BanditRLlib
Lean gate passed before this site build; local proof declarations are shown as compiled.

Lean module · ETC

BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence

# ETC centered reward-difference independence transfer This module transfers time-coordinate independence of a stochastic reward trace through the deterministic centered pairwise reward-difference transform used by the ETC wrong-commit tail route.

Module map

Teaching chapter
3. Explore-Then-Commit
Declarations
1
Placeholders
0

Imports

BanditRLProof.Algorithms.ETCSumRewardsDiff

Imported by

BanditRLProof, BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian

Declarations

Open an item to read its exact compact statement and source link. Detailed teaching notes are linked when registered.

theorem BanditRLProof.ETC.iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward Compiled

If the reward trace coordinates are independent across time, then the centered pairwise reward-difference summands are independent across time for every non-best arm. This is the `ETC-CENTERED-DIFF-INDEPENDENCE-WITNESS` leaf. It only proves the deterministic-transform part of the reward-law independence obligation. It does not prove the reward trace coordinates are independent from a kernel or environment model, and it does not prove any sub-Gaussian witness.

theorem iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu