This thesis consists of three parts: statistical inference for additive monotone models, sharp regret bounds for multi-armed bandits and a review of instance-optimality of online reinforcement learning. First, we investigate statistical inference for least squares estimators of additive monotone models. We prove local limiting distribution theory for least squares estimators under a fixed lattice design, and establish the joint asymptotic independence. We also derive pivotal limiting distributions for constructing tuning-free confidence intervals. Second, we analyze multi-armed bandits, focusing on upper confidence bound (UCB) allocation rules. We establish sharp non-asymptotic regret bounds for Lai's UCB under Gaussian reward assumptions and propose a new proof schema based on nonlinear renewal theory that yields improved regret bounds under weaker distributional assumptions. Finally, we review recent progress on instance optimality in online reinforcement learning, unifying asymptotic lower bounds for both infinite-horizon average-reward and finite-horizon episodic Markov decision processes. We review information lower bounds and instance-optimal algorithms, and highlight the common algorithmic designs and the core proof techniques for establishing instance optimality. Our theoretical contributions advance the understanding of nonparametric shape-constrained inference, sharpen regret guarantees for multi-armed bandits, and offer a systematic perspective on instance-optimal reinforcement learning.
Huachen Ren (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: