Advances in Functional Data Analysis and Reinforcement Learning
Open AccessIn the era of Big Data, functional data analysis (FDA) and reinforcement learning (RL) have received increasing attentions in many scientific and engineering fields. FDA handles data in the form of curves, surfaces, volumes, etc. while RL deals with data-driven sequential decision making. Despite their rapid developments in the past few decades, some fundamental problems remain unsolved, including how to test the independence between two random functions without the risk of model misspecification and how to learn the evaluate a decision when the Markovian assumption in RL is violated. In this dissertation, we study the following two topics: (1) an independence test for bivariate functional data, and (2) off-policy evaluation for episodic partially observable Markov decision processes under non-parametric models.Part I. Measuring and testing the dependency between multiple random functions is often an important task in FDA. In the literature, a model-based method relies on a model which is subject to the risk of model misspecification, while a model-free method only provides a correlation measure which is inadequate to test independence. In this work, we adopt the Hilbert-Schmidt Independence Criterion (HSIC) to measure the dependency between two random functions. We develop a two-step procedure by first pre-smoothing each function based on its discrete and noisy measurements and then applying the HSIC to recovered functions. To ensure the compatibility between the two steps such that the effect of the pre-smoothing error on the subsequent HSIC is asymptotically negligible, we propose to use wavelet soft-thresholding for pre-smoothing and Besov-norm-induced kernels for HSIC. We also provide the corresponding asymptotic analysis. The superior numerical performance of the proposed method over existing ones is demonstrated in a simulation study. Moreover, in an magnetoencephalography data application, the functional connectivity patterns identified by the proposed method are more anatomically interpretable than those by existing methods. Part II. We study the problem of off-policy evaluation (OPE) for episodic Partially Observable Markov Decision Processes (POMDPs) with continuous states. Motivated by the recently proposed proximal causal inference framework, we develop a nonparametric identification result for estimating the policy value via a sequence of so-called V-bridge functions with the help of time-dependent proxy variables. We then develop a fitted-Q-evaluation-type algorithm to estimate V-bridge functions recursively, where a non-parametric instrumental variable (NPIV) problem is solved at each step. By analyzing this challenging sequential NPIV estimation, we establish the finite-sample error bounds for estimating the V-bridge functions and accordingly that for evaluating the policy value, in terms of the sample size, length of horizon and so-called (local) measure of ill-posedness at each step. To the best of our knowledge, this is the first finite-sample error bound for OPE in POMDPs under non-parametric models.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.