Electronic Thesis/Dissertation
 

Quantile Regression for Censored Data and Spatial Data

Open Access Deposited

For a random variable with a cumulative distribution function, the quantile at a given level represents the smallest value at which the distribution reaches that probability level. Quantile regression models these conditional quantiles given a set of covariates. Unlike mean regression, quantile regression is inherently robust to heavy-tailed distributions, outliers, and departures from parametric assumptions, as it characterizes conditional distributional features beyond the mean. Moreover, quantile regression is particularly well suited for settings in which tail heterogeneity is present, allowing covariate effects to vary across different parts of the conditional distribution rather than being summarized by a single average effect. For example, when studying the relationship between years of education and income, conventional mean regression methods such as ordinary least squares may yield estimates similar to those from median regression. However, education may have little association with income among low-income individuals while exerting a much stronger effect in the upper tail of the income distribution, reflecting pronounced heterogeneity between lower and upper quantiles that cannot be captured by mean-based approaches. In biomedical studies, observational data frequently involve early dropout, leading to nonignorable missingness. In such settings, inference based on conditional mean regression is fundamentally limited, since the conditional mean is generally not identifiable without strong parametric assumptions on either the outcome distribution or the missingness mechanism. Beyond capturing tail heterogeneity, quantile regression provides a robust alternative for analyzing missing data, as conditional quantiles remain identifiable for a range of quantile levels under censoring and are therefore less sensitive to missing observations. For many biomedical applications, it is of interest to estimate bent-line regression models, in which the target relationship is piecewise linear and continuous across regions defined by unknown change points. However, censoring can induce substantial bias in both slope and change-point estimation for existing mean-based approaches. Motivated by an experimental autoimmune myasthenia gravis dataset, Chapter 1 proposes a censored bent-line quantile regression framework that formulates the problem directly in terms of identifiable quantile-based functionals and avoids parametric assumptions on the missingness mechanism. The resulting estimator is consistent, asymptotically normal, and easy to implement using an informative subset of the data. Simulation studies and real-data analysis demonstrate substantial bias reduction under nonignorable missingness. In addition to tail heterogeneity, spatial data often exhibit substantial spatial heterogeneity. Existing spatial methods frequently struggle to capture heterogeneous patterns over complex domains or ignore distributional heterogeneity in the tails of the response. Chapter 2 introduces a quantile spatial modeling framework that accommodates both spatial nonstationarity and tail heterogeneity through constant and spatially varying coefficients. A smoothed quantile bivariate triangulation method is proposed based on penalized splines on triangulations combined with smoothing of the quantile loss. The method effectively captures spatial nonstationarity while preserving important data features such as smoothness and shape over complex and irregular domains. Under mild regularity conditions, the estimator achieves optimal convergence rates. A Bahadur representation is established, enabling asymptotic normality for the constant coefficient estimator and facilitating the construction of confidence intervals. A wild bootstrap procedure is also developed to improve finite-sample performance. Simulation studies demonstrate the numerical and computational advantages of the proposed method, and application to United States mortality data reveals how socioeconomic factors influence mortality rates differently across spatial regions and distributional tails. Modern spatial datasets are often massive in both volume and resolution, frequently exceeding the memory capacity of a single machine and posing substantial challenges for scalable statistical modeling. While the proposed method in Chapter 2 performs well for spatial data of moderate size, it does not scale efficiently to large datasets. Chapter 3 proposes a distributed inference framework for quantile spatial modeling that accommodates both constant and spatially varying coefficients through triangulation-based smoothing. The framework is built on a surrogate loss induced by smoothed and decorrelated scores and employs a communication-efficient multi-round aggregation strategy. It achieves the global convergence rate for constant coefficients without requiring large local sample sizes and remains robust even when spatially varying components are imperfectly estimated on local machines. Asymptotic normality is established for the constant coefficient estimator, along with corresponding inference procedures. Simulation studies highlight the numerical performance of the method, and its practical utility is demonstrated through an application to large-scale United States census-tract-level data on coronary heart disease prevalence.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Lin_gwu_0075A_17844.pdf Lin_gwu_0075A_17844.pdf 2026-06-24 Open Access