Advances in Subgroup Identification and Expected Shortfall Regression
Open AccessThis dissertation focuses on developing new statistical methods in three related directionsto capture heterogeneity and tail features from data. In the first project, we develop a generalized quantile tree method for subgroup identification. The method employs quantile rank score tests for split variable selection and utilizes composite quantile loss for associated split point estimation. The proposed split rule is free of variable selection bias and robust against outliers and heavy-tailed distributions. In addition, we introduce a generalized quantile treatment effect test for the selection and confirmation of predictive subgroups. Numerical studies suggest that the proposed method gives more accurate subgroup identification than existing methods for cases with heteroscedastic or heavy-tailed errors. The second and third projects focus on statistical inference for regression of expected shortfall, a commonly used risk measure that is related to quantiles while with distinguished features. In the second project, we consider the joint modeling of linear conditional quantile and expected shortfall at the same probability level. A two-step estimation procedure is proposed to reduce the computational effort. We show that the resulting two-step estimator is asymptotically equivalent to the joint estimator, but the former is numerically more efficient. We further develop a score-type inference method for hypothesis testing and confidence interval construction. The proposed score-type method is superior to the Wald-type method in finite samples, especially for cases with a large number of confounding factors and heterogeneous errors. We extend the joint regression framework to partially linear varying coefficient (PLVC) models for the analysis of longitudinal data in the third project. The functional coefficients are estimated by B-spline approximations. We adapt the two-step estimation method introduced in Section 3.2.2 and consider two types of loss functions for the joint modeling of conditional quantile and expected shortfall. The estimation procedure is easy to implement, and it requires no specification of the error distributions or intra-subject dependence structure. The asymptotic properties of the proposed estimators are established for both varying and constant coefficients. We develop rank score-type tests for hypotheses on the constant coefficients, and for testing the constancy of a subset of varying coefficients. We assess the finite sample performance of the proposed methods by Monte Carlo simulation studies, and demonstrate their practical value by analyzing an income data set.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.