Electronic Thesis/Dissertation
 

Three Essays on Offline Learning, In-Context Learning Calibration, and Explainable AI

Open Access Deposited

an explanation is faithful if an independent interpreter can reconstruct the black-box model’s predictive distribution from the input and the explanation. We operationalize this idea with a judge LLM and a Kullback--Leibler (KL)-consistent objective that measures agreement between the black-box distribution and the judge-implied distribution, and we train an explainer LLM using a GRPO-style reinforcement learning procedure. Preliminary experiments show that the KL-based signal provides a principled foundation for faithfulness in language space, but can induce reward-hacking when optimized in isolation, motivating auxiliary regularization and an ongoing extension based on mechanistic interpretability. Taken together, the three essays develop a coherent perspective on reliable data-driven decision making under real-world informational constraints. More broadly, the dissertation shows how classical ideas from dynamic programming and statistical learning theory can be adapted to contemporary problems in offline learning, LLM calibration, and explainable AI, yielding methods that are not only accurate, but also structurally valid, robust to practical limitations, and aligned with how humans evaluate and use model outputs.

Data-driven decision making increasingly relies on models and algorithms that are powerful, adaptive, and scalable, yet often difficult to design, evaluate, and trust in real-world settings. Across many applications, decision makers must act under incomplete information and operational constraints while producing outputs that are useful and reliable for human stakeholders. A recurring difficulty is that the information directly observed in practice is often not the same as the information that should ideally guide learning, optimization, or evaluation. In the three essays of this dissertation, that gap appears in different forms

realized sales can hide true demand under stockouts, in-context predictions of large language models (LLMs) can be unstable and systematically biased despite strong few-shot performance, and natural-language explanations of black-box machine learning models can be plausible without faithfully reflecting the model's actual decision making process. This dissertation develops methodological frameworks that make these gaps explicit and improve performance and reliability under the resulting informational constraints. The first essay studies offline sequential feature-based joint pricing and inventory control under censored and dependent demand. In lost-sales environments, firms observe realized sales but not excess demand during stockouts, and this missing information is especially consequential when demand is temporally dependent. We show that censoring creates missing reward information and can destroy the Markov property in the observed state, even when the underlying uncensored system is Markovian. To address this, we develop a general modeling framework that separates the underlying decision process from the observed censored process and show that the observed process inherits a finite high-order Markov structure indexed by consecutive censoring. Building on this result, we derive specialized Bellman equations, propose two offline learning algorithms---Censored Fitted Q-Iteration (C-FQI) and Pessimistic Censored Fitted Q-Iteration (PC-FQI)---and establish finite-sample regret guarantees with interpretable components. Numerical experiments illustrate the effectiveness of the proposed methods and the value of censoring-aware policy learning. The second essay studies calibration of in-context learning (ICL) predictions in LLMs. Although LLMs exhibit strong few-shot classification performance, their predictions often suffer from systematic biases, leading to unstable performance in classification. While calibration techniques are proposed to mitigate these biases, we show that many of these methods can be understood, in logit space, as shifting the base LLM’s decision boundary without changing its orientation, which limits correction under severe misalignment. Motivated by this, we propose Supervised Calibration (SC), a loss-minimization framework that learns per-class affine transformations of LLM logits using surrogate training data generated from the in-context examples. The framework incorporates context-invariance and directional trust-region regularization, and supports ensembling across context sizes and sub-contexts. Theoretical analysis clarifies how SC generalizes prior calibration methods, and empirical results across multiple LLMs and benchmark classification tasks show substantial and consistent gains. The third essay studies faithful natural-language explanations for black-box machine learning models. One of the central challenges in natural-language based explainable AI is that explanations can be convincing and plausible while failing to be faithful. In particular, they might not necessarily reflect the actual decision making process of the black-box model. We address this by proposing a distribution-preserving notion of faithfulness

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Gundem_gwu_0075A_17804.pdf Gundem_gwu_0075A_17804.pdf 2026-06-24 Open Access