Electronic Thesis/Dissertation
 

Understanding Reasoning Mechanisms in Large Language Models Through Direction Learning

Open Access Deposited

This thesis explores directional learning as a framework to understand and manipulate reasoning mechanisms within large language models (LLMs). Current LLMs demonstrate impressive linguistic proficiency but often lack robust conceptual reasoning, particularly in ethical decision-making. This study hypothesizes that such reasoning can be represented as multidimensional vectors within the activation space of LLMs, termed "directions," which encode complex behaviors like ethical refusal, qualitative reasoning, and quantitative problem-solving. Leveraging techniques such as Singular Value Decomposition (SVD), this research proposes a computational framework to identify, analyze, and manipulate these directions. Experimental results across various LLM architectures reveal that directional learning enables scalable and efficient characterization of refusal behaviors and reasoning patterns. The study demonstrates that refusal behaviors are dynamically encoded and that reasoning tasks are hierarchically structured, with intermediate model layers playing a critical role in abstract conceptual encoding. Contributions include: (1) a scalable SVD-based method for extracting key directional representations, (2) insights into the temporal dynamics of refusal mechanisms, (3) robust classification of qualitative and quantitative reasoning tasks, and (4) implications for improving interpretability, alignment, and robustness in LLMs. This work advances the understanding of LLM reasoning mechanisms, offering pathways for enhancing AI safety and cognitive capabilities. Future directions involve exploring adversarial robustness, transferability of directional representations, and dynamic interventions for real-time ethical alignment.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of AitHou_gwu_0075M_17155.pdf AitHou_gwu_0075M_17155.pdf 2025-04-09 Open Access