Electronic Thesis/Dissertation
 

A Case Study of TripAdvisor Reviews in Juneau, Alaska

Open Access Deposited

Investigating The Impact of LLM-Generated Data on Broad-Based Sentiment Classification

10% zero-shot, 20% one-shot, and 70% few-shot. In this case study, we find that integrating a limited proportion of synthetic data reviews (roughly 20%) into the training set leads to marginal improvement in precision, recall, and F1-scores for unseen real-world data in the minority classes. However, as we increase the amount of LLM-generated data – with reviews of low linguistic diversity – the performance on the unseen real-world data in the minority classes degrades over time (with the lowest performance at 100% LLM-generated data). These results demonstrate the trade-off between data augmentation and data quality and highlight that linguistic diversity is a key factor in determining the utility of synthetic data and the classification performance on real-world data. We use this analysis to show that this is a proof-of-concept study that demonstrates that LLM-generated data holds significant potential to augment and strengthen training datasets in the tourism & hospitality industries. In the context of tourism and hospitality, this work has direct implications for advancing the detection of tourist feedback, enabling the growth of actionable insights that key decision makers can leverage to improve businesses’ performance. By doing so, this work contributes to a core research question in the tourism and hospitality industries – how digital tools, such as user-generated content and tourism analytics, can be enhanced by LLM-generated data to influence tourist analytics, decision-making, and co-creation.

In this study, we examine the ability of large language model (LLM)-generated synthetic data (artificial data reviews that mimic real-world data) to act as a partial substitute for real-world data in tourism sentiment analysis. Specifically, we investigate how varying proportions of synthetic data affect the evaluation of unseen real-world data in supervised broad-based sentiment classification for the minority classes. We further examine the role of linguistic diversity (complex linguistic features) in synthetic data and analyze how in-context learning (e.g., zero-shot, one-shot, few-shot) affects the diversity and quality of synthetic data reviews. We use the Llama3.1 8B and Claude Opus 4.6 large language models to generate synthetic data. We employ a mixture of in-context learning techniques at the following ratio

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Sharda_gwu_0075M_17934.pdf Sharda_gwu_0075M_17934.pdf 2026-06-24 Open Access