Electronic Thesis/Dissertation
 

Turning Visual Designs into Code with Generative AI and Multimodal Language Models

Open Access Deposited

A critical but resource-intensive process in software development is the transformation of visual designs into functional, maintainable code. Despite significant advancements in artificial intelligence (AI) and multimodal large language models (LLMs), the currently available solutions and existing approaches to automating the design-to-code process struggle with generalization, accuracy, structural fidelity, and computational efficiency and remain inadequate. This praxis presents a comprehensive multimodal generative framework specifically developed to automate the conversion of high-fidelity visual designs into industry-standard frontend code. The central focus of this research is the creation and validation of TruthfulPixel, a novel benchmark dataset comprising 7621 high-quality, annotated design-code pairs. The TruthfulPixel dataset integrates and enhances existing resources, such as WebSight and Design2Code, to provide robust, diverse, and representative scenarios for training and benchmarking multimodal generative models. The framework proposed in this research leverages state-of-the-art multimodal generative AI techniques that specifically target fine-tuning methodologies utilizing LoRA adapters and parameter-efficient training strategies. The models evaluated included GPT-4o-mini, llama-3.2-Vision, llama-3b, llava-v1.6-mistral-7b, and qwen2.5vl, which were selected for their balanced complexity and feasibility within the resource constraints. A comprehensive evaluation framework was developed that encompasses critical metrics, such as syntax correctness, maintainability indices, structural fidelity (TreeBLEU and SSIM), and real-world computational efficiency (Page Load Time and Largest Contentful Paint). Empirical results demonstrate substantial improvements in accuracy, maintainability, structural fidelity, and computational efficiency through fine-tuning with TruthfulPixel data. Specifically, the fine-tuned models exhibited reductions in syntax errors, measurable enhancements in maintainability, and consistent performance across various UI complexity levels. Furthermore, incorporating image modality alongside textual inputs enhanced the overall code quality, validating the robustness and efficacy of multimodal integration. The practical implications of this research are extensive, offering the potential for fully automated, efficient, and accurate design-to-code translation workflows. Additionally, the established dataset and evaluation framework contribute to future research by promoting reproducibility and transparency method. Despite its notable strengths, this research acknowledges several limitations, including dataset generalizability, computational resource constraints, metric comprehensiveness, and the exclusion of direct human-generated code comparisons. Future research directions include enhancing dataset diversity, exploring additional computational optimizations, integrating user-centric evaluations, and innovating multimodal approaches. This praxis advances the theoretical understanding and practical application of multimodal generative AI in software development, laying a robust foundation for continued innovation and adoption in the field of automated front-end engineering.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Zhang_gwu_0075A_17400.pdf Zhang_gwu_0075A_17400.pdf 2025-12-12 Open Access