Secure Data Handling for Large Language Models (LLMs) and Artificial Intelligence (AI)

Authors

  • Jacob Thomas

Abstract

The development of Large Language Models (LLMs), particularly the generative models and AI technologies, has changed many fields, paving the way with far-reaching effects in enhancing natural language understanding and generation, empowering many organizations to transform from traditional to learning organizations. But such developments also pose a plethora of challenges, especially to the secure handling of the data used for training, considering the sensitive nature of the data itself and the privacy threats posed to AI training and deployment across geographies and domains. This dissertation investigates the principles, methodologies, and technologies required for secure data management in respect of LLMs, AI systems, among the rational agents currently available in the market. This work investigates training data from the perspective of data privacy and governance mechanism in place with the help of secondary research and recommends an inclusive framework of training data being recorded in Blockchain for guaranteeing secure data management and identifying data leakage through the entire lifecycle of AI learning and output generation. Most of these training data is obtained from webpages scrapings of the crawlers from internet and from online data sources. Secure data handling is essential and crucial to keep its users' privacy protected, preserve the trust, and adhere to regulatory requirements like GDPR and HIPAA and unfortunately, this aspect is lacking when learning happens by defining the LLMs on the market based on their exposure against the datasets
This dissertation also explores if DLT (Distributed Ledger Technology) and hashing which is a backbone of Blockchain technology can be a solution to the workaround of Data leakage identification and fine tuning of the LLM model. Concept of Digital Twin and the possibility of Digital Twin used for feedback training / fine tuning has been visited in this work. This dissertation is intended to prove that training data entry in Blockchain ledger is a solution to data leakage and can significantly capture instances of data leakage by inherently identifying the trained data in a hashed format and act as a backward loop in fine tuning the model to eradicate the occurrence of generating data in an unintentional manner.

Downloads

Published

2026-06-19

How to Cite

Thomas, J. (2026). Secure Data Handling for Large Language Models (LLMs) and Artificial Intelligence (AI). Digital Repository of Theses. Retrieved from https://repository.learn-portal.org/index.php/rps/article/view/1291