Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 1 - 30 of 129

Full-Text Articles in Databases and Information Systems

Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev Dec 2025

Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev

Theses and Dissertations

DataOps has been coined as a novel term that emerged as a synthesis of data management practices with software engineering concepts, such as DevOps and Agile, with the goal of improving data quality and governance in enterprises. The proliferation of scratch table use and transformation tools, such as dbt, has led to an exponential increase in the number of data models, which complicates standardization efforts and increases maintenance overhead. Although the market is saturated with various flavors of text-to-SQL engines that promote increased productivity and self-service use in organizations, there are limited tools available to optimize individual queries, enforce consistency, …


Blockchain-Enabled Master Data Management, Shakhawat Hossain May 2025

Blockchain-Enabled Master Data Management, Shakhawat Hossain

Theses and Dissertations

Master Data Management (MDM) is essential for maintaining data quality, accuracy, consistency, and governance within organizations. However, traditional centralized MDM systems continue to face challenges related to data integrity, security, and scalability. This research presents a blockchain-enabled MDM framework designed to overcome these limitations by leveraging blockchain’s decentralized, immutable, and secure architecture. The study aims to identify and address the shortcomings of conventional MDM practices, examine the applicability of blockchain technology in enhancing these systems, and develop a functional prototype to validate the proposed model. The framework incorporates decentralized review mechanisms that improve auditability and ensure trusted data verification by …


Qlorax: Heuristic-Guided Fine-Tuning Of Llama-2 For Domain Adaptation In Entrepreneurship, Gaurob Saha Jan 2025

Qlorax: Heuristic-Guided Fine-Tuning Of Llama-2 For Domain Adaptation In Entrepreneurship, Gaurob Saha

Theses and Dissertations

This thesis presents a study on the fine-tuning of large language models (LLMs) for domain-specific applications using limited data. We fine-tuned the LLaMA-2 (7B) model on a curated entrepreneurial dataset containing 3,545 human-written question-answer pairs, of which 3,095 were used for training and 450 were reserved for evaluation. A complete fine-tuning and evaluation pipeline was developed, which included clustering human-written answers, generating centroid-based summaries for each cluster, and evaluating the model's generated responses through cosine similarity.

Training was carried out over five epochs, with model performance evaluated after each epoch. The fine-tuned model demonstrated strong semantic alignment with human-written content, …


Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn Nov 2024

Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn

Theses and Dissertations

In the ever-expanding landscape of digital technologies, the exponential growth of data presents both challenges and opportunities, demanding innovative approaches to data curation. Effective data curation is pivotal for extracting meaningful insights from vast and complex datasets. This study explores the integration of spectral clustering and Shannon Entropy within the Data Washing Machine (DWM), a novel tool designed to streamline unsupervised data curation processes. The DWM incorporates Shannon Entropy into its clustering process, allowing for adaptive refinement of clustering strategies based on entropy levels observed within data clusters. Spectral clustering, known for its ability to handle complex and non-linearly separable …


Increasing The Robustness Of Machine Learning By Adversarial Attacks, Gourab Mukhopadhyay Jul 2024

Increasing The Robustness Of Machine Learning By Adversarial Attacks, Gourab Mukhopadhyay

Theses and Dissertations

By perturbation or physical attacks any machine can be fooled into predicting something else other than the intended output. There are training data based on which the model is trained to predict unknown things. The objective was to create noises and shades of different levels on the images and do experiments for measuring accuracy and making the model classify the traffic signs. When it comes to adding shades to the pictures, pixels were modified for three different layers of the pictures. The experiment also shows that with the shadows getting deeper, the accuracies drop significantly. Here, some changes in pixels …


Deep Learning In Indus Valley Script Digitization, Deva Munikanta Reddy Atturu May 2024

Deep Learning In Indus Valley Script Digitization, Deva Munikanta Reddy Atturu

Theses and Dissertations

This research introduces ASR-net(Ancient Script Recognition), a groundbreaking system that automatically digitizes ancient Indus seals by converting them into coded text, similar to Optical Character Recognition for modern languages. ASR-net, with an 95% success rate in identifying individual symbols, aims to address the crucial need for automated techniques in deciphering the enigmatic Indus script. Initially Yolov3 is utilized to create the bounding boxes around each graphemes present in the Indus Valley Seal. In addition to that we created M-net(Mahadevan) model to encode the graphemes. Beyond digitization, the paper proposes a new research challenge called the Motif Identification Problem (MIP) related …


Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen May 2024

Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen

Theses and Dissertations

This dissertation explores applications of representation learning and generative models to challenges in healthcare, astronautics, and aviation.

The first part investigates the use of Generative Adversarial Networks (GANs) to synthesize realistic electronic health record (EHR) data. An initial attempt at training a GAN on the MIMIC-IV dataset encountered stability and convergence issues, motivating a deeper study of 1-Lipschitz regularization techniques for Auxiliary Classifier GANs (AC-GANs). An extensive ablation study on the CIFAR-10 dataset found that Spectral Normalization is key for AC-GAN stability and performance, while Weight Clipping fails to converge without Spectral Normalization. Analysis of the training dynamics provided further …


Modeling & Engineering Of Usmepcom Business Intelligence Data, Merrick A. Bedford Mar 2024

Modeling & Engineering Of Usmepcom Business Intelligence Data, Merrick A. Bedford

Theses and Dissertations

This thesis investigates the USMEPCOM’s issue of modeling and engineering Business Intelligence data centralized around the MEPS of Excellence (MOE) program. MEPS around the US conduct military personnel in-processing and in doing so have a vested interest in the standardization and application of the data associated with such processes. There are currently 65 MEPS stations and one RPS that handle military personnel onboarding paperwork and make determinations for military eligibility. This topic is important due to the MEPS cloud data processing system modernization efforts requiring data processing adaptation to ensure relevant and meaningful usage of current data.


Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua Jan 2024

Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua

Theses and Dissertations

Entity resolution (ER) plays a pivotal role across domains by enabling data integration and quality improvement. This dissertation delves into the evolving landscape of ER, introducing innovative approaches that redefine this fundamental task. The first contribution is TriBERTa, a novel representation learning model tailored for ER. TriBERTa sets new benchmarks in entity matching and demonstrates versatility across ER processes like data blocking and resolution. Empirical evaluations on diverse datasets showcase TriBERTa’s superior performance over existing representations, including from large language models. The second contribution explores the use generative language models like GPT-3.5 and Dolly 2.0 for cross-domain entity matching using …


A Conceptual Decentralized Identity Solution For State Government, Martin Duclos Dec 2023

A Conceptual Decentralized Identity Solution For State Government, Martin Duclos

Theses and Dissertations

In recent years, state governments, exemplified by Mississippi, have significantly expanded their online service offerings to reduce costs and improve efficiency. However, this shift has led to challenges in managing digital identities effectively, with multiple fragmented solutions in use. This paper proposes a Self-Sovereign Identity (SSI) framework based on distributed ledger technology. SSI grants individuals control over their digital identities, enhancing privacy and security without relying on a centralized authority. The contributions of this research include increased efficiency, improved privacy and security, enhanced user satisfaction, and reduced costs in state government digital identity management. The paper provides background on digital …


Pattern-Of-Life Modeling With Automatic Dependent Surveillance-Broadcast (Ads-B), Sarah J. Bolton Sep 2023

Pattern-Of-Life Modeling With Automatic Dependent Surveillance-Broadcast (Ads-B), Sarah J. Bolton

Theses and Dissertations

This dissertation and research were sponsored by the Air Force Research Laboratory Layered Sensing Exploitation Branch (AFRL/RYA) to investigate the utility of using the data found within aircraft secondary radar to make predictions about aircraft characteristics and intent. The research focuses on making predictions on aircraft characteristics using only the kinetic data within one type of secondary radar, Automatic Dependent Surveillance-Broadcast (ADS-B), as a surrogate for primary radar. The results from this research provide a means to reduce the reliance on a type of aircraft tracking that is vulnerable to cyber attack and other integrity concerns.


Analysis And Optimization Of Contract Data Schema, Franklin Sun Mar 2023

Analysis And Optimization Of Contract Data Schema, Franklin Sun

Theses and Dissertations

agement, development, and growth of U.S Air Force assets demand extensive organizational communication and structuring. These interactions yield substantial amounts of contracting and administrative information. Over 4 million such contracts as a means towards obtaining valuable insights on Department of Defense resource usage. This set of contracting data is largely not optimized for backend service in an analytics environment. To this end, the following research evaluates the efficiency and performance of various data structuring methods. Evaluated designs include a baseline unstructured schema, a Data Mart schema, and a snowflake schema. Overall design success metrics include ease of use by end …


A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson Feb 2023

A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson

Theses and Dissertations

The Data Washing Machine (DWM) is a known and documented open-source Python Jupyter Notebook project that is the foundation for an Unsupervised Data Curation process. The DWM ingests reference data without a prior data cleansing activity and ultimately runs Entity Resolution (ER) on acceptable entity data to cluster duplicate references within the dataset. The DWM currently has 17 modifiable parameters that are used to help tokenize, cleanse, organize, link and cluster like references. With such a large number of parameters, some type of beginning settings as optimal as possible are needed for the DWM process for it to be useful …


The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique Jan 2023

The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique

Theses and Dissertations

The main objective of Entity resolution (ER) is to find duplicate records within the same data table from the same source or different data tables from various sources. A traditional pair-wise supervised entity resolution matching depends on pre-built rules for finding matched records. On the other hand, unsupervised or semisupervised also relies on pair-wise matching. In the maximum case, group membership is left behind for consideration. In this dissertation, I have discussed the design, implementation and evaluation of a graph-based entity resolution for group membership to enhance the pair-wise matching ER system. I have designed and implemented a pipeline for …


Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed Jul 2022

Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed

Theses and Dissertations

In an era, where data is being used by organizations in achieving their business goals and driving their business decisions, it is key to ensure the quality of data. For an organization, the most important data assets include master data assets such as product master data and customer master data, supplier master data, employee master data which can be generalized as party master data. There has been significant growth and variety in data in recent years because of which the traditional rules-based approach and dependency on data experts for assessing data quality is no longer working. This dissertation evaluates and …


Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira Jun 2022

Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira

Theses and Dissertations

For the last two decades, the cloud computing ecosystems has become a major defining force for America, because of their unique economic, social and national importance. These ecosystems have now taken their place alongside the nation’s other infrastructure such as the food/agricultural, energy healthcare, and roads/highways. Scientists, engineers, and researchers over the same period have experienced a tremendous growth in the need for resources to that support assorted research. Cloud computing enables these communities to undertake wide-ranging research efforts, while requiring no maintenance, management or significant invest of local resources. The genius of this project examines how create on-ramps for …


The Applications Of The Internet Of Things In The Medical Field, Cody Repass May 2022

The Applications Of The Internet Of Things In The Medical Field, Cody Repass

Theses and Dissertations

The Internet of Things (IoT) paradigm promises to make “things” include a more generic set of entities such as smart devices, sensors, human beings, and any other IoT objects to be accessible at anytime and anywhere. IoT varies widely in its applications, and one of its most beneficial uses is in the medical field. However, the large attack surface and vulnerabilities of IoT systems needs to be secured and protected. Security is a requirement for IoT systems in the medical field where the Health Insurance Portability and Accountability Act (HIPAA) applies.

This work investigates various applications of IoT in healthcare …


Bayesian Convolutional Neural Network With Prediction Smoothing And Adversarial Class Thresholds, Noah M. Miller Mar 2022

Bayesian Convolutional Neural Network With Prediction Smoothing And Adversarial Class Thresholds, Noah M. Miller

Theses and Dissertations

Using convolutional neural networks (CNNs) for image classification for each frame in a video is a very common technique. Unfortunately, CNNs are very brittle and have a tendency to be over confident in their predictions. This can lead to what we will refer to as “flickering,” which is when the predictions between frames jump back and forth between classes. In this paper, new methods are proposed to combat these shortcomings. This paper utilizes a Bayesian CNN which allows for a distribution of outputs on each data point instead of just a point estimate. These distributions are then smoothed over multiple …


A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye Sep 2021

A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye

Theses and Dissertations

Data governance is mission critical. Organizations start to realize the data and information are very important assets. Data governance program could minimize the risk of data lose and maximize the data values. As existing best practices to implement data governance program, most of these practices rely on management strategies, which focus on data governance literacy education, enforcing data policy, standard and processes, etc. However, IT infrastructure strategy has not changed too much to support data governance achievement. Data governance requirements for system design usually is the last concern on the list. Most of the data governance programs collect metadata or …


Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi Apr 2021

Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi

Theses and Dissertations

This dissertation describes a first attempt to build a data washing machine, a system able to take dirty data and through an unsupervised process, output clean data. The washing machine design described here focuses on two main aspects of the data curation process, token correction and data redundancy. It aims to simplify and automate the preparation of data used to create information products. In this approach, all these steps would be automated, thus saving the time and effort of the data analysts who ordinarily perform these actions. In other words, this is the opposite of the current approach to first …


Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy Jan 2021

Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy

Theses and Dissertations

Social media plays an important role in the propagation and dissemination of ideas and thoughts. Compared to other social media platforms, blogs provide a convenient platform for users to post detailed information, engage in active discussions and share the content on other social media sites, such as Facebook and Twitter. Thus, the blogosphere has been an enormous and ever-growing part of the open-source intelligence. In order to track and monitor online social behavior particularly from blogs, the first challenging part is to mine the vast pool of unstructured data. To scale up this process and cope with the continuously changing …


Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon Jan 2021

Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon

Theses and Dissertations

Machine learning models for chemical property predictions are high dimension design challenges spanning multiple disciplines. Free and open-source software libraries have streamlined the model implementation process, but the design complexity remains. In order better navigate and understand the machine learning design space, model information needs to be organized and contextualized. In this work, instances of chemical property models and their associated parameters were stored in a Neo4j property graph database. Machine learning model instances were created with permutations of dataset, learning algorithm, molecular featurization, data scaling, data splitting, hyperparameters, and hyperparameter optimization techniques. The resulting graph contains over 83,000 nodes …


Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai Dec 2020

Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai

Theses and Dissertations

The development of web technologies has progressed rapidly in the past few decades. As a result, these technologies have increasingly gained popularity among researchers as an instrument for data analysis and visualization to aid their work. The purpose of this project is to pursue a new method for mining time-based data to perform curve fitting, using web client technology. This requires comparing the advantages and disadvantages of different JavaScript software designs to determine which is best. Based on the results, we emphasize developing a client-side web application and demonstrate its use through fitting a curve for a new metric to …


Barriers And Drivers Influencing The Growth Of E-Commerce In Uzbekistan, Madinakhon Tursunboeva Nov 2020

Barriers And Drivers Influencing The Growth Of E-Commerce In Uzbekistan, Madinakhon Tursunboeva

Theses and Dissertations

Electronic commerce (e-commerce) has become a major retail channel for businesses in developed countries. However, it is still considered an innovation in developing countries. Specifically, e-commerce in Uzbekistan is in the early stages of emergence despite its advance in recent years in terms of Internet penetration, a strong retail sector, new national regulations, and a young population. The study aimed to identify barriers and drivers influencing e-commerce growth in Uzbekistan. A Delphi research design was utilized to answer the research questions of the study, which categorized and ranked factors that Uzbekistani entrepreneurs are facing when engaging in e-commerce processes. A …


A Framework To Determine The Impact Of Poor Data Quality On The Reliability Of Iot Sensor-Based Real-Time Decision-Making, Arnold Rego Nov 2020

A Framework To Determine The Impact Of Poor Data Quality On The Reliability Of Iot Sensor-Based Real-Time Decision-Making, Arnold Rego

Theses and Dissertations

Today, more and more systems make real-time decisions based on the data received from sensors. But it is inevitable that sensors will fail, or send bad data – leading to faulty real-time decision-making. Using this poor-quality data in real-time decision-making can lead to deadly consequences. Traditional Data Quality (DQ) foundations call for “fixing” the poor DQ before the data can be used. However, time is often of the essence in a real-time decision-making environment, leaving little or no time to “fix” poor-quality data before it can be used in the decision-making process. The primary objective in real-time decision-making is not …


Privacy And The Digital Divide: Investigating Strategies For Digital Safety By People Of Color, Denavious Hoover Oct 2020

Privacy And The Digital Divide: Investigating Strategies For Digital Safety By People Of Color, Denavious Hoover

Theses and Dissertations

People of color are becoming increasingly concerned with digital privacy. They are concerned about the obfuscated data collection and sharing practices of major social media plat- forms and the strong entitlement of other users in the online space to their content. This study examines how people of color conceptualize and behave to produce safety in the online space, or, in other words, digital privacy. This study challenges notions that people are not purposeful about privacy in the online space and highlights the voices of people of color, whom are not of- ten included in theorizing or decision making about the …


A Methodology To Identify Alternative Suitable Nosql Data Models Via Observation Of Relational Database Interactions, Paul M. Beach Sep 2020

A Methodology To Identify Alternative Suitable Nosql Data Models Via Observation Of Relational Database Interactions, Paul M. Beach

Theses and Dissertations

The effectiveness and performance of data-intensive applications are influenced by the suitability of the data models upon which they are built. The relational data model has been the de facto data model underlying most database systems since the 1970’s. However, the recent emergence of NoSQL data models have provided users with alternative ways of storing and manipulating data. Previous research has demonstrated the potential value in applying NoSQL data models in non-distributed environments. However, knowing when to apply these data models has generally required inputs from system subject matter experts to make this determination. This research, sponsored by the Air …


Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li Aug 2020

Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li

Theses and Dissertations

As a fundamental task in data integration and data quality, Entity Resolution (ER) has been investigated for decades in various domains. The emerging volume of heterogeneously structured data, and even unstructured data, poses a challenge to traditional ER methods. This research is to explore machine learning and deep learning approach to address the challenge from unstructured references data. This research starts with pairwise matching, the core function of all ER tasks. Based on the similarity score vector derived from our designed similarity measurement tool, scoring matrix, machine leaning enhances the performance significantly compared to the manually threshold method. Without similarity …


Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney Aug 2020

Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney

Theses and Dissertations

This thesis describes the creation of a natural-language dataset from a corpus of legislation passed in the State of Arkansas between 2001 and 2019. The dataset also includes metadata about individual acts of legislation and the lawmakers who sponsored them. This thesis describes the creation of the dataset, including the transformation of raw textual input using various natural language processing techniques such as sentiment analysis and topic modeling. Finally, this thesis examines a use case for identifying corrupt lawmakers using machine learning tools trained on transformations of the dataset.


Snow-Albedo Feedback In Northern Alaska: How Vegetation Influences Snowmelt, Lucas C. Reckhaus Aug 2020

Snow-Albedo Feedback In Northern Alaska: How Vegetation Influences Snowmelt, Lucas C. Reckhaus

Theses and Dissertations

This paper investigates how the snow-albedo feedback mechanism of the arctic is changing in response to rising climate temperatures. Specifically, the interplay of vegetation and snowmelt, and how these two variables can be correlated. This has the potential to refine climate modelling of the spring transition season. Research was conducted at the ecoregion scale in northern Alaska from 2000 to 2020. Each ecoregion is defined by distinct topographic and ecological conditions, allowing for meaningful contrast between the patterns of spring albedo transition across surface conditions and vegetation types. The five most northerly ecoregions of Alaska are chosen as they encompass …