Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (17)
- Business (11)
- Information Security (10)
- Social and Behavioral Sciences (6)
- Electrical and Computer Engineering (5)
-
- Artificial Intelligence and Robotics (4)
- Management Information Systems (4)
- Numerical Analysis and Scientific Computing (4)
- Operations Research, Systems Engineering and Industrial Engineering (4)
- Data Science (3)
- Geography (3)
- Systems Architecture (3)
- Aerospace Engineering (2)
- Applied Mathematics (2)
- Aviation (2)
- Computer Engineering (2)
- Controls and Control Theory (2)
- Cybersecurity (2)
- Finance and Financial Management (2)
- Geographic Information Sciences (2)
- Graphics and Human Computer Interfaces (2)
- Life Sciences (2)
- Longitudinal Data Analysis and Time Series (2)
- OS and Networks (2)
- Oceanography and Atmospheric Sciences and Meteorology (2)
- Operational Research (2)
- Physical and Environmental Geography (2)
- Institution
- Keyword
-
- Entity Resolution (9)
- Object-oriented databases (9)
- Machine learning (8)
- Data quality (6)
- Entity resolution (5)
-
- Databases (4)
- Deep Learning (4)
- Distributed databases (4)
- Information Quality (4)
- Natural Language Processing (4)
- Center_CCR (3)
- Data Washing Machine (3)
- Data management (3)
- Database management (3)
- Entity Identity Information Management (3)
- Information retrieval (3)
- Organizational learning (3)
- #antcenter (2)
- Automation (2)
- Big Data (2)
- Clerical Review (2)
- Data Mining (2)
- Data Quality (2)
- Data bases (2)
- Data governance (2)
- Data integration (2)
- Data mining (2)
- Database (2)
- Database design (2)
- Distributed data processing (2)
Articles 1 - 30 of 129
Full-Text Articles in Databases and Information Systems
Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev
Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev
Theses and Dissertations
DataOps has been coined as a novel term that emerged as a synthesis of data management practices with software engineering concepts, such as DevOps and Agile, with the goal of improving data quality and governance in enterprises. The proliferation of scratch table use and transformation tools, such as dbt, has led to an exponential increase in the number of data models, which complicates standardization efforts and increases maintenance overhead. Although the market is saturated with various flavors of text-to-SQL engines that promote increased productivity and self-service use in organizations, there are limited tools available to optimize individual queries, enforce consistency, …
Blockchain-Enabled Master Data Management, Shakhawat Hossain
Blockchain-Enabled Master Data Management, Shakhawat Hossain
Theses and Dissertations
Master Data Management (MDM) is essential for maintaining data quality, accuracy, consistency, and governance within organizations. However, traditional centralized MDM systems continue to face challenges related to data integrity, security, and scalability. This research presents a blockchain-enabled MDM framework designed to overcome these limitations by leveraging blockchain’s decentralized, immutable, and secure architecture. The study aims to identify and address the shortcomings of conventional MDM practices, examine the applicability of blockchain technology in enhancing these systems, and develop a functional prototype to validate the proposed model. The framework incorporates decentralized review mechanisms that improve auditability and ensure trusted data verification by …
Qlorax: Heuristic-Guided Fine-Tuning Of Llama-2 For Domain Adaptation In Entrepreneurship, Gaurob Saha
Qlorax: Heuristic-Guided Fine-Tuning Of Llama-2 For Domain Adaptation In Entrepreneurship, Gaurob Saha
Theses and Dissertations
This thesis presents a study on the fine-tuning of large language models (LLMs) for domain-specific applications using limited data. We fine-tuned the LLaMA-2 (7B) model on a curated entrepreneurial dataset containing 3,545 human-written question-answer pairs, of which 3,095 were used for training and 450 were reserved for evaluation. A complete fine-tuning and evaluation pipeline was developed, which included clustering human-written answers, generating centroid-based summaries for each cluster, and evaluating the model's generated responses through cosine similarity.
Training was carried out over five epochs, with model performance evaluated after each epoch. The fine-tuned model demonstrated strong semantic alignment with human-written content, …
Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn
Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn
Theses and Dissertations
In the ever-expanding landscape of digital technologies, the exponential growth of data presents both challenges and opportunities, demanding innovative approaches to data curation. Effective data curation is pivotal for extracting meaningful insights from vast and complex datasets. This study explores the integration of spectral clustering and Shannon Entropy within the Data Washing Machine (DWM), a novel tool designed to streamline unsupervised data curation processes. The DWM incorporates Shannon Entropy into its clustering process, allowing for adaptive refinement of clustering strategies based on entropy levels observed within data clusters. Spectral clustering, known for its ability to handle complex and non-linearly separable …
Increasing The Robustness Of Machine Learning By Adversarial Attacks, Gourab Mukhopadhyay
Increasing The Robustness Of Machine Learning By Adversarial Attacks, Gourab Mukhopadhyay
Theses and Dissertations
By perturbation or physical attacks any machine can be fooled into predicting something else other than the intended output. There are training data based on which the model is trained to predict unknown things. The objective was to create noises and shades of different levels on the images and do experiments for measuring accuracy and making the model classify the traffic signs. When it comes to adding shades to the pictures, pixels were modified for three different layers of the pictures. The experiment also shows that with the shadows getting deeper, the accuracies drop significantly. Here, some changes in pixels …
Deep Learning In Indus Valley Script Digitization, Deva Munikanta Reddy Atturu
Deep Learning In Indus Valley Script Digitization, Deva Munikanta Reddy Atturu
Theses and Dissertations
This research introduces ASR-net(Ancient Script Recognition), a groundbreaking system that automatically digitizes ancient Indus seals by converting them into coded text, similar to Optical Character Recognition for modern languages. ASR-net, with an 95% success rate in identifying individual symbols, aims to address the crucial need for automated techniques in deciphering the enigmatic Indus script. Initially Yolov3 is utilized to create the bounding boxes around each graphemes present in the Indus Valley Seal. In addition to that we created M-net(Mahadevan) model to encode the graphemes. Beyond digitization, the paper proposes a new research challenge called the Motif Identification Problem (MIP) related …
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Theses and Dissertations
This dissertation explores applications of representation learning and generative models to challenges in healthcare, astronautics, and aviation.
The first part investigates the use of Generative Adversarial Networks (GANs) to synthesize realistic electronic health record (EHR) data. An initial attempt at training a GAN on the MIMIC-IV dataset encountered stability and convergence issues, motivating a deeper study of 1-Lipschitz regularization techniques for Auxiliary Classifier GANs (AC-GANs). An extensive ablation study on the CIFAR-10 dataset found that Spectral Normalization is key for AC-GAN stability and performance, while Weight Clipping fails to converge without Spectral Normalization. Analysis of the training dynamics provided further …
Modeling & Engineering Of Usmepcom Business Intelligence Data, Merrick A. Bedford
Modeling & Engineering Of Usmepcom Business Intelligence Data, Merrick A. Bedford
Theses and Dissertations
This thesis investigates the USMEPCOM’s issue of modeling and engineering Business Intelligence data centralized around the MEPS of Excellence (MOE) program. MEPS around the US conduct military personnel in-processing and in doing so have a vested interest in the standardization and application of the data associated with such processes. There are currently 65 MEPS stations and one RPS that handle military personnel onboarding paperwork and make determinations for military eligibility. This topic is important due to the MEPS cloud data processing system modernization efforts requiring data processing adaptation to ensure relevant and meaningful usage of current data.
Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua
Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua
Theses and Dissertations
Entity resolution (ER) plays a pivotal role across domains by enabling data integration and quality improvement. This dissertation delves into the evolving landscape of ER, introducing innovative approaches that redefine this fundamental task. The first contribution is TriBERTa, a novel representation learning model tailored for ER. TriBERTa sets new benchmarks in entity matching and demonstrates versatility across ER processes like data blocking and resolution. Empirical evaluations on diverse datasets showcase TriBERTa’s superior performance over existing representations, including from large language models. The second contribution explores the use generative language models like GPT-3.5 and Dolly 2.0 for cross-domain entity matching using …
A Conceptual Decentralized Identity Solution For State Government, Martin Duclos
A Conceptual Decentralized Identity Solution For State Government, Martin Duclos
Theses and Dissertations
In recent years, state governments, exemplified by Mississippi, have significantly expanded their online service offerings to reduce costs and improve efficiency. However, this shift has led to challenges in managing digital identities effectively, with multiple fragmented solutions in use. This paper proposes a Self-Sovereign Identity (SSI) framework based on distributed ledger technology. SSI grants individuals control over their digital identities, enhancing privacy and security without relying on a centralized authority. The contributions of this research include increased efficiency, improved privacy and security, enhanced user satisfaction, and reduced costs in state government digital identity management. The paper provides background on digital …
Pattern-Of-Life Modeling With Automatic Dependent Surveillance-Broadcast (Ads-B), Sarah J. Bolton
Pattern-Of-Life Modeling With Automatic Dependent Surveillance-Broadcast (Ads-B), Sarah J. Bolton
Theses and Dissertations
This dissertation and research were sponsored by the Air Force Research Laboratory Layered Sensing Exploitation Branch (AFRL/RYA) to investigate the utility of using the data found within aircraft secondary radar to make predictions about aircraft characteristics and intent. The research focuses on making predictions on aircraft characteristics using only the kinetic data within one type of secondary radar, Automatic Dependent Surveillance-Broadcast (ADS-B), as a surrogate for primary radar. The results from this research provide a means to reduce the reliance on a type of aircraft tracking that is vulnerable to cyber attack and other integrity concerns.
Analysis And Optimization Of Contract Data Schema, Franklin Sun
Analysis And Optimization Of Contract Data Schema, Franklin Sun
Theses and Dissertations
agement, development, and growth of U.S Air Force assets demand extensive organizational communication and structuring. These interactions yield substantial amounts of contracting and administrative information. Over 4 million such contracts as a means towards obtaining valuable insights on Department of Defense resource usage. This set of contracting data is largely not optimized for backend service in an analytics environment. To this end, the following research evaluates the efficiency and performance of various data structuring methods. Evaluated designs include a baseline unstructured schema, a Data Mart schema, and a snowflake schema. Overall design success metrics include ease of use by end …
A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson
A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson
Theses and Dissertations
The Data Washing Machine (DWM) is a known and documented open-source Python Jupyter Notebook project that is the foundation for an Unsupervised Data Curation process. The DWM ingests reference data without a prior data cleansing activity and ultimately runs Entity Resolution (ER) on acceptable entity data to cluster duplicate references within the dataset. The DWM currently has 17 modifiable parameters that are used to help tokenize, cleanse, organize, link and cluster like references. With such a large number of parameters, some type of beginning settings as optimal as possible are needed for the DWM process for it to be useful …
The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique
The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique
Theses and Dissertations
The main objective of Entity resolution (ER) is to find duplicate records within the same data table from the same source or different data tables from various sources. A traditional pair-wise supervised entity resolution matching depends on pre-built rules for finding matched records. On the other hand, unsupervised or semisupervised also relies on pair-wise matching. In the maximum case, group membership is left behind for consideration. In this dissertation, I have discussed the design, implementation and evaluation of a graph-based entity resolution for group membership to enhance the pair-wise matching ER system. I have designed and implemented a pipeline for …
Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed
Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed
Theses and Dissertations
In an era, where data is being used by organizations in achieving their business goals and driving their business decisions, it is key to ensure the quality of data. For an organization, the most important data assets include master data assets such as product master data and customer master data, supplier master data, employee master data which can be generalized as party master data. There has been significant growth and variety in data in recent years because of which the traditional rules-based approach and dependency on data experts for assessing data quality is no longer working. This dissertation evaluates and …
Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira
Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira
Theses and Dissertations
For the last two decades, the cloud computing ecosystems has become a major defining force for America, because of their unique economic, social and national importance. These ecosystems have now taken their place alongside the nation’s other infrastructure such as the food/agricultural, energy healthcare, and roads/highways. Scientists, engineers, and researchers over the same period have experienced a tremendous growth in the need for resources to that support assorted research. Cloud computing enables these communities to undertake wide-ranging research efforts, while requiring no maintenance, management or significant invest of local resources. The genius of this project examines how create on-ramps for …
The Applications Of The Internet Of Things In The Medical Field, Cody Repass
The Applications Of The Internet Of Things In The Medical Field, Cody Repass
Theses and Dissertations
The Internet of Things (IoT) paradigm promises to make “things” include a more generic set of entities such as smart devices, sensors, human beings, and any other IoT objects to be accessible at anytime and anywhere. IoT varies widely in its applications, and one of its most beneficial uses is in the medical field. However, the large attack surface and vulnerabilities of IoT systems needs to be secured and protected. Security is a requirement for IoT systems in the medical field where the Health Insurance Portability and Accountability Act (HIPAA) applies.
This work investigates various applications of IoT in healthcare …
Bayesian Convolutional Neural Network With Prediction Smoothing And Adversarial Class Thresholds, Noah M. Miller
Bayesian Convolutional Neural Network With Prediction Smoothing And Adversarial Class Thresholds, Noah M. Miller
Theses and Dissertations
Using convolutional neural networks (CNNs) for image classification for each frame in a video is a very common technique. Unfortunately, CNNs are very brittle and have a tendency to be over confident in their predictions. This can lead to what we will refer to as “flickering,” which is when the predictions between frames jump back and forth between classes. In this paper, new methods are proposed to combat these shortcomings. This paper utilizes a Bayesian CNN which allows for a distribution of outputs on each data point instead of just a point estimate. These distributions are then smoothed over multiple …
A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye
A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye
Theses and Dissertations
Data governance is mission critical. Organizations start to realize the data and information are very important assets. Data governance program could minimize the risk of data lose and maximize the data values. As existing best practices to implement data governance program, most of these practices rely on management strategies, which focus on data governance literacy education, enforcing data policy, standard and processes, etc. However, IT infrastructure strategy has not changed too much to support data governance achievement. Data governance requirements for system design usually is the last concern on the list. Most of the data governance programs collect metadata or …
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Theses and Dissertations
This dissertation describes a first attempt to build a data washing machine, a system able to take dirty data and through an unsupervised process, output clean data. The washing machine design described here focuses on two main aspects of the data curation process, token correction and data redundancy. It aims to simplify and automate the preparation of data used to create information products. In this approach, all these steps would be automated, thus saving the time and effort of the data analysts who ordinarily perform these actions. In other words, this is the opposite of the current approach to first …
Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy
Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy
Theses and Dissertations
Social media plays an important role in the propagation and dissemination of ideas and thoughts. Compared to other social media platforms, blogs provide a convenient platform for users to post detailed information, engage in active discussions and share the content on other social media sites, such as Facebook and Twitter. Thus, the blogosphere has been an enormous and ever-growing part of the open-source intelligence. In order to track and monitor online social behavior particularly from blogs, the first challenging part is to mine the vast pool of unstructured data. To scale up this process and cope with the continuously changing …
Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon
Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon
Theses and Dissertations
Machine learning models for chemical property predictions are high dimension design challenges spanning multiple disciplines. Free and open-source software libraries have streamlined the model implementation process, but the design complexity remains. In order better navigate and understand the machine learning design space, model information needs to be organized and contextualized. In this work, instances of chemical property models and their associated parameters were stored in a Neo4j property graph database. Machine learning model instances were created with permutations of dataset, learning algorithm, molecular featurization, data scaling, data splitting, hyperparameters, and hyperparameter optimization techniques. The resulting graph contains over 83,000 nodes …
Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai
Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai
Theses and Dissertations
The development of web technologies has progressed rapidly in the past few decades. As a result, these technologies have increasingly gained popularity among researchers as an instrument for data analysis and visualization to aid their work. The purpose of this project is to pursue a new method for mining time-based data to perform curve fitting, using web client technology. This requires comparing the advantages and disadvantages of different JavaScript software designs to determine which is best. Based on the results, we emphasize developing a client-side web application and demonstrate its use through fitting a curve for a new metric to …
Barriers And Drivers Influencing The Growth Of E-Commerce In Uzbekistan, Madinakhon Tursunboeva
Barriers And Drivers Influencing The Growth Of E-Commerce In Uzbekistan, Madinakhon Tursunboeva
Theses and Dissertations
Electronic commerce (e-commerce) has become a major retail channel for businesses in developed countries. However, it is still considered an innovation in developing countries. Specifically, e-commerce in Uzbekistan is in the early stages of emergence despite its advance in recent years in terms of Internet penetration, a strong retail sector, new national regulations, and a young population. The study aimed to identify barriers and drivers influencing e-commerce growth in Uzbekistan. A Delphi research design was utilized to answer the research questions of the study, which categorized and ranked factors that Uzbekistani entrepreneurs are facing when engaging in e-commerce processes. A …
A Framework To Determine The Impact Of Poor Data Quality On The Reliability Of Iot Sensor-Based Real-Time Decision-Making, Arnold Rego
Theses and Dissertations
Today, more and more systems make real-time decisions based on the data received from sensors. But it is inevitable that sensors will fail, or send bad data – leading to faulty real-time decision-making. Using this poor-quality data in real-time decision-making can lead to deadly consequences. Traditional Data Quality (DQ) foundations call for “fixing” the poor DQ before the data can be used. However, time is often of the essence in a real-time decision-making environment, leaving little or no time to “fix” poor-quality data before it can be used in the decision-making process. The primary objective in real-time decision-making is not …
Privacy And The Digital Divide: Investigating Strategies For Digital Safety By People Of Color, Denavious Hoover
Privacy And The Digital Divide: Investigating Strategies For Digital Safety By People Of Color, Denavious Hoover
Theses and Dissertations
People of color are becoming increasingly concerned with digital privacy. They are concerned about the obfuscated data collection and sharing practices of major social media plat- forms and the strong entitlement of other users in the online space to their content. This study examines how people of color conceptualize and behave to produce safety in the online space, or, in other words, digital privacy. This study challenges notions that people are not purposeful about privacy in the online space and highlights the voices of people of color, whom are not of- ten included in theorizing or decision making about the …
A Methodology To Identify Alternative Suitable Nosql Data Models Via Observation Of Relational Database Interactions, Paul M. Beach
A Methodology To Identify Alternative Suitable Nosql Data Models Via Observation Of Relational Database Interactions, Paul M. Beach
Theses and Dissertations
The effectiveness and performance of data-intensive applications are influenced by the suitability of the data models upon which they are built. The relational data model has been the de facto data model underlying most database systems since the 1970’s. However, the recent emergence of NoSQL data models have provided users with alternative ways of storing and manipulating data. Previous research has demonstrated the potential value in applying NoSQL data models in non-distributed environments. However, knowing when to apply these data models has generally required inputs from system subject matter experts to make this determination. This research, sponsored by the Air …
Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li
Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li
Theses and Dissertations
As a fundamental task in data integration and data quality, Entity Resolution (ER) has been investigated for decades in various domains. The emerging volume of heterogeneously structured data, and even unstructured data, poses a challenge to traditional ER methods. This research is to explore machine learning and deep learning approach to address the challenge from unstructured references data. This research starts with pairwise matching, the core function of all ER tasks. Based on the similarity score vector derived from our designed similarity measurement tool, scoring matrix, machine leaning enhances the performance significantly compared to the manually threshold method. Without similarity …
Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney
Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney
Theses and Dissertations
This thesis describes the creation of a natural-language dataset from a corpus of legislation passed in the State of Arkansas between 2001 and 2019. The dataset also includes metadata about individual acts of legislation and the lawmakers who sponsored them. This thesis describes the creation of the dataset, including the transformation of raw textual input using various natural language processing techniques such as sentiment analysis and topic modeling. Finally, this thesis examines a use case for identifying corrupt lawmakers using machine learning tools trained on transformations of the dataset.
Snow-Albedo Feedback In Northern Alaska: How Vegetation Influences Snowmelt, Lucas C. Reckhaus
Snow-Albedo Feedback In Northern Alaska: How Vegetation Influences Snowmelt, Lucas C. Reckhaus
Theses and Dissertations
This paper investigates how the snow-albedo feedback mechanism of the arctic is changing in response to rising climate temperatures. Specifically, the interplay of vegetation and snowmelt, and how these two variables can be correlated. This has the potential to refine climate modelling of the spring transition season. Research was conducted at the ecoregion scale in northern Alaska from 2000 to 2020. Each ecoregion is defined by distinct topographic and ecological conditions, allowing for meaningful contrast between the patterns of spring albedo transition across surface conditions and vegetation types. The five most northerly ecoregions of Alaska are chosen as they encompass …