Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Keyword
-
- Entity Resolution (9)
- Data quality (6)
- Entity resolution (5)
- Information Quality (4)
- Machine learning (4)
-
- Data Washing Machine (3)
- Entity Identity Information Management (3)
- Automation (2)
- Clerical Review (2)
- Data Quality (2)
- Data governance (2)
- Data integration (2)
- Master Data Management (2)
- Natural language processing (2)
- Record linkage (2)
- API (1)
- Ad-Hoc Retrieval (1)
- Approximate string matching (1)
- Assertion (1)
- Big Data (1)
- Blockchain Technology (1)
- Blocking (1)
- Blog crawling (1)
- Blog posts (1)
- Blogs (1)
- Boolean Rule (1)
- Cloud computing (1)
- Cloud infrastructure (1)
- Cohort graduation rate (1)
- Confidence Rating (1)
Articles 1 - 30 of 33
Full-Text Articles in Databases and Information Systems
Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev
Implementing Dataops: A Scalable Framework For Modern Data Warehousing, Dmytro Valiaiev
Theses and Dissertations
DataOps has been coined as a novel term that emerged as a synthesis of data management practices with software engineering concepts, such as DevOps and Agile, with the goal of improving data quality and governance in enterprises. The proliferation of scratch table use and transformation tools, such as dbt, has led to an exponential increase in the number of data models, which complicates standardization efforts and increases maintenance overhead. Although the market is saturated with various flavors of text-to-SQL engines that promote increased productivity and self-service use in organizations, there are limited tools available to optimize individual queries, enforce consistency, …
Blockchain-Enabled Master Data Management, Shakhawat Hossain
Blockchain-Enabled Master Data Management, Shakhawat Hossain
Theses and Dissertations
Master Data Management (MDM) is essential for maintaining data quality, accuracy, consistency, and governance within organizations. However, traditional centralized MDM systems continue to face challenges related to data integrity, security, and scalability. This research presents a blockchain-enabled MDM framework designed to overcome these limitations by leveraging blockchain’s decentralized, immutable, and secure architecture. The study aims to identify and address the shortcomings of conventional MDM practices, examine the applicability of blockchain technology in enhancing these systems, and develop a functional prototype to validate the proposed model. The framework incorporates decentralized review mechanisms that improve auditability and ensure trusted data verification by …
Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn
Improving Data Curation With Spectral Clustering And Shannon Entropy: An Unsupervised Approach Within The Data Washing Machine, Erin Hathorn
Theses and Dissertations
In the ever-expanding landscape of digital technologies, the exponential growth of data presents both challenges and opportunities, demanding innovative approaches to data curation. Effective data curation is pivotal for extracting meaningful insights from vast and complex datasets. This study explores the integration of spectral clustering and Shannon Entropy within the Data Washing Machine (DWM), a novel tool designed to streamline unsupervised data curation processes. The DWM incorporates Shannon Entropy into its clustering process, allowing for adaptive refinement of clustering strategies based on entropy levels observed within data clusters. Spectral clustering, known for its ability to handle complex and non-linearly separable …
Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua
Triberta And Beyond: Redefining Entity Resolution With Large Language Models, Bi Foua
Theses and Dissertations
Entity resolution (ER) plays a pivotal role across domains by enabling data integration and quality improvement. This dissertation delves into the evolving landscape of ER, introducing innovative approaches that redefine this fundamental task. The first contribution is TriBERTa, a novel representation learning model tailored for ER. TriBERTa sets new benchmarks in entity matching and demonstrates versatility across ER processes like data blocking and resolution. Empirical evaluations on diverse datasets showcase TriBERTa’s superior performance over existing representations, including from large language models. The second contribution explores the use generative language models like GPT-3.5 and Dolly 2.0 for cross-domain entity matching using …
A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson
A Parameter Discovery Process For The Data Washing Machine Created For Unsupervised Data Curation, Kris E. Anderson
Theses and Dissertations
The Data Washing Machine (DWM) is a known and documented open-source Python Jupyter Notebook project that is the foundation for an Unsupervised Data Curation process. The DWM ingests reference data without a prior data cleansing activity and ultimately runs Entity Resolution (ER) on acceptable entity data to cluster duplicate references within the dataset. The DWM currently has 17 modifiable parameters that are used to help tokenize, cleanse, organize, link and cluster like references. With such a large number of parameters, some type of beginning settings as optimal as possible are needed for the DWM process for it to be useful …
The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique
The Application Of Graph Technology For Improving Entity Resolution Results In The Context Of Group Membership, Md Abdus Salam Siddique
Theses and Dissertations
The main objective of Entity resolution (ER) is to find duplicate records within the same data table from the same source or different data tables from various sources. A traditional pair-wise supervised entity resolution matching depends on pre-built rules for finding matched records. On the other hand, unsupervised or semisupervised also relies on pair-wise matching. In the maximum case, group membership is left behind for consideration. In this dissertation, I have discussed the design, implementation and evaluation of a graph-based entity resolution for group membership to enhance the pair-wise matching ER system. I have designed and implemented a pipeline for …
Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed
Evaluation Of Automation Techniques For Data Quality Assessment For Party And Product Master Data, Mahmood Mohammed
Theses and Dissertations
In an era, where data is being used by organizations in achieving their business goals and driving their business decisions, it is key to ensure the quality of data. For an organization, the most important data assets include master data assets such as product master data and customer master data, supplier master data, employee master data which can be generalized as party master data. There has been significant growth and variety in data in recent years because of which the traditional rules-based approach and dependency on data experts for assessing data quality is no longer working. This dissertation evaluates and …
Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira
Data As A Service Ecosystem For Data-Driven Research, Leonardo Vieira
Theses and Dissertations
For the last two decades, the cloud computing ecosystems has become a major defining force for America, because of their unique economic, social and national importance. These ecosystems have now taken their place alongside the nation’s other infrastructure such as the food/agricultural, energy healthcare, and roads/highways. Scientists, engineers, and researchers over the same period have experienced a tremendous growth in the need for resources to that support assorted research. Cloud computing enables these communities to undertake wide-ranging research efforts, while requiring no maintenance, management or significant invest of local resources. The genius of this project examines how create on-ramps for …
A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye
A Positive Data Control System For The Automation Of Data Governance Functions, Yanbin Ye
Theses and Dissertations
Data governance is mission critical. Organizations start to realize the data and information are very important assets. Data governance program could minimize the risk of data lose and maximize the data values. As existing best practices to implement data governance program, most of these practices rely on management strategies, which focus on data governance literacy education, enforcing data policy, standard and processes, etc. However, IT infrastructure strategy has not changed too much to support data governance achievement. Data governance requirements for system design usually is the last concern on the list. Most of the data governance programs collect metadata or …
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Theses and Dissertations
This dissertation describes a first attempt to build a data washing machine, a system able to take dirty data and through an unsupervised process, output clean data. The washing machine design described here focuses on two main aspects of the data curation process, token correction and data redundancy. It aims to simplify and automate the preparation of data used to create information products. In this approach, all these steps would be automated, thus saving the time and effort of the data analysts who ordinarily perform these actions. In other words, this is the opposite of the current approach to first …
Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy
Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy
Theses and Dissertations
Social media plays an important role in the propagation and dissemination of ideas and thoughts. Compared to other social media platforms, blogs provide a convenient platform for users to post detailed information, engage in active discussions and share the content on other social media sites, such as Facebook and Twitter. Thus, the blogosphere has been an enormous and ever-growing part of the open-source intelligence. In order to track and monitor online social behavior particularly from blogs, the first challenging part is to mine the vast pool of unstructured data. To scale up this process and cope with the continuously changing …
Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai
Applying Web Technologies For Data Mining Of Models Of Advancement Over Time In Space Exploration Technology, Peng-Hung Tsai
Theses and Dissertations
The development of web technologies has progressed rapidly in the past few decades. As a result, these technologies have increasingly gained popularity among researchers as an instrument for data analysis and visualization to aid their work. The purpose of this project is to pursue a new method for mining time-based data to perform curve fitting, using web client technology. This requires comparing the advantages and disadvantages of different JavaScript software designs to determine which is best. Based on the results, we emphasize developing a client-side web application and demonstrate its use through fitting a curve for a new metric to …
A Framework To Determine The Impact Of Poor Data Quality On The Reliability Of Iot Sensor-Based Real-Time Decision-Making, Arnold Rego
Theses and Dissertations
Today, more and more systems make real-time decisions based on the data received from sensors. But it is inevitable that sensors will fail, or send bad data – leading to faulty real-time decision-making. Using this poor-quality data in real-time decision-making can lead to deadly consequences. Traditional Data Quality (DQ) foundations call for “fixing” the poor DQ before the data can be used. However, time is often of the essence in a real-time decision-making environment, leaving little or no time to “fix” poor-quality data before it can be used in the decision-making process. The primary objective in real-time decision-making is not …
Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li
Machine Learning And Deep Learning Based Entity Resolution Approaches For Unstructured References, Xinming Li
Theses and Dissertations
As a fundamental task in data integration and data quality, Entity Resolution (ER) has been investigated for decades in various domains. The emerging volume of heterogeneously structured data, and even unstructured data, poses a challenge to traditional ER methods. This research is to explore machine learning and deep learning approach to address the challenge from unstructured references data. This research starts with pairwise matching, the core function of all ER tasks. Based on the similarity score vector derived from our designed similarity measurement tool, scoring matrix, machine leaning enhances the performance significantly compared to the manually threshold method. Without similarity …
Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney
Arlegislation: An R Package Of Arkansas Legislation Data And An Exploratory Use Case For Using Machine Learning To Identify Public Corruption, Nathan P. Chaney
Theses and Dissertations
This thesis describes the creation of a natural-language dataset from a corpus of legislation passed in the State of Arkansas between 2001 and 2019. The dataset also includes metadata about individual acts of legislation and the lawmakers who sponsored them. This thesis describes the creation of the dataset, including the transformation of raw textual input using various natural language processing techniques such as sentiment analysis and topic modeling. Finally, this thesis examines a use case for identifying corrupt lawmakers using machine learning tools trained on transformations of the dataset.
An Investigation Into The Optimal Use Of Frequency-Based Weights To Improve The Performance Of Entity Resolution, Bingyi Zhong
An Investigation Into The Optimal Use Of Frequency-Based Weights To Improve The Performance Of Entity Resolution, Bingyi Zhong
Theses and Dissertations
Using a weight-based match score has been studied as a way to improve the accuracy of record linking (entity resolution) since Fellegi and Sunter first described the idea of probabilistic agreement and disagreement weights in their seminal work “A Theory of Record Linking.” However, the original work only described weight associated with an entity attribute. Later researchers such as Herzog et al suggested the weighting scheme could be extended to apply to frequently-occurring attribute values (frequency-based weights) instead of just the attribute as in the Fellegi-Sunter scheme. However, there has been little definitive research as to how frequency-based weights should …
Multi-Resolution Models For Learning Multilevel Abstract Representation With Application To Information Retrieval, Tolgahan Cakaloglu
Multi-Resolution Models For Learning Multilevel Abstract Representation With Application To Information Retrieval, Tolgahan Cakaloglu
Theses and Dissertations
Deep language models learning a hierarchical representation proved to be a powerful tool for natural language processing, text mining, and information retrieval tasks. However, more specifically, representations that perform well for ad-hoc retrieval must capture semantic meaning at different levels of abstraction or context-scopes. The primary goal of ad-hoc retrieval is to find relevant documents satisfying the information need posted in a natural language query. It requires a good understanding of the query and all the documents in a corpus, which is difficult because the meaning of natural language texts depends on the context, syntax, and semantics. In this dissertation, …
A System For Stratified Sampling Of Entity Resolution Results To Assess And Improve Accuracy With Minimal Clerical Review Effort, Daniel L. Pullen
A System For Stratified Sampling Of Entity Resolution Results To Assess And Improve Accuracy With Minimal Clerical Review Effort, Daniel L. Pullen
Theses and Dissertations
Organizations across many industries from banking to medicine depend on master data that describe their customers, products and services. Data integration and unique representations of master data are supported by Entity Resolution processes to link records for the same master entity. Master data is key to the core operations of these organizations. Unfortunately, many of these organizations do not use a systematic method, or in some cases, use no method for measuring and improving the quality of record linkage provided by Entity Resolution. This research presents an implementation and proposed methodology for routinely and systematically measuring the quality of Entity …
Stream Data Quality Assessment Based On Distributed Computing Platforms, Wei Dai
Stream Data Quality Assessment Based On Distributed Computing Platforms, Wei Dai
Theses and Dissertations
In this era of big data, data quality will be increasingly important because people need high quality data to make decisions, analyze patterns, and discover knowledge. So, measuring data quality is a vital mission. In this thesis, Chapter 1 is the introduction, Chapter 2 is a literature review, Chapter 3 illustrates how to discover potentially important data based on a reference algorithm, a frequency algorithm, and an entropy algorithm, in Chapter 4, the author offers a concise five-layer data quality framework to measure stream data quality scorecards, in Chapter 5, the author shows how to visualize data quality scorecards through …
Blocking Strategies For Performing Entity Resolution In A Distributed Computing Environment, Pei Wang
Blocking Strategies For Performing Entity Resolution In A Distributed Computing Environment, Pei Wang
Theses and Dissertations
Entity resolution (ER) is an O(n2) problem where n is the number of records to be processed. The pair-wise nature of ER makes it impractical to perform on large datasets without the use of a technique called blocking. In blocking the records are separated into groups (called blocks) in such a way the records most likely to match are within the same block. The ER system only compares pairs of records within the same block, thus reducing the total number of pairs to match. Traditionally, blocking algorithms build inverted indices in memory to quickly locate potential matches. With the advent …
Inferred Error Rates For Entity Resolution, Melody Lynn Penning
Inferred Error Rates For Entity Resolution, Melody Lynn Penning
Theses and Dissertations
This dissertation is focused on the methodology of determining the quality of the results of entity resolution. Entity resolution methodologies results in different success rates. Until now these success rates have been measured by counting all of the correctly and incorrectly matched results. This is quickly becoming an intractable task as datasets grow in size. In order to address this problem this research tests and describes the results of count based measures as compared with inferred measures borrowed from the information retrieval community. The key contribution of this research is a proof of concept in a controlled environment that demonstrates …
A Framework For Collecting, Extracting And Managing Event Identity Information From Textual Content In Social Media, Debanjan Mahata
A Framework For Collecting, Extracting And Managing Event Identity Information From Textual Content In Social Media, Debanjan Mahata
Theses and Dissertations
With the popularity of social media platforms such as Facebook, Twitter and Google Plus, there has been voluminous growth in the digital footprints of real-life events on the Internet. The user-generated colloquial and concise textual content related to different types of real-life events, available in these websites, acts as an extremely useful source for researchers and organizations for extracting valuable and insightful information. There has been significant improvement in natural language processing techniques for mining formal and long textual content commonly found in newspapers. It is still a challenging task to mine textual information from the social media channels producing …
A Design For An Identity Resolution Service As An Extension Of The Entity Identity Information Management Model, Fumiko Kobayashi
A Design For An Identity Resolution Service As An Extension Of The Entity Identity Information Management Model, Fumiko Kobayashi
Theses and Dissertations
This research describes the design of an identity resolution service (IRS), which provides extensions and enhancements to the current EIIM model. The IRS provides a set of application programming interfaces (API) which allow identity resolution (IR) to be performed interactively. Interactive IR is a logical extension to the existing EIIM model. It will separate the IR function from the EIIM batch update process. The new (IR) function: 1) Operates interactively 2) Has its own probabilistic matching rules defined separately from the identity rules used in the EIIM update process 3) Has the ability to return a confidence rating alongside a …
A System To Support Clerical Review, Correction, And Confirmation Assertions In Entity Identity Information Management, Cheng Chen
Theses and Dissertations
Clerical review of Entity Resolution(ER) is crucial for maintaining the entity identity integrity of an Entity Identity Information Management (EIIM) system. However, the clerical review process presents several problems. These problems include Entity Identity Structures (EIS) that are difficult to read and interpret, excessive time and effort to review large Identity Knowledgebase (IKB), and the duplication of effort in repeatedly reviewing the same EIS in same EIIM review cycle or across multiple review cycles. Although the original EIIM model envisioned and demonstrated the value of correction assertions, these are applied to correct errors after they have been found. The original …
Context-Sensitive Entity Resolution, William C. Decker
Context-Sensitive Entity Resolution, William C. Decker
Theses and Dissertations
This dissertation proposes an entity resolution approach that is context-sensitive, meaning it relies less on high-risk information, such as social security numbers, to discern whether records from a data set belong to the same real-world individual or to different real-world individuals. The research follows an iterative process of assessing the quality, or fitness of use, of identity data housed in an educational institution and then processing the data with the proposed context-sensitive entity resolution (ER) rule set. The efficacy of this process is demonstrated through calculation of the four-year adjusted cohort graduation rate, a prevalent longitudinal data analysis challenge for …
The Information Value Methodology: How Users Assess The Quality Of Web Information, Marilou Haines
The Information Value Methodology: How Users Assess The Quality Of Web Information, Marilou Haines
Theses and Dissertations
The Internet is a self-regulating complex system in which users decide what is relevant through their actions. Since the burden of locating and evaluating information depends on knowledge, experience, and skill, this study rigorously and holistically investigates the Web users' experience. The study began with a multidisciplinary literature review to determine how academics assess IQ on the Web. While individual models are narrowly focused, in aggregate, they delineate the boundaries of IQ needs, and identify the dimensions that are significant to the Web community. These dimensions, categorized into six knowledge domains, are the core of the information value methodology (IVM). …
Modeling And Analysis Of Information Product Maps, Christopher Harris Heien
Modeling And Analysis Of Information Product Maps, Christopher Harris Heien
Theses and Dissertations
Information Product Maps are visual diagrams used to represent the inputs, processing, and outputs of data within an Information Manufacturing System. A data unit, drawn as an edge, symbolizes a grouping of raw data as it travels through this system. Processes, drawn as vertices, transform each data unit input into various forms prior to delivery to consumers. These visual representations act as documentation for all processes, and data unit transformations, leading to the final construction and delivery of an Information Product. The science of Information Quality strives to measure the fitness for use of these products, expressed by consumer satisfaction. …
The ∑Iq Methodology: An Information Quality Perspective On Oil Data, Yusuf Yiliyasi
The ∑Iq Methodology: An Information Quality Perspective On Oil Data, Yusuf Yiliyasi
Theses and Dissertations
Information quality (IQ) theories and frameworks have been increasingly studied and applied in various organizations to assess, improve and monitor the quality of their information products. Yet information quality problems remain pervasive in many organizations, industries and government institutions. For example various governmental as well as non-governmental organizations, institutions, and companies collect, compile and distribute information products in order to satisfy the needs of information consumers in the energy industry. What are the qualities of their information products? How do they collect or disseminate their data products? Why do information quality problems exist and how are the problems created? How …
Design And Construction Of An Entity Resolution System That Supports Entity Identity Information Management And Asserted Resolution, Eric D. Nelson
Design And Construction Of An Entity Resolution System That Supports Entity Identity Information Management And Asserted Resolution, Eric D. Nelson
Theses and Dissertations
This work describes the design and construction of an open source, entity resolution system that enables users to assign and maintain persistent identifiers for master data items. Two key features of this system that are not available in current ER systems and that make persistent identification possible are (1) The capture and management of entity identity information (2) Support for user-directed asserted resolution to complement automated direct matching and transitive closure Another important feature of the design is that the system can be easily configured at run-time into any one of four types of entity resolution architectures including * Traditional …
A Data-Intensive Approach To Named Entity Recognition Using Domain And Language Independent Methods, Olukayode Isaac Osesina
A Data-Intensive Approach To Named Entity Recognition Using Domain And Language Independent Methods, Olukayode Isaac Osesina
Theses and Dissertations
In this dissertation, I proposed a novel approach to Named Entity Recognition (NER) in which the contextual and intrinsic indicators are used for locating named entities and their semantic meanings in unstructured textual information (UTI). Named entity is the process of locating a word or a phrase that references a particular entity within a text. The data-intensive approach introduced in this dissertation departs from the traditional Natural Language Processing used in NER tasks in that it does not apply linguistic rules or knowledge in the entity recognition process. It leverages the wide availability of huge amounts of data as well …