Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2641 - 2670 of 3244

Full-Text Articles in Data Science

Moving Ethnography: Infrastructuring Doubletakes And Switchbacks In Experimental Collaborative Methods, Aalok Khandekar, Brandon Costelloe-Kuehn, Lindsay Poirier, Alli Morgan, Alison Kenner, Kim Fortun, Mike Fortun Jan 2021

Moving Ethnography: Infrastructuring Doubletakes And Switchbacks In Experimental Collaborative Methods, Aalok Khandekar, Brandon Costelloe-Kuehn, Lindsay Poirier, Alli Morgan, Alison Kenner, Kim Fortun, Mike Fortun

Statistical and Data Sciences: Faculty Publications

In this article, we describe how our work at a particular nexus of STS, ethnography, and critical theory—informed by experimental sensibilities in both the arts and sciences—transformed as we built and learned to use collaborative workflows and supporting digital infrastructure. Responding to the call of this special issue to be “ethnographic about ethnography,” we describe what we have learned about our own methods and collaborative practices through building digital infrastructure to support them. Supporting and accounting for how experimental ethnographic projects move—through different points in a research workflow, with many switchbacks, with project designs constantly changing as the research develops—was …


Searching Harder, Localizing Better, Classifying Faster: Optimizing Fast Radio Burst Detection And Analysis, Kshitij Aggarwal Jan 2021

Searching Harder, Localizing Better, Classifying Faster: Optimizing Fast Radio Burst Detection And Analysis, Kshitij Aggarwal

Graduate Theses, Dissertations, and Problem Reports (ETD)

Fast Radio Bursts (or FRBs) are millisecond-duration transients of extragalactic origin. They exhibit dispersion caused by propagation through an ionized medium, and quantified by Dispersion Measure (DM). Around 800 FRBs (24 repeaters) have been discovered; so far, 24 FRBs have been confidently associated with a host galaxy. In this thesis, we discuss multiple new FRB search and analysis techniques and the corresponding tools that enable us to search for FRBs harder, localize them better, and classify candidates faster.

We discuss five open-source software suites that can be used in FRB analysis. These suites are used to distinguish between FRBs and …


Topic Modeling And Cultural Nature Of Citations, Marie Coraline Dumaz Jan 2021

Topic Modeling And Cultural Nature Of Citations, Marie Coraline Dumaz

Graduate Theses, Dissertations, and Problem Reports (ETD)

Ever since the beginning of research journals, the number of academic publications has been increasing steadily. Nowadays, especially, with the new importance of online open-access journals and databases, research papers are more easily available to read and share. It also becomes harder to keep up with novelties and grasp an idea of the general impact of a given researcher, institution, journal, or field. For this reason, different bibliometric indicators are now routinely used to classify and evaluate the impact or significance of individual researchers, conferences, journals, or entire scientific communities. In this thesis, we provide tools to study trends in …


Analysis And Classification Of Software Fault-Proneness And Vulnerabilities, Mohammad Jamil Ahmad Jan 2021

Analysis And Classification Of Software Fault-Proneness And Vulnerabilities, Mohammad Jamil Ahmad

Graduate Theses, Dissertations, and Problem Reports (ETD)

Software bugs are expensive to fix and can lead to catastrophic consequences. Therefore, their analysis and the use of machine learning for prediction are of the utmost importance. Many prediction models have been proposed and different factors affecting the prediction performance have been extensively studied. This work addresses four topics in two areas in software engineering: software fault-proneness prediction and analysis and classification of security-related bug reports. The first topic focuses on the effect of the learning approach (i.e., the way software fault-proneness prediction models are trained and tested) on the performance of software fault-proneness prediction which lacks extensive research …


A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui Jan 2021

A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui

Faculty, Staff and Student Publications

Uncovering and fixing errors in biomedical terminologies is essential so that they provide accurate knowledge to downstream applications that rely on them. Non-lattice-based methods have been applied to identify various kinds of inconsistencies in different biomedical terminologies. In previous work, we have introduced two inference-based approaches that were applied in an exhaustive manner to audit hierarchical relations in the Gene Ontology: (1) Lexical-based inference framework, and (2) Subsumption-based sub-term inference framework. However, it is unclear how effective these exhaustive approaches perform compared with their corresponding non-lattice-based approaches. Therefore, in this paper, we implement the non-lattice versions of these two exhaustive …


A Review And Evaluation Of Techniques For Improved Feature Detection In Mass Spectrometry Data, Annika R. Tostengard, Rob Smith Jan 2021

A Review And Evaluation Of Techniques For Improved Feature Detection In Mass Spectrometry Data, Annika R. Tostengard, Rob Smith

Graduate Student Theses, Dissertations, & Professional Papers

Mass spectrometry (MS) is used in analysis of chemical samples to identify the molecules present and their quantities. This analytical technique has applications in many fields, from pharmacology to space exploration. Its impacts on medicine are particularly significant, since MS aids in the identification of molecules associated with disease; for instance, in proteomics, MS allows researchers to identify proteins that are associated with autoimmune disorders, cancers, and other conditions. Since the applications are so wide-ranging and the tool is ubiquitous across so many fields, it is critical that the analytical methods used to collect data are sound.

Data analysis in …


Inference Of Surface Velocities From Oblique Time Lapse Photos And Terrestrial Based Lidar At The Helheim Glacier, Franklyn T. Dunbar Ii Jan 2021

Inference Of Surface Velocities From Oblique Time Lapse Photos And Terrestrial Based Lidar At The Helheim Glacier, Franklyn T. Dunbar Ii

Graduate Student Theses, Dissertations, & Professional Papers

Using time dependent observations derived from terrestrial LiDAR and oblique
time-lapse imagery, we demonstrate that a Bayesian approach to glacial motion es-
timation provides a concise way to incorporate multiple data products into a single
motion estimation procedure effectively producing surface velocity estimates with
an associated uncertainty. This approach brings both improved computational effi-
ciency, and greater scalability across observational time-frames when compared to
existing methods. To gauge efficacy, we apply these methods to a set of observa-
tions from the Helheim Glacier, a critical actor in contemporary mass loss trends
observed in the Greenland Ice Sheet. We find that …


Clustering Data To Classify Hearthstone Decks, Tim Inzitari Jan 2021

Clustering Data To Classify Hearthstone Decks, Tim Inzitari

Williams Honors College, Honors Research Projects

The esports game of "Hearthstone" is a collectible card game with a competitive format that has every team submit 4 decks of 30 cards each. Using K-Means clustering an adaptable way to group data for classifying can be made that works well in every update of the game. This system will take in a list of decks and cluster them to easily classify large amounts of information in a timely fashion. This system will be able to be used by the Universities esports department for years to come to aid the preparation of "Hearthstone" matches. This model uses qualities about …


The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr Jan 2021

The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr

Statistical and Data Sciences: Faculty Publications

We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.


Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse, Omar El-Gayar, Abdullah Wahbeh, Tareq Nasralah, Ahmed El Noshokaty, Mohammad A. Al-Ramahi Jan 2021

Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse, Omar El-Gayar, Abdullah Wahbeh, Tareq Nasralah, Ahmed El Noshokaty, Mohammad A. Al-Ramahi

Computer Information Systems Faculty Publications (Archived)

This study analyzed Twitter discourse to understand the association of the COVID-19 pandemic with mental health. The study compared tweets’ volume over time, tweets’ volume per mental health category, emotions, and the top hashtags on mental health before and after November 2019, the month on which the first COVID-19 case was reported. We analyzed a total of 273 million English tweets on mental health collected from 56 million unique users. Results and analysis showed a significant shift in trend for the volume of tweets on mental health over time. There was also a notable increase in the volume of tweets …


Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv Jan 2021

Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv

Theses and Dissertations

With data becoming a new form of currency, its analysis has become a top priority in both academia and industry, furthering advancements in high-performance computing and machine learning. However, these large, real-world datasets come with additional complications such as noise and class overlap. Problems are magnified when with multi-class data is presented, especially since many of the popular algorithms were originally designed for binary data. Another challenge arises when the number of examples are not evenly distributed across all classes in a dataset. This often causes classifiers to favor the majority class over the minority classes, leading to undesirable results …


Goes-R Supervised Machine Learning, Ronald Adomako Jan 2021

Goes-R Supervised Machine Learning, Ronald Adomako

Dissertations and Theses

The GOES-R series is a product line of four satellite, with two currently on-orbit (GOES-16 “East” and GOES-17 “West”). GOES-17 is susceptible to a Loop-Heat-Pipe (LHP) phenomenon where during Fall and Spring seasons, there are times of day where some of the infrared bands records inaccurate readings from the Advanced Baseline Imager (ABI). This occurs from joint astronomical behavior and position of the GOES-17. This calibration issue occurs when the LHP instrument fails to radiate the heat of the sun out of ABI. Predictive Calibration (pCal) is an algorithm developed by instrument vendors for the National Oceanic Atmospheric Agency (NOAA) …


Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee Jan 2021

Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee

CMC Senior Theses

This paper attempts to quantify predictive power of social media sentiment and financial data in stock prediction by utilizing a comprehensive set of stock-related fundamental and technical variables and social media sentiments. For conducting sentiment analysis, this study employs a pretrained finBERT model that provides three different sentiment classifications and respective softmax scores. Hence, the significance of these variables is evaluated with XGBoost regression and Shapley Additive exPlanations (SHAP) frameworks. Through investigating feature importance, this study finds that statistical properties of sentiment variables provide a stronger predictive power than a weighted sentiment score and that it is possible to quantify …


Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard Jan 2021

Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard

CMC Senior Theses

Using a Twitter API, I gather and analyze tweets by performing sentiment analysis to solve the GOAT debate among professional athletes with the primary focus on comparing Michael Jordan and LeBron James. Athletes from the National Football League (NFL), the National Basketball Association (NBA), Major League Baseball (MLB), and the National Collegiate Athletic Association (NCAA) Division 1 Men's and Women's Basketball were selected to compare how sentiment polarity varies across sports. Sentiment polarity is measured by labeling text as "positive", "neutral", or "negative" which allows us to determine which athlete/sport is highly favored among the Twitter community when it comes …


An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill Jan 2021

An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill

Articles

This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble's annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to …


Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems, Rui Wang, Yusheng Xu Jan 2021

Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems, Rui Wang, Yusheng Xu

Mathematics & Statistics Faculty Publications

Learning a function from a finite number of sampled data points (measurements) is a fundamental problem in science and engineering. This is often formulated as a minimum norm interpolation (MNI) problem, a regularized learning problem or, in general, a semi discrete inverse problem (SDIP), in either Hilbert spaces or Banach spaces. The goal of this paper is to systematically study solutions of these problems in Banach spaces. We aim at obtaining explicit representer theorems for their solutions, on which convenient solution methods can then be developed. For the MNI problem, the explicit representer theorems enable us to express the infimum …


A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments, Roger Sayre, Kevin Butler, Keith Van Graafeiland, Sean Breyer, Dawn Wright, Charlie Frye, Deniz Karagulle, Madeline Martin, Jill Cress, Tom Allen, Rebecca J. Allee, Rost Parsons, Bjorn Nyberg, Mark J. Costello, Peter Harris, Frank E. Muller-Karger Jan 2021

A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments, Roger Sayre, Kevin Butler, Keith Van Graafeiland, Sean Breyer, Dawn Wright, Charlie Frye, Deniz Karagulle, Madeline Martin, Jill Cress, Tom Allen, Rebecca J. Allee, Rost Parsons, Bjorn Nyberg, Mark J. Costello, Peter Harris, Frank E. Muller-Karger

Political Science & Geography Faculty Publications

A new data layer provides Coastal and Marine Ecological Classification Standard (CMECS) labels for global coastal segments at 1 km or shorter resolution. These characteristics are summarized for six US Marine Biodiversity Observation Network (MBON) sites and one MBON Pole to Pole of the Americas site in Argentina. The global coastlines CMECS classifications were produced from a partitioning of a 30 m Landsat-derived shoreline vector that was segmented into 4 million 1 km or shorter segments. Each segment was attributed with values from 10 variables that represent the ecological settings in which the coastline occurs, including properties of the adjacent …


Ensemble Encoder-Decoder Models For Predicting Land Transformation, Pariya Pourmohammadi Jan 2021

Ensemble Encoder-Decoder Models For Predicting Land Transformation, Pariya Pourmohammadi

Graduate Theses, Dissertations, and Problem Reports (ETD)

In studying dynamic and complex processes which are influenced by a system of inter-connected driving variables, it is crucial to apply models that can learn the complexity of the interactions. Land transformation is one of such complex processes, prediction of which can help to mitigate severe climate situations and improve the resiliency of communities. In this study, a multi-spectral set of data cubes is used to capture various characteristics of a geographic region. Based on the data cube, a feature space is constructed using socio-economic attributes, terrain characteristics, and landscape traits of the study region. Two-dimensional and three-dimensional convolutional neural …


Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries, Chibuzo Iloabuchi Jan 2021

Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries, Chibuzo Iloabuchi

Graduate Theses, Dissertations, and Problem Reports (ETD)

In the United States (US), 25% of healthcare spending is considered wasteful because it is spent reimbursing low-value care. Low-value care is the utilization of healthcare services, medical tests, and procedures that have unclear or no clinical benefit to patients but still exposes them to risk. World-wide, low-value care imposes a significant economic burden on patients, payers, governments, and society. Cancer care among older adults > 65 years is one of the biggest drivers of healthcare expenditure in the US and accounts for nearly 40% of all spending, and low-value care among cancer patients is prevalent and contributes to the financial …


Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington Jan 2021

Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington

Graduate Theses, Dissertations, and Problem Reports (ETD)

The necessary materials for most human activities are water and energy. Integrated analysis to accurately forecast water and energy consumption enables the implementation of efficient short and long-term resource management planning as well as expanding policy and research possibilities for the supportive infrastructure. However, the integral relationship between water and energy (water-energy nexus) poses a difficult problem for modeling. The accessibility and physical overlay of data sets related to water-energy nexus is another main issue for a reliable water-energy consumption forecast. The framework of urban metabolism (UM) uses several types of data to build a global view and highlight issues …


A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh Jan 2021

A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh

Dartmouth College Master’s Theses

Computational methods for predicting binding interfaces between antigens and antibodies (epitopes and paratopes) are faster and cheaper than traditional experimental structure determination methods. A sufficiently reliable computational predictor that could scale to large sets of available antibody sequence data could thus inform and expedite many biomedical pursuits, such as better understanding immune responses to vaccination and natural infection and developing better drugs and vaccines. However, current state-of-the-art predictors produce discontiguous predictions, e.g., predicting the epitope in many different spots on an antigen, even though in reality they typically comprise a single localized region. We seek to produce contiguous predicted epitopes, …


Why We Need Better Corporate Governance Data, Jens Frankenreiter, Cathy Hwang, Yaron Nili, Eric L. Talley Jan 2021

Why We Need Better Corporate Governance Data, Jens Frankenreiter, Cathy Hwang, Yaron Nili, Eric L. Talley

Scholarship@WashULaw

Three decades of finance, economics, and legal studies in corporate governance have been built substantially on data sets with nearly unknown provenance. A new paper sets to correct this fatal flaw of contemporary corporate governance research by debuting a brand new resource—the Cleaning Corporate Governance database.


Research Data Curation And Management Bibliography, Charles W. Bailey Jr. Jan 2021

Research Data Curation And Management Bibliography, Charles W. Bailey Jr.

Copyright, Fair Use, Scholarly Communication, etc.

Preface

The Research Data Curation and Management Bibliography includes over 800 selected English-language articles and books that are useful in understanding the curation of digital research data in academic and other research institutions.

The "digital curation" concept is still evolving. In "Digital Curation and Trusted Repositories: Steps toward Success," Christopher A. Lee and Helen R. Tibbo define digital curation as follows:

Digital curation involves selection and appraisal by creators and archivists; evolving provision of intellectual access; redundant storage; data transformations; and, for some materials, a commitment to long-term preservation. Digital curation is stewardship that provides for the reproducibility and re-use …


Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam Jan 2021

Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic triggered a large-scale work-from-home trend globally in recent months. In this paper, we study the phenomenon of “work-from-home” (WFH) by performing social listening. We propose an analytics pipeline designed to crawl social media data and perform text mining analyzes on textual data from tweets scrapped based on hashtags related to WFH in COVID-19 situation. We apply text mining and NLP techniques to analyze the tweets for extracting the WFH themes and sentiments (positive and negative). Our Twitter theme analysis adds further value by summarizing the common key topics, allowing employers to gain more insights on areas of …


Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee Jan 2021

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee

Senior Independent Study Theses

According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.

We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …


Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang Jan 2021

Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang

Faculty, Staff and Student Publications

Disordered sleep is associated with poor cognitive function and cognitive decline. However, little is known regarding the association of sleep-related factors with cognitive function in underrepresented cohorts such as the Hispanic/Latino population. Leveraging the National Sleep Research Resource, one of the most comprehensive collections of sleep studies, we identified a Hispanic/Latino cohort of 1,031 lower cognitive function cases and 2,062 normal controls. We developed a novel dual random forest (DRF) approach to discriminate cases against controls for estimating the potential impact of sleep-related variables related to the decline of cognitive function. Several important sleep-related factors were identified which may be …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke Jan 2021

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue Jan 2021

Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue

Graduate Student Theses, Dissertations, & Professional Papers

Steady improvements to the instruments used in remote sensing has led to much higher resolution data, often contemporaneous with lower resolution instruments that continue to collect data. There is a clear opportunity to reconcile recent high resolution satellite data with the lower resolution data of the past. Super-resolution (SR) imaging is a technique that increases the spatial resolution of image data by training statistical methods on simultaneously occurring lower and higher resolution data sets. The special sensor microwave/imager (SSMI) and advanced microwave scanning radiometer (AMSR2) brightness temperature data products are well suited to super-resolution imaging, and SR can be used …


Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle Jan 2021

Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle

All Master's Theses

Machine learning algorithms often produce models considered as complex black-box models by both end users and developers. Such algorithms fail to explain the model in terms of the domain they are designed for. The proposed Iterative Visual Logical Classifier (IVLC) is an interpretable machine learning algorithm that allows end users to design a model and classify data with more confidence and without having to compromise on the accuracy. Such technique is especially helpful when dealing with sensitive and crucial data like cancer data in the medical domain with high cost of errors. With the help of the proposed interactive and …


K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant Jan 2021

K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant

Theses and Dissertations

Traditional density-based clustering approaches rely on a distance-based parameter to define data connectivity and density. However, an appropriate value of this parameter can be difficult to determine as it is highly dependent on the underlying distribution of the data. In particular, distribution parameters affect the scale of inter-group distances (e.g., variance); this dependence leads to a well-known inability to simultaneously detect clusters at varying levels of density. In this work, connectivity and density are defined according to the rank-order induced by the distance metric (i.e., invariant to the expected scale of the distances). Connectivity by k-nearest neighbors and density by …