Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 134 of 155.

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee 2021 The College of Wooster

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee

Senior Independent Study Theses

According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.

We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …


Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang 2021 The Texas Medical Center Library

Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang

Faculty, Staff and Student Publications

Disordered sleep is associated with poor cognitive function and cognitive decline. However, little is known regarding the association of sleep-related factors with cognitive function in underrepresented cohorts such as the Hispanic/Latino population. Leveraging the National Sleep Research Resource, one of the most comprehensive collections of sleep studies, we identified a Hispanic/Latino cohort of 1,031 lower cognitive function cases and 2,062 normal controls. We developed a novel dual random forest (DRF) approach to discriminate cases against controls for estimating the potential impact of sleep-related variables related to the decline of cognitive function. Several important sleep-related factors were identified which may be …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke 2021 University of Montana, Missoula

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue 2021 University of Montana

Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue

Graduate Student Theses, Dissertations, & Professional Papers

Steady improvements to the instruments used in remote sensing has led to much higher resolution data, often contemporaneous with lower resolution instruments that continue to collect data. There is a clear opportunity to reconcile recent high resolution satellite data with the lower resolution data of the past. Super-resolution (SR) imaging is a technique that increases the spatial resolution of image data by training statistical methods on simultaneously occurring lower and higher resolution data sets. The special sensor microwave/imager (SSMI) and advanced microwave scanning radiometer (AMSR2) brightness temperature data products are well suited to super-resolution imaging, and SR can be used …


Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle 2021 Central Washington University

Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle

All Master's Theses

Machine learning algorithms often produce models considered as complex black-box models by both end users and developers. Such algorithms fail to explain the model in terms of the domain they are designed for. The proposed Iterative Visual Logical Classifier (IVLC) is an interpretable machine learning algorithm that allows end users to design a model and classify data with more confidence and without having to compromise on the accuracy. Such technique is especially helpful when dealing with sensitive and crucial data like cancer data in the medical domain with high cost of errors. With the help of the proposed interactive and …


K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant 2021 Virginia Commonwealth University

K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant

Theses and Dissertations

Traditional density-based clustering approaches rely on a distance-based parameter to define data connectivity and density. However, an appropriate value of this parameter can be difficult to determine as it is highly dependent on the underlying distribution of the data. In particular, distribution parameters affect the scale of inter-group distances (e.g., variance); this dependence leads to a well-known inability to simultaneously detect clusters at varying levels of density. In this work, connectivity and density are defined according to the rank-order induced by the distance metric (i.e., invariant to the expected scale of the distances). Connectivity by k-nearest neighbors and density by …


Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini 2021 Virginia Commonwealth University

Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini

Theses and Dissertations

Multi-label data streams are sequences of multi-label instances arriving over time to a multi-label classifier. The properties of the data stream may continuously change due to concept drift. Therefore, algorithms must adapt constantly to the new data distributions. In this paper we propose a novel ensemble method for multi-label drifting streams named Homogeneous Ensemble of Self-Adjusting Nearest Neighbors (HESAkNN). It leverages a self-adjusting kNN as a base classifier with the advantages of ensembles to adapt to concept drift in the multi-label environment. To promote diverse knowledge within the ensemble, each base classifier is given a unique subset of features and …


Research Data Curation And Management Bibliography, Charles W. Bailey Jr. 2021 Digital Scholarship

Research Data Curation And Management Bibliography, Charles W. Bailey Jr.

Copyright, Fair Use, Scholarly Communication, etc.

Preface

The Research Data Curation and Management Bibliography includes over 800 selected English-language articles and books that are useful in understanding the curation of digital research data in academic and other research institutions.

The "digital curation" concept is still evolving. In "Digital Curation and Trusted Repositories: Steps toward Success," Christopher A. Lee and Helen R. Tibbo define digital curation as follows:

Digital curation involves selection and appraisal by creators and archivists; evolving provision of intellectual access; redundant storage; data transformations; and, for some materials, a commitment to long-term preservation. Digital curation is stewardship that provides for the reproducibility and re-use …


Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna GOTTIPATI, Kyong Jin SHIM, Hui Hian TEO, Karthik NITYANAND, Shreyansh SHIVAM 2021 Singapore Management University

Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic triggered a large-scale work-from-home trend globally in recent months. In this paper, we study the phenomenon of “work-from-home” (WFH) by performing social listening. We propose an analytics pipeline designed to crawl social media data and perform text mining analyzes on textual data from tweets scrapped based on hashtags related to WFH in COVID-19 situation. We apply text mining and NLP techniques to analyze the tweets for extracting the WFH themes and sentiments (positive and negative). Our Twitter theme analysis adds further value by summarizing the common key topics, allowing employers to gain more insights on areas of …


Change Detection And Landscape Similarity Comparison Using Computer Vision Methods, Karim Malik Wilfrid Laurier University 2021 Wilfrid Laurier University

Change Detection And Landscape Similarity Comparison Using Computer Vision Methods, Karim Malik Wilfrid Laurier University

Theses and Dissertations (Comprehensive)

Human-induced disturbances of terrestrial and aquatic ecosystems continue at alarming rates. With the advent of both raw sensor and analysis-ready datasets, the need to monitor ecosystem disturbances is now more imperative than ever; yet the task is becoming increasingly complex with increasing sources and varieties of earth observation data. In this research, computer vision methods and tools are interrogated to understand their capability for comparing spatial patterns. A critical survey of literature provides evidence that computer vision methods are relatively robust to scale and highlights issues involved in parameterization of computer vision models for characterizing significant pattern information in a …


Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani 2021 Wilfrid Laurier University

Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani

Theses and Dissertations (Comprehensive)

Hawkes process was evolved so that the past events contribute to the occurrence time of future events by self-exciting or mutually exciting. However, many real-world data do not follow the Hawkes process's assumptions (i.e., positivity, additivity, and exponential decay) and become more complex to be modeled by the traditional Hawkes processes, so the neural Hawkes process was developed to tackle the challenges. However, Recurrent Neural Networks (RNN) fail to capture long-term dependencies among multiple point processes, and Transformer Hawkes processes only address temporal characteristics of Hawkes processes. In this thesis, we proposed a combination of neural networks and Hawkes processes …


K-Nearest Neighbour Classifiers - A Tutorial, Padraig Cunningham, Sarah Jane Delany 2021 University College Dublin

K-Nearest Neighbour Classifiers - A Tutorial, Padraig Cunningham, Sarah Jane Delany

Conference papers

Perhaps the most straightforward classifier in the arsenal or Machine Learning techniques is the Nearest Neighbour Classifier – classification is achieved by identifying the nearest neighbours to a query example and using those neighbours to determine the class of the query. This approach to classification is of particular importance because issues of poor run-time performance is not such a problem these days with the computational power that is available. This paper presents an overview of techniques for Nearest Neighbour classification focusing on; mechanisms for assessing similarity (distance), computational issues in identifying nearest neighbours and mechanisms for reducing the dimension of …


Binary Black Widow Optimization Algorithm For Feature Selection Problems, Ahmed Al-Saedi 2021 Wilfrid Laurier University

Binary Black Widow Optimization Algorithm For Feature Selection Problems, Ahmed Al-Saedi

Theses and Dissertations (Comprehensive)

This thesis addresses feature selection (FS) problems, which is a primary stage in data mining. FS is a significant pre-processing stage to enhance the performance of the process with regards to computation cost and accuracy to offer a better comprehension of stored data by removing the unnecessary and irrelevant features from the basic dataset. However, because of the size of the problem, FS is known to be very challenging and has been classified as an NP-hard problem. Traditional methods can only be used to solve small problems. Therefore, metaheuristic algorithms (MAs) are becoming powerful methods for addressing the FS problems. …


Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang 2021 Wilfrid Laurier University

Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang

Theses and Dissertations (Comprehensive)

Terrorism becomes more rampant in recent years because of separatism and extreme nationalism, which brings a serious threat to the national security of many countries in the world. The analysis of spatial and temporal patterns of terror data is significant in containing terrorism. This thesis focuses on building and applying a temporal point process called self-exciting point process to fit the terror data from 1970 to 2018 of 10 countries. The data come from the Global Terrorism database. Further, an application in predicting the number of terror events based on the self-exciting model is another main innovative idea, in which …


Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott 2020 Louisiana State University and Agricultural and Mechanical College

Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott

LSU Doctoral Dissertations

Modern human-machine systems such as microservices rely upon agile engineering practices which require changes to be tested and released more frequently than classically engineered systems. A critical step in the testing of such systems is the generation of realistic workloads or load testing. Generated workload emulates the expected behaviors of users and machines within a system under test in order to find potentially unknown failure states. Typical testing tools rely on static testing artifacts to generate realistic workload conditions. Such artifacts can be cumbersome and costly to maintain; however, even model-based alternatives can prevent adaptation to changes in a system …


Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao 2020 Technological University Dublin

Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao

Articles

It is often the case with new technologies that it is very hard to predict their long-term impacts and as a result, although new technology may be beneficial in the short term, it can still cause problems in the longer term. This is what happened with oil by-products in different areas: the use of plastic as a disposable material did not take into account the hundreds of years necessary for its decomposition and its related long-term environmental damage. Data is said to be the new oil. The message to be conveyed is associated with its intrinsic value. But as in …


Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar 2020 Kennesaw State University

Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar

Master of Science in Computer Science Theses

Under partial shading conditions, photovoltaic (PV) modules in a solar array experience varying irradiance. A Global Maximum (GM) and multiple Local Maximums (LMs) can originate on the Power-Voltage (P-V) curve under nonuniform irradiance conditions. There are many maximum power point tracking (MPPT) algorithms developed to detect the true maximum power point (MPP) of a PV array. However, in the real-world environment, limited samples of power-voltage (P-V) data might be available to quickly and accurately predict the position of the global maximum point. Since the change of environmental conditions are dynamic, limited time is available to locate the global peak. Machine …


Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, GUANSHEN WANG 2020 Southern Methodist University

Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang

Statistical Science Theses and Dissertations

This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.

In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …


Analysis Of Github Pull Requests, Canon Ellis 2020 Southern Methodist University

Analysis Of Github Pull Requests, Canon Ellis

Computer Science and Engineering Theses and Dissertations

The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …


Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran 2020 LSU New Orleans

Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran

LSU New Orleans Theses and Dissertations

This work is concerned with the viability of Machine Learning (ML) in training models for predicting global bathymetry, and whether there is a best fit model for predicting that bathymetry. The desired result is an investigation of the ability for ML to be used in future prediction models and to experiment with multiple trained models to determine an optimum selection. Ocean features were aggregated from a set of external studies and placed into two minute spatial grids representing the earth's oceans. A set of regression models, classification models, and a novel classification model were then fit to this data and …


Digital Commons powered by bepress