Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2671 - 2700 of 3244

Full-Text Articles in Data Science

Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini Jan 2021

Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini

Theses and Dissertations

Multi-label data streams are sequences of multi-label instances arriving over time to a multi-label classifier. The properties of the data stream may continuously change due to concept drift. Therefore, algorithms must adapt constantly to the new data distributions. In this paper we propose a novel ensemble method for multi-label drifting streams named Homogeneous Ensemble of Self-Adjusting Nearest Neighbors (HESAkNN). It leverages a self-adjusting kNN as a base classifier with the advantages of ensembles to adapt to concept drift in the multi-label environment. To promote diverse knowledge within the ensemble, each base classifier is given a unique subset of features and …


Change Detection And Landscape Similarity Comparison Using Computer Vision Methods, Karim Malik Wilfrid Laurier University Jan 2021

Change Detection And Landscape Similarity Comparison Using Computer Vision Methods, Karim Malik Wilfrid Laurier University

Theses and Dissertations (Comprehensive)

Human-induced disturbances of terrestrial and aquatic ecosystems continue at alarming rates. With the advent of both raw sensor and analysis-ready datasets, the need to monitor ecosystem disturbances is now more imperative than ever; yet the task is becoming increasingly complex with increasing sources and varieties of earth observation data. In this research, computer vision methods and tools are interrogated to understand their capability for comparing spatial patterns. A critical survey of literature provides evidence that computer vision methods are relatively robust to scale and highlights issues involved in parameterization of computer vision models for characterizing significant pattern information in a …


Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani Jan 2021

Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani

Theses and Dissertations (Comprehensive)

Hawkes process was evolved so that the past events contribute to the occurrence time of future events by self-exciting or mutually exciting. However, many real-world data do not follow the Hawkes process's assumptions (i.e., positivity, additivity, and exponential decay) and become more complex to be modeled by the traditional Hawkes processes, so the neural Hawkes process was developed to tackle the challenges. However, Recurrent Neural Networks (RNN) fail to capture long-term dependencies among multiple point processes, and Transformer Hawkes processes only address temporal characteristics of Hawkes processes. In this thesis, we proposed a combination of neural networks and Hawkes processes …


Binary Black Widow Optimization Algorithm For Feature Selection Problems, Ahmed Al-Saedi Jan 2021

Binary Black Widow Optimization Algorithm For Feature Selection Problems, Ahmed Al-Saedi

Theses and Dissertations (Comprehensive)

This thesis addresses feature selection (FS) problems, which is a primary stage in data mining. FS is a significant pre-processing stage to enhance the performance of the process with regards to computation cost and accuracy to offer a better comprehension of stored data by removing the unnecessary and irrelevant features from the basic dataset. However, because of the size of the problem, FS is known to be very challenging and has been classified as an NP-hard problem. Traditional methods can only be used to solve small problems. Therefore, metaheuristic algorithms (MAs) are becoming powerful methods for addressing the FS problems. …


Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang Jan 2021

Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang

Graduate Theses, Dissertations, and Problem Reports (ETD)

Automated single-pulse search approaches are necessary as ever-increasing amount of observed data makes the manual inspection impractical. Detecting radio pulsars using single-pulse searches, however, is a challenging problem for machine learning because pul- sar signals often vary significantly in brightness, width, and shape and are only detected in a small fraction of observed data.

The research work presented in this dissertation is focused on development of ma- chine learning algorithms and approaches for single-pulse searches in the time domain. Specifically, (1) We developed a two-stage single-pulse search approach, named Single- Pulse Event Group IDentification (SPEGID), which automatically identifies and clas- …


Plant Species Identification In The Wild Based On Images Of Organs, Meghana Kovur Jan 2021

Plant Species Identification In The Wild Based On Images Of Organs, Meghana Kovur

Graduate Theses, Dissertations, and Problem Reports (ETD)

Image-based plant species identification in the wild is a difficult problem for several reasons. First, the input data is subject to a very high degree of variability because it is captured under fully unconstrained conditions. The same plant species may look very different in different images, while different species can often appear very similar, challenging even the recognition skills of human experts in the field. The large intra-class and small inter-class image variability makes this a fine-grained visual classification problem. One way to cope with this variability and to reduce image background noise is to predict species based on the …


Estimating The Azimuthal Mode Structure Of Ultra Low Frequency Waves And Its Effects On The Radial Diffusion Of Radiation Belt Electrons, Mohammad Barani Jan 2021

Estimating The Azimuthal Mode Structure Of Ultra Low Frequency Waves And Its Effects On The Radial Diffusion Of Radiation Belt Electrons, Mohammad Barani

Graduate Theses, Dissertations, and Problem Reports (ETD)

Characterizing the azimuthal mode number �� of Ultra Low Frequency (ULF) waves is critical to quantifying the radial diffusion of radiation belt electrons. A Wavelet cross-spectral technique is applied to the compressional ULF waves observed by multiple pairs of GOES and MMS satellites to estimate the mode structure of ULF waves. A more realistic distribution of mode numbers is achieved by inclusion of the modes corresponding to different wave propagation directions as well as at �� higher than fundamental mode number. For the event study of a geomagnetic storm using GOES data, ULF wave power is found to dominate at …


K-Nearest Neighbour Classifiers - A Tutorial, Padraig Cunningham, Sarah Jane Delany Jan 2021

K-Nearest Neighbour Classifiers - A Tutorial, Padraig Cunningham, Sarah Jane Delany

Conference papers

Perhaps the most straightforward classifier in the arsenal or Machine Learning techniques is the Nearest Neighbour Classifier – classification is achieved by identifying the nearest neighbours to a query example and using those neighbours to determine the class of the query. This approach to classification is of particular importance because issues of poor run-time performance is not such a problem these days with the computational power that is available. This paper presents an overview of techniques for Nearest Neighbour classification focusing on; mechanisms for assessing similarity (distance), computational issues in identifying nearest neighbours and mechanisms for reducing the dimension of …


A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity, Adewale Aremu Adesanya Jan 2021

A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity, Adewale Aremu Adesanya

Dissertations, Master's Theses and Master's Reports

The transition to using clean, affordable, and reliable electrical energy is critical for enhancing human opportunities and capabilities. In the United States, many states and localities are engaging in this transition despite the lack of ambitious federal policy support. This research builds on the theoretical framework of the multilevel perspective (MLP) of sociotechnical transitions as well as the concept of energy justice to investigate potential pathways to 100 percent renewable energy (RE) for electricity provision in the U.S. This research seeks to answer the question: what are the technical, policy, and perceptual pathways, barriers, and opportunities for just transition to …


Explainable Feature- And Decision-Level Fusion, Siva Krishna Kakula Jan 2021

Explainable Feature- And Decision-Level Fusion, Siva Krishna Kakula

Dissertations, Master's Theses and Master's Reports

Information fusion is the process of aggregating knowledge from multiple data sources to produce more consistent, accurate, and useful information than any one individual source can provide. In general, there are three primary sources of data/information: humans, algorithms, and sensors. Typically, objective data---e.g., measurements---arise from sensors. Using these data sources, applications such as computer vision and remote sensing have long been applying fusion at different "levels" (signal, feature, decision, etc.). Furthermore, the daily advancement in engineering technologies like smart cars, which operate in complex and dynamic environments using multiple sensors, are raising both the demand for and complexity of fusion. …


Using Text Mining And Machine Learning Classifiers To Analyze Stack Overflow, Taylor Morris Jan 2021

Using Text Mining And Machine Learning Classifiers To Analyze Stack Overflow, Taylor Morris

Dissertations, Master's Theses and Master's Reports

StackOverflow is an extensively used platform for programming questions. In this report, text mining and machine learning classifiers such as decision trees and Naive Bayes are used to evaluate whether a given question posted on StackOverflow will be closed or answered. While multiple models were used in the analysis, the performance for the models was no better than the majority classifier. Future work to develop better performing classifiers to understand why a question is closed or answered will require additional natural language processing or methods to address the imbalanced data.


Regional Impacts Of Invasive Species And Climate Change On Black Ash Wetlands, Joseph Shannon Jan 2021

Regional Impacts Of Invasive Species And Climate Change On Black Ash Wetlands, Joseph Shannon

Dissertations, Master's Theses and Master's Reports

For more than a decade intensive research on the ecohydrology of black ash wetland ecosystems has been performed to understand these systems before they are drastically altered by the invasive species, emerald ash borer (EAB). In that time there has been little research aimed at the scale and persistence of the alterations. Three distinct but related research articles will be presented to demonstrate a method for moderate resolution mapping of black ash across its entire range, understand the relative impacts of EAB and climate change on probable future wetland conditions, and develop an experimental and modeling approach to quantify and …


Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang Jan 2021

Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang

Theses and Dissertations (Comprehensive)

Terrorism becomes more rampant in recent years because of separatism and extreme nationalism, which brings a serious threat to the national security of many countries in the world. The analysis of spatial and temporal patterns of terror data is significant in containing terrorism. This thesis focuses on building and applying a temporal point process called self-exciting point process to fit the terror data from 1970 to 2018 of 10 countries. The data come from the Global Terrorism database. Further, an application in predicting the number of terror events based on the self-exciting model is another main innovative idea, in which …


Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott Dec 2020

Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott

LSU Doctoral Dissertations

Modern human-machine systems such as microservices rely upon agile engineering practices which require changes to be tested and released more frequently than classically engineered systems. A critical step in the testing of such systems is the generation of realistic workloads or load testing. Generated workload emulates the expected behaviors of users and machines within a system under test in order to find potentially unknown failure states. Typical testing tools rely on static testing artifacts to generate realistic workload conditions. Such artifacts can be cumbersome and costly to maintain; however, even model-based alternatives can prevent adaptation to changes in a system …


Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao Dec 2020

Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao

Articles

It is often the case with new technologies that it is very hard to predict their long-term impacts and as a result, although new technology may be beneficial in the short term, it can still cause problems in the longer term. This is what happened with oil by-products in different areas: the use of plastic as a disposable material did not take into account the hundreds of years necessary for its decomposition and its related long-term environmental damage. Data is said to be the new oil. The message to be conveyed is associated with its intrinsic value. But as in …


Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar Dec 2020

Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar

Master of Science in Computer Science Theses

Under partial shading conditions, photovoltaic (PV) modules in a solar array experience varying irradiance. A Global Maximum (GM) and multiple Local Maximums (LMs) can originate on the Power-Voltage (P-V) curve under nonuniform irradiance conditions. There are many maximum power point tracking (MPPT) algorithms developed to detect the true maximum power point (MPP) of a PV array. However, in the real-world environment, limited samples of power-voltage (P-V) data might be available to quickly and accurately predict the position of the global maximum point. Since the change of environmental conditions are dynamic, limited time is available to locate the global peak. Machine …


Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang Dec 2020

Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang

Statistical Science Theses and Dissertations

This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.

In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …


Analysis Of Github Pull Requests, Canon Ellis Dec 2020

Analysis Of Github Pull Requests, Canon Ellis

Computer Science and Engineering Theses and Dissertations

The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …


Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran Dec 2020

Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran

LSU New Orleans Theses and Dissertations

This work is concerned with the viability of Machine Learning (ML) in training models for predicting global bathymetry, and whether there is a best fit model for predicting that bathymetry. The desired result is an investigation of the ability for ML to be used in future prediction models and to experiment with multiple trained models to determine an optimum selection. Ocean features were aggregated from a set of external studies and placed into two minute spatial grids representing the earth's oceans. A set of regression models, classification models, and a novel classification model were then fit to this data and …


Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner Dec 2020

Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner

Master of Science in Computer Science Theses

Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …


Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud Dec 2020

Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud

SMU Data Science Review

A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …


Data Science In The Time Of Covid-19, Tony Breitzman Dec 2020

Data Science In The Time Of Covid-19, Tony Breitzman

College of Science & Mathematics Departmental Research

No abstract provided.


Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi Dec 2020

Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi

Publications and Research

In Spring 2020, I did a project, "Decision Tree Predicting the Party of Legislators," and construct a decision tree model to predict legislators' parties' based on their votes. We also use this model to identify legislators who frequently voted against their parties. We used the legislators' roll call votes, Office of Clerk U.S. House of Representatives Data Sets (Categorical values) collected in 2018 and 2019. In this new project, We study the 2018 and 2019 vote data using Principal Component Analysis (PCA). The goal is to find a (compressed) model using unsupervised learning to distinguish the legislators' parties, and PCA …


Introduction To Data Science Lti 110, Joanna Burkhardt Dec 2020

Introduction To Data Science Lti 110, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Introduction To Data Science, Joanna Burkhardt Dec 2020

Introduction To Data Science, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi Dec 2020

Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi

Student Scholar Symposium Abstracts and Posters

Visual processing in humans is done by integrating and updating multiple streams of global and local sensory input. Interaction between these two systems can be disrupted in individuals with ASD and other learning disabilities. When this integration is not done smoothly, it becomes difficult to see the “big picture”, which has been found to have implications on emotion recognition, social skills, and conversation skills. An example of this phenomenon is local interference, which is when local details are prioritized over the global features. Previous research in this field has aimed to decrease local interference by developing and evaluating a filter …


Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke Dec 2020

Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke

Library Philosophy and Practice (e-journal)

Computer science is a burgeoning research field and has the potential to accelerate the rate of industrialisation and subsequently, economic development. Using bibliometric data obtained from Scopus, this study employed a 15-year bibliometric analysis to highlight Nigeria’s productivity and impact trends in the computer science research landscape. Our findings are summarised as follows: First, Nigeria’s computer science research contribution and citations are meager in comparison to the global output. Secondly, international collaboration is generally weak as most collaborations are national in scope. Third, Nigeria’s computer science-related research is published in low-quality outlets, as Scopus has discontinued the indexing of most …


Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz Dec 2020

Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz

Dartmouth Scholarship

Recent advances in wearable sensor technologies have led to a variety of approaches for detecting physiological stress. Even with over a decade of research in the domain, there still exist many significant challenges, including a near-total lack of reproducibility across studies. Researchers often use some physiological sensors (custom-made or off-the-shelf), conduct a study to collect data, and build machine-learning models to detect stress. There is little effort to test the applicability of the model with similar physiological data collected from different devices, or the efficacy of the model on data collected from different studies, populations, or demographics.

This paper takes …


Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker Dec 2020

Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker

Publications and Research

Ocean observing systems are well-recognized as platforms for long-term monitoring of near-shore and remote locations in the global ocean. High-quality observatory data is freely available and accessible to all members of the global oceanographic community—a democratization of data that is particularly useful for early career scientists (ECS), enabling ECS to conduct research independent of traditional funding models or access to laboratory and field equipment. The concurrent collection of distinct data types with relevance for oceanographic disciplines including physics, chemistry, biology, and geology yields a unique incubator for cutting-edge, timely, interdisciplinary research. These data are both an opportunity and an incentive …


Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang Dec 2020

Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang

School of Continuing and Professional Studies Student Papers

No abstract provided.