Clustering Data To Classify Hearthstone Decks,
2021
The University of Akron
Clustering Data To Classify Hearthstone Decks, Tim Inzitari
Williams Honors College, Honors Research Projects
The esports game of "Hearthstone" is a collectible card game with a competitive format that has every team submit 4 decks of 30 cards each. Using K-Means clustering an adaptable way to group data for classifying can be made that works well in every update of the game. This system will take in a list of decks and cluster them to easily classify large amounts of information in a timely fashion. This system will be able to be used by the Universities esports department for years to come to aid the preparation of "Hearthstone" matches. This model uses qualities about …
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students,
2021
Amherst College
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr
Statistical and Data Sciences: Faculty Publications
We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.
Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse,
2021
Dakota State University
Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse, Omar El-Gayar, Abdullah Wahbeh, Tareq Nasralah, Ahmed El Noshokaty, Mohammad A. Al-Ramahi
Computer Information Systems Faculty Publications (Archived)
This study analyzed Twitter discourse to understand the association of the COVID-19 pandemic with mental health. The study compared tweets’ volume over time, tweets’ volume per mental health category, emotions, and the top hashtags on mental health before and after November 2019, the month on which the first COVID-19 case was reported. We analyzed a total of 273 million English tweets on mental health collected from 56 million unique users. Results and analysis showed a significant shift in trend for the volume of tweets on mental health over time. There was also a notable increase in the volume of tweets …
Learning From Multi-Class Imbalanced Big Data With Apache Spark,
2021
Virginia Commonwealth University
Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv
Theses and Dissertations
With data becoming a new form of currency, its analysis has become a top priority in both academia and industry, furthering advancements in high-performance computing and machine learning. However, these large, real-world datasets come with additional complications such as noise and class overlap. Problems are magnified when with multi-class data is presented, especially since many of the popular algorithms were originally designed for binary data. Another challenge arises when the number of examples are not evenly distributed across all classes in a dataset. This often causes classifiers to favor the majority class over the minority classes, leading to undesirable results …
A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity,
2021
Michigan Technological University
A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity, Adewale Aremu Adesanya
Dissertations, Master's Theses and Master's Reports
The transition to using clean, affordable, and reliable electrical energy is critical for enhancing human opportunities and capabilities. In the United States, many states and localities are engaging in this transition despite the lack of ambitious federal policy support. This research builds on the theoretical framework of the multilevel perspective (MLP) of sociotechnical transitions as well as the concept of energy justice to investigate potential pathways to 100 percent renewable energy (RE) for electricity provision in the U.S. This research seeks to answer the question: what are the technical, policy, and perceptual pathways, barriers, and opportunities for just transition to …
Explainable Feature- And Decision-Level Fusion,
2021
Michigan Technological University
Explainable Feature- And Decision-Level Fusion, Siva Krishna Kakula
Dissertations, Master's Theses and Master's Reports
Information fusion is the process of aggregating knowledge from multiple data sources to produce more consistent, accurate, and useful information than any one individual source can provide. In general, there are three primary sources of data/information: humans, algorithms, and sensors. Typically, objective data---e.g., measurements---arise from sensors. Using these data sources, applications such as computer vision and remote sensing have long been applying fusion at different "levels" (signal, feature, decision, etc.). Furthermore, the daily advancement in engineering technologies like smart cars, which operate in complex and dynamic environments using multiple sensors, are raising both the demand for and complexity of fusion. …
Using Text Mining And Machine Learning Classifiers To Analyze Stack Overflow,
2021
Michigan Technological University
Using Text Mining And Machine Learning Classifiers To Analyze Stack Overflow, Taylor Morris
Dissertations, Master's Theses and Master's Reports
StackOverflow is an extensively used platform for programming questions. In this report, text mining and machine learning classifiers such as decision trees and Naive Bayes are used to evaluate whether a given question posted on StackOverflow will be closed or answered. While multiple models were used in the analysis, the performance for the models was no better than the majority classifier. Future work to develop better performing classifiers to understand why a question is closed or answered will require additional natural language processing or methods to address the imbalanced data.
Regional Impacts Of Invasive Species And Climate Change On Black Ash Wetlands,
2021
Michigan Technological University
Regional Impacts Of Invasive Species And Climate Change On Black Ash Wetlands, Joseph Shannon
Dissertations, Master's Theses and Master's Reports
For more than a decade intensive research on the ecohydrology of black ash wetland ecosystems has been performed to understand these systems before they are drastically altered by the invasive species, emerald ash borer (EAB). In that time there has been little research aimed at the scale and persistence of the alterations. Three distinct but related research articles will be presented to demonstrate a method for moderate resolution mapping of black ash across its entire range, understand the relative impacts of EAB and climate change on probable future wetland conditions, and develop an experimental and modeling approach to quantify and …
Goes-R Supervised Machine Learning,
2021
CUNY City College
Goes-R Supervised Machine Learning, Ronald Adomako
Dissertations and Theses
The GOES-R series is a product line of four satellite, with two currently on-orbit (GOES-16 “East” and GOES-17 “West”). GOES-17 is susceptible to a Loop-Heat-Pipe (LHP) phenomenon where during Fall and Spring seasons, there are times of day where some of the infrared bands records inaccurate readings from the Advanced Baseline Imager (ABI). This occurs from joint astronomical behavior and position of the GOES-17. This calibration issue occurs when the LHP instrument fails to radiate the heat of the sun out of ABI. Predictive Calibration (pCal) is an algorithm developed by instrument vendors for the National Oceanic Atmospheric Agency (NOAA) …
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification,
2021
Claremont McKenna College
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee
CMC Senior Theses
This paper attempts to quantify predictive power of social media sentiment and financial data in stock prediction by utilizing a comprehensive set of stock-related fundamental and technical variables and social media sentiments. For conducting sentiment analysis, this study employs a pretrained finBERT model that provides three different sentiment classifications and respective softmax scores. Hence, the significance of these variables is evaluated with XGBoost regression and Shapley Additive exPlanations (SHAP) frameworks. Through investigating feature importance, this study finds that statistical properties of sentiment variables provide a stronger predictive power than a weighted sentiment score and that it is possible to quantify …
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James,
2021
Claremont Colleges
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard
CMC Senior Theses
Using a Twitter API, I gather and analyze tweets by performing sentiment analysis to solve the GOAT debate among professional athletes with the primary focus on comparing Michael Jordan and LeBron James. Athletes from the National Football League (NFL), the National Basketball Association (NBA), Major League Baseball (MLB), and the National Collegiate Athletic Association (NCAA) Division 1 Men's and Women's Basketball were selected to compare how sentiment polarity varies across sports. Sentiment polarity is measured by labeling text as "positive", "neutral", or "negative" which allows us to determine which athlete/sport is highly favored among the Twitter community when it comes …
An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags,
2021
Bowling Green State University
An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill
Articles
This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble's annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to …
Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems,
2021
Old Dominion University
Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems, Rui Wang, Yusheng Xu
Mathematics & Statistics Faculty Publications
Learning a function from a finite number of sampled data points (measurements) is a fundamental problem in science and engineering. This is often formulated as a minimum norm interpolation (MNI) problem, a regularized learning problem or, in general, a semi discrete inverse problem (SDIP), in either Hilbert spaces or Banach spaces. The goal of this paper is to systematically study solutions of these problems in Banach spaces. We aim at obtaining explicit representer theorems for their solutions, on which convenient solution methods can then be developed. For the MNI problem, the explicit representer theorems enable us to express the infimum …
Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries,
2021
West Virginia University
Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries, Chibuzo Iloabuchi
Graduate Theses, Dissertations, and Problem Reports (ETD)
In the United States (US), 25% of healthcare spending is considered wasteful because it is spent reimbursing low-value care. Low-value care is the utilization of healthcare services, medical tests, and procedures that have unclear or no clinical benefit to patients but still exposes them to risk. World-wide, low-value care imposes a significant economic burden on patients, payers, governments, and society. Cancer care among older adults > 65 years is one of the biggest drivers of healthcare expenditure in the US and accounts for nearly 40% of all spending, and low-value care among cancer patients is prevalent and contributes to the financial …
A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments,
2021
Old Dominion University
A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments, Roger Sayre, Kevin Butler, Keith Van Graafeiland, Sean Breyer, Dawn Wright, Charlie Frye, Deniz Karagulle, Madeline Martin, Jill Cress, Tom Allen, Rebecca J. Allee, Rost Parsons, Bjorn Nyberg, Mark J. Costello, Peter Harris, Frank E. Muller-Karger
Political Science & Geography Faculty Publications
A new data layer provides Coastal and Marine Ecological Classification Standard (CMECS) labels for global coastal segments at 1 km or shorter resolution. These characteristics are summarized for six US Marine Biodiversity Observation Network (MBON) sites and one MBON Pole to Pole of the Americas site in Argentina. The global coastlines CMECS classifications were produced from a partitioning of a 30 m Landsat-derived shoreline vector that was segmented into 4 million 1 km or shorter segments. Each segment was attributed with values from 10 variables that represent the ecological settings in which the coastline occurs, including properties of the adjacent …
Plant Species Identification In The Wild Based On Images Of Organs,
2021
West Virginia University
Plant Species Identification In The Wild Based On Images Of Organs, Meghana Kovur
Graduate Theses, Dissertations, and Problem Reports (ETD)
Image-based plant species identification in the wild is a difficult problem for several reasons. First, the input data is subject to a very high degree of variability because it is captured under fully unconstrained conditions. The same plant species may look very different in different images, while different species can often appear very similar, challenging even the recognition skills of human experts in the field. The large intra-class and small inter-class image variability makes this a fine-grained visual classification problem. One way to cope with this variability and to reduce image background noise is to predict species based on the …
Ensemble Encoder-Decoder Models For Predicting Land Transformation,
2021
West Virginia University
Ensemble Encoder-Decoder Models For Predicting Land Transformation, Pariya Pourmohammadi
Graduate Theses, Dissertations, and Problem Reports (ETD)
In studying dynamic and complex processes which are influenced by a system of inter-connected driving variables, it is crucial to apply models that can learn the complexity of the interactions. Land transformation is one of such complex processes, prediction of which can help to mitigate severe climate situations and improve the resiliency of communities. In this study, a multi-spectral set of data cubes is used to capture various characteristics of a geographic region. Based on the data cube, a feature space is constructed using socio-economic attributes, terrain characteristics, and landscape traits of the study region. Two-dimensional and three-dimensional convolutional neural …
Identification And Classification Of Radio Pulsar Signals Using Machine Learning,
2021
West Virginia University
Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang
Graduate Theses, Dissertations, and Problem Reports (ETD)
Automated single-pulse search approaches are necessary as ever-increasing amount of observed data makes the manual inspection impractical. Detecting radio pulsars using single-pulse searches, however, is a challenging problem for machine learning because pul- sar signals often vary significantly in brightness, width, and shape and are only detected in a small fraction of observed data.
The research work presented in this dissertation is focused on development of ma- chine learning algorithms and approaches for single-pulse searches in the time domain. Specifically, (1) We developed a two-stage single-pulse search approach, named Single- Pulse Event Group IDentification (SPEGID), which automatically identifies and clas- …
Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model,
2021
West Virginia University
Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington
Graduate Theses, Dissertations, and Problem Reports (ETD)
The necessary materials for most human activities are water and energy. Integrated analysis to accurately forecast water and energy consumption enables the implementation of efficient short and long-term resource management planning as well as expanding policy and research possibilities for the supportive infrastructure. However, the integral relationship between water and energy (water-energy nexus) poses a difficult problem for modeling. The accessibility and physical overlay of data sets related to water-energy nexus is another main issue for a reliable water-energy consumption forecast. The framework of urban metabolism (UM) uses several types of data to build a global view and highlight issues …
Why We Need Better Corporate Governance Data,
2021
Washington University in St. Louis School of Law
Why We Need Better Corporate Governance Data, Jens Frankenreiter, Cathy Hwang, Yaron Nili, Eric L. Talley
Scholarship@WashULaw
Three decades of finance, economics, and legal studies in corporate governance have been built substantially on data sets with nearly unknown provenance. A new paper sets to correct this fatal flaw of contemporary corporate governance research by debuting a brand new resource—the Cleaning Corporate Governance database.
