Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 1 - 26 of 26

Full-Text Articles in Categorical Data Analysis

Adaptive Ensemble Of Classifiers With Regularization For Imbalanced Data Classification, Chen Wang, Chengyuan Deng, Zhoulu Yu, Dafeng Hui, Xiaofeng Gong, Ruisen Luo Dec 2020

Adaptive Ensemble Of Classifiers With Regularization For Imbalanced Data Classification, Chen Wang, Chengyuan Deng, Zhoulu Yu, Dafeng Hui, Xiaofeng Gong, Ruisen Luo

Biology Faculty Research

The dynamic ensemble selection of classifiers is an effective approach for processing label-imbalanced data classifications. However, such a technique is prone to overfitting, owing to the lack of regularization methods and the dependence on local geometry of data. In this study, focusing on binary imbalanced data classification, a novel dynamic ensemble method, namely adaptive ensemble of classifiers with regularization (AER), is proposed, to overcome the stated limitations. The method solves the overfitting problem through a new perspective of implicit regularization. Specifically, it leverages the properties of stochastic gradient descent to obtain the solution with the minimum norm, thereby achieving regularization; …


Can Statcast Variables Explain The Variation In Weighted Runs Created Plus?, Ryan Kupiec Dec 2020

Can Statcast Variables Explain The Variation In Weighted Runs Created Plus?, Ryan Kupiec

Student Research

The release of Statcast data in 2015 was revolutionary for data analysis in the game of baseball. Many analysts have begun using this data regularly, but none have used it exclusively. Often older, less reliable statistics (on-base percentage) are still used in favor of the newer statistics (weighted runs created plus). In this paper, we attempt to explain the variation in weighted runs created plus (wRC+) using Statcast variables such as exit velocity and launch angle. We find that exit velocity along with other Statcast variables, can explain as much as 70% of the variation in wRC+. Launch angle can …


Developing A Tourism Opportunity Index Regarding The Prospective Of Overtourism In Nepal, Susan Phuyal Dec 2020

Developing A Tourism Opportunity Index Regarding The Prospective Of Overtourism In Nepal, Susan Phuyal

Graduate Theses/Dissertations

This research explores Nepal's overtourism scenario based on the capacity of a locality to manage sustainable tourism practices. Environmental degradation, local infrastructure degradation, negative tourist experience and local resident responses regarding visitors are the four main variables used in this study to analyze overtourism. In order to analyze the case study of overtourism, we select the three top touristic cities of Nepal, Kathmandu, Pokhara, and Chitwan based on the number of annual visitors. Nepal's case analysis of overtourism conditions reviews the overall threat of over-tourism and establishes a metric by which tourism can be viewed as potentially detrimental to sustainability. …


Direct Questioning Of Sensitive Topics In Public Health Studies: A Simulation Study, Jessica K. Fox, Evrim Oral Nov 2020

Direct Questioning Of Sensitive Topics In Public Health Studies: A Simulation Study, Jessica K. Fox, Evrim Oral

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Predicting Postoperative Delirium Risk For Intracranial Surgery: A Statistical Machine Learning Approach, Juliet Aygun, Alaina Bartfeld, Sahana Rayan Aug 2020

Predicting Postoperative Delirium Risk For Intracranial Surgery: A Statistical Machine Learning Approach, Juliet Aygun, Alaina Bartfeld, Sahana Rayan

The Journal of Purdue Undergraduate Research

No abstract provided.


Evaluation Of China Shipping Hub-And-Spoke Network Based On Herfindahl-Hirschmann Index (Hhi), Wenjin Sun Aug 2020

Evaluation Of China Shipping Hub-And-Spoke Network Based On Herfindahl-Hirschmann Index (Hhi), Wenjin Sun

World Maritime University Dissertations

No abstract provided.


“Playing The Whole Game”: A Data Collection And Analysis Exercise With Google Calendar, Albert Y. Kim, Johanna Hardin Aug 2020

“Playing The Whole Game”: A Data Collection And Analysis Exercise With Google Calendar, Albert Y. Kim, Johanna Hardin

Statistical and Data Sciences: Faculty Publications

We provide a computational exercise suitable for early introduction in an undergraduate statistics or data science course that allows students to “play the whole game” of data science: performing both data collection and data analysis. While many teaching resources exist for data analysis, such resources are not as abundant for data collection given the inherent difficulty of the task. Our proposed exercise centers around student use of Google Calendar to collect data with the goal of answering the question “How do I spend my time?” On the one hand, the exercise involves answering a question with near universal appeal, but …


Analyzing The Fractal Dimension Of Various Musical Pieces, Nathan Clark Aug 2020

Analyzing The Fractal Dimension Of Various Musical Pieces, Nathan Clark

Industrial Engineering Undergraduate Honors Theses

One of the most common tools for evaluating data is regression. This technique, widely used by industrial engineers, explores linear relationships between predictors and the response. Each observation of the response is a fixed linear combination of the predictors with an added error element. The method is built on the assumption that this error is normally distributed across all observations and has a mean of zero. In some cases, it has been found that the inherent variation is not the result of a random variable, but is instead the result of self-symmetric properties of the observations. For data with these …


Integrating Data Science Ethics Into An Undergraduate Major, Benjamin Baumer, Randi L. Garcia, Albert Y. Kim, Katherine M. Kinnaird, Miles Q. Ott Jul 2020

Integrating Data Science Ethics Into An Undergraduate Major, Benjamin Baumer, Randi L. Garcia, Albert Y. Kim, Katherine M. Kinnaird, Miles Q. Ott

Statistical and Data Sciences: Faculty Publications

We present a programmatic approach to incorporating ethics into an undergraduate major in statistical and data sciences. We discuss departmental-level initiatives designed to meet the National Academy of Sciences recommendation for weaving ethics into the curriculum from top-to-bottom as our majors progress from our introductory courses to our senior capstone course, as well as from side-to-side through co-curricular programming. We also provide six examples of data science ethics modules used in five different courses at our liberal arts college, each focusing on a different ethical consideration. The modules are designed to be portable such that they can be flexibly incorporated …


Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett Jul 2020

Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett

HCA Healthcare Journal of Medicine

Electronic health records (EHRs) are an excellent source for secondary data analysis. Studies based on EHR-derived data, if designed properly, can answer previously unanswerable clinical research questions. In this paper we will highlight the benefits of large retrospective studies from secondary sources such as EHRs, examine retrospective cohort and case-control study design challenges, as well as methodological and statistical adjustment that can be made to overcome some of the inherent design limitations, in order to increase the generalizability, validity and reliability of the results obtained from these studies.


Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker Jul 2020

Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker

Graduate Theses and Dissertations

We study the use of distance correlation for statistical inference on categorical data, especially the induction of probability networks. Szekely et al. first defined distance correlation for continuous variables in [42], and Zhang translated the concept into the categorical setting in [57] by defining dCor(X,Y) for categorical variables X = (x1,...,xI) and Y = (y1,...,yJ) where P(X=xi)=[pi]i and P(Y=yi)=[pi]j with the formula [Please open the document]

Part I of the dissertation covers the background we need to understand this formula, and prepares us to analyze the properties and performance of its applications.

Part II then presents the main results of …


Next-Term Grade Prediction: A Machine Learning Approach, Audrey Tedja Widjaja, Lei Wang, Nghia Truong Trong, Aldy Gunawan, Ee-Peng Lim Jul 2020

Next-Term Grade Prediction: A Machine Learning Approach, Audrey Tedja Widjaja, Lei Wang, Nghia Truong Trong, Aldy Gunawan, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

As students progress in their university programs, they have to face many course choices. It is important for them to receive guidance based on not only their interest, but also the "predicted" course performance so as to improve learning experience and optimise academic performance. In this paper, we propose the next-term grade prediction task as a useful course selection guidance. We propose a machine learning framework to predict course grades in a specific program term using the historical student-course data. In this framework, we develop the prediction model using Factorization Machine (FM) and Long Short Term Memory combined with FM …


Analysis Of Gameplay Strategies In Hearthstone: A Data Science Approach, Connor W. Watson May 2020

Analysis Of Gameplay Strategies In Hearthstone: A Data Science Approach, Connor W. Watson

Theses

In recent years, games have been a popular test bed for AI research, and the presence of Collectible Card Games (CCGs) in that space is still increasing. One such CCG for both competitive/casual play and AI research is Hearthstone, a two-player adversarial game where players seeks to implement one of several gameplay strategies to defeat their opponent and decrease all of their Health points to zero. Although some open source simulators exist, some of their methodologies for simulated agents create opponents with a relatively low skill level. Using evolutionary algorithms, this thesis seeks to evolve agents with a higher skill …


Decision Tree For Predicting The Party Of Legislators, Afsana Mimi May 2020

Decision Tree For Predicting The Party Of Legislators, Afsana Mimi

Publications and Research

The motivation of the project is to identify the legislators who voted frequently against their party in terms of their roll call votes using Office of Clerk U.S. House of Representatives Data Sets collected in 2018 and 2019. We construct a model to predict the parties of legislators based on their votes. The method we used is Decision Tree from Data Mining. Python was used to collect raw data from internet, SAS was used to clean data, and all other calculations and graphical presentations are performed using the R software.


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya May 2020

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


Using Alteryx Designer In Audit, Nolan Asiala Apr 2020

Using Alteryx Designer In Audit, Nolan Asiala

Honors Projects

My senior project was built around data analysis and how it relates to the auditing profession. Initially, I was planning on attending a data analytics competition, but that was canceled due to the events of COVID-19. This project utilized the Alteryx Designer program to demonstrate how it can be used during an audit engagement. By creating a workflow in Alteryx Designer, a report from a client can be cleaned and reformatted into a working dataset. My project includes two Excel files, a Microsoft Word document that serves as a brief introduction to the program, and a video describing the workflow …


How Data Is Changing The World Of Healthcare, Cameron Marous Apr 2020

How Data Is Changing The World Of Healthcare, Cameron Marous

Honors Capstone Enhancement Presentations

No abstract provided.


Preparing For The Future: The Effects Of Financial Literacy On Financial Planning For Young Professionals, Tanay Singh Apr 2020

Preparing For The Future: The Effects Of Financial Literacy On Financial Planning For Young Professionals, Tanay Singh

Senior Theses

Purpose – Many people between the age of 20 and 34 have not considered planning financially for the future in any significant capacity and in doing so, they limit their potential savings. The purpose of this study is to examine what financial expectations are for people in the early stages of their career and determine if improving financial literacy and revealing financial realities helps to produce more accurate or realistic expectations. Ultimately, the goal is to better prepare participants in the study for the working world and increased responsibilities outside of the college/university environment by getting them to start thinking …


A Permutation Test And Spatial Cross-Validation Approach To Assess Models Of Interspecific Competition Between Trees, David Allen, Albert Y. Kim Mar 2020

A Permutation Test And Spatial Cross-Validation Approach To Assess Models Of Interspecific Competition Between Trees, David Allen, Albert Y. Kim

Statistical and Data Sciences: Faculty Publications

Measuring species-specific competitive interactions is key to understanding plant communities. Repeat censused large forest dynamics plots offer an ideal setting to measure these interactions by estimating the species-specific competitive effect on neighboring tree growth. Estimating these interaction values can be difficult, however, because the number of them grows with the square of the number of species. Furthermore, confidence in the estimates can be overestimated if any spatial structure of model errors is not considered. Here we measured these interactions in a forest dynamics plot in a transitional oak-hickory forest. We analytically fit Bayesian linear regression models of annual tree radial …


Evaluation Of Text Mining Techniques Using Twitter Data For Hurricane Disaster Resilience, Joshua Eason, Sathish Kumar Feb 2020

Evaluation Of Text Mining Techniques Using Twitter Data For Hurricane Disaster Resilience, Joshua Eason, Sathish Kumar

SDSU Data Science Symposium

Data obtained from social media microblogging websites such as Twitter provide the unique ability to collect and analyze conversations of the public in order to gain perspective on the thoughts and feelings of the general public. Sentiment and volume analysis techniques were applied to the dataset in order to gain an understanding of the amount and level of sentiment associated with certain disaster-related tweets, including a topical analysis of specific terms. This study showed that disaster-type events such as a hurricane can cause some strong negative sentiment in the period of time directly preceding the event, but ultimately returns quickly …


Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock Feb 2020

Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock

SMU Journal of Undergraduate Research

Professional development in its most traditional form is a classroom setting with a lecturer and an overwhelming amount of information. It is no surprise, then, that informal professional development away from institutions and on the teacher's own terms is a growing phenomenon due to an increased presence of educators on social media. These communities of educators use hashtags to broadcast to each other, with general hashtags such as #edchat having the broadest audience. However, many math educators usethe hashtags #ITeachMath and #MTBoS, communities I was interested in learning more about. I built a python script that used Tweepy to connect …


Mapping Relationships And Positions Of Objects In Images Using Mask And Bounding Box Data, Jaime M. Villanueva Jr, Anantharam Subramanian, Vishal Ahir, Andrew Pollock Jan 2020

Mapping Relationships And Positions Of Objects In Images Using Mask And Bounding Box Data, Jaime M. Villanueva Jr, Anantharam Subramanian, Vishal Ahir, Andrew Pollock

SMU Data Science Review

In this paper we present novel methods for automatically annotating images with relationship and position tags that are derived using mask and bounding box data. A Mask Region-based Convolutional Neural Network (Mask R-CNN) is used as the foundation for the ob- ject detection process. The relationships are found by manipulating the bounding box and mask segmentation outputs of a Mask R-CNN. The absolute positions, the positions of the objects relative to the image, and the relative positions, the positions of objects relative to the other objects, are then associated with the images as annotations that are out- put in order …


Wage-Productivity Analysis Of U.S. Domestic Airlines, Jesse Lucas, Khairul Azuar, Justin Tan, Syed Ilyas Jan 2020

Wage-Productivity Analysis Of U.S. Domestic Airlines, Jesse Lucas, Khairul Azuar, Justin Tan, Syed Ilyas

Introduction to Research Methods RSCH 202

This study examines the impact of wages on productivity by examining US domestic airlines.

Current literature places emphasis on jobs conducted in-flight, specifically pilots and cabin crew. This paper considers all job titles involved in the operations of the airline, including executives and management. Existing research focuses on factors such as governance, domestic economic level, and personal attributes such as intrinsic motivation, gender, and age. There is insufficient research regarding the relationship between wage and productivity. Thus, it is uncertain if high wage leads to high productivity. Preliminary findings suggest higher wage equates to higher productivity.


Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui Jan 2020

Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui

Theses and Dissertations--Statistics

In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.

In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …


Teaching Introductory Statistics With Datacamp, Benjamin Baumer, Andrew P. Bray, Mine Çetinkaya-Rundel, Johanna S. Hardin Jan 2020

Teaching Introductory Statistics With Datacamp, Benjamin Baumer, Andrew P. Bray, Mine Çetinkaya-Rundel, Johanna S. Hardin

Statistical and Data Sciences: Faculty Publications

We designed a sequence of courses for the DataCamp online learning platform that approximates the content of a typical introductory statistics course. We discuss the design and implementation of these courses and illustrate how they can be successfully integrated into a brick-and-mortar class. We reflect on the process of creating content for online consumers, ruminate on the pedagogical considerations we faced, and describe an R package for statistical inference that became a by-product of this development process. We discuss the pros and cons of creating the course sequence and express our view that some aspects were particularly problematic. The issues …


Inventory Models For Perishable Items Under Markdown Policy, Nurzahara Atika Kamaruzaman Jan 2020

Inventory Models For Perishable Items Under Markdown Policy, Nurzahara Atika Kamaruzaman

Student Works (2020-2029)

As expected, the demand for a fresh product depends on how fresh it is, therefore, it is important to take expiration date into consideration. Based on marketing and economic theory, several factors such as price, inventory level and advertisement play a crucial role in influencing the demand. Hence, we study the effect of these factors in influencing the demand in the inventory model. Since the demand for perishable product declines over time, markdown policy is offered to increase the demand and profit while reducing the inventory. Salvage value is incorporated to the deteriorating units. In this research, we extend previous …