Predicting Diabetes Diagnoses,
2020
Misericordia University
Predicting Diabetes Diagnoses, Sarah Netchert
Student Research Poster Presentations 2020
This study explored the traits and health state of African Americans in central Virginia in order to determine what traits put people at a higher probability of being diagnosed with diabetes. We also want to know which traits will generate the highest probability a person will be diagnosed with diabetes. Traits that were included and used in this study were cholesterol, stabilized glucose, high density lipoprotein levels, age(years), gender, height(inches), weight(pounds), systolic blood pressure, diastolic blood pressure, waist size(inches), and hip size(inches). There were 403 individuals included in study since they were only ones screened for diabetes out of 1,046 …
An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines,
2020
Western Kentucky University
An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines, Jeron Russell
Mahurin Honors College Capstone Experience/Thesis Projects
The purpose of this research is to look at the relationship that market-specific, economic, and demographic variables have with the success of farmers markets in Kentucky. It additionally seeks to build a tool for predicting farmers market success that could be used by policy makers to aid in decision-making processes concerning farmers markets. Logistic regression and Support Vector Machines (SVMs) are used on data acquired from the Kentucky Department of Agriculture and the American Community Survey in order to analyze the data in a traditional statistical approach as well as a machine learning approach. The results included an SVM model …
Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering,
2020
Minnesota State University, Mankato
Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering, Daniel Silva
All Graduate Theses, Dissertations, and Other Capstone Projects
The purpose of this paper is to develop the theory of principal components analysis succinctly from the fundamentals of matrix algebra and multivariate statistics. Principal components analysis is sometimes used as a descriptive technique to explain the variance-covariance or correlation structure of a dataset. However, most often, it is used as a dimensionality reduction technique to visualize a high dimensional dataset in a lower dimensional space. Principal components analysis accomplishes this by using the first few principal components, provided that they account for a substantial proportion of variation in the original dataset. In the same way, the first few principal …
Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis,
2020
Murray State University
Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown
Murray State Theses and Dissertations
Data and algorithmic modeling are two different approaches used in predictive analytics. The models discussed from these two approaches include the proportional odds logit model (POLR), the vector generalized linear model (VGLM), the classification and regression tree model (CART), and the random forests model (RF). Patterns in the data were analyzed using trigonometric polynomial approximations and Fast Fourier Transforms. Predictive modeling is used frequently in statistics and data science to find the relationship between the explanatory (input) variables and a response (output) variable. Both approaches prove advantageous in different cases depending on the data set. In our case, the data …
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer,
2020
Virginia Commonwealth University
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Theses and Dissertations
Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …
Unitary And Symmetric Structure In Deep Neural Networks,
2020
University of Kentucky
Unitary And Symmetric Structure In Deep Neural Networks, Kehelwala Dewage Gayan Maduranga
Theses and Dissertations--Mathematics
Recurrent neural networks (RNNs) have been successfully used on a wide range of sequential data problems. A well-known difficulty in using RNNs is the vanishing or exploding gradient problem. Recently, there have been several different RNN architectures that try to mitigate this issue by maintaining an orthogonal or unitary recurrent weight matrix. One such architecture is the scaled Cayley orthogonal recurrent neural network (scoRNN), which parameterizes the orthogonal recurrent weight matrix through a scaled Cayley transform. This parametrization contains a diagonal scaling matrix consisting of positive or negative one entries that can not be optimized by gradient descent. Thus the …
Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks,
2020
University of Kentucky
Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich
Theses and Dissertations--Mathematics
Despite the recent success of various machine learning techniques, there are still numerous obstacles that must be overcome. One obstacle is known as the vanishing/exploding gradient problem. This problem refers to gradients that either become zero or unbounded. This is a well known problem that commonly occurs in Recurrent Neural Networks (RNNs). In this work we describe how this problem can be mitigated, establish three different architectures that are designed to avoid this issue, and derive update schemes for each architecture. Another portion of this work focuses on the often used technique of batch normalization. Although found to be successful …
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups,
2020
University of Kentucky
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Theses and Dissertations--Statistics
Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …
Nonparametric Analysis Of Clustered And Multivariate Data,
2020
University of Kentucky
Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui
Theses and Dissertations--Statistics
In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.
In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …
Nonparametric Tests Of Lack Of Fit For Multivariate Data,
2020
University of Kentucky
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Theses and Dissertations--Statistics
A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data,
2020
University of Kentucky
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Theses and Dissertations--Statistics
Kinetic modeling of the time dependence of metabolite concentrations including the unstable isotope labeled species is an important approach to simulate metabolic pathway dynamics. It is also essential for quantitative metabolic flux analysis using tracer data. However, as the metabolic networks are complex including extensive compartmentation and interconnections, the parameter estimation for enzymes that catalyze individual reactions needed for kinetic modeling is challenging. As the pa- rameter space is large and multi-dimensional while kinetic data are comparatively sparse, the estimation procedure (especially the point estimation methods) often en- counters multiple local maximum such that standard maximum likelihood methods may yield …
Statistical Intervals For Various Distributions Based On Different Inference Methods,
2020
University of Kentucky
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Theses and Dissertations--Statistics
Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …
Measuring Change: Prediction Of Early Onset Sepsis,
2020
University of Kentucky
Measuring Change: Prediction Of Early Onset Sepsis, Aric Schadler
Theses and Dissertations--Statistics
Sepsis occurs in a patient when an infection enters into the blood stream and spreads throughout the body causing a cascading response from the immune system. Sepsis is one of the leading causes of morbidity and mortality in today’s hospitals. This is despite published and accepted guidelines for timely and appropriate interventions for septic patients. The largest barrier to applying these interventions is the early identification of septic patients. Early identification and treatment leads to better outcomes, shorter lengths of stay, and financial savings for healthcare institutions. In order to increase the lead time in recognizing patients trending towards septicemia …
How We Can Extend The Standard Deviation Notion With Neutrosophic Interval And Quadruple Neutrosophic Numbers,
2020
University of New Mexico
How We Can Extend The Standard Deviation Notion With Neutrosophic Interval And Quadruple Neutrosophic Numbers, Victor Christianto, Florentin Smarandache, Muhammad Aslam
Branch Mathematics and Statistics Faculty and Staff Publications
During scientific demonstrating of genuine specialized framework we can meet any sort and rate model vulnerability. Its reasons can be incognizance of modelers or information mistake. In this way, characterization of vulnerabilities, as for their sources, recognizes aleatory and epistemic ones. The aleatory vulnerability is an inalienable information variety related with the researched framework or its condition. Epistemic one is a vulnerability that is because of an absence of information on amounts or procedures of the framework or the earth [7]. Right now, we examine fourfold neutrosophic numbers and their potential application for practical displaying of physical frameworks, particularly in …
Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals,
2020
Old Dominion University
Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals, Sihai Dave Zhao, Yet Tian Nguyen
Mathematics & Statistics Faculty Publications
It is frequently of interest to identify simultaneous signals, defined as features that exhibit statistical significance across each of several independent experiments. For example, genes that are consistently differentially expressed across experiments in different animal species can reveal evolutionarily conserved biological mechanisms. However, in some problems the test statistics corresponding to these features can have complicated or unknown null distributions. This paper proposes a novel nonparametric false discovery rate control procedure that can identify simultaneous signals even without knowing these null distributions. The method is shown, theoretically and in simulations, to asymptotically control the false discovery rate. It was also …
Generalization Of Kullback-Leibler Divergence For Multi-Stage Diseases: Application To Diagnostic Test Accuracy And Optimal Cut-Points Selection Criterion,
2020
Georgia Southern University
Generalization Of Kullback-Leibler Divergence For Multi-Stage Diseases: Application To Diagnostic Test Accuracy And Optimal Cut-Points Selection Criterion, Chen Mo
College of Graduate Studies: Theses & Dissertations
The Kullback-Leibler divergence (KL), which captures the disparity between two distributions, has been considered as a measure for determining the diagnostic performance of an ordinal diagnostic test. This study applies KL and further generalizes it to comprehensively measure the diagnostic accuracy test for multi-stage (K > 2) diseases, named generalized total Kullback-Leibler divergence (GTKL). Also, GTKL is proposed as an optimal cut-points selection criterion for discriminating subjects among different disease stages. Moreover, the study investigates a variety of applications of GTKL on measuring the rule-in/out potentials in the single-stage and multi-stage levels. Intensive simulation studies are conducted to compare the performance …
The Role Of Topography, Soil, And Remotely Sensed Vegetation Condition Towards Predicting Crop Yield,
2020
University of Nebraska-Lincoln
The Role Of Topography, Soil, And Remotely Sensed Vegetation Condition Towards Predicting Crop Yield, Trenton E. Franz, Sayli Pokal, Justin P. Gibson, Yuzhen Zhou, Hamed Gholizadeh, Fatima Amor Tenorio, Daran Rudnick, Derek M. Heeren, Matthew F. Mccabe, Matteo Ziliani, Zhenong Jin, Kaiyu Guan, Ming Pan, John Gates, Brian Wardlow
School of Natural Resources: Faculty Publications
Foreknowledge of the spatiotemporal drivers of crop yield would provide a valuable source of information to optimize on-farm inputs and maximize profitability. In recent years, an abundance of spatial data providing information on soils, topography, and vegetation condition have become available from both proximal and remote sensing platforms. Given the wide range of data costs (between USD $0−50/ha), it is important to understand where often limited financial resources should be directed to optimize field production. Two key questions arise. First, will these data actually aid in better fine-resolution yield prediction to help optimize crop management and farm economics? Second, what …
Distribution Of Human Exposure To Ozone During Commuting Hours In Connecticut Using The Cellular Device Network,
2020
Bucknell University
Distribution Of Human Exposure To Ozone During Commuting Hours In Connecticut Using The Cellular Device Network, Owais Gilani, Simon Urbanek, Michael J. Kane
Faculty Journal Articles
Epidemiologic studies have established associations between various air pollutants and adverse health outcomes for adults and children. Due to high costs of monitoring air pollutant concentrations for subjects enrolled in a study, statisticians predict exposure concentrations from spatial models that are developed using concentrations monitored at a few sites. In the absence of detailed information on when and where subjects move during the study window, researchers typically assume that the subjects spend their entire day at home, school, or work. This assumption can potentially lead to large exposure assignment bias. In this study, we aim to determine the distribution of …
Tropical Cyclone Hazards In Relation To Propagation Speed,
2020
CUNY City College
Tropical Cyclone Hazards In Relation To Propagation Speed, Jiehao Huang
Dissertations and Theses
As the population and infrastructure along the US East Coast increase, it becomes increasingly important to study the characteristics of tropical cyclones that can impact the coast. A recent study shows that the propagation speed of tropical cyclones has slowed over the past 60 years, which can lead to greater accumulation of precipitation and greater storm surge impacts. The study presented herein is meant to examine and analyze the relationships that exist between the propagation speed of tropical cyclones, their surface wind strength, displacement angles, and cyclone averaged winds. This analysis is focused on tropical cyclones spanning from 1950-2015 in …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation,
2020
Claremont Colleges
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
