Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2018

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 541 - 570 of 596

Full-Text Articles in Statistics and Probability

A Comparison Of Discriminant Function Analysis And Logistic Regression By Categorizing The Incarcerated Mentally Ill, Mona King Jan 2018

A Comparison Of Discriminant Function Analysis And Logistic Regression By Categorizing The Incarcerated Mentally Ill, Mona King

Wayne State University Dissertations

Both discriminant function analysis (DFA) and logistic regression (LR) are used to classify subjects into a category/group based upon several explanatory variables (Liong & Foo, 2013). Although the two procedures are generally related, there is no clear advice in the statistical literature on when to use DFA vs. LR, although LR appears to be preferred due to the claim that its underlying assumptions are more easily met (Liong & Foo, 2013). Although DFA and LR use different methods to accomplish their objectives, they can answer the same research questions (Antonogeorgos et al., 2009). This facilitates a practical comparison of their …


Facilitators And Outcomes Of Stem-Education Groups Working Toward Disciplinary Integration, Anna E. Bargagliotti Jan 2018

Facilitators And Outcomes Of Stem-Education Groups Working Toward Disciplinary Integration, Anna E. Bargagliotti

Mathematics, Statistics and Data Science Faculty Works

There is a growing societal recognition of the need for transdisciplinary scholarly collaboration which can enhance undergraduate physics, science, and engineering education. A regional conference/network with 100 university education researchers in physics and other STEM fields was formed to address three themes (problemsolving, computational thinking, and equity) with multiple goals including to strive for transdisciplinary publications. As part of an ongoing participant observation study, phone interviews were conducted 3-4 months later. One year later, publications that were completed as a result of the conference were analyzed for their disciplinary integration. The papers showed evidence of interdispliciplanry collaboration but transdiciplinary collaboration …


Decision Trees: Predicting Future Losses For Insurance Data, Amanda Lahrmann Jan 2018

Decision Trees: Predicting Future Losses For Insurance Data, Amanda Lahrmann

Williams Honors College, Honors Research Projects

Big data is a term that has come to the spotlight for companies within recent years. Data analysis and business intelligence have become prominent sectors of companies and agencies. But what is big data? How has it impacted large companies and agencies? Why must it be embraced?

The best way to approach utilizing a big data set is to establish a question to answer. For this data set, the question that must be answered is “What variables cause a loss to occur?” To answer this question, first, we must understand what is meant by a “loss”, and take a look …


Catastrophe Modeling With Financial Applications, Jeremy Gensel Jan 2018

Catastrophe Modeling With Financial Applications, Jeremy Gensel

Williams Honors College, Honors Research Projects

Catastrophe modeling is used to prepare for losses caused by natural catastrophes such as earthquakes, hurricanes, or tornadoes and man-made catastrophes such as terrorism. Modeled data can be used to create a comprehensive distribution of possible disasters. The distribution gives probabilities of potential catastrophes of different severities occurring over a certain time frame. Calculating potential losses and probability of those losses occurring allows insurance companies to plan and reserve enough money to protect themselves from catastrophic events. Using a catastrophe case study posted online from the Casualty Actuarial Society and R software, this paper shows the use of statistical techniques …


Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry Jan 2018

Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry

Theses and Dissertations

The Brisbane Longitudinal Twin Study (BLTS) was being conducted in Australia and was funded by the US National Institute on Drug Abuse (NIDA). Adolescent twins were sampled as a part of this study and surveyed about their substance use as part of the Pathways to Cannabis Use, Abuse and Dependence project. The methods developed in this dissertation were designed for the purpose of analyzing a subset of the Pathways data that includes demographics, cannabis use metrics, personality measures, and imputed genotypes (SNPs) for 493 complete twin pairs (986 subjects.) The primary goal was to determine what combination of SNPs and …


Joint Analysis Of Multiple Phenotypes In Association Studies, Xiaoyu Liang Jan 2018

Joint Analysis Of Multiple Phenotypes In Association Studies, Xiaoyu Liang

Dissertations, Master's Theses and Master's Reports

Genome-wide association studies (GWAS) have become a very effective research tool to identify genetic variants of underlying various complex diseases. In spite of the success of GWAS in identifying thousands of reproducible associations between genetic variants and complex disease, in general, the association between genetic variants and a single phenotype is usually weak. It is increasingly recognized that joint analysis of multiple phenotypes can be potentially more powerful than the univariate analysis, and can shed new light on underlying biological mechanisms of complex diseases. Therefore, developing statistical methods to test for genetic association with multiple phenotypes has become increasingly important. …


The Kumaraswamy Weibull Geometric Distribution With Applications, Mahdi Rasekhi, Morad Alizadeh, Gholamhossein G. Hamedani Jan 2018

The Kumaraswamy Weibull Geometric Distribution With Applications, Mahdi Rasekhi, Morad Alizadeh, Gholamhossein G. Hamedani

Mathematical and Statistical Science Faculty Research and Publications

In this work, we study the kumaraswamy weibull geometric (Kw-WG) distribution which includes as special cases, several models such as the kumaraswamy weibull distribution, kumaraswamy exponential distribution, weibull geometric distribution, exponential geometric distribution, to name a few. This distribution was monotone and non-monotone hazard rate functions, which are useful in lifetime data analysis and reliability. We derive some basic properties of the Kw-WG distribution including non-central rth-moments, skewness, kurtosis, generating functions, mean deviations, mean residual life, entropy, order statistics and certain characterizations of our distribution. The method of maximum likelihood is used for estimating the model parameters and a simulation …


Goodness Of Fit Via Residual Plots In Item Response Theory, Bryonna Bowen Jan 2018

Goodness Of Fit Via Residual Plots In Item Response Theory, Bryonna Bowen

Theses and Dissertations

Goodness-of-fit criteria developed for the evaluation of item response functions have been examined by many scholars using different theories and criteria. A number of potential graphical analysis approaches, such as residual plots, have been described in literature, but have received little attention from researchers. While many tests of goodness-of-fit are available, those that incorporate the analysis of residuals may be most useful. The unmistakable presence of a pattern in the residual plot for the logistic model item response functions even when we know the model fits raises a red flag up and calls for greater analysis. This study explores different …


Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh Jan 2018

Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh

Theses and Dissertations

Any counting system is prone to recording errors including underreporting and overreporting. Ignoring the misreporting pattern in count data can give rise to bias in the estimation of model parameters. Accordingly, Poisson, negative binomial and generalized Poisson regression have been expanded in some instances to capture reporting biases. However, to our knowledge, no program has been developed to allow users to apply all of these models when needed. In the first part of the dissertation, we review the available models for underreported counts and develop a Stata command to estimate Poisson, negative binomial and generalized Poisson regression models for underreported …


Mlb Rule Iv Draft: Valuing Draft Pick Slots, Anthony Cacchione Jan 2018

Mlb Rule Iv Draft: Valuing Draft Pick Slots, Anthony Cacchione

Dissertations and Theses

This study explored the Net Present Value (NPV) in dollar terms of draft pick slots in the Major League Rule IV Draft. In order to accomplish this, the cumulative performance of players selected in each slot within the draft was evaluated and brought to the Present Value of the time they were selected using a discount rate. The performance of the players was determined using the baseball-reference Wins Above Replacement (WAR) metric. It is intuitive that earlier draft picks are the most valuable; however, it is unclear how quickly the value of draft picks decline. This research demonstrates that the …


Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar Jan 2018

Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar

CMC Senior Theses

Understanding what factors influence wildlife movement allows landscape planners to make informed decisions that benefit both animals and humans. New quantitative methods, such as step-selection functions, provide valuable objective analyses of wildlife connectivity. This paper provides a framework for creating a step-selection function and demonstrates its use in a case study. The first section provides a general introduction about wildlife connectivity research. The second section explains the math behind the step-selection function using a simple example. The last section gives the results of a step-selection model for African buffalo in the Kavango Zambezi Transfrontier Conservation Area. Buffalo were found to …


Predictive Golf Analytics Versus The Daily Fantasy Sports Market, John O'Malley Jan 2018

Predictive Golf Analytics Versus The Daily Fantasy Sports Market, John O'Malley

CMC Senior Theses

This study examines the different skills necessary for PGA tour players to succeed at specific annual tournaments, in order to create a predictive model for DraftKings PGA contests. The model takes into account data from the PGA Tour ShotLink Intelligence Program. The predictive model is created each week based on past results from the specific tournament in question, with the hope of predicting a group of twenty-five players who should be successful based on their statistical profile. The results of the model are detailed in this paper, which covers the first nine weeks of the 2017 PGA Tour season, with …


Spatiotemporal Variations In Coexisting Multiple Causes Of Death And The Associated Factors, Emmanuel Oluwatobi Salawu Jan 2018

Spatiotemporal Variations In Coexisting Multiple Causes Of Death And The Associated Factors, Emmanuel Oluwatobi Salawu

Walden Dissertations and Doctoral Studies

The study and practice of epidemiology and public health benefit from the use of mortality statistics, such as mortality rates, which are frequently used as key health indicators. Furthermore, multiple causes of death (MCOD) data offer important information that could not possibly be gathered from other mortality data. This study aimed to describe the interrelationships between various causes of death in the United States in order to improve the understanding of the coexistence of MCOD and thereby improve public health and enhance longevity. The social support theory was used as a framework, and multivariate linear regression analyses were conducted to …


Development Of Biclustering Techniques For Gene Expression Data Modeling And Mining, Juan Xie Jan 2018

Development Of Biclustering Techniques For Gene Expression Data Modeling And Mining, Juan Xie

Electronic Theses and Dissertations

The next-generation sequencing technologies can generate large-scale biological data with higher resolution, better accuracy, and lower technical variation than the arraybased counterparts. RNA sequencing (RNA-Seq) can generate genome-scale gene expression data in biological samples at a given moment, facilitating a better understanding of cell functions at genetic and cellular levels. The abundance of gene expression datasets provides an opportunity to identify genes with similar expression patterns across multiple conditions, i.e., co-expression gene modules (CEMs). Genomescale identification of CEMs can be modeled and solved by biclustering, a twodimensional data mining technique that allows clustering of rows and columns in a gene …


Rock Paper Scissors And Evolutionary Game Theory, Christian Cordova, Rudolf Jovero, Evan Thomas Jan 2018

Rock Paper Scissors And Evolutionary Game Theory, Christian Cordova, Rudolf Jovero, Evan Thomas

Math 365 Class Projects

In Rock Paper Scissors (RPS), three different "species" compete, but no single species has a dominating strategy. In evolutionary game theory, replicator equations model population densities over time. When a mutation is introduced, they are called "replicator-mutator" equations. Using the replicator-mutator equation in [1] we have shown how population density of three species change.


Understanding The Novice Decision-Making Process In Forensic Footwear Examinations: Accuracy And Decision Rules, Madonna A. Nobel Jan 2018

Understanding The Novice Decision-Making Process In Forensic Footwear Examinations: Accuracy And Decision Rules, Madonna A. Nobel

Graduate Theses, Dissertations, and Problem Reports (ETD)

The reproducibility of experienced-based forensic pattern interpretation is founded on the notion that domain-specific knowledge can be successfully distributed and applied among experts within a group. This assumption persists, even when the examination is complicated by variations in case circumstances, such as impression clarity and totality, as well as media, substrate, collection mechanism and enhancement. While it is further theorized that many of these factors (as well as additional confounding factors) are at play during an examination, the manner and extent to which these sources of variability affect the examination of footwear evidence remain unclear. In order to explore this …


Developing, Piloting, And Factor Analysis Of A Brief Survey Tool For Evaluating Food And Composting Behaviors: The Short Composting Survey, Jennie Norton Jan 2018

Developing, Piloting, And Factor Analysis Of A Brief Survey Tool For Evaluating Food And Composting Behaviors: The Short Composting Survey, Jennie Norton

All Master's Theses

Composting on a university campus may take a variety of forms. Sustainable approaches to waste management can be taught and supported through educational programs, peer-to-peer behavior modeling, and composting program interventions. Although peer-reviewed research on composting interventions is somewhat lacking, student interest in the topic is demonstrated by a range of exploratory senior projects and pilot interventions conducted at colleges across the United States and abroad. The purpose of this study was twofold: conduct an educational compost intervention pilot study and develop a survey tool to measure participant attitudes surrounding food behaviors and composting. The Compost Project pilot study focused …


Stat 380: Statistics And Applications, Yumou Qiu Jan 2018

Stat 380: Statistics And Applications, Yumou Qiu

UNL Faculty Course Portfolios

This portfolio prepares for the introduction level undergraduate statistic course (Stat 380) for the students with calculus background, which includes the course goals, contents as well as teaching methods for large session classes. Analysis and evaluation of student learning are also included. In-class activities and discussion helping students’ engagement are introduced. The homework is assigned by Canvas, which summaries the students’ solution and provides analysis for each student.


Geographic Variations In Antenatal Care Services In Sierra Leone, Eunice Nyambura Chege Jan 2018

Geographic Variations In Antenatal Care Services In Sierra Leone, Eunice Nyambura Chege

Walden Dissertations and Doctoral Studies

Despite antenatal care presenting opportunities to identify and monitor women at risk, use of recommended antenatal care services remains. Barriers preventing use of antenatal services vary between countries, and limited knowledge exists about the link between geographical settings and antenatal service use. The objective of this cross-sectional quantitative study was to explore geographical variations and investigate how social demographic characteristics affect use of antenatal care for women in Sierra Leone using the Andersen behavioral model. The data used were from the 2016 maternal death surveillance report of the whole counrty (N =706). Logistic regression analysis was used to determine the …


Effect Of Neuromodulation Of Short-Term Plasticity On Information Processing In Hippocampal Interneuron Synapses, Elham Bayat Mokhtari Jan 2018

Effect Of Neuromodulation Of Short-Term Plasticity On Information Processing In Hippocampal Interneuron Synapses, Elham Bayat Mokhtari

Graduate Student Theses, Dissertations, & Professional Papers

Neurons convey information about the complex dynamic environment in the form of signals. Computational neuroscience provides a theoretical foundation toward enhancing our understanding of nervous system. The aim of this dissertation is to present techniques to study the brain and how it processes information in particular neurons in hippocampus.

We begin with a brief review of the history of neuroscience and biological background of basic neurons. To appreciate the importance of information theory, familiarity with the information theoretic basics is required, these basics are presented in Chapter 2. In Chapter 3, we use information theory to estimate the amount of …


Multiclass Classification Using Support Vector Machines, Duleep Prasanna W. Rathgamage Don Jan 2018

Multiclass Classification Using Support Vector Machines, Duleep Prasanna W. Rathgamage Don

College of Graduate Studies: Theses & Dissertations

In this thesis, we discuss different SVM methods for multiclass classification and introduce the Divide and Conquer Support Vector Machine (DCSVM) algorithm which relies on data sparsity in high dimensional space and performs a smart partitioning of the whole training data set into disjoint subsets that are easily separable. A single prediction performed between two partitions eliminates one or more classes in a single partition, leaving only a reduced number of candidate classes for subsequent steps. The algorithm continues recursively, reducing the number of classes at each step until a final binary decision is made between the last two classes …


On Modeling Quantities For Insurer Solvency Against Catastrophe Under Some Markovian Assumptions, Daniel Jefferson Geiger Jan 2018

On Modeling Quantities For Insurer Solvency Against Catastrophe Under Some Markovian Assumptions, Daniel Jefferson Geiger

Doctoral Dissertations

"Insurance companies sometimes face catastrophic losses, yet they must remain solvent enough to meet the legal obligation of covering all claims. Catastrophes can result in large damages to the policyholders, causing the arrival of numerous claims to insurance companies at once. Furthermore, the severity of an event could impact the time until the next occurrence. An insurer needs certain levels of startup capital to meet all claims, and then must have adequate reserves on a continual basis, even more so when catastrophes occur. This work examines two facets of these matters: for an infinite time horizon, we extend and develop …


Some New And Generalized Distributions Via Exponentiation, Gamma And Marshall-Olkin Generators With Applications, Hameed Abiodun Jimoh Jan 2018

Some New And Generalized Distributions Via Exponentiation, Gamma And Marshall-Olkin Generators With Applications, Hameed Abiodun Jimoh

College of Graduate Studies: Theses & Dissertations

Three new generalized distributions developed via completing risk, gamma generator, Marshall-Olkin generator and exponentiation techniques are proposed and studied. Structural properties including quantile functions, hazard rate functions, moment, conditional moments, mean deviations, R\'enyi entropy, distribution of order statistics and maximum likelihood estimates are presented. Monte Carlo simulation is employed to examine the performance of the proposed distributions. Applications of the generalized distributions to real lifetime data are presented to illustrate the usefulness of the models.


New Developments Of Dimension Reduction, Lei Huo Jan 2018

New Developments Of Dimension Reduction, Lei Huo

Doctoral Dissertations

"Variable selection becomes more crucial than before, since high dimensional data are frequently seen in many research areas. Many model-based variable selection methods have been developed. However, the performance might be poor when the model is mis-specified. Sufficient dimension reduction (SDR, Li 1991; Cook 1998) provides a general framework for model-free variable selection methods.

In this thesis, we first propose a novel model-free variable selection method to deal with multi-population data by incorporating the grouping information. Theoretical properties of our proposed method are also presented. Simulation studies show that our new method significantly improves the selection performance compared with those …


The South Carolina Safety Belt Study: Large-Scale Location Sampling, Stephanie Jones Jan 2018

The South Carolina Safety Belt Study: Large-Scale Location Sampling, Stephanie Jones

Theses and Dissertations

The South Carolina Safety Belt Study is a statewide survey completed yearly to assess the prevalence of safety belt usage on of South Carolina roads through observations from different locations across the state. Every five years the sites for observation are resampled. This thesis breaks down the most recent sampling done for the years of 2018 through 2022. Both the methodology of large scale location sampling and the mathematical idea behind the strategy employed are covered. Further, three different software packages were utilized: R, SAS, and ArcGIS. The steps that were taken and the written function code run for each …


Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard Jan 2018

Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard

Theses and Dissertations

Linear regression is a widely used method for analysis that is well understood across a wide variety of disciplines. In order to use linear regression, a number of assumptions must be met. These assumptions, specifically normality and homoscedasticity of the error distribution can at best be met only approximately with real data. Quantile regression requires fewer assumptions, which offers a potential advantage over linear regression. In this simulation study, we compare the performance of linear (least squares) regression to quantile regression when these assumptions are violated, in order to investigate under what conditions quantile regression becomes the more advantageous method …


Semiparametric Regression In The Presence Of Measurement Error, Xiang Li Jan 2018

Semiparametric Regression In The Presence Of Measurement Error, Xiang Li

Theses and Dissertations

The error-in-covariates problem has received great attention among researchers who study semiparametric and nonparametric inference for regression models over the past two decades. Without correcting for the measurement error in covariates, estimators for covariate effect usually contain bias. To account for measurement error, much research have been done in mean regression (Liang et al., 1999; Fuller, 2009; Carroll et al., 2006) and quantile regression (He and Liang, 2000; Hardle et al., 2000; Wei and Carroll, 2009). In contrast, there is little research in mode regression and this motivates us to propose semiparametric methods to address this error-incovariates problem in Chapters …


Bayesian Semiparametric Methods For Analyzing Panel Count Data, Jianhong Wang Jan 2018

Bayesian Semiparametric Methods For Analyzing Panel Count Data, Jianhong Wang

Theses and Dissertations

Panel count data commonly arise in epidemiological, social science, medical studies, in which subjects have repeated measurements on the recurrent events of interest at different observation times. Since the subjects are not under continuous monitoring, the exact times of those recurrent events are not observed but the counts of such events within the adjacent observation times are known. Panel count data can be considered as a special type of longitudinal data with a count response variable in the literature. Compared to the frequentist literature, very limited Bayesian approaches have been developed to analyze panel count data. In this dissertation, several …


How Often Does The Best Team Win? A Unified Approach To Understanding Randomness In North American Sport, Michael J. Lopez, Gregory J. Matthews, Benjamin S. Baumer Jan 2018

How Often Does The Best Team Win? A Unified Approach To Understanding Randomness In North American Sport, Michael J. Lopez, Gregory J. Matthews, Benjamin S. Baumer

Mathematics and Statistics: Faculty Publications and Other Works

Statistical applications in sports have long centered on how to best separate signal (e.g., team talent) from random noise. However, most of this work has concentrated on a single sport, and the development of meaningful cross-sport comparisons has been impeded by the difficulty of translating luck from one sport to another. In this manuscript we develop Bayesian state-space models using betting market data that can be uniformly applied across sporting organizations to better understand the role of randomness in game outcomes. These models can be used to extract estimates of team strength, the between-season, within-season and game-to-game variability of team …


How Often Does The Best Team Win? A Unified Approach To Understanding Randomness In North American Sport, Michael J. Lopez, Gregory J. Matthews, Benjamin S. Baumer Jan 2018

How Often Does The Best Team Win? A Unified Approach To Understanding Randomness In North American Sport, Michael J. Lopez, Gregory J. Matthews, Benjamin S. Baumer

Mathematics and Statistics: Faculty Publications and Other Works

Statistical applications in sports have long centered on how to best separate signal (e.g. team talent) from random noise. However, most of this work has concentrated on a single sport, and the development of meaningful cross-sport comparisons has been impeded by the difficulty of translating luck from one sport to another. In this manuscript, we develop Bayesian state-space models using betting market data that can be uniformly applied across sporting organizations to better understand the role of randomness in game outcomes. These models can be used to extract estimates of team strength, the between-season, within-season, and game-to-game variability of team …