Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Statistical Methodology

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 841 - 870 of 1562

Full-Text Articles in Statistics and Probability

Are There Predictors Of A Running Back’S Success?, Joshua Price Aug 2021

Are There Predictors Of A Running Back’S Success?, Joshua Price

Symposium of Student Scholars

People who analyze football have concentrated in the past on a running back’s 40-yard dash, shuffle, broad jump, vertical jump, and bench press measures. My research will test if the following variables can predict a running back’s success in the NFL: height, weight, conference, offensive line ranking for their team, the running back’s total yards for the season, their average yards for each attempt, the number of times the running back has entered the end zone for a touchdown that season, the running back’s time average time behind the line of scrimmage (TLOS), the percentage of times the running back …


Sources And Aftermaths Of Pipeline Related Leaks And Spills, Justin Smith Aug 2021

Sources And Aftermaths Of Pipeline Related Leaks And Spills, Justin Smith

Symposium of Student Scholars

The escape of oil and other hazardous materials have been shown to pollute and destroy ecosystems. As an aspiring chemist, I am adamant about the secure handling and transportation of oil and other hazardous materials. In the past, researchers have concentrated on oil’s high viscosity. Oil’s high viscosity physically smothers wildlife, affecting their ability to continue critical functions such as respiration, feeding, and thermoregulation. My research focuses on the source of these oil spills, as well as natural gas leaks, for the purpose of risk assessment. In addition, I compare recovery efforts based on the cause of the leak/spill, the …


On The Front Lines Of Fire: How Do We Save Their Lives?, Cathrine Jatta Aug 2021

On The Front Lines Of Fire: How Do We Save Their Lives?, Cathrine Jatta

Symposium of Student Scholars

The National Institute for Occupational Safety and Health (NIOSH) reports that the United States depends on about 1.1 million firefighters to protect its citizens and property from fire. NIOSH adds that approximately 336,000 are career firefighters; 812,000 are volunteers; and 80 to 100 die in the line of duty each year. NIOSH investigates each fatality individually for the cause and prevention. In contrast, my research will look at a complete dataset of 2005 firefighter fatalities and see if any of the following variables may predict firefighter death: age, cause of death, property type, type of duty (e.g. on-duty, training), and …


Cervical Cancer: Are There Ways To Reduce The Risks?, Madelyn Dorn Aug 2021

Cervical Cancer: Are There Ways To Reduce The Risks?, Madelyn Dorn

Symposium of Student Scholars

History has shown us that when caught early, cervical cancer is curable. Past research has found that the sexually transmitted diseases (STDs), herpes and human papillomavirus (HPV), have been associated with cervical cancer. In contrast, my dataset on 859 women has many more STDs and lifestyle choices compiled on 36 variables. The diagnoses in the dataset are many: cervical condylomatosis, vaginal condylomatosis, vulvo-perineral condylomatosis, syphilis, pelvic inflammatory disease, genital herpes, molluscum contagiosum, acquired immune deficiency syndrome (AIDS), human immunodeficiency virus (HIV), hepatitis B, HPV, and cervical cancer. In addition to the demographic variable on age, there are many lifestyle choice …


Marijuana Arrests In Toronto Canada: A Look Into The Canadian Criminal Justice System, Steven Tully Aug 2021

Marijuana Arrests In Toronto Canada: A Look Into The Canadian Criminal Justice System, Steven Tully

Symposium of Student Scholars

Marijuana related drug offenses made up fifty-eight percent of all Controlled Drugs and Substances Act offenses in Canada in 2016. On October 17, 2018, Canada legalized marijuana. As part of the efforts to legalize marijuana, descriptive statistics of single variables, like the age of the arrestees and the number of people arrested per year, were reported by the Toronto Star newspaper. The dataset analyzed in this research predates the legalization of marijuana and was collected from 1997 to 2002 on 5,226 individuals arrested in Toronto, Canada for simple possession of small quantities of marijuana. When an offender was arrested for …


Who Is Next? Evaluating Factors That May Contribute To Heart Failure, Davon Broadwater Aug 2021

Who Is Next? Evaluating Factors That May Contribute To Heart Failure, Davon Broadwater

Symposium of Student Scholars

Cardiovascular diseases are the number one causes of death globally, and for African Americans those risks are even higher. As an African American university student studying Biology, I am passionate about researching the diseases that affect my race. Current research states that behavioral factors such as obesity, tobacco use, unhealthy diet, and harmful use of alcohol should be avoided. I have chosen to research predictors of what helps patients survive if they already have heart failure. Heart failure develops gradually, where the heart becomes weaker over time and has trouble pumping blood to nourish the cells in the body. Data …


Eradicating Zebra Mussels: What Works?, Elijah Davies Aug 2021

Eradicating Zebra Mussels: What Works?, Elijah Davies

Symposium of Student Scholars

The invasion of U.S lakes and rivers by the invasive species of zebra mussels called Dreissena polymorpha has caused catastrophic harm to the local ecosystem by reproducing and outcompeting native mussel species as well as harm to pipes leading into water sources by binding to surfaces and reproducing to the point that the mussels clog pipes. In addition, recreation areas must be closed due to the sharp shells making areas unusable. In the past, research has focused on individual molluscicides and their eradication of zebra mussels, as well as their effect on native flora and fauna. My research will contrast …


Bias In Police Shootings: Is It Just An Opinion?, Phuong Ho Aug 2021

Bias In Police Shootings: Is It Just An Opinion?, Phuong Ho

Symposium of Student Scholars

The claims of racism have drawn public attention toward police brutality and its impact on minorities. Is this just an opinion or is there any statistical evidence? Recent studies from The Atlantic have investigated the average age and ethnicity of victims from police killings in 2015-2016. As an Asian-American, I am motivated to examine the issue of police killings among races and other demographics to find any bias that is present. Using the dataset of 2,204 victims of police killings (2015-2016) collected by The Guardian, I will examine the following variables for bias: age, cause of death, armed/unarmed, race/ethnicity, and …


Do Environmental Toxins Predict Violent Crimes?, Tyler Stahl Aug 2021

Do Environmental Toxins Predict Violent Crimes?, Tyler Stahl

Symposium of Student Scholars

Do chemical pollutants that persistent in the environment and bioaccumulate in the body affect human health and behavior? Could these Persistent, Bioaccumulative, and Toxic (PBT) chemicals play a role in the cause of violent crimes due to deterioration of mental and cognitive functions? In the past, Mercury, a PBT chemical, has been shown in salmon to be associated with aggression. Could similar aggression occur in humans exposed to mercury through a toxic spill? Two sources of data are utilized in this analysis. The Environmental Protection Agency’s (EPA) Annual Toxic Release Inventory publishes data on toxic releases into the environment and …


From Mathematics To Medicine: A Practical Primer On Topological Data Analysis (Tda) And The Development Of Related Analytic Tools For The Functional Discovery Of Latent Structure In Fmri Data, Andrew Salch, Adam Regalski, Hassan Abdallah, Raviteja Suryadevara, Michael J. Catanzaro, Vaibhav A. Diwadkar Aug 2021

From Mathematics To Medicine: A Practical Primer On Topological Data Analysis (Tda) And The Development Of Related Analytic Tools For The Functional Discovery Of Latent Structure In Fmri Data, Andrew Salch, Adam Regalski, Hassan Abdallah, Raviteja Suryadevara, Michael J. Catanzaro, Vaibhav A. Diwadkar

Mathematics Faculty Research Publications

fMRI is the preeminent method for collecting signals from the human brain in vivo, for using these signals in the service of functional discovery, and relating these discoveries to anatomical structure. Numerous computational and mathematical techniques have been deployed to extract information from the fMRI signal. Yet, the application of Topological Data Analyses (TDA) remain limited to certain sub-areas such as connectomics (that is, with summarized versions of fMRI data). While connectomics is a natural and important area of application of TDA, applications of TDA in the service of extracting structure from the (non-summarized) fMRI data itself are heretofore nonexistent. …


Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin Aug 2021

Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin

Electronic Theses and Dissertations

In this work, we seek to develop a variable screening and selection method for Bayesian mixture models with longitudinal data. To develop this method, we consider data from the Health and Retirement Survey (HRS) conducted by University of Michigan. Considering yearly out-of-pocket expenditures as the longitudinal response variable, we consider a Bayesian mixture model with $K$ components. The data consist of a large collection of demographic, financial, and health-related baseline characteristics, and we wish to find a subset of these that impact cluster membership. An initial mixture model without any cluster-level predictors is fit to the data through an MCMC …


On The Use Of Minimum Penalties In Statistical Learning, Ben Sherwood, Bradley S. Price Jul 2021

On The Use Of Minimum Penalties In Statistical Learning, Ben Sherwood, Bradley S. Price

Faculty & Staff Scholarship

Modern multivariate machine learning and statistical methodologies estimate parameters of interest while leveraging prior knowledge of the association between outcome variables. The methods that do allow for estimation of relationships do so typically through an error covariance matrix in multivariate regression which does not scale to other types of models. In this article we proposed the MinPEN framework to simultaneously estimate regression coefficients associated with the multivariate regression model and the relationships between outcome variables using mild assumptions. The MinPen framework utilizes a novel penalty based on the minimum function to exploit detected relationships between responses. An iterative algorithm that …


Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon Jul 2021

Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon

Graduate Theses and Dissertations

Markov chain Monte Carlo (MCMC) is a simulation technique that produces a Markov chain designed to converge to a stationary distribution. In Bayesian statistics, MCMC is used to obtain samples from a posterior distribution for inference. To ensure the accuracy of estimates using MCMC samples, the convergence to the stationary distribution of an MCMC algorithm has to be checked. As computation time is a resource, optimizing the efficiency of an MCMC algorithm in terms of effective sample size (ESS) per time unit is an important goal for statisticians. In this paper, we use simulation studies to demonstrate how the Gibbs …


Statistical Modeling For High-Dimensional Compositional Data With Applications To The Human Microbiome, Thy Dao Jul 2021

Statistical Modeling For High-Dimensional Compositional Data With Applications To The Human Microbiome, Thy Dao

Graduate Theses and Dissertations

Compositional data refer to the data that lie on a simplex, which are common in many scientific domains such as genomics, geology, and economics. As the components in a composition must sum to one, traditional tests based on unconstrained data become inappropriate, and new statistical methods are needed to analyze this special type of data. This dissertation is motivated by some statistical problems arising in the analysis of compositional data. In particular, we focus on the high-dimensional and over-dispersed setting, where the dimensionality of compositions is greater than the sample size and the dispersion parameter is moderate or large. In …


Pivot Points In Bivariate Linear Regression, David L. Farnsworth, Carl V. Lutzer Jun 2021

Pivot Points In Bivariate Linear Regression, David L. Farnsworth, Carl V. Lutzer

Articles

There are little-noticed points in the plane, which are artifacts of linear regression. The points, which are called pivot points, are the intersections of sets of regression lines. We derive the coordinates of the pivot point and explain its sources. We show how a pivot point arises in a certain notable data set, which has been analyzed often for points of high leverage. We obtain the application of pivot points that shortens calculations when updating a set of bivariate observations by adding a new point.


A Geometric Approach To Conditioning And The Search For Minimum Variance Unbiased Estimators, David L. Farnsworth, James E. Marengo Jun 2021

A Geometric Approach To Conditioning And The Search For Minimum Variance Unbiased Estimators, David L. Farnsworth, James E. Marengo

Articles

Our purpose is twofold: to present a prototypical example of the conditioning technique to obtain the best estimator of a parameter and to show that this technique resides in the structure of an inner product space. The technique uses conditioning of an unbiased estimator on a sufficient statistic. This procedure is founded upon the conditional variance formula, which leads to an inner product space and a geometric interpretation. The example clearly illustrates the dependence on the sampling methodology. These advantages show the power and centrality of this process.


Modeling And Solving The Outsourcing Risk Management Problem In Multi-Echelon Supply Chains, Arian A. Nahangi Jun 2021

Modeling And Solving The Outsourcing Risk Management Problem In Multi-Echelon Supply Chains, Arian A. Nahangi

Master's Theses

Worldwide globalization has made supply chains more vulnerable to risk factors, increasing the associated costs of outsourcing goods. Outsourcing is highly beneficial for any company that values building upon its core competencies, but the emergence of the COVID-19 pandemic and other crises have exposed significant vulnerabilities within supply chains. These disruptions forced a shift in the production of goods from outsourcing to domestic methods.

This paper considers a multi-echelon supply chain model with global and domestic raw material suppliers, manufacturing plants, warehouses, and markets. All levels within the supply chain network are evaluated from a holistic perspective, calculating a total …


Compare And Contrast Maximum Likelihood Method And Inverse Probability Weighting Method In Missing Data Analysis, Scott Sun May 2021

Compare And Contrast Maximum Likelihood Method And Inverse Probability Weighting Method In Missing Data Analysis, Scott Sun

Mathematical Sciences Technical Reports (MSTR)

Data can be lost for different reasons, but sometimes the missingness is a part of the data collection process. Unbiased and efficient estimation of the parameters governing the response mean model requires the missing data to be appropriately addressed. This paper compares and contrasts the Maximum Likelihood and Inverse Probability Weighting estimators in an Outcome-Dependendent Sampling design that deliberately generates incomplete observations. WE demonstrate the comparison through numerical simulations under varied conditions: different coefficient of determination, and whether or not the mean model is misspecified.


Characterizing The Northern Hemisphere Circumpolar Vortex Through Space And Time, Nazla Bushra May 2021

Characterizing The Northern Hemisphere Circumpolar Vortex Through Space And Time, Nazla Bushra

LSU Doctoral Dissertations

This hemispheric-scale, steering atmospheric circulation represented by the circumpolar vortices (CPVs) are the middle- and upper-tropospheric wind belts circumnavigating the poles. Variability in the CPV area, shape, and position are important topics in geoenvironmental sciences because of the many links to environmental features. However, a means of characterizing the CPV has remained elusive. The goal of this research is to (i) identify the Northern Hemisphere CPV (NHCPV) and its morphometric characteristics, (ii) understand the daily characteristics of NHCPV area and circularity over time, (iii) identify and analyze spatiotemporal variability in the NHCPV’s centroid, and (iv) analyze how CPV features relate …


Guidelines For Regression Analysis In Sas And R: A Case Study, Sarah Milligan May 2021

Guidelines For Regression Analysis In Sas And R: A Case Study, Sarah Milligan

Honors Program Theses and Projects

When a player is a free agent, an individual who is able to sign to any team, one wonders what their best option is. Will signing with Team A or Team B provide them with the largest salary? What factors will affect their salary the most? Does last year’s statistics have a strong impact on next year’s salary? These questions can be answered by performing a regression analysis on previous years data. The primary focus of this project is to determine the most important variables related to an NBA salary. Likewise, the statistical programs SAS and R will be compared …


Cointegration And Statistical Arbitrage Of Precious Metals, Judge Van Horn May 2021

Cointegration And Statistical Arbitrage Of Precious Metals, Judge Van Horn

Finance Undergraduate Honors Theses

When talking about financial instruments correlation is often thrown around as a measure of the relation between two securities. An often more useful or tradeable measure is cointegration. Cointegration is the measure of two securities tendency to revert to an average price over time. In other words, cointegration ignores directionality and only cares about the distance between two securities. For a mean reversion strategy such as statistical arbitrage cointegration proves to be a far more reliable statistical measure of mean reversion, and while it is more reliable than correlation it still has its own problems. One thing to consider is …


Markov Chains And Their Applications, Fariha Mahfuz Apr 2021

Markov Chains And Their Applications, Fariha Mahfuz

Math Theses

Markov chain is a stochastic model that is used to predict future events. Markov chain is relatively simple since it only requires the information of the present state to predict the future states. In this paper we will go over the basic concepts of Markov Chain and several of its applications including Google PageRank algorithm, weather prediction and gamblers ruin.

We examine on how the Google PageRank algorithm works efficiently to provide PageRank for a Google search result. We also show how can we use Markov chain to predict weather by creating a model from real life data.


Comparing Radiation Shielding Potential Of Liquid Propellants To Water For Application In Space, John Czaplewski Mar 2021

Comparing Radiation Shielding Potential Of Liquid Propellants To Water For Application In Space, John Czaplewski

Master's Theses

The radiation environment in space is a threat that engineers and astronauts need to mitigate as exploration into the solar system expands. Passive shielding involves placing as much material between critical components and the radiation environment as possible. However, with mass and size budgets, it is important to select efficient materials to provide shielding. Currently, NASA and other space agencies plan on using water as a shield against radiation since it is already necessary for human missions. Water has been tested thoroughly and has been proven to be effective. Liquid propellants are needed for every mission and also share similar …


Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels Jan 2021

Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels

SMU Data Science Review

Understanding diagnostic tests and examining important features of novel coronavirus (COVID-19) infection are essential steps for controlling the current pandemic of 2020. In this paper, we study the relationship between clinical diagnosis and analytical features of patient blood panels from the US, Mexico, and Brazil. Our analysis confirms that among adults, the risk of severe illness from COVID-19 increases with pre-existing conditions such as diabetes and immunosuppression. Although more than eight months into pandemic, more data have become available to indicate that more young adults were getting infected. In addition, we expand on the definition of COVID-19 test and discuss …


Evaluation Of The Effect Of The Clinical-Decision-Support Systems On Diabetes Management: A Multivariate Meta-Analysis Comparison With Univariate Meta-Analysis, Abdelfattah Elbarsha Jan 2021

Evaluation Of The Effect Of The Clinical-Decision-Support Systems On Diabetes Management: A Multivariate Meta-Analysis Comparison With Univariate Meta-Analysis, Abdelfattah Elbarsha

Electronic Theses and Dissertations

The advantage of using meta-analysis lies in its ability in providing a quantitative summary of the findings from multiple studies. The aim of this dissertation was first to conduct a simulation study in order to understand what factors (sample size, between-study correlation, and percent of missing data) have a significant effect on meta-analysis estimates and whether using univariate or multivariate meta-analysis would produce different estimates.

The second goal of this study was to evaluate the effect of clinical decision support systems CDSS on diabetes care management by conducting three separate univariate meta-analyses and one multivariate meta-analysis. CDSS are health information …


The Combined Impact Of Continuous And Ordinal Auxiliary Variables On Missing Data Imputation In Sem, Salina Wu Whitaker Jan 2021

The Combined Impact Of Continuous And Ordinal Auxiliary Variables On Missing Data Imputation In Sem, Salina Wu Whitaker

Electronic Theses and Dissertations

“Modern” methods of addressing missing data using full-information maximum-likelihood (FIML) have become mainstays in SEM analyses. FIML allows the inclusion of auxiliary variables which carry information that is related to missing values and can reduce bias in parameter estimates. Past research has illustrated the benefits of auxiliary variable inclusion under different missingness conditions (MCAR and MNAR; e.g., Enders, 2008), missingness proportions (e.g., Collins et al., 2001), and although limited, missingness patterns (e.g., Yoo, 2009) in FIML analyses. While past studies have focused on the effects of either continuous or ordinal auxiliary variables, no study has included both types in their …


Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao Jan 2021

Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao

Theses and Dissertations

Drug addiction can lead to many health-related problems and social concerns. Functional connectivity obtained from functional magnetic resonance imaging (fMRI) data promotes a variety of fundamental understandings in such association. Due to its complex correlation structure and large dimensionality, the modeling and analysis of the functional connectivity from neuroimage are challenging. By proposing a spatio-temporal model for multi-subject neuroimage data, we incorporate voxel-level spatio-temporal dependencies of whole-brain measurements to improve the accuracy of statistical inference. To tackle large-scale spatio-temporal neuroimage data, we develop a computationally efficient algorithm to estimate the parameters. Our method is used to identify functional connectivity and …


Investigations Into The Genetics Of Mixed Pathologies In Dementia, Adam Dugan Jan 2021

Investigations Into The Genetics Of Mixed Pathologies In Dementia, Adam Dugan

Theses and Dissertations--Epidemiology and Biostatistics

Alzheimer’s disease (AD) is an irreversible, progressive brain disorder that leads to a loss of memory and thinking skills. While tremendous progress has been made in our understanding of the genetics underlying AD, currently known genetic variants explain only approximately 30% of the heritable risk of developing AD. One hurdle to AD research is that it can only be definitively diagnosed at autopsy, making cruder, clinic-based diagnoses more common. In recent years, several brain pathologies that mimic AD’s clinical presentation have been identified including brain arteriolosclerosis, hippocampal sclerosis (HS), and, most recently, limbic-predominant age-related TDP-43 encephalopathy (LATE). It has become …


Dimension Reduction Techniques In Regression, Pei Wang Jan 2021

Dimension Reduction Techniques In Regression, Pei Wang

Theses and Dissertations--Statistics

Because of the advances of modern technology, the size of the collected data nowadays is larger and the structure is more complex. To deal with such kinds of data, sufficient dimension reduction (SDR) and reduced rank (RR) regression are two powerful tools. This dissertation focuses on these two tools and it is composed of three projects. In the first project, we introduce a new SDR method through a novel approach of feature filter to recover the central mean subspace exhaustively along with a method to determine the dimension, two variable selection methods, and extensions to multivariate response and large p …


Novel Nonparametric Testing Approaches For Multivariate Growth Curve Data: Finite-Sample, Resampling And Rank-Based Methods, Ting Zeng Jan 2021

Novel Nonparametric Testing Approaches For Multivariate Growth Curve Data: Finite-Sample, Resampling And Rank-Based Methods, Ting Zeng

Theses and Dissertations--Statistics

Multivariate growth curve data naturally arise in various fields, for example, biomedical science, public health, agriculture, social science and so on. For data of this type, the classical approach is to conduct multivariate analysis of variance (MANOVA) based on Wilks' Lambda and other multivariate statistics, which require the assumptions of multivariate normality and homogeneity of within-cell covariance matrices. However, data being analyzed nowadays show marked departure from multivariate normal distribution and homoscedasticity. In this dissertation, we investigate nonparametric testing approaches for multivariate growth curve data from three aspects, i.e., finite-sample, resampling and rank-based methods.

The first project proposes an approximate …