Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (23)
- Statistical Methodology (18)
- Biostatistics (16)
- Social and Behavioral Sciences (11)
- Statistical Theory (9)
-
- Applied Mathematics (8)
- Multivariate Analysis (8)
- Longitudinal Data Analysis and Time Series (7)
- Mathematics (7)
- Data Science (6)
- Life Sciences (6)
- Other Statistics and Probability (6)
- Probability (6)
- Sociology (6)
- Survival Analysis (6)
- Business (5)
- Computer Sciences (5)
- Education (5)
- Quantitative, Qualitative, Comparative, and Historical Methodologies (5)
- Genetics and Genomics (4)
- Medicine and Health Sciences (4)
- Other Applied Mathematics (4)
- Bioinformatics (3)
- Clinical Trials (3)
- Geography (3)
- Microarrays (3)
- Other Mathematics (3)
- Institution
- Keyword
-
- Morgridge College of Education (7)
- Research Methods and Information Science (6)
- Research Methods and Statistics (6)
- Regression (3)
- Statistics (3)
-
- Survival Analysis (3)
- Bootstrap (2)
- Breast cancer (2)
- Confidence Interval (2)
- Data mining (2)
- Machine learning (2)
- Meta-analysis (2)
- Mixed data (2)
- Nonparametric (2)
- Observational data (2)
- Propensity scores (2)
- Robust (2)
- Simulation (2)
- Variation (2)
- ANNs (1)
- AUC (1)
- Actuarial (1)
- Adaptive Design (1)
- Adaptive Designs (1)
- Adolescent (1)
- Alpha Spending Functions (1)
- Anomaly Detection (1)
- Applied statistics (1)
- Asset risk (1)
- Autoencoder (1)
Articles 31 - 58 of 58
Full-Text Articles in Applied Statistics
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Electronic Theses and Dissertations
Recently, gene set analysis has become the first choice for gaining insights into the underlying complex biology of diseases through high-throughput genomic studies, such as Microarrays, bulk RNA-Sequencing, single cell RNA-Sequencing, etc. It also reduces the complexity of statistical analysis and enhances the explanatory power of the obtained results. Further, the statistical structure and steps common to these approaches have not yet been comprehensively discussed, which limits their utility. Hence, a comprehensive overview of the available gene set analysis approaches used for different high-throughput genomic studies is provided. The analysis of gene sets is usually carried out based on …
Aspects Of Causal Inference., John A. Craycroft
Aspects Of Causal Inference., John A. Craycroft
Electronic Theses and Dissertations
Observational studies differ from experimental studies in that assignment of subjects to treatments is not randomized but rather occurs due to natural mechanisms, which are usually hidden from researchers. Yet objectives of the two studies are frequently the same: identify the causal – rather than merely associational – relationship between some treatment or exposure and an outcome. The statistical issues that arise in properly analyzing observational data for this goal are numerous and fascinating, and these issues are encompassed in the domain of causal inference. The research presented in this dissertation explores several distinct aspects of causal inference. This dissertation …
Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder
Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder
Electronic Theses and Dissertations
In data sets where there are a small number of observations but a large number of variables observed for each observation, ordinary least squares estimation cannot be used for regression models. There are many alternative including stepwise regression, penalized methods such as ridge regression and the LASSO, and methods based on derived inputs such as principal components regression and partial least squares regression. In this thesis, these five methods are described. K-fold cross validation is also discussed as a way for determining regularization parameters for each method. The performance of these methods in estimation and prediction is also examined through …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden
Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden
Electronic Theses and Dissertations
Many tools exist to assist law enforcement agencies in mitigating criminal activity. For centuries, academics used statistics in the study of crime and criminals, and more recently, police departments make use of spatial statistics and geographic information systems in that pursuit. Clustering and hot spot methods of analysis are popular in this application for their relative simplicity of interpretation and ease of process. With recent advancements in geospatial technology, it is easier than ever to publicly share data through visual communication tools like web applications and dashboards. Sharing data and results of analyses boosts transparency and the public image of …
Function Space Tensor Decomposition And Its Application In Sports Analytics, Justin Reising
Function Space Tensor Decomposition And Its Application In Sports Analytics, Justin Reising
Electronic Theses and Dissertations
Recent advancements in sports information and technology systems have ushered in a new age of applications of both supervised and unsupervised analytical techniques in the sports domain. These automated systems capture large volumes of data points about competitors during live competition. As a result, multi-relational analyses are gaining popularity in the field of Sports Analytics. We review two case studies of dimensionality reduction with Principal Component Analysis and latent factor analysis with Non-Negative Matrix Factorization applied in sports. Also, we provide a review of a framework for extending these techniques for higher order data structures. The primary scope of this …
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Electronic Theses and Dissertations
Premature birth has been identified as the single greatest cause of death worldwide in children under the age of five. This thesis will implement binary logistic regression and proportional odds ordinal logistic regression to predict different levels of premature birth and identify associated risk factors. The models will be built from the Center for Disease Control and Prevention's 2014 Vital Statistics Natality Birth Data containing nearly 4 million live births within the United States. Odds ratios and confidence intervals on risk factors were produced utilizing binary logistic regression.
A Comparison Of Standard Denoising Methods For Peptide Identification, Skylar Carpenter
A Comparison Of Standard Denoising Methods For Peptide Identification, Skylar Carpenter
Electronic Theses and Dissertations
Peptide identification using tandem mass spectrometry depends on matching the observed spectrum with the theoretical spectrum. The raw data from tandem mass spectrometry, however, is often not optimal because it may contain noise or measurement errors. Denoising this data can improve alignment between observed and theoretical spectra and reduce the number of peaks. The method used by Lewis et. al (2018) uses a combined constant and moving threshold to denoise spectra. We compare the effects of using the standard preprocessing methods baseline removal, wavelet smoothing, and binning on spectra with Lewis et. al’s threshold method. We consider individual methods and …
Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor
Electronic Theses and Dissertations
Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Electronic Theses and Dissertations
Sorting out data into partitions is increasing becoming complex as the constituents of data is growing outward everyday. Mixed data comprises continuous, categorical, directional functional and other types of variables. Clustering mixed data is based on special dissimilarities of the variables. Some data types may influence the clustering solution. Assigning appropriate weight to the functional data may improve the performance of the clustering algorithm. In this paper we use the extension of the Gower coefficient with judciously chosen weight for the L2 to cluster mixed data.The benefits of weighting are demonstrated both in in applications to the Buoy data set …
Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen
Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen
Electronic Theses and Dissertations
The bootstrap procedure is widely used in nonparametric statistics to generate an empirical sampling distribution from a given sample data set for a statistic of interest. Generally, the results are good for location parameters such as population mean, median, and even for estimating a population correlation. However, the results for a population variance, which is a spread parameter, are not as good due to the resampling nature of the bootstrap method. Bootstrap samples are constructed using sampling with replacement; consequently, groups of observations with zero variance manifest in these samples. As a result, a bootstrap variance estimator will carry a …
Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek
Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek
Electronic Theses and Dissertations
ABSTRACT
Examination and Comparison of the Performance of Common Non-Parametric and Robust Regression Models
By
Gregory Frank Malek
Stephen F. Austin State University, Masters in Statistics Program,
Nacogdoches, Texas, U.S.A.
This work investigated common alternatives to the least-squares regression method in the presence of non-normally distributed errors. An initial literature review identified a variety of alternative methods, including Theil Regression, Wilcoxon Regression, Iteratively Re-Weighted Least Squares, Bounded-Influence Regression, and Bootstrapping methods. These methods were evaluated using a simple simulated example data set, as well as various real data sets, including math proficiency data, Belgian telephone call data, and faculty …
Uses Of The Hypergeometric Distribution For Determining Survival Or Complete Representation Of Subpopulations In Sequential Sampling, Brooke Busbee
Uses Of The Hypergeometric Distribution For Determining Survival Or Complete Representation Of Subpopulations In Sequential Sampling, Brooke Busbee
Electronic Theses and Dissertations
This thesis will explore the hypergeometric probability distribution by looking at many different aspects of the distribution. These include, and are not limited to: history and origin, derivation and elementary applications, properties, relationships to other probability models, kindred hypergeometric distributions and elements of statistical inference associated with the hypergeometric distribution. Once the above are established, an investigation into and furthering of work done by Walton (1986) and Charlambides (2005) will be done. Here, we apply the hypergeometric distribution to sequential sampling in order to determine a surviving subcategory as well as study the problem of and complete representation of the …
A Cross-Sectional Exploration Of Household Financial Reactions And Homebuyer Awareness Of Registered Sex Offenders In A Rural, Suburban, And Urban County., John Charles Navarro
A Cross-Sectional Exploration Of Household Financial Reactions And Homebuyer Awareness Of Registered Sex Offenders In A Rural, Suburban, And Urban County., John Charles Navarro
Electronic Theses and Dissertations
As stigmatized persons, registered sex offenders betoken instability in communities. Depressed home sale values are associated with the presence of registered sex offenders even though the public is largely unaware of the presence of registered sex offenders. Using a spatial multilevel approach, the current study examines the role registered sex offenders influence sale values of homes sold in 2015 for three U.S. counties (rural, suburban, and urban) located in Illinois and Kentucky within the social disorganization framework. Homebuyers were surveyed to examine whether awareness of local registered sex offenders and the homebuyer’s community type operate as moderators between home selling …
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch
Electronic Theses and Dissertations
Missing data is one of the challenges we are facing today in modeling valid statistical models. It reduces the representativeness of the data samples. Hence, population estimates, and model parameters estimated from such data are likely to be biased.
However, the missing data problem is an area under study, and alternative better statistical procedures have been presented to mitigate its shortcomings. In this paper, we review causes of missing data, and various methods of handling missing data. Our main focus is evaluating various multiple imputation (MI) methods from the multiple imputation of chained equation (MICE) package in the statistical software …
Denoising Tandem Mass Spectrometry Data, Felix Offei
Denoising Tandem Mass Spectrometry Data, Felix Offei
Electronic Theses and Dissertations
Protein identification using tandem mass spectrometry (MS/MS) has proven to be an effective way to identify proteins in a biological sample. An observed spectrum is constructed from the data produced by the tandem mass spectrometer. A protein can be identified if the observed spectrum aligns with the theoretical spectrum. However, data generated by the tandem mass spectrometer are affected by errors thus making protein identification challenging in the field of proteomics. Some of these errors include wrong calibration of the instrument, instrument distortion and noise. In this thesis, we present a pre-processing method, which focuses on the removal of noisy …
Meta-Analyses Of The Relationship Between Depression And Nine Dimensions Of Perfectionism, Gabriel Lynn Hottinger
Meta-Analyses Of The Relationship Between Depression And Nine Dimensions Of Perfectionism, Gabriel Lynn Hottinger
Electronic Theses and Dissertations
Perfectionism has been shown to be related to depression, but perfectionism is multidimensional. Some dimensions are related to positive psychological characteristics and outcomes and other dimensions are related to negative psychological characteristics and outcomes. This study reports results of nine meta-analyses performed to investigate the association between each of nine subscales of perfectionism and depression to determine which dimensions of perfectionism are most strongly associated with depression. The two subscales that were used from the Hewitt and Flett (1991b) Multidimensional Perfectionism scale were Self-Oriented Perfectionism (SOP) and Socially-Prescribed Perfectionism (SPP). The five subscales that were used from the Frost et …
A Multi-Indexed Logistic Model For Time Series, Xiang Liu
A Multi-Indexed Logistic Model For Time Series, Xiang Liu
Electronic Theses and Dissertations
In this thesis, we explore a multi-indexed logistic regression (MILR) model, with particular emphasis given to its application to time series. MILR includes simple logistic regression (SLR) as a special case, and the hope is that it will in some instances also produce significantly better results. To motivate the development of MILR, we consider its application to the analysis of both simulated sine wave data and stock data. We looked at well-studied SLR and its application in the analysis of time series data. Using a more sophisticated representation of sequential data, we then detail the implementation of MILR. We compare …
Multilevel Models For Longitudinal Data, Aastha Khatiwada
Multilevel Models For Longitudinal Data, Aastha Khatiwada
Electronic Theses and Dissertations
Longitudinal data arise when individuals are measured several times during an ob- servation period and thus the data for each individual are not independent. There are several ways of analyzing longitudinal data when different treatments are com- pared. Multilevel models are used to analyze data that are clustered in some way. In this work, multilevel models are used to analyze longitudinal data from a case study. Results from other more commonly used methods are compared to multilevel models. Also, comparison in output between two software, SAS and R, is done. Finally a method consisting of fitting individual models for each …
Spatio-Temporal Analysis Of Point Patterns, Abdul-Nasah Soale
Spatio-Temporal Analysis Of Point Patterns, Abdul-Nasah Soale
Electronic Theses and Dissertations
In this thesis, the basic tools of spatial statistics and time series analysis are applied to the case study of the earthquakes in a certain geographical region and time frame. Then some of the existing methods for joint analysis of time and space are described and applied. Finally, additional research questions about the spatial-temporal distribution of the earthquakes are posed and explored using statistical plots and models. The focus in the last section is in the relationship between number of events per year and maximum magnitude and its effect on how clustered the spatial distribution is and the relationship between …
Newsvendor Models With Monte Carlo Sampling, Ijeoma W. Ekwegh
Newsvendor Models With Monte Carlo Sampling, Ijeoma W. Ekwegh
Electronic Theses and Dissertations
Newsvendor Models with Monte Carlo Sampling by Ijeoma Winifred Ekwegh The newsvendor model is used in solving inventory problems in which demand is random. In this thesis, we will focus on a method of using Monte Carlo sampling to estimate the order quantity that will either maximizes revenue or minimizes cost given that demand is uncertain. Given data, the Monte Carlo approach will be used in sampling data over scenarios and also estimating the probability density function. A bootstrapping process yields an empirical distribution for the order quantity that will maximize the expected profit. Finally, this method will be used …
Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft
Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft
Electronic Theses and Dissertations
Observational data presents unique challenges for analysis that are not encountered with experimental data resulting from carefully designed randomized controlled trials. Selection bias and unbalanced treatment assignments can obscure estimations of treatment effects, making the process of causal inference from observational data highly problematic. In 1983, Paul Rosenbaum and Donald Rubin formalized an approach for analyzing observational data that adjusts treatment effect estimates for the set of non-treatment variables that are measured at baseline. The propensity score is the conditional probability of assignment to a treatment group given the covariates. Using this score, one may balance the covariates across treatment …
Analyses Of 2002-2013 China’S Stock Market Using The Shared Frailty Model, Chao Tang
Analyses Of 2002-2013 China’S Stock Market Using The Shared Frailty Model, Chao Tang
Electronic Theses and Dissertations
This thesis adopts a survival model to analyze China’s stock market. The data used are the capitalization-weighted stock market index (CSI 300) and the 300 stocks for creating the index. We define the recurrent events using the daily return of the selected stocks and the index. A shared frailty model which incorporates the random effects is then used for analyses since the survival times of individual stocks are correlated. Maximization of penalized likelihood is presented to estimate the parameters in the model. The covariates are selected using the Akaike information criterion (AIC) and the variance inflation factor (VIF) to avoid …
Are Highly Dispersed Variables More Extreme? The Case Of Distributions With Compact Support, Benedict E. Adjogah
Are Highly Dispersed Variables More Extreme? The Case Of Distributions With Compact Support, Benedict E. Adjogah
Electronic Theses and Dissertations
We consider discrete and continuous symmetric random variables X taking values in [0; 1], and thus having expected value 1/2. The main thrust of this investigation is to study the correlation between the variance, Var(X) of X and the value of the expected maximum E(Mn) = E(X1,...,Xn) of n independent and identically distributed random variables X1,X2,...,Xn, each distributed as X. Many special cases are studied, some leading to very interesting alternating sums, and some progress is made towards a general theory.
Solving The Differential Equation For The Probit Function Using A Variant Of The Carleman Embedding Technique., Kelechukwu Iroajanma Alu
Solving The Differential Equation For The Probit Function Using A Variant Of The Carleman Embedding Technique., Kelechukwu Iroajanma Alu
Electronic Theses and Dissertations
The probit function is the inverse of the cumulative distribution function associated with the standard normal distribution. It is of great utility in statistical modelling. The Carleman embedding technique has been shown to be effective in solving first order and, less efficiently, second order nonlinear differential equations. In this thesis, we show that solutions to the second order nonlinear differential equation for the probit function can be approximated efficiently using a variant of the Carleman embedding technique.
Early Stopping Of A Neural Network Via The Receiver Operating Curve., Daoping Yu
Early Stopping Of A Neural Network Via The Receiver Operating Curve., Daoping Yu
Electronic Theses and Dissertations
This thesis presents the area under the ROC (Receiver Operating Characteristics) curve, or abbreviated AUC, as an alternate measure for evaluating the predictive performance of ANNs (Artificial Neural Networks) classifiers. Conventionally, neural networks are trained to have total error converge to zero which may give rise to over-fitting problems. To ensure that they do not over fit the training data and then fail to generalize well in new data, it appears effective to stop training as early as possible once getting AUC sufficiently large via integrating ROC/AUC analysis into the training process. In order to reduce learning costs involving the …
Examining Significant Differences Of Gunshot Residue Patterns Using Same Make And Model Of Firearms In Forensic Distance Determination Tests., Heather Lewey
Electronic Theses and Dissertations
In many cases of crimes involving a firearm, police investigators need to know how far the firearm was held from the victim when it was discharged. Knowing this distance, vital questions regarding the re-construction of the crime scene can be known. Often, the original firearm used in commission of a suspected crime is not available for testing or is damaged. Crime laboratories require the original firearm in order to conduct distance determination tests. However, no empirical research has ever been conducted to determine if same make and model firearms produce different results in distance determination testing. It was the purpose …
Building A Model Of Campus Diversity: An Organization-As-The-Individual Based Approach, Ian Spencer Ray
Building A Model Of Campus Diversity: An Organization-As-The-Individual Based Approach, Ian Spencer Ray
Electronic Theses and Dissertations
Diversity is a vital concept when engaging with modern social science research and among higher education practitioners. Despite the pervasive discourse surrounding diversity, equity, and inclusion (DEI) initiatives, quantification of diversity remains a challenge across the social sciences. This study proposes a new conceptual framework, applies the framework to the quantification of diversity and measurement of its impacts. Further, this study demonstrates the robustness of the new conceptual framework alongside the positive impacts of diversity on institutional outcomes.
The Organization-As-The-Individual (OATI) framework arises from an integration of the social and biological sciences. The framework argues that humans, as biological organisms, …