Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Electronic Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 121 - 150 of 261

Full-Text Articles in Statistics and Probability

Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang Jan 2020

Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang

Electronic Theses and Dissertations

Measurement invariance is crucial for an effective and valid measure of a construct. Invariance holds when the latent trait varies consistently across subgroups; in other words, the mean differences among subgroups are only due to true latent ability differences. Differential item functioning (DIF) occurs when measurement invariance is violated. There are two kinds of traditional tools for DIF detection: non-parametric methods and parametric methods. Mantel Haenszel (MH), SIBTEST, and standardization are examples of non-parametric DIF detection methods. The majority of parametric DIF detection methods are item response theory (IRT) based. Both non-parametric methods and parametric methods compare differences among subgroups …


How 6-12th Grade Staff Support Students With Depression: A Pilot Study To Develop Measures Of Implicit Associations, Explicit Attitudes And Helping Behavior, Paul M. Thompson Jan 2020

How 6-12th Grade Staff Support Students With Depression: A Pilot Study To Develop Measures Of Implicit Associations, Explicit Attitudes And Helping Behavior, Paul M. Thompson

Electronic Theses and Dissertations

Students with emotional disabilities are disproportionately suspended and expelled in K-12 schools. Attribution theory suggests individuals are less likely to provide assistance to others if they believe the individuals are responsible for their own difficulties. To test attribution theory, this study created new measures of explicit attitudes and implicit associations of licensed 6-12th grade staff regarding students with depression as well as a helping behavior measure of staff toward students with depression. The survey was distributed within a single school district in the western United States. A majority of the sample (N = 52) held a mental health license (60%), …


Statistical Methods For Estimating And Testing Treatment Effect For Multiple Treatment Groups In Observational Studies., Xiaofang Yan Dec 2019

Statistical Methods For Estimating And Testing Treatment Effect For Multiple Treatment Groups In Observational Studies., Xiaofang Yan

Electronic Theses and Dissertations

Note: Abstract would not save due to an issue with some of the characters.


Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden Dec 2019

Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden

Electronic Theses and Dissertations

Many tools exist to assist law enforcement agencies in mitigating criminal activity. For centuries, academics used statistics in the study of crime and criminals, and more recently, police departments make use of spatial statistics and geographic information systems in that pursuit. Clustering and hot spot methods of analysis are popular in this application for their relative simplicity of interpretation and ease of process. With recent advancements in geospatial technology, it is easier than ever to publicly share data through visual communication tools like web applications and dashboards. Sharing data and results of analyses boosts transparency and the public image of …


Function Space Tensor Decomposition And Its Application In Sports Analytics, Justin Reising Dec 2019

Function Space Tensor Decomposition And Its Application In Sports Analytics, Justin Reising

Electronic Theses and Dissertations

Recent advancements in sports information and technology systems have ushered in a new age of applications of both supervised and unsupervised analytical techniques in the sports domain. These automated systems capture large volumes of data points about competitors during live competition. As a result, multi-relational analyses are gaining popularity in the field of Sports Analytics. We review two case studies of dimensionality reduction with Principal Component Analysis and latent factor analysis with Non-Negative Matrix Factorization applied in sports. Also, we provide a review of a framework for extending these techniques for higher order data structures. The primary scope of this …


Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood Aug 2019

Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood

Electronic Theses and Dissertations

Premature birth has been identified as the single greatest cause of death worldwide in children under the age of five. This thesis will implement binary logistic regression and proportional odds ordinal logistic regression to predict different levels of premature birth and identify associated risk factors. The models will be built from the Center for Disease Control and Prevention's 2014 Vital Statistics Natality Birth Data containing nearly 4 million live births within the United States. Odds ratios and confidence intervals on risk factors were produced utilizing binary logistic regression.


Designing And Sample Size Calculation In Presence Of Heterogeneity In Biological Studies Involving High-Throughput Data., Sudhir Srivastava Aug 2019

Designing And Sample Size Calculation In Presence Of Heterogeneity In Biological Studies Involving High-Throughput Data., Sudhir Srivastava

Electronic Theses and Dissertations

The designing and determination of sample size are important for conducting high-throughput biological experiments such as proteomics experiments and RNA-Seq expression studies, thus leading to better understanding of complex mechanisms underlying various biological processes. The variations in the biological data or technical approaches to data collection lead to heterogeneity for the samples under study. We critically worked on the issues of technical and biological heterogeneity. The quantitative measurements based on liquid chromatography (LC) coupled with mass spectrometry (MS) often suffer from the problem of missing values (MVs) and data heterogeneity. We considered a proteomics data set generated from human kidney …


Novel Bayesian Methodology In Multivariate Problems., Debamita Kundu Aug 2019

Novel Bayesian Methodology In Multivariate Problems., Debamita Kundu

Electronic Theses and Dissertations

This dissertation involves developing novel Bayesian methodology for multivariate problems. In particular, it focuses on two contexts: shrinkage based variable selection in multivariate regression and simultaneous covariance estimation of multiple groups. Both these projects are centered around fully Bayesian inference schemes based on hierarchical modeling to capture context-specific features of the data and the development of computationally efficient estimation algorithm. Variable selection over a potentially large set of covariates in a linear model is quite popular. In the Bayesian context, common prior choices can lead to a posterior expectation of the regression coefficients that is a sparse (or nearly sparse) …


Robustness Of Semi-Parametric Survival Model: Simulation Studies And Application To Clinical Data, Isaac Nwi-Mozu Aug 2019

Robustness Of Semi-Parametric Survival Model: Simulation Studies And Application To Clinical Data, Isaac Nwi-Mozu

Electronic Theses and Dissertations

An efficient way of analyzing survival clinical data such as cancer data is a great concern to health experts. In this study, we investigate and propose an efficient way of handling survival clinical data. Simulation studies were conducted to compare performances of various forms of survival model techniques using an R package ``survsim". Models performance was conducted with varying sample sizes as small ($n5000$). For small and mild samples, the performance of the semi-parametric outperform or approximate the performance of the parametric model. However, for large samples, the parametric model outperforms the semi-parametric model. We compared the effectiveness and reliability …


Is Corequisite Developmental Math Effective At East Tennessee State University?, Christine Padden Aug 2019

Is Corequisite Developmental Math Effective At East Tennessee State University?, Christine Padden

Electronic Theses and Dissertations

This thesis looks at the corequisite developmental math program at East Tennessee State University (ETSU) and compares the effectiveness to the previous developmental math program by comparing the student outcomes in MATH 1530. MATH 1530 is a non-calculus based statistic and probability course that satisfies most majors’ general education math requirements. ETSU sees approximately 1,000 students a year pass through MATH 1530 which is around 6.7% of the total enrollment at ETSU[9]. We are interested in the last five years of the developmental math program before it was changed to corequisite developmental math and the first five years of corequisite …


A Systematic Assessment Of Socio-Economic Impacts Of Prolonged Episodic Volcano Crises, Justin Peers May 2019

A Systematic Assessment Of Socio-Economic Impacts Of Prolonged Episodic Volcano Crises, Justin Peers

Electronic Theses and Dissertations

Uncertainty surrounding volcanic activity can lead to socio-economic crises with or without an eruption as demonstrated by the post-1978 response to unrest of Long Valley Caldera (LVC), CA. Extensive research in physical sciences provides a foundation on which to assess direct impacts of hazards, but fewer resources have been dedicated towards understanding human responses to volcanic risk. To evaluate natural hazard risk issues at LVC, a multi-hazard, mail-based, household survey was conducted to compare perceptions of volcanic, seismic, and wildfire hazards. Impacts of volcanic activity on housing prices and businesses were examined at the county-level for three volcanoes with a …


Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt May 2019

Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt

Electronic Theses and Dissertations

A statistician's job is to produce statistical models. When these models are precise and unbiased, we can relate them to new data appropriately. However, when data sets have missing values, assumptions to statistical methods are violated and produce biased results. The statistician's objective is to implement methods that produce unbiased and accurate results. Research in missing data is becoming popular as modern methods that produce unbiased and accurate results are emerging, such as MICE in R, a statistical software. Using real data, we compare four common imputation methods, in the MICE package in R, at different levels of missingness. The …


Generalizations Of The Arcsine Distribution, Rebecca Rasnick May 2019

Generalizations Of The Arcsine Distribution, Rebecca Rasnick

Electronic Theses and Dissertations

The arcsine distribution looks at the fraction of time one player is winning in a fair coin toss game and has been studied for over a hundred years. There has been little further work on how the distribution changes when the coin tosses are not fair or when a player has already won the initial coin tosses or, equivalently, starts with a lead. This thesis will first cover a proof of the arcsine distribution. Then, we explore how the distribution changes when the coin the is unfair. Finally, we will explore the distribution when one person has won the first …


A Comparison Of Standard Denoising Methods For Peptide Identification, Skylar Carpenter May 2019

A Comparison Of Standard Denoising Methods For Peptide Identification, Skylar Carpenter

Electronic Theses and Dissertations

Peptide identification using tandem mass spectrometry depends on matching the observed spectrum with the theoretical spectrum. The raw data from tandem mass spectrometry, however, is often not optimal because it may contain noise or measurement errors. Denoising this data can improve alignment between observed and theoretical spectra and reduce the number of peaks. The method used by Lewis et. al (2018) uses a combined constant and moving threshold to denoise spectra. We compare the effects of using the standard preprocessing methods baseline removal, wavelet smoothing, and binning on spectra with Lewis et. al’s threshold method. We consider individual methods and …


A Comparison Of Bayesian Estimation Techniques In A Multidimensional Two-Parameter Partial Credit Item Response Model, Peiyan Liu Jan 2019

A Comparison Of Bayesian Estimation Techniques In A Multidimensional Two-Parameter Partial Credit Item Response Model, Peiyan Liu

Electronic Theses and Dissertations

Bayesian estimation methods have shown better performance than the traditional Marginal Maximum Likelihood (MML) estimation method for parameter estimation in relatively simple item response models. However, extant literature is lacking on the investigation of Bayesian parameter estimation approaches for a multidimensional two parameter partial credit (M2PPC) model, therefore this simulation study investigated the performance of two Bayesian Markov Chain Monte Carlo (MCMC) algorithms: Gibbs Sampler and Hamiltonian Monte Carlo-No-U-Turn-Sampler (HMC-NUTS) for M2PPC models' parameter estimation. It compared the estimation accuracy and computing speed in different combinations of situations, including prior choices, test lengths, and the relationships between dimensions.

The datasets …


Finite Mixture Of Regression Models For Complex Survey Data, Abdelbaset Abdalla Jan 2019

Finite Mixture Of Regression Models For Complex Survey Data, Abdelbaset Abdalla

Electronic Theses and Dissertations

Over time, survey data has become an essential source of information for modern society. However, to be effective, the structures of survey data require sampling designs that are more complex than simple random sampling. The complex sampling data collected from enormous national surveys via these complex designs ideally include sample weights that allow analysis to take account of complicated population structures. When the target of inference is the parameters of a regression model, it is crucial to know whether these weights should be incorporated into the sampling weight when fitting the model to the survey data. The finite mixture models …


Cramer Type Moderate Deviations For Random Fields And Mutual Information Estimation For Mixed-Pair Random Variables, Aleksandr Beknazaryan Jan 2019

Cramer Type Moderate Deviations For Random Fields And Mutual Information Estimation For Mixed-Pair Random Variables, Aleksandr Beknazaryan

Electronic Theses and Dissertations

In this dissertation we first study Cramer type moderate deviation for partial sums of random fields by applying the conjugate method. In 1938 Cramer published his results on large deviations of sums of i.i.d. random variables after which a lot of research has been done on establishing Cramer type moderate and large deviation theorems for different types of random variables and for various statistics. In particular results have been obtained for independent non-identically distributed random variables for the sum of independent random to estimate the mutual information between two random variables. The estimates enjoy a central limit theorem under some …


Development Of A Data-Driven Patient Engagement Score Using Finite Mixture Models, Eric Bae Jan 2019

Development Of A Data-Driven Patient Engagement Score Using Finite Mixture Models, Eric Bae

Electronic Theses and Dissertations

Patient activation measure (PAM) is widely adopted by health care providers to access individual's knowledge, skill, and confidence for managing one's health and healthcare. Patient activation measure (PAM), licensed by Insignia Health, is widely adopted by health care providers to access individual's knowledge, skill, and confidence for managing one's health and healthcare. Multiple studies corroborate the effectiveness of activation measure in predicting most health behaviors, including preventive behaviors, healthy behaviors, self-management behaviors, and health information seeking. However, PAM is heavily dependent on subjective patient-reported data, which are often incomplete. The purpose of this study is to develop an objective statistical …


Mefenamic Acid – Hpmc As Hg Amorphous Solid Dispersions: Dissolution Enhancement Using Hot Melt Extrusion Technology, Ashay Shukla Jan 2019

Mefenamic Acid – Hpmc As Hg Amorphous Solid Dispersions: Dissolution Enhancement Using Hot Melt Extrusion Technology, Ashay Shukla

Electronic Theses and Dissertations

Mefenamic acid, a BCS class II drug, displays high permeability and low solubility, thereby exhibiting a poor dissolution profile. Hence to improve the solubility and dissolution rate of Mefenamic acid, Hot Melt Extrusion (HME) technique was employed. The amorphous solid dispersion matrix exhibited enhanced dissolution with desired release characteristics. Hydroxypropylmethylcellulose acetate succinate (AquaSolve™ HPMC-AS HG) was used as a carrier with the poloxamer (Kolliphor P407). The drug load was varied from 20% to 40% within the blend. Drug and polymers were blended using a twin shell V-blender for 10 minutes and extruded using an 11mm twin-screw co-rotating extruder (ThermoFisher Scientific, …


Applying Machine Learning Algorithms For The Analysis Of Biological Sequences And Medical Records, Shaopeng Gu Jan 2019

Applying Machine Learning Algorithms For The Analysis Of Biological Sequences And Medical Records, Shaopeng Gu

Electronic Theses and Dissertations

The modern sequencing technology revolutionizes the genomic research and triggers explosive growth of DNA, RNA, and protein sequences. How to infer the structure and function from biological sequences is a fundamentally important task in genomics and proteomics fields. With the development of statistical and machine learning methods, an integrated and user-friendly tool containing the state-of-the-art data mining methods are needed. Here, we propose SeqFea-Learn, a comprehensive Python pipeline that integrating multiple steps: feature extraction, dimensionality reduction, feature selection, predicting model constructions based on machine learning and deep learning approaches to analyze sequences. We used enhancers, RNA N6- methyladenosine sites and …


Effectiveness Of Prescribed Fire On Meeting Fuel Load And Wildlife Habitat Management Objectives In East Texas National Forests, Trey Wall Dec 2018

Effectiveness Of Prescribed Fire On Meeting Fuel Load And Wildlife Habitat Management Objectives In East Texas National Forests, Trey Wall

Electronic Theses and Dissertations

Using standardized methodology outlined by the United States Forest Service and the National Forests and Grasslands in Texas’ Fire Monitoring Program for data collection, the efficacy of current Forest Service prescribed burn regimes were analyzed for 24 study sites in East Texas National Forests. Study sites were located within Sam Houston, Davy Crockett, and Angelina/Sabine National Forests. Efficacy was determined by comparing defined management objectives established by the Forest Service to the data collected at the study sites. The results conclude that most objectives, as outlined by the Forest Service, are not being met with the current practices. Re-visitation of …


Innate Immunity, The Hepatic Extracellular Matrix, And Liver Injury: Mathematical Modeling Of Metastatic Potential And Tumor Development In Alcoholic Liver Disease., Shanice V. Hudson Dec 2018

Innate Immunity, The Hepatic Extracellular Matrix, And Liver Injury: Mathematical Modeling Of Metastatic Potential And Tumor Development In Alcoholic Liver Disease., Shanice V. Hudson

Electronic Theses and Dissertations

The overarching goals of the current work are to fill key gaps in the current understanding of alcohol consumption and the risk of metastasis to the liver. Considering the evidence this research group has compiled confirming that the hepatic matrisome responds dynamically to injury, an altered extracellular matrix (ECM) profile appears to be a key feature of pre-fibrotic inflammatory injury in the liver. This group has demonstrated that the hepatic ECM responds dynamically to alcohol exposure, in particular, sensitizing the liver to LPS-induced inflammatory damage. Although the study of alcohol in its role as a contributing factor to oncogenesis and …


Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage Aug 2018

Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage

Electronic Theses and Dissertations

This thesis is based on a Poisson model that uses both error-free data and error-prone data subject to misclassification in the form of false-negative and false-positive counts. We present maximum likelihood estimators (MLEs), Fisher's Information, and Wald statistics for Poisson rate parameter and the two misclassification parameters. Next, we invert the Wald statistics to get asymptotic confidence intervals for Poisson rate parameter and false-negative rate parameter. The coverage and width properties for various sample size and parameter configurations are studied via a simulation study. Finally, we apply the MLEs and confidence intervals to one real data set and another realistic …


Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor Aug 2018

Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor

Electronic Theses and Dissertations

Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …


Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal Aug 2018

Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal

Electronic Theses and Dissertations

This dissertation consists of three projects and can be categorized in two broad research areas: generalized spatiotemporal modeling and causal inference based on observational data. In the first project, I introduce a Bayesian hierarchical mixed effect hurdle model with a nested random effect structure to model the count for primary care providers and understand their spatial and temporal variation. This study further enables us to identify the health professional shortage areas and the possible impacting factors. In the second project, I have unified popular parametric and nonparametric propensity score-based methods to assess the treatment effect of multiple groups for ordinal …


Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister Aug 2018

Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister

Electronic Theses and Dissertations

A probability distribution is a statistical function that describes the probability of possible outcomes in an experiment or occurrence. There are many different probability distributions that give the probability of an event happening, given some sample size n. An important question in statistics is to determine the distribution of the sum of independent random variables when the sample size n is fixed. For example, it is known that the sum of n independent Bernoulli random variables with success probability p is a Binomial distribution with parameters n and p: However, this is not true when the sample size …


The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh Aug 2018

The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh

Electronic Theses and Dissertations

Previous work by Flaxman (2004) and Biers-Ariel et al. (2018) focused on the number of distinct words embedded in a string of words of length n. In this thesis, we will extend this work to permutations, focusing on the maximum number of distinct permutations contained in a permutation on [n] = {1,2,...,n} and on the expected number of distinct permutations contained in a random permutation on [n]. We further considered the problem where repetition of subsequences are as a result of the occurrence of (Type A and/or Type B) replications. Our method of enumerating the Type A replications causes double …


Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong Aug 2018

Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong

Electronic Theses and Dissertations

Sorting out data into partitions is increasing becoming complex as the constituents of data is growing outward everyday. Mixed data comprises continuous, categorical, directional functional and other types of variables. Clustering mixed data is based on special dissimilarities of the variables. Some data types may influence the clustering solution. Assigning appropriate weight to the functional data may improve the performance of the clustering algorithm. In this paper we use the extension of the Gower coefficient with judciously chosen weight for the L2 to cluster mixed data.The benefits of weighting are demonstrated both in in applications to the Buoy data set …


Designing A Calibration Set In Spectral Space For Efficient Development Of An Nir Method For Tablet Analysis, Md Anik Alam May 2018

Designing A Calibration Set In Spectral Space For Efficient Development Of An Nir Method For Tablet Analysis, Md Anik Alam

Electronic Theses and Dissertations

Designing a calibration set is the first step in developing a spectroscopic calibration method for quantitative analysis of pharmaceutical tablets. This step is critical because successful model development depends on the suitability of the calibration data. For spectroscopic-based methods, traditional concentration based techniques for designing calibration sets are prone to have redundant information while simultaneously lacking necessary information for a successful calibration model. The traditional method also follows the same design approach for different spectroscopic techniques and different formulations, thereby lacks the optimizing capability to be technique and formulation specific.

A method for designing a calibration set in the Near …


Re-Evaluating Performance Measurement: New Mathematical Methods To Address Common Performance Measurement Challenges, Jordan David Benis May 2018

Re-Evaluating Performance Measurement: New Mathematical Methods To Address Common Performance Measurement Challenges, Jordan David Benis

Electronic Theses and Dissertations

Performance Measurement is an essential discipline for any business. Robust and reliable performance metrics for people, processes, and technologies enable a business to identify and address deficiencies to improve performance and profitability. The complexity of modern operating environments presents real challenges to developing equitable and accurate performance metrics. This thesis explores and develops two new methods to address common challenges encountered in businesses across the world. The first method addresses the challenge of estimating the relative complexity of various tasks by utilizing the Pearson Correlation Coefficient to identify potentially over weighted and under weighted tasks. The second method addresses the …