Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Electronic Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 61 - 90 of 261

Full-Text Articles in Statistics and Probability

Penalized Bayesian Exponential Random Graph Models., Vicki Modisette Aug 2023

Penalized Bayesian Exponential Random Graph Models., Vicki Modisette

Electronic Theses and Dissertations

Networks have the critical ability to represent the complex interconnectedness of social relationships, biological processes, and the spread of diseases and information. Exponential random graph models (ERGM) are one of the popular statistical methods for analyzing network data. ERGM, however, struggle with computational challenges and degeneracy issues, further exacerbated by their inability to handle high-dimensional network data. Bayesian techniques provide a promising avenue to overcome these two problems. This paper considers penalized Bayesian exponential random graph models with adaptive lasso and adaptive ridge penalties to perform variable selection and reduce multicollinearity on a variety of networks. The experimental results demonstrate …


An Analysis Of All-Cause Mortality On Patients With Sickle Cell Disease And Kidney Disease Using Propensity Score Matching, Adam Garrison May 2023

An Analysis Of All-Cause Mortality On Patients With Sickle Cell Disease And Kidney Disease Using Propensity Score Matching, Adam Garrison

Electronic Theses and Dissertations

In this work, we provide an overview of the Cox proportional hazards model for time to event or survival analysis and the notion of propensity score matching to deal with confounding factors. A full analysis is reported in Chapter 2 concerning mortality for in-center dialysis patients with sickle cell disease to demonstrate the application of a general analysis strategy that has some logistical benefits over more traditional approaches to accounting for confounding variables. We also provide some insight and discussions on the challenges and future research questions that will emerge when trying to implement this strategy as a monitoring tool …


Identifying And Analyzing Multi-Star Systems Among Tess Planetary Candidates Using Gaia, Katie E. Bailey May 2023

Identifying And Analyzing Multi-Star Systems Among Tess Planetary Candidates Using Gaia, Katie E. Bailey

Electronic Theses and Dissertations

Exoplanets represent a young, rapidly advancing subfield of astrophysics where much is still unknown. It is therefore important to analyze trends among their parameters to learn more about these systems. More complexity is added to these systems with the presence of additional stellar companions. To study these complex systems, one can employ programming languages such as Python to parse databases such as those constructed by TESS and Gaia to bridge the gap between exoplanets and stellar companions. Data can then be analyzed for trends in these multi-star exoplanet systems and in juxtaposition to their single-star counterparts. This research was able …


Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham May 2023

Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham

Electronic Theses and Dissertations

The abundance, accessibility, and scale of data have engendered an era where machine learning can quickly and accurately solve complex problems, identify complicated patterns, and uncover intricate trends. One research area where many have applied these techniques is the stock market. Yet, financial domains are influenced by many factors and are notoriously difficult to predict due to their volatile and multivariate behavior. However, the literature indicates that public sentiment data may exhibit significant predictive qualities and improve a model’s ability to predict intricate trends. In this study, momentum SVM classification accuracy was compared between datasets that did and did not …


Non-Destructive Imaging Of Phytosulfokine Trafficking In Plants Using Fiber-Optic Fluorescence Microscopy, Bernard Abakah May 2023

Non-Destructive Imaging Of Phytosulfokine Trafficking In Plants Using Fiber-Optic Fluorescence Microscopy, Bernard Abakah

Electronic Theses and Dissertations

Plants secrete peptide ligands and use receptor signaling to respond to stress and control development. Understanding these phenomena is key to improving plant health and productivity for food, fiber, and energy applications. Phytosulfokine (PSK), a sulfated peptide hormone, regulates plant cell division, growth, and stress tolerance via specific phytosulfokine receptors (PSKRs). This study uses fiber-optic fluorescence microscopy to elucidate trafficking of PSK in live plants. The microscope features two-color optics and an objective lens connected to a 1-m coherent imaging fiber mounted on either a conventional upright microscope body or 5-axis positioning system (X–Y–Z plus pitch and yaw). PSK and …


Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg Jan 2023

Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg

Electronic Theses and Dissertations

Using a multi-level ensemble design, we forecast international stock market returns with a novel high-dimensional data set of aggregated cross sectional firm-level predictors. The method includes considerations of model uncertainty, parameter instability, model density and non-linearities with machine learning, shrinkage and model averaging. We provide evidence that it is important to systematically focus on all four sources of forecast failure, shed light on the sparsity/density debate in the stock return forecasting dialogue and contribute interesting findings on the efficacy dimensionality reduction with principal components analysis and partial least squares. The robustness of the approach is demonstrated through applications in four …


A Class Of Regression Models For Pairwise Comparisons Of Forensic Handwriting Comparison Systems, Cami M. Fuglsby Jan 2023

A Class Of Regression Models For Pairwise Comparisons Of Forensic Handwriting Comparison Systems, Cami M. Fuglsby

Electronic Theses and Dissertations

Handwriting analysis is a complex field largely living in forensic science and the legal realm. One task of a forensic document examiner (FDE) may be to determine the writer(s) of handwritten documents. Automated identification systems (AIS) were built to aid FDEs in their examinations. Part of the uses of these AIS (such as FISH[5] [7],WANDA [6], CEDAR-FOX [17], and FLASHID®2) are tomeasure features about a handwriting sample and to provide the user with a numeric value of the evidence. These systems use their own algorithms and definitions of features to quantify the writing and can be considered a black-box. The …


A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani Dec 2022

A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani

Electronic Theses and Dissertations

Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.

Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance …


Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury Dec 2022

Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury

Electronic Theses and Dissertations

Graphical models determine associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models, where the relationships are formalized by non-null entries of the precision matrix. However, in high-dimensional cases, covariance estimates are typically unstable. Moreover, it is natural to expect only a few significant associations to be present in many realistic applications. This necessitates the injection of sparsity techniques into the estimation method. Classical frequentist methods, like GLASSO, use penalization techniques for this purpose. Fully Bayesian methods, on the contrary, are slow because they require iteratively sampling over a quadratic …


Mathematical Models Yield Insights Into Cnns: Applications In Natural Image Restoration And Population Genetics, Ryan Cecil Aug 2022

Mathematical Models Yield Insights Into Cnns: Applications In Natural Image Restoration And Population Genetics, Ryan Cecil

Electronic Theses and Dissertations

Due to a rise in computational power, machine learning (ML) methods have become the state-of-the-art in a variety of fields. Known to be black-box approaches, however, these methods are oftentimes not well understood. In this work, we utilize our understanding of model-based approaches to derive insights into Convolutional Neural Networks (CNNs). In the field of Natural Image Restoration, we focus on the image denoising problem. Recent work have demonstrated the potential of mathematically motivated CNN architectures that learn both `geometric' and nonlinear higher order features and corresponding regularizers. We extend this work by showing that not only can geometric features …


Computer Aided Diagnosis System For Breast Cancer Using Deep Learning., Asma Baccouche Aug 2022

Computer Aided Diagnosis System For Breast Cancer Using Deep Learning., Asma Baccouche

Electronic Theses and Dissertations

The recent rise of big data technology surrounding the electronic systems and developed toolkits gave birth to new promises for Artificial Intelligence (AI). With the continuous use of data-centric systems and machines in our lives, such as social media, surveys, emails, reports, etc., there is no doubt that data has gained the center of attention by scientists and motivated them to provide more decision-making and operational support systems across multiple domains. With the recent breakthroughs in artificial intelligence, the use of machine learning and deep learning models have achieved remarkable advances in computer vision, ecommerce, cybersecurity, and healthcare. Particularly, numerous …


Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage Aug 2022

Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage

Electronic Theses and Dissertations

Due to the wide availability of functional data from multiple disciplines, the studies of functional data analysis have become popular in the recent literature. However, the related development in censored survival data has been relatively sparse. In Chapter 2, we consider the problem of analyzing time-to-event data in the presence of functional predictors. We develop a conditional generalized Kaplan Meier (KM) estimator that incorporates functional predictors using kernel weights and rigorously establishes its asymptotic properties. In addition, we propose to select the optimal bandwidth based on a time-dependent Brier score. We then carry out extensive numerical studies to examine the …


Statistical Methods For Assessing Drug Interactions And Identifying Effect Modifiers Using Observational Data., Qian Xu May 2022

Statistical Methods For Assessing Drug Interactions And Identifying Effect Modifiers Using Observational Data., Qian Xu

Electronic Theses and Dissertations

This dissertation consists of three projects related to causal inference based on observational data. In the first project, we propose a double robust to identify the effect modifiers and estimate optimal treatment. Observational studies differ from experimental studies in that assignment of subjects to treatments is not randomized but rather occurs due to natural mechanisms, which are usually hidden from the researchers. Many statistical methods to identify the treatment effect and select the optimal personalized treatment for experimental studies may not be suitable for observational studies any more. In this project, we propose a exible outcome model to select the …


Finding A Representative Distribution For The Tail Index Alpha, Α, For Stock Return Data From The New York Stock Exchange, Jett Burns May 2022

Finding A Representative Distribution For The Tail Index Alpha, Α, For Stock Return Data From The New York Stock Exchange, Jett Burns

Electronic Theses and Dissertations

Statistical inference is a tool for creating models that can accurately display real-world events. Special importance is given to the financial methods that model risk and large price movements. A parameter that describes tail heaviness, and risk overall, is α. This research finds a representative distribution that models α. The absolute value of standardized stock returns from the Center for Research on Security Prices are used in this research. The inference is performed using R. Approximations for α are found using the ptsuite package. The GAMLSS package employs maximum likelihood estimation to estimate distribution parameters using the CRSP data. The …


Development And Initial Validation Of An Acculturation Measure For Arab Expatriates In Qatar Using Mixed Methods, Sara Ali Ahmed Zikri Jan 2022

Development And Initial Validation Of An Acculturation Measure For Arab Expatriates In Qatar Using Mixed Methods, Sara Ali Ahmed Zikri

Electronic Theses and Dissertations

Prior research on acculturation has predominantly focused on specific immigrant and ethnic groups in the United States, Europe, and Asia, and there is a lack of publicly available scales that measure acculturation among expatriates in countries of the Gulf Cooperation Council. The purpose of this study was to use a mixed methods research design to develop a theoretical framework of acculturation for white-collar Arab expatriates in Qatar as well as a culturally appropriate self-report acculturation measure. The acculturation stories and experiences of white-collar Arab expatriates in Qatar were qualitatively explored and analyzed, and their findings were used to construct an …


Expanding The Network Evaluation Toolkit: Combining Social Network Analysis & Qualitative Comparative Analysis, Debbie Gowensmith Jan 2022

Expanding The Network Evaluation Toolkit: Combining Social Network Analysis & Qualitative Comparative Analysis, Debbie Gowensmith

Electronic Theses and Dissertations

Collective action networks are complex systems of interrelated individuals or groups that come together for a common social change purpose (Ernstson, 2011). Researchers have used social network analysis (SNA) to examine the relationship structures and characteristics of collective action networks. However, determining whether collective action networking produces outcomes has been challenging because networks are complex, affected by context, and produce interdependent data. I addressed these challenges by pairing SNA with qualitative comparative analysis (QCA), a configurational comparative method. Using QCA, researchers can tease out which conditions are necessary or sufficient to produce an outcome. I analyzed a collective action network …


Using The Fraction Of Missing Information (Fmi) In Selecting Auxiliary Variables To Impute Missingness In Confirmatory Factor Analysis (Cfa), Dareen Taha Alzahrani Jan 2022

Using The Fraction Of Missing Information (Fmi) In Selecting Auxiliary Variables To Impute Missingness In Confirmatory Factor Analysis (Cfa), Dareen Taha Alzahrani

Electronic Theses and Dissertations

This study aimed to investigate the effectiveness of using the fraction of missing information (FMI) to select auxiliary variables in imputing missing data in confirmatory factor analysis (CFA). This was done by conducting two studies (a simulation study and an empirical study). A Monte Carlo simulation technique was used to compare the performance and the effect of the restrictive strategy based on FMI and the inclusive strategy on parameter estimate bias and parameter estimate efficiency. The missing data mechanisms, missing data proportion, correlation strength between the analysis variables and auxiliary variables, and the inclusive and restrictive strategies were assessed in …


Examining The Credibility Of Story-Based Causal Methodologies, Megan E. Kauffmann Jan 2022

Examining The Credibility Of Story-Based Causal Methodologies, Megan E. Kauffmann

Electronic Theses and Dissertations

The purpose of this study was to explore how evaluators justify using story-based methodologies when examining causality. The two primary research questions of the study included: 1) what arguments are made by evaluators to justify the credibility of story-based causal methodologies to evaluation stakeholders; and 2) from the perspective of evaluators, how do contextual factors influence whether story-based causal methodologies are perceived as credible by evaluation stakeholders? A case study was conducted to examine the cases of four evaluators who had experience implementing a story-based methodology in an evaluation. Data collection procedures included two interviews with each participant and a …


Mis-Specification Of Functional Forms In Growth Mixture Modeling: A Monte Carlo Simulation, Richa Ghevarghese Jan 2022

Mis-Specification Of Functional Forms In Growth Mixture Modeling: A Monte Carlo Simulation, Richa Ghevarghese

Electronic Theses and Dissertations

Growth mixture modeling (GMM) is a methodological tool used to represent heterogeneity in longitudinal datasets through the identification of unobserved subgroups following qualitatively and quantitatively distinct trajectories in a population. These growth trajectories or functional forms are informed by the underlying developmental theory, are distinct to each subgroup, and form the core assumptions of the model. Therefore, the accuracy of the assumed functional forms of growth strongly influences substantive research and theories of growth. While there is evidence of mis-specified functional forms of growth in GMM literature, the weight of this violation has been largely overlooked. Current solutions to circumvent …


Application Of An Organizational Evaluation Capacity Assessment In A Multinational Ngo: A Case Study To Support Applied Practice, Ryan James Smyth Jan 2022

Application Of An Organizational Evaluation Capacity Assessment In A Multinational Ngo: A Case Study To Support Applied Practice, Ryan James Smyth

Electronic Theses and Dissertations

As evaluation capacity building (ECB) has rapidly emerged as a practice in human service organizations and as a field of academic inquiry, attention has focused on methods of evaluation capacity building while assessment of organizational evaluation capacity (EC) has lagged behind. To examine the practice of organizational evaluation capacity assessment, this dissertation presents two separate but related studies. In sub-study 1, I present a qualitative evidence synthesis of the research theorizing organizational evaluation capacity models. In sub-study 2, I support the implementation of one of the tools from the evidence-synthesis at a multinational human service organization. I use a concurrent …


Investigaion Of The Gamma Hurdle Model For A Single Population Mean, Alissa Jacobs Jan 2022

Investigaion Of The Gamma Hurdle Model For A Single Population Mean, Alissa Jacobs

Electronic Theses and Dissertations

A common issue in some statistical inference problems is dealing with a high frequency of zeroes in a sample of data. For many distributions such as the gamma, optimal inference procedures do not allow for zeroes to be present. In practice, however, it is natural to observe real data sets where nonnegative distributions would make sense to model but naturally zeroes will occur. One example of this is in the analysis of cost in insurance claim studies. One common approach to deal with the presence of zeroes is using a hurdle model. Most literary work on hurdle models will focus …


Using Deep Neural Networks To Analyze Precision Agriculture Data, Stephanie Liebl Jan 2022

Using Deep Neural Networks To Analyze Precision Agriculture Data, Stephanie Liebl

Electronic Theses and Dissertations

As the population of the Earth increases, there is a growing need for food to feed the inhabitants. Precision agriculture offers techniques and tools that can be used to help accommodate the growing population. One specific precision agriculture tool is remote sensing data, which can be used to image fields as an effort to better predict or understand the crops. In this thesis, deep neural networks are used to evaluate various spatial, spectral, and temporal resolutions of three different satellite images to determine which best predicts corn yield. The main metrics we used to evaluate the models were R-squared (R2), …


Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel Dec 2021

Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel

Electronic Theses and Dissertations

Statistical inference for the mean of a beta distribution has become increasingly popular in various fields of academic research. In this study, we developed a novel statistical model from likelihood-based techniques to evaluate various confidence interval techniques for the mean of a beta distribution. Simulation studies will be implemented to compare the performance of the confidence intervals. In addition to the development and study involving confidence intervals, we will also apply the confidence intervals to real biological data that was gathered by the Department of Biology at Stephen F. Austin State University and provide recommendations on the best practice.


Estimating Treatment Effect On Medical Cost And Examining Medical Cost Trajectory Using Splines And Change Point Techniques., Indranil Ghosh Dec 2021

Estimating Treatment Effect On Medical Cost And Examining Medical Cost Trajectory Using Splines And Change Point Techniques., Indranil Ghosh

Electronic Theses and Dissertations

In the world of growing medical needs, other than the clinical outcomes, the cost of healthcare is one of the important aspects to evaluate. The cost of treatment could act as a decisive factor on which one to choose from two equally likely effective treatment options. In literature, the most used quantity for the cost of treatment is cumulative lifetime cost since the diagnosis of a disease. While it provides a bird' eye view of the treatment cost, it fails to capture the underlying pattern of the treatment cost trajectory. We developed a marginal structural functional model (MSFM) using an …


Functional Mixed Data Clustering With Fourier Basis Smoothing, Ishmael Amartey Dec 2021

Functional Mixed Data Clustering With Fourier Basis Smoothing, Ishmael Amartey

Electronic Theses and Dissertations

Clustering is an important analytical technique that has proven to affect human life positively through its application in cancer research, market segmentation, city planning etc. In this time of growing technological systems, mixed data has seen another face of longitudinal, directional and functional attributes which is worth paying attention to and analyzing. Previous research works on clustering relied largely on the inverse weight technique and B-spline in smoothing data and assessing the performance of various clustering algorithms. In 1971, Gower proposed a method of clustering for mixed variable types which has been extended to include functional and directional variables by …


Prediction Intervals: The Effects And Identification Of Sparse Regions For Nonparametric Regression Methods, Jackson Faires Aug 2021

Prediction Intervals: The Effects And Identification Of Sparse Regions For Nonparametric Regression Methods, Jackson Faires

Electronic Theses and Dissertations

In this work, we provide an overview of different nonparametric methods for prediction interval estimation and investigate how well they perform when making predictions in sparse regions of the predictor space. This sparsity is an extension to the more common concept of extrapolation in linear regression settings. Using simulation studies, we show that coverage probabilities using prediction intervals from quantile k-nearest neighbors and quantile random forest can be biased to low or too high from the nominal level under various situations of sparsity. We also introduce a test that can be used to see if a new data point lies …


Predictive Modeling Of Clinical Outcomes For Hospitalized Covid-19 Patients Utilizing Cytof And Clinical Data., Onajia Stubblefield Aug 2021

Predictive Modeling Of Clinical Outcomes For Hospitalized Covid-19 Patients Utilizing Cytof And Clinical Data., Onajia Stubblefield

Electronic Theses and Dissertations

In December 2019, an outbreak of a novel coronavirus initiated a global pandemic. Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is a virus that causes the disease coronavirus disease 2019 (COVID-19). Symptoms of infection with COVID-19 vary widely between individuals. While some infected individuals are asymptomatic, others need more extensive care and require hospitalization. Indeed, the COVID-19 pandemic was characterized by a shortage of hospital beds which presented additional complications in providing adequate care for patients. In this study, we used a combination of T cell population data collected from mass cytometry analysis and clinical markers to form a predictive …


Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin Aug 2021

Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin

Electronic Theses and Dissertations

In this work, we seek to develop a variable screening and selection method for Bayesian mixture models with longitudinal data. To develop this method, we consider data from the Health and Retirement Survey (HRS) conducted by University of Michigan. Considering yearly out-of-pocket expenditures as the longitudinal response variable, we consider a Bayesian mixture model with $K$ components. The data consist of a large collection of demographic, financial, and health-related baseline characteristics, and we wish to find a subset of these that impact cluster membership. An initial mixture model without any cluster-level predictors is fit to the data through an MCMC …


Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari Aug 2021

Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari

Electronic Theses and Dissertations

One of the concerns in the field of statistics is the presence of missing data, which leads to bias in parameter estimation and inaccurate results. However, the multiple imputation procedure is a remedy for handling missing data. This study looked at the best multiple imputation methods used to handle mixed variable datasets with different sample sizes and variability along with different levels of missingness. The study employed the predictive mean matching, classification and regression trees, and the random forest imputation methods. For each dataset, the multiple regression parameter estimates for the complete datasets were compared to the multiple regression parameter …


Applying Deep Learning To The Ice Cream Vendor Problem: An Extension Of The Newsvendor Problem, Gaffar Solihu Aug 2021

Applying Deep Learning To The Ice Cream Vendor Problem: An Extension Of The Newsvendor Problem, Gaffar Solihu

Electronic Theses and Dissertations

The Newsvendor problem is a classical supply chain problem used to develop strategies for inventory optimization. The goal of the newsvendor problem is to predict the optimal order quantity of a product to meet an uncertain demand in the future, given that the demand distribution itself is known. The Ice Cream Vendor Problem extends the classical newsvendor problem to an uncertain demand with unknown distribution, albeit a distribution that is known to depend on exogenous features. The goal is thus to estimate the order quantity that minimizes the total cost when demand does not follow any known statistical distribution. The …