Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (23)
- Biostatistics (15)
- Statistical Methodology (12)
- Statistical Theory (10)
- Mathematics (8)
-
- Social and Behavioral Sciences (8)
- Multivariate Analysis (7)
- Other Statistics and Probability (7)
- Applied Mathematics (6)
- Data Science (6)
- Longitudinal Data Analysis and Time Series (6)
- Probability (6)
- Education (5)
- Computer Sciences (4)
- Educational Assessment, Evaluation, and Research (4)
- Medicine and Health Sciences (4)
- Survival Analysis (4)
- Categorical Data Analysis (3)
- Life Sciences (3)
- Microarrays (3)
- Other Mathematics (3)
- Psychology (3)
- Algebra (2)
- Biomedical (2)
- Business (2)
- Earth Sciences (2)
- Electrical and Computer Engineering (2)
- Institution
- Keyword
-
- Morgridge College of Education (8)
- Research Methods and Information Science (8)
- Research Methods and Statistics (8)
- Regression (3)
- Breast cancer (2)
-
- Confirmatory factor analysis (2)
- Dimension Reduction (2)
- Features Extraction (2)
- Item response theory (2)
- Machine learning (2)
- Parameter estimation (2)
- Propensity scores (2)
- Statistics (2)
- Takens Vectors (2)
- Actuarial (1)
- Ant colony optimization (1)
- Antibiotic overuse (1)
- Applied statistics (1)
- Artificial intelligence (1)
- Asset risk (1)
- Asymptotic Confidence Intervals (1)
- Asymptotics (1)
- Bayesian (1)
- Bayesian analysis (1)
- Bayesian inference (1)
- Bayesian inferences (1)
- Bayesian regularization (1)
- Bayesian shrinkage priors (1)
- Beta Distribution (1)
- Bootstrap resampling. (1)
Articles 1 - 30 of 48
Full-Text Articles in Statistical Models
Modeling Mean And Variability Of Anxiety In Ecological Momentary Assessment Data Using Mixed-Effects Location–Scale Models, Trenzy Odero
Modeling Mean And Variability Of Anxiety In Ecological Momentary Assessment Data Using Mixed-Effects Location–Scale Models, Trenzy Odero
Electronic Theses and Dissertations
Ecological Momentary Assessment is a method of collecting repeated measures of people in real time within natural environments. This results in hierarchical data that has a significant amount of variation at the person level. The traditional linear mixedeffects models assume that the residual variance is constant, which might not be true when the residual variance varies among individuals as well as in time. This thesis uses mixed-effects location-scale (MELS) models to model the mean and variance of an EMA outcome together. By introducing the possibility of variability in residual variance within and across individuals and with covariates, the MELS framework …
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
Electronic Theses and Dissertations
This thesis develops a discrete stochastic linear systems interpretation of age–stage demographic evolution grounded in Leslie operators and realized in a discrete-event simulation implemented with salabim. The central claim is that one annual cycle of the simulation constitutes a cone-preserving, stochastic affine transformation on a high- dimensional population state vector indexed by age, sex, marital status, household type, employment, and education, and that the composition of yearly operators yields a random matrix product whose top Lyapunov exponent is the stochastic counterpart of the Perron–Frobenius growth rate (Caswell, 2001; Tuljapurkar, 1997)[1, 2]. The actuarial bridge is constructed by mapping simulated survival …
Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman
Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman
Electronic Theses and Dissertations
Compositional data analysis (CoDA) addresses multivariate data constrained to a constant sum, such as proportions or percentages. Originating from early warnings regarding misinterpretation by Pearson (1897), the field was formalized by John Aitchison in 1986, whose foundational work remains highly influential. Over time, new modeling techniques and visualization tools have advanced the field, as noted by Greenacre et al. More recently, Turner et al. proposed an approach based on the Nested Dirichlet Distribution (NDD), which accommodates more flexible dependence structures than the standard Dirichlet model. This thesis builds on the methodology of Turner et al. Chapter 1 introduces the nature …
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Electronic Theses and Dissertations
Networks are powerful tools for modeling the complexity of social interactions, biological systems, and information spread. A leading statistical frameworks for analyzing network data are Exponential Random Graph Models (ERGMs), which provide a principled approach to capturing structural dependencies. However, ERGMs remain challenging to estimate, especially in sparse or high-dimensional settings where models suffer from degeneracy and unstable parameter inference. This paper proposes a penalized Bayesian approach to ERGMs that utilizes the horseshoe prior, a sparsity-inducing global-local shrinkage prior. This prior offers robust regularization while preserving important signals, improving estimation by shrinking irrelevant parameters and reducing the impact of extreme …
Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson
Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson
Electronic Theses and Dissertations
Most work surrounding changepoint analysis focuses on linear models. This dissertation explores changepoint detection in nonlinear power function models, specifically focusing on models where the constant multiplier and power are the parameters to be estimated in addition to the changepoint parameter. The study assumes an asymptotic framework as the number of observations approaches infinity. The study explores various model fitting algorithms, and decides to employ the Newton-Raphson method for parameter estimation, with a custom implementation developed to optimize the process. The research first establishes the strong consistency of estimators for the model without a changepoint. Building on this result, consistency …
Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han
Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han
Electronic Theses and Dissertations
This dissertation consists of two projects. The first one involves nonparametric methods on Continuous Time Markov Chains (CTMCs). The second one is centered around Bayesian shrinkage models for detecting prognostic and predictive biomarkers in high-dimensional clinical data. Both these projects build on methods from across the frequentist and Bayesian paradigm to offer novel solutions. In the first project, we aim to model the nonlinear effects of continuous variables within multistate framework in a non-parametrically by appealing to the rich mathematical framework of Reproducing Kernel Hilbert Spaces (RKHS). Then we adapted the classical Representer Theorem to penalized (squared norm) log-likelihood which …
Exploring The Diagnostic Potential Of Radiomics-Based Pet Image Analysis For T-Stage Tumor Diagnosis, Victor Aderanti
Exploring The Diagnostic Potential Of Radiomics-Based Pet Image Analysis For T-Stage Tumor Diagnosis, Victor Aderanti
Electronic Theses and Dissertations
Cancer is a leading cause of death globally, and early detection is crucial for better
outcomes. This research aims to improve Region Of Interest (ROI) segmentation
and feature extraction in medical image analysis using Radiomics techniques
with 3D Slicer, Pyradiomics, and Python. Dimension reduction methods, including
PCA, K-means, t-SNE, ISOMAP, and Hierarchical Clustering, were applied to highdimensional features to enhance interpretability and efficiency. The study assessed the ability of the reduced feature set to predict T-staging, an essential component of the TNM system for cancer diagnosis. Multinomial logistic regression models were developed and evaluated using MSE, AIC, BIC, and Deviance …
Capturing Latent Abilities And Latent Capacities Of Professional Golfers Using Nonlinear Mixed Effects Growth Modeling, Mac Wetherbee
Capturing Latent Abilities And Latent Capacities Of Professional Golfers Using Nonlinear Mixed Effects Growth Modeling, Mac Wetherbee
Electronic Theses and Dissertations
This study demonstrates an effective and innovative approach to measuring the latent athletic abilities and capacities of professional golfers. I used nonlinear mixed effects growth modeling (e.g., Dynamic Measurement Modeling) to measure professional golfers’ ability levels and capacities for improvement. I accomplished this using a two-stage modeling approach. First, a crossed linear mixed effects model estimated each player’s ability level in each year. In the second stage, I used the results from the first stage to estimate several candidate nonlinear growth trajectories for players’ abilities over time. The quadratic growth trajectory was the best-fitting of these trajectories and was used …
Extending The Utility Of Ant Colony Optimization Through The Incorporation Of An Intraclass Correlation Coefficient To Assess For Rater Consistency, Mark Leveling
Electronic Theses and Dissertations
Ant Colony Optimization (ACO) is a flexible algorithm designed to solve complex combinatorial problems. While the method was derived from the behavior of ants by researchers in the field of computer science, its application to solving complex combinatorial problems is widespread in a growing number of fields in behavioral science, including psychometrics. Over the last two decades, psychometricians have adapted ACO to measurement model specification problems with the intention of generating measurement models that express measurement model fit and reliability within the standards of what is considered acceptable. Additionally, psychometricians have used ACO to generate shortened versions of existing measures …
Advancement Of Iterative Optimization Technology Algorithms Toward Calibration-Free Process Analytical Technology Applications, Adam Rish
Electronic Theses and Dissertations
The expansion of spectroscopic process analytical technology (PAT) tools within the pharmaceutical industry has the potential to elevate the current state-of-the-art of pharmaceutical manufacturing by offering opportunities for reduced quality testing times, enhanced process control, and greater production flexibility. Spectroscopic PAT tools are dependent on multivariate models to extract the relevant information from the spectral outputs. However, there is a substantial calibration burden for developing and maintaining these multivariate models that discourages the application of PAT, despite the encouragement from regulators. This has led to an interest in calibration-free methods such as iterative optimization technology (IOT) for spectroscopic PAT that …
Landslide Susceptibility And Tree Ring Eccentricity Analysis Along Unstable Slopes Of The New River Watershed, Anderson And Morgan Counties, Tn, Megan Palmer
Electronic Theses and Dissertations
Landslides are mass movements that affect infrastructure across East Tennessee, causing problems for the Tennessee Department of Transportation (TDOT). An assessment of conditions and locations of unstable slopes can aid TDOT in infrastructure management. Landslide susceptibility was evaluated for Anderson and Morgan counties, TN, off State Route 116 in the New River watershed. Susceptibility maps used a landslide inventory and six factors: elevation, slope, geology, distance from stream, rainfall, and curvature, input in forest-based classification and logistic regression models. Additionally, affected trees along these unstable slopes in Anderson and Morgan counties were cored to analyze mass movement impacts on tree …
Assessment Of Method Effects Of Keying And Wording In Instruments: A Mixed-Methods Explanatory Sequential Study, Lin Ma
Electronic Theses and Dissertations
This dissertation presents an innovative approach to examining the keying method, wording method, and construct validity on psychometric instruments. By employing a mixed methods explanatory sequential design, the effects of keying and wording in two psychometric assessments were examined and validated. Those two self-report psychometric assessments were the Effortful Control assessment (Ellis & Rothbart, 2001) and the Grit assessment (Duckworth & Quinn, 2009). Moreover, the quantitative phase utilized structural equation modeling to analyze 2,104 students’ responses and assess the construct of keying and wording. Various hypothetical models were investigated and evaluated. The reliability of each construct in each method was …
Bayesian Strategies For Propensity Score Estimation In Causal Inference., Uthpala I. Wanigasekara
Bayesian Strategies For Propensity Score Estimation In Causal Inference., Uthpala I. Wanigasekara
Electronic Theses and Dissertations
Causal inference is a method used in various fields to draw causal conclusions based on data. It involves using assumptions, study designs, and estimation strategies to minimize the impact of confounding variables. Propensity scores are used to estimate outcome effects, through matching methods, stratification, weighting methods, and the Covariate Balancing Propensity Score method. However, they can be sensitive to estimation techniques and can lead to unstable findings. Researchers have proposed integrating weighing with regression adjustment in parametric models to improve causal inference validity. The first project focuses on Bayesian joint and two-stage methods for propensity score analysis. Propensity score modeling …
Proteomics And Machine Learning For Pulmonary Embolism Risk With Protein Markers, Yaa Amankwah Awuah
Proteomics And Machine Learning For Pulmonary Embolism Risk With Protein Markers, Yaa Amankwah Awuah
Electronic Theses and Dissertations
This thesis investigates protein markers linked to pulmonary embolism risk using proteomics and statistical methods, employing unsupervised and supervised machine learning techniques. The research analyzes existing datasets, identifies significant features, and observes gender differences through MANOVA. Principal Component Analysis reduces variables from 378 to 59, and Random Forest achieves 70% accuracy. These findings contribute to our understanding of pulmonary embolism and may lead to diagnostic biomarkers. MANOVA reveals significant gender differences, and applying proteomics holds promise for clinical practice and research.
The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña
The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña
Electronic Theses and Dissertations
Since the late 1970s, multiple linear regression has been the preferred method for identifying discrimination in pay. An empirical study on this topic was conducted using quantitative critical methods. A literature review first examined conflicting views on using multiple linear regression in pay equity studies. The review found that multiple linear regression is used so prevalently in pay equity studies because the courts and practitioners have widely accepted it and because of its simplicity and ability to parse multiple sources of variance simultaneously. Commentaries in the literature cautioned about errors in model specification, the use of tainted variables, and the …
Indirect Aggression And Victimization: Investigating Instrument Psychometrics, Gender Differences, And Its Relationship To Social Information Processing, Taylor Steeves
Electronic Theses and Dissertations
The study of indirect bullying behaviors, relational aggression and social aggression, has been of theoretical importance and interest to researchers and psychologists within the last few decades. In this investigation, using a convenience sample of 451 late adolescents attending a private university in the mid-Atlantic U.S., I examined the factor structure of two measures of indirect bullying, the Young Adult Social Behavior Scale – Victim (YASB-V) and the Young Adult Social Behavior Scale – Perpetrator (YASB-P). Using confirmatory factor analysis (CFA), I found that the YASB-V comprised a four-factor model, differing from the model that had been identified in the …
Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman
Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman
Electronic Theses and Dissertations
This dissertation consists of three research projects on cancer screening probability modeling. In these projects, the three key modeling parameters (sensitivity, sojourn time, transition density) for cancer screening were estimated, along with the long-term outcomes (including overdiagnosis as one outcome), the optimal screening time/age, the lead time distribution, and the probability of overdiagnosis at the future screening time were simulated to provide a statistical perspective on the effectiveness of cancer screening programs. In the first part of this dissertation, a statistical inference was conducted for male and female smokers using the National Lung Screening Trial (NLST) chest X-ray data. A …
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Electronic Theses and Dissertations
Graphical models determine associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models, where the relationships are formalized by non-null entries of the precision matrix. However, in high-dimensional cases, covariance estimates are typically unstable. Moreover, it is natural to expect only a few significant associations to be present in many realistic applications. This necessitates the injection of sparsity techniques into the estimation method. Classical frequentist methods, like GLASSO, use penalization techniques for this purpose. Fully Bayesian methods, on the contrary, are slow because they require iteratively sampling over a quadratic …
Computer Aided Diagnosis System For Breast Cancer Using Deep Learning., Asma Baccouche
Computer Aided Diagnosis System For Breast Cancer Using Deep Learning., Asma Baccouche
Electronic Theses and Dissertations
The recent rise of big data technology surrounding the electronic systems and developed toolkits gave birth to new promises for Artificial Intelligence (AI). With the continuous use of data-centric systems and machines in our lives, such as social media, surveys, emails, reports, etc., there is no doubt that data has gained the center of attention by scientists and motivated them to provide more decision-making and operational support systems across multiple domains. With the recent breakthroughs in artificial intelligence, the use of machine learning and deep learning models have achieved remarkable advances in computer vision, ecommerce, cybersecurity, and healthcare. Particularly, numerous …
Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage
Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage
Electronic Theses and Dissertations
Due to the wide availability of functional data from multiple disciplines, the studies of functional data analysis have become popular in the recent literature. However, the related development in censored survival data has been relatively sparse. In Chapter 2, we consider the problem of analyzing time-to-event data in the presence of functional predictors. We develop a conditional generalized Kaplan Meier (KM) estimator that incorporates functional predictors using kernel weights and rigorously establishes its asymptotic properties. In addition, we propose to select the optimal bandwidth based on a time-dependent Brier score. We then carry out extensive numerical studies to examine the …
Finding A Representative Distribution For The Tail Index Alpha, Α, For Stock Return Data From The New York Stock Exchange, Jett Burns
Electronic Theses and Dissertations
Statistical inference is a tool for creating models that can accurately display real-world events. Special importance is given to the financial methods that model risk and large price movements. A parameter that describes tail heaviness, and risk overall, is α. This research finds a representative distribution that models α. The absolute value of standardized stock returns from the Center for Research on Security Prices are used in this research. The inference is performed using R. Approximations for α are found using the ptsuite package. The GAMLSS package employs maximum likelihood estimation to estimate distribution parameters using the CRSP data. The …
Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel
Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel
Electronic Theses and Dissertations
Statistical inference for the mean of a beta distribution has become increasingly popular in various fields of academic research. In this study, we developed a novel statistical model from likelihood-based techniques to evaluate various confidence interval techniques for the mean of a beta distribution. Simulation studies will be implemented to compare the performance of the confidence intervals. In addition to the development and study involving confidence intervals, we will also apply the confidence intervals to real biological data that was gathered by the Department of Biology at Stephen F. Austin State University and provide recommendations on the best practice.
Predictive Modeling Of Clinical Outcomes For Hospitalized Covid-19 Patients Utilizing Cytof And Clinical Data., Onajia Stubblefield
Predictive Modeling Of Clinical Outcomes For Hospitalized Covid-19 Patients Utilizing Cytof And Clinical Data., Onajia Stubblefield
Electronic Theses and Dissertations
In December 2019, an outbreak of a novel coronavirus initiated a global pandemic. Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is a virus that causes the disease coronavirus disease 2019 (COVID-19). Symptoms of infection with COVID-19 vary widely between individuals. While some infected individuals are asymptomatic, others need more extensive care and require hospitalization. Indeed, the COVID-19 pandemic was characterized by a shortage of hospital beds which presented additional complications in providing adequate care for patients. In this study, we used a combination of T cell population data collected from mass cytometry analysis and clinical markers to form a predictive …
Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin
Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin
Electronic Theses and Dissertations
In this work, we seek to develop a variable screening and selection method for Bayesian mixture models with longitudinal data. To develop this method, we consider data from the Health and Retirement Survey (HRS) conducted by University of Michigan. Considering yearly out-of-pocket expenditures as the longitudinal response variable, we consider a Bayesian mixture model with $K$ components. The data consist of a large collection of demographic, financial, and health-related baseline characteristics, and we wish to find a subset of these that impact cluster membership. An initial mixture model without any cluster-level predictors is fit to the data through an MCMC …
Assessing The Variations Of Educational Attainment At National And Subnational Levels Using Hierarchical Linear Models, Bingxin Qi
Electronic Theses and Dissertations
Education is a human right, and equal access to education is not only crucial for an individual’s well-being, but also essential for eradicating poverty, ensuring long-term prosperity for all, transforming the society, and achieving sustainable development. Measuring education development, especially the variations of educational attainment, in a timely and accurate manner can help educators, practitioners, scientists, and policymakers compare and evaluate various education indicators at both subnational and national levels. This research presents an approach that combines multi-source and multidimensional data including population distribution, human settlement, and education data to assess and explore educational attainment trajectories at both national and …
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Electronic Theses and Dissertations
Recently, gene set analysis has become the first choice for gaining insights into the underlying complex biology of diseases through high-throughput genomic studies, such as Microarrays, bulk RNA-Sequencing, single cell RNA-Sequencing, etc. It also reduces the complexity of statistical analysis and enhances the explanatory power of the obtained results. Further, the statistical structure and steps common to these approaches have not yet been comprehensively discussed, which limits their utility. Hence, a comprehensive overview of the available gene set analysis approaches used for different high-throughput genomic studies is provided. The analysis of gene sets is usually carried out based on …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Measuring The Connective Action Of Black Lives Matter Activists: A Psychometric Investigation Into Twitter Data, Paige Alfonzo
Measuring The Connective Action Of Black Lives Matter Activists: A Psychometric Investigation Into Twitter Data, Paige Alfonzo
Electronic Theses and Dissertations
Many protest movements from the last twenty-first century have become increasingly networked and personalized. Several scholars have tapped into this change coining terms such as participatory action, digitally mediated action, computer-mediated communication, issue-based organization, and what I focus on in this project, connective action. Building on the ideas percolating across the literary landscape at the time, Bennett and Segerberg (2012) introduced the logic of connective action based on emergent characteristics they observed in post-2010 large-scale social movements. Both the logic of connective action and related work have become deeply ingrained in today's social movement scholarship. As such, I felt it …
Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang
Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang
Electronic Theses and Dissertations
Measurement invariance is crucial for an effective and valid measure of a construct. Invariance holds when the latent trait varies consistently across subgroups; in other words, the mean differences among subgroups are only due to true latent ability differences. Differential item functioning (DIF) occurs when measurement invariance is violated. There are two kinds of traditional tools for DIF detection: non-parametric methods and parametric methods. Mantel Haenszel (MH), SIBTEST, and standardization are examples of non-parametric DIF detection methods. The majority of parametric DIF detection methods are item response theory (IRT) based. Both non-parametric methods and parametric methods compare differences among subgroups …
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Electronic Theses and Dissertations
Premature birth has been identified as the single greatest cause of death worldwide in children under the age of five. This thesis will implement binary logistic regression and proportional odds ordinal logistic regression to predict different levels of premature birth and identify associated risk factors. The models will be built from the Center for Disease Control and Prevention's 2014 Vital Statistics Natality Birth Data containing nearly 4 million live births within the United States. Odds ratios and confidence intervals on risk factors were produced utilizing binary logistic regression.