Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Logistic regression

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 69

Full-Text Articles in Statistics and Probability

A Variance Decomposition Approach To Inconclusives In Forensic Black Box Studies, Amanda Luby, Joseph B. Kadane Jan 2026

A Variance Decomposition Approach To Inconclusives In Forensic Black Box Studies, Amanda Luby, Joseph B. Kadane

Mathematics and Statistics Faculty Work

In the USA, ‘black box’ studies are increasingly being used to estimate the error rate of forensic disciplines. A sample of forensic examiner participants is asked to evaluate a set of items whose source is known to the researchers but not to the participants. Participants are asked to make a source determination (typically an identification, exclusion, or some kind of inconclusive). We study inconclusives in two black box studies, one on fingerprints and one on bullets. Rather than treating all inconclusive responses as functionally correct (as is the practice in reported error rates in the two studies we address), irrelevant …


Identifying The Factors Affecting The Survival Of Trauma Patients Using Logistic Regression Analysis, Maggie Smith Apr 2025

Identifying The Factors Affecting The Survival Of Trauma Patients Using Logistic Regression Analysis, Maggie Smith

Honors College Theses

There is a broad interest among researchers and clinicians in identifying factors affecting clinical outcomes of patients with physical trauma. Numerous factors affect Hospital Discharge Status (HDS), one of the main binary outcome variables of trauma patients. Logistic regression is one of the widely used methods to analyze relationships between a set of predictors with a binary outcome. In this study, a logistic regression model is built for HDS. Predictors include arrival time, age, trauma level, injury severity score, arrival heart rate, arrival blood pressure, length of hospital stay, time from injury to arrival at Billings Clinic (BC), patient transfer …


Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma Mar 2025

Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma

Research Symposium

Background: With a rapid development of data collection technology, high dimensional data, whose model dimension k may be growing or much larger than the sample size n, is becoming increasingly prevalent in different fields of study, such as ecology, genetics, among others. This data deluge is introducing new challenges to traditional statistical procedures and theories and is thus generating a renewed interest in the problems of variable selection and classification in high dimensional regression models. In large k, small n settings, variable selection is usually the first step for dimension reduction to uncover significant covariates, which contribute to …


Utilizing Machine Learning Techniques For Accurate Diagnosis Of Breast Cancer And Comprehensive Statistical Analysis Of Clinical Data, Myat Ei Ei Phyo Mar 2024

Utilizing Machine Learning Techniques For Accurate Diagnosis Of Breast Cancer And Comprehensive Statistical Analysis Of Clinical Data, Myat Ei Ei Phyo

USF Tampa Graduate Theses and Dissertations

Breast cancer represents a formidable malignancy, presenting a substantial threat to global health and individual well-being. Conventionally, it is widely held that the prognosis for breast cancer patients hinges predominantly upon the timing of diagnosis and the extent of cancer progression, typically delineated by its stage. However, emerging evidence from robust regression and machine learning analyses challenges this prevailing notion. The results indicate that survival months cannot be solely attributed to diagnosis and socio-economic factors. Instead, additional variables such as existing diseases and treatment complexities may contribute to the intricate landscape of breast cancer outcomes.

This research aims to delve …


Effects Of Maternal Anthropometry On Infant Anthropometry: A Cross-Sectional Study At Public Hospital X In Ternate, Indonesia, Yuni Nurwati, Hardinsyah Hardinsyah, Sri Anna Marliyati, Budi Iman Santoso, Dewi Anggraini Feb 2024

Effects Of Maternal Anthropometry On Infant Anthropometry: A Cross-Sectional Study At Public Hospital X In Ternate, Indonesia, Yuni Nurwati, Hardinsyah Hardinsyah, Sri Anna Marliyati, Budi Iman Santoso, Dewi Anggraini

Kesmas

Infant anthropometry is an indicator of neonatal survival. This study aimed to determine the effects of maternal anthropometry on estimating infant anthropom­etry. This cross-sectional study on 173 pregnant women at Public Hospital X in Ternate, Indonesia, was conducted from August 2018 to March 2023. The el­igible criteria were pregnant women aged ≥18 years, single pregnancy, and antenatal care (ANC) visits to the same hospital. The variables used included ma­ternal anthropometric measurements (body weight, body height, third-trimester weight (TTW)), gestational weight gain (GWG), education, age, ANC visits, and gestational age at delivery (GAD). A logistic regression model was employed to estimate …


Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang Feb 2024

Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang

School of Mathematical & Statistical Sciences Faculty Publications

In this paper, we propose a sparse Bayesian procedure with global and local(GL) shrinkage priors for the problems of variable selection and classification in high-dimensional logistic regression models. In particular, we consider two types of GL shrinkage priors for the regression coefficients, the horseshoe (HS)prior and the normal-gamma (NG) prior, and then specify a correlated prior for the binary vector to distinguish models with the same size. The GL priors are then combined with mixture representations of logistic distribution to construct a hierarchical Bayes model that allows efficient implementation of a Markov chain Monte Carlo (MCMC) to generate samples from …


Classification In Supervised Statistical Learning With The New Weighted Newton-Raphson Method, Toma Debnath Jan 2024

Classification In Supervised Statistical Learning With The New Weighted Newton-Raphson Method, Toma Debnath

College of Graduate Studies: Theses & Dissertations

In this thesis, the Weighted Newton-Raphson Method (WNRM), an innovative optimization technique, is introduced in statistical supervised learning for categorization and applied to a diabetes predictive model, to find maximum likelihood estimates. The iterative optimization method solves nonlinear systems of equations with singular Jacobian matrices and is a modification of the ordinary Newton-Raphson algorithm. The quadratic convergence of the WNRM, and high efficiency for optimizing nonlinear likelihood functions, whenever singularity in the Jacobians occur allow for an easy inclusion to classical categorization and generalized linear models such as the Logistic Regression model in supervised learning. The WNRM is thoroughly investigated …


Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre Dec 2023

Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre

SMU Data Science Review

Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …


Approaches To Detecting And Modeling Over-And Underdispersion In Alternative Count Data Distributions And An Application Of Logistic Regression And Random Forest Modeling To Improve Screening Tools For Tic Disorders In Children, Rebecca C. Wardrop Jul 2023

Approaches To Detecting And Modeling Over-And Underdispersion In Alternative Count Data Distributions And An Application Of Logistic Regression And Random Forest Modeling To Improve Screening Tools For Tic Disorders In Children, Rebecca C. Wardrop

Theses and Dissertations

This dissertation focuses on theory and application of discrete data methods, particularly approaches to over- and underdispersion relative to the Poisson distribution and an application of random forest and logistic regression modeling. The first chapter derives a score test for over- and underdispersion in the heaped generalized Poisson distribution. Equi-, over-, and underdispersed heaped generalized Poisson and heaped negative binomial data are simulated to evaluate the performance of the score test by comparing the power it achieves to that of Wald and likelihood ratio tests. We find that the score test we derive performs comparably to both the Wald and …


Does The Three Point Shot Affect Winning Percentage, Marcamus Winn Apr 2023

Does The Three Point Shot Affect Winning Percentage, Marcamus Winn

Mathematics Senior Capstone Papers

The three-point shot, introduced in the late 1970s, is a shot that occurs typically 24 feet away from the basket at the professional level. Strategically the game of basketball was originally based on two-point field goals. Recently, there has been a noticeable trend in the popularity of the three-point shot amongst professional teams. Nowadays, three point shot attempts account for more than a third of average NBA shot selection. Statistical analysis is becoming integral to athletics. Statistics has become a critical component to the development of not only on court basketball strategies, but also team structure as well. There are …


A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani Dec 2022

A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani

Electronic Theses and Dissertations

Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.

Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance …


Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair Jan 2022

Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair

Graduate Student Theses, Dissertations, & Professional Papers

Landslides are a globally pervasive problem with the potential to cause significant fatalities and economic losses. Although landslides are widespread, many at-risk regions may not have the high-quality data or resources used in most landslide susceptibility analyses. This study aims to develop regional susceptibility relationships that are versatile and use publicly available data and open-sourced software. Logistic Regression and Frequency Ratio susceptibility relationships were developed in 23 regions in Washington, Utah, North Carolina, and Kentucky, with a region referring to a unique area and data combination. Regions were diverse in their geology, morphology, climate, and nature and quality of their …


Smoking, Alcohol Consumption, And Depression In Association With Incidence Of Type 2 Diabetes Among Mexican Americans In Starr County, Texas, Gabriela Rubannelsonkumar Dec 2021

Smoking, Alcohol Consumption, And Depression In Association With Incidence Of Type 2 Diabetes Among Mexican Americans In Starr County, Texas, Gabriela Rubannelsonkumar

Honors Program Theses and Research Projects

Previous studies on conditions like obesity, hypertension, and type 2 diabetes mellitus (T2DM) have explored the correlations between them and various other human conditions, including aortic stiffness, left ventricular hypertrophy and sleep apnea, as they predict possibilities of developing certain diseases in Mexican Americans. This study aims to observe the correlation between lifestyle decisions that could relate to the onset of the depression in normal, prediabetic, and diabetic individuals. These include smoking habits and alcohol consumption. Many papers have previously conducted research on these lifestyle habits as they relate to obesity, hypertension, diabetes, however, have done so in a singular …


Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao Jul 2021

Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao

Graduate Theses and Dissertations

Nowadays industries are collecting a massive and exponentially growing amount of data that can be utilized to extract useful insights for improving various aspects of our life. Data analytics (e.g., via the use of machine learning) has been extensively applied to make important decisions in various real world applications. However, it is challenging for resource-limited clients to analyze their data in an efficient way when its scale is large. Additionally, the data resources are increasingly distributed among different owners. Nonetheless, users' data may contain private information that needs to be protected.

Cloud computing has become more and more popular in …


An Examination Of Civilian Retention In The United States Air Force, William F. Wilson Mar 2021

An Examination Of Civilian Retention In The United States Air Force, William F. Wilson

Theses and Dissertations

The backbone of the United States Air Force is undoubtedly the large civilian workforce that supplements the great work that is accomplished. Many research studies have been conducted on officer and enlisted personnel to ensure that the career fields are properly developed and managed to meet the ever growing demands of the military's varied missions, but no recent studies have focused on the civilian workforce. Striking a balance between new and experienced employees is paramount to success given the ever-changing economic and political landscapes where we find ourselves. The first part of the research uses logistic regression to determine the …


Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee Jan 2021

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee

Senior Independent Study Theses

According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.

We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …


Logistic Regression Under Sparse Data Conditions, David A. Walker, Thomas J. Smith Sep 2020

Logistic Regression Under Sparse Data Conditions, David A. Walker, Thomas J. Smith

Journal of Modern Applied Statistical Methods

The impact of sparse data conditions was examined among one or more predictor variables in logistic regression and assessed the effectiveness of the Firth (1993) procedure in reducing potential parameter estimation bias. Results indicated sparseness in binary predictors introduces bias that is substantial with small sample sizes, and the Firth procedure can effectively correct this bias.


Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox Jun 2020

Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox

Journal of Modern Applied Statistical Methods

For a binary random variable Y, let p(x) = P(Y = 1 | X = x) for some covariate X. The goal of computing a confidence interval for p(x) is considered. In the logistic regression model, even a slight departure difficult to detect via a goodness-of-fit test can yield inaccurate results. The accuracy of a confidence interval can deteriorate as the sample size increases. The goal is to suggest an alternative approach based on a smoother, which provides a more flexible approximation of p(x).


Investigating The Performance Of Propensity Score Approaches For Differential Item Functioning Analysis, Yan Liu, Chanmin Kim, Amrey D. Wu, Paul Gustafson, Edward Kroc, Bruno D. Zumbo Apr 2020

Investigating The Performance Of Propensity Score Approaches For Differential Item Functioning Analysis, Yan Liu, Chanmin Kim, Amrey D. Wu, Paul Gustafson, Edward Kroc, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

To evaluate the performance of propensity score approaches for differential item functioning analysis, this simulation study was conducted to assess bias, mean square error, Type I error, and power under different levels of effect size and a variety of model misspecification conditions, including different types and missing patterns of covariates.


An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone Jan 2020

An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone

Published and Grey Literature from PhD Candidates

Data mining techniques have numerous applications in bankcard response modeling. Logistic regression has been used as the standard modeling tool in the financial industry because of its almost always desirable performance and its interpretability. In this paper, we propose a hybrid bankcard response model, which integrates decision tree-based chi-square automatic interaction detection (CHAID) into logistic regression. In the first stage of the hybrid model, CHAID analysis is used to detect the possible potential variable interactions. Then in the second stage, these potential interactions are served as the additional input variables in logistic regression. The motivation of the proposed hybrid model …


Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu Jan 2020

Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu

College of Graduate Studies: Theses & Dissertations

The misclassification simulation extrapolation (MC-SIMEX) method proposed by Küchenho et al. is a general method of handling categorical data with measurement error. It consists of two steps, the simulation and extrapolation steps. In the simulation step, it simulates observations with varying degrees of measurement error. Then parameter estimators for varying degrees of measurement error are obtained based on these observations. In the extrapolation step, it uses a parametric extrapolation function to obtain the parameter estimators for data with no measurement error. However, as shown in many studies, the parameter estimators are still biased as a result of the parametric extrapolation …


Longitudinal Analysis With Modes Of Operation For Aes, Dana Geislinger, Cory Thigpen, Daniel W. Engels Aug 2019

Longitudinal Analysis With Modes Of Operation For Aes, Dana Geislinger, Cory Thigpen, Daniel W. Engels

SMU Data Science Review

In this paper, we present an empirical evaluation of the randomness of the ciphertext blocks generated by the Advanced Encryption Standard (AES) cipher in Counter (CTR) mode and in Cipher Block Chaining (CBC) mode. Vulnerabilities have been found in the AES cipher that may lead to a reduction in the randomness of the generated ciphertext blocks that can result in a practical attack on the cipher. We evaluate the randomness of the AES ciphertext using the standard key length and NIST randomness tests. We evaluate the randomness through a longitudinal analysis on 200 billion ciphertext blocks using logistic regression and …


Prediction Of High School Graduation With Decision Trees, Andrea M. Lee Aug 2019

Prediction Of High School Graduation With Decision Trees, Andrea M. Lee

Graduate Theses/Dissertations

While working as an educator for the past fourteen years, we are always looking at data and determining ways to help our students. Graduation status is one area of interest. I wanted to apply statistical methods to try and find early indicators of those students who may drop out, thus being able to provide early intervention to those students. With early intervention, we may be able to lower our dropout rate. While studying different methods of pattern recognition, I found that the decision tree method in machine learning was the best for the data that I had collected. Decision trees …


The Price Is Right: Analyzing Bidding Behavior On Contestants’ Row, Paul Kvam May 2019

The Price Is Right: Analyzing Bidding Behavior On Contestants’ Row, Paul Kvam

Department of Math & Statistics Faculty Publications

The TV game show “The Price is Right” features a bidding auction called Contestant’s Row that rewards the player (out of four) who bids closest to an item’s value without overbidding. By exploring 903 game outcomes from the 2000–2001 season, we show how player strategies are significantly inefficient, and compare the empirical results to probability outcomes for optimal bid strategies found in a recent study. Findings show that the last bidder would do better using the naïve strategy of bidding a dollar more than the highest of the three bids. We apply the EM algorithm in a novel way to …


Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley May 2019

Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley

SMU Data Science Review

In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …


Random Forest Vs Logistic Regression: Binary Classification For Heterogeneous Datasets, Kaitlin Kirasich, Trace Smith, Bivin Sadler Aug 2018

Random Forest Vs Logistic Regression: Binary Classification For Heterogeneous Datasets, Kaitlin Kirasich, Trace Smith, Bivin Sadler

SMU Data Science Review

Selecting a learning algorithm to implement for a particular application on the basis of performance still remains an ad-hoc process using fundamental benchmarks such as evaluating a classifier’s overall loss function and misclassification metrics. In this paper we address the difficulty of model selection by evaluating the overall classification performance between random forest and logistic regression for datasets comprised of various underlying structures: (1) increasing the variance in the explanatory and noise variables, (2) increasing the number of noise variables, (3) increasing the number of explanatory variables, (4) increasing the number of observations. We developed a model evaluation tool capable …


Fitting The Rasch Model Under The Logistic Regression Framework To Reduce Estimation Bias, Tianshu Pan Jun 2018

Fitting The Rasch Model Under The Logistic Regression Framework To Reduce Estimation Bias, Tianshu Pan

Journal of Modern Applied Statistical Methods

This article showed how and why the Rasch model can be fitted under the logistic regression framework. Then a penalized maximum likelihood (Firth 1993) for logistic regression models can also be used to reduce ML biases when fitting the Rasch model. These conclusions are supported by a simulation study.


The Impact Of Changing Requirements, James C. Ellis Mar 2018

The Impact Of Changing Requirements, James C. Ellis

Theses and Dissertations

The fundamental purpose of an Engineering Change Proposal (ECP) is to change the requirements of a contract. To build in flexibility, the acquisition practice is to estimate a dollar value to hold in reserve after the contract is awarded. There appears to be no empirical-based method for estimating this ECP withhold in the literature. Using the Cost Assessment Data Enterprise (CADE) database, 533 contracts were randomly selected to build two regression models: one to predict the likelihood of a contract experiencing an ECP, and the other to determine the expected median percent increase in baseline contract cost if an ECP …


Preference Probability Based On Ranks - A New Approach Using Logistic Regression With Zero Intercept, Oluwagbenga David Agboola Jan 2018

Preference Probability Based On Ranks - A New Approach Using Logistic Regression With Zero Intercept, Oluwagbenga David Agboola

Theses, Dissertations and Capstones

Many probability models have been proposed to describe rankings. One of these is the BradleyTerry model, which is based on observed pairwise preferences. For this study, we reverse the case and propose a new approach for estimating pairwise preference probabilities based on observed rankings. The new approach uses logistic regression with zero intercept as the statistical model that fits this situation. In order to implement the model, we first estimate the parameter using maximum likelihood estimation. Then we evaluate this estimation using numerical approximation procedures. We consider three such procedures: bisection method, Newton-Raphson method, and improved Newton’s method. Using simulated …


The Use Of Item Response Theory In Survey Methodology: Application In Seat Belt Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter Jan 2018

The Use Of Item Response Theory In Survey Methodology: Application In Seat Belt Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter

Mathematics & Statistics Faculty Publications

Problem: Several approaches to analyze survey data have been proposed in the literature. One method that is not popular in survey research methodology is the use of item response theory (IRT). Since accurate methods to make prediction behaviors are based upon observed data, the design model must overcome computation challenges, but also consideration towards calibration and proficiency estimation. The IRT model deems to be offered those latter options. We review that model and apply it to an observational survey data. We then compare the findings with the more popular weighted logistic regression. Method: Apply IRT model to the observed data …