Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Logistic regression

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 69

Full-Text Articles in Statistics and Probability

Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell Dec 2017

Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Movement and habitat selection by Greater Sage-grouse (Centrocercus uropasianus) is of great interest to wildlife managers tasked with applying conservation measures for this iconic western species. Current technology has created small and lightweight GPS (Global Positioning Systems) transmitters that can be attached to sage-grouse. Using GIS software and statistical programs such as Program R, land managers can analyze GPS location data to assess how sage-grouse are geospatially interacting with their habitats. Within the Panguitch Sage-Grouse Management Area (SGMA) thousands of acres of land have been restored or manipulated to enhance sage-grouse habitat; this usually involves removal of pinyon pine …


Exact Approaches For Bias Detection And Avoidance With Small, Sparse, Or Correlated Categorical Data, Sarah E. Schwartz Dec 2017

Exact Approaches For Bias Detection And Avoidance With Small, Sparse, Or Correlated Categorical Data, Sarah E. Schwartz

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Every day, traditional statistical methodology are used world wide to study a variety of topics and provides insight regarding countless subjects. Each technique is based on a distinct set of assumptions to ensure valid results. Additionally, many statistical approaches rely on large sample behavior and may collapse or degenerate in the presence of small, spare, or correlated data. This dissertation details several advancements to detect these conditions, avoid their consequences, and analyze data in a different way to yield trustworthy results.

One of the most commonly used modeling techniques for outcomes with only two possible categorical values (eg. live/die, pass/fail, …


Supervised Classification Using Finite Mixture Copula, Sumen Sen, Norou Diawara Aug 2017

Supervised Classification Using Finite Mixture Copula, Sumen Sen, Norou Diawara

Mathematics & Statistics Faculty Publications

Use of copula for statistical classification is recent and gaining popularity. For example, statistical classification using copula has been proposed for automatic character recognition, medical diagnostic and most recently in data mining. Classical discrimination rules assume normality. But in this data age time, this assumption is often questionable. In fact features of data could be a mixture of discrete and continues random variables. In this paper, mixture copula densities are used to model class conditional distributions. Such types of densities are useful when the marginal densities of the vector of features are not normally distributed and are of a mixed …


Application Of Support Vector Machine Modeling And Graph Theory Metrics For Disease Classification, Jessica M. Rudd Jul 2017

Application Of Support Vector Machine Modeling And Graph Theory Metrics For Disease Classification, Jessica M. Rudd

Published and Grey Literature from PhD Candidates

Disease classification is a crucial element of biomedical research. Recent studies have demonstrated that machine learning techniques, such as Support Vector Machine (SVM) modeling, produce similar or improved predictive capabilities in comparison to the traditional method of Logistic Regression. In addition, it has been found that social network metrics can provide useful predictive information for disease modeling. In this study, we combine simulated social network metrics with SVM to predict diabetes in a sample of data from the Behavioral Risk Factor Surveillance System. In this dataset, Logistic Regression outperformed SVM with ROC index of 81.8 and 81.7 for models with …


Binary Classification On Past Due Of Service Accounts Using Logistic Regression And Decision Tree, Yan Wang, Jennifer L. Priestley Jan 2017

Binary Classification On Past Due Of Service Accounts Using Logistic Regression And Decision Tree, Yan Wang, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

This paper aims at predicting businesses’ past due in service accounts as well as determining the variables that impact the likelihood of repayment. Two binary classification approaches, logistic regression and the decision tree, were conducted and compared. Both approaches have very good performances with respect to the accuracy. However, the decision tree only uses 10 predictors and reaches an accuracy of 96.69% on the validation set while logistic regression includes 14 predictors and reaches an accuracy of 94.58%. Due to the large concern of false negatives in financial industry, the decision tree technique is a better option than logistic regression …


Logistic Ensemble Models, Bob Vanderheyden, Jennifer L. Priestley Jan 2017

Logistic Ensemble Models, Bob Vanderheyden, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

Predictive models that are developed in a regulated industry or a regulated application, like determination of credit worthiness must be interpretable and “rational” (e.g., improvements in basic credit behavior must result in improved credit worthiness scores). Machine Learning technologies provide very good performance with minimal analyst intervention, so they are well suited to a high volume analytic environment but the majority are “black box” tools that provide very limited insight or interpretability into key drivers of model performance or predicted model output values. This paper presents a methodology that blends one of the most popular predictive statistical modeling methods with …


Inference Using Bhattacharyya Distance To Model Interaction Effects When The Number Of Predictors Far Exceeds The Sample Size, Sarah A. Janse Jan 2017

Inference Using Bhattacharyya Distance To Model Interaction Effects When The Number Of Predictors Far Exceeds The Sample Size, Sarah A. Janse

Theses and Dissertations--Statistics

In recent years, statistical analyses, algorithms, and modeling of big data have been constrained due to computational complexity. Further, the added complexity of relationships among response and explanatory variables, such as higher-order interaction effects, make identifying predictors using standard statistical techniques difficult. These difficulties are only exacerbated in the case of small sample sizes in some studies. Recent analyses have targeted the identification of interaction effects in big data, but the development of methods to identify higher-order interaction effects has been limited by computational concerns. One recently studied method is the Feasible Solutions Algorithm (FSA), a fast, flexible method that …


A Comparison Of Decision Tree With Logistic Regression Model For Prediction Of Worst Non-Financial Payment Status In Commercial Credit, Jessica M. Rudd Mph, Gstat, Jennifer L. Priestley Jan 2017

A Comparison Of Decision Tree With Logistic Regression Model For Prediction Of Worst Non-Financial Payment Status In Commercial Credit, Jessica M. Rudd Mph, Gstat, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

Credit risk prediction is an important problem in the financial services domain. While machine learning techniques such as Support Vector Machines and Neural Networks have been used for improved predictive modeling, the outcomes of such models are not readily explainable and, therefore, difficult to apply within financial regulations. In contrast, Decision Trees are easy to explain, and provide an easy to interpret visualization of model decisions. The aim of this paper is to predict worst non-financial payment status among businesses, and evaluate decision tree model performance against traditional Logistic Regression model for this task. The dataset for analysis is provided …


What Affects Parents’ Choice Of Milk? An Application Of Bayesian Model Averaging, Yingzhe Cheng Dec 2016

What Affects Parents’ Choice Of Milk? An Application Of Bayesian Model Averaging, Yingzhe Cheng

Mathematics & Statistics ETDs

This study identifies the factors that influence parents’ choice of milk for their children, using data from a unique survey administered in 2013 in Hunan province, China. In this survey, we identified two brands of milk, which differ in their prices and safety claims by the producer. Data were collected on parents’ choice of milk between the two brands, demographics, attitude towards food safety and behaviors related to food. Stepwise model selection and Bayesian model averaging (BMA) are used to search for influential factors. The two approaches consistently select the same factors suggested by an economic theoretical model, including price …


A Multi-Indexed Logistic Model For Time Series, Xiang Liu Dec 2016

A Multi-Indexed Logistic Model For Time Series, Xiang Liu

Electronic Theses and Dissertations

In this thesis, we explore a multi-indexed logistic regression (MILR) model, with particular emphasis given to its application to time series. MILR includes simple logistic regression (SLR) as a special case, and the hope is that it will in some instances also produce significantly better results. To motivate the development of MILR, we consider its application to the analysis of both simulated sine wave data and stock data. We looked at well-studied SLR and its application in the analysis of time series data. Using a more sophisticated representation of sequential data, we then detail the implementation of MILR. We compare …


Prevalence Of And Risk Factors For Adolescent Obesity In Tennessee Using The 2010 Youth Risk Behavior Survey (Yrbs) Data: An Analysis Using Weighted Hierarchical Logistic Regression, Shimin Zheng, Nicole Holt, Jodi L. Southerland, Yan Cao, Trevor Taylor, Deborah L. Slawson, Mark Bloodworth Oct 2016

Prevalence Of And Risk Factors For Adolescent Obesity In Tennessee Using The 2010 Youth Risk Behavior Survey (Yrbs) Data: An Analysis Using Weighted Hierarchical Logistic Regression, Shimin Zheng, Nicole Holt, Jodi L. Southerland, Yan Cao, Trevor Taylor, Deborah L. Slawson, Mark Bloodworth

ETSU Faculty Works

Background: The rate of adolescent overweight and obesity has more than quadrupled over the past few decades, and has become a major public health problem [1]. In 2011, 55% of 12-19 year olds in the United States (U.S.) were overweight or obese [2]. Adolescence is a pivotal time in which many health risk behaviors such as tobacco, alcohol, and drug use are initiated. Such health risk behaviors have been significantly associated with overweight and obesity among adolescents.

Objective: The purpose of this study is to evaluate the relationship between obesity and the health risk behaviors most commonly associated with premature …


Exploring New Models For Seatbelt Use In Survey Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter Oct 2016

Exploring New Models For Seatbelt Use In Survey Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter

Virginia Journal of Science

Problem: Several approaches to analyze seatbelt use have been proposed in the literature. Two methods that have not been explored are the use of unweighted and weighted logistic regression models and the use of item response theory (IRT) or the Rasch model. Since accurate methods to predict seatbelt use behavior based upon observed data must include a built-in design method and model and overcome computation challenges, weighted and IRT methods deem to be other options for an observational survey of seatbelt use in the state of Virginia.

Method: The data observed from 136 sites within the Commonwealth of …


Liu-Type Logistic Estimators With Optimal Shrinkage Parameter, Yasin Asar May 2016

Liu-Type Logistic Estimators With Optimal Shrinkage Parameter, Yasin Asar

Journal of Modern Applied Statistical Methods

Multicollinearity in logistic regression affects the variance of the maximum likelihood estimator negatively. In this study, Liu-type estimators are used to reduce the variance and overcome the multicollinearity by applying some existing ridge regression estimators to the case of logistic regression model. A Monte Carlo simulation is given to evaluate the performances of these estimators when the optimal shrinkage parameter is used in the Liu-type estimators, along with an application of real case data.


Separation Of Points And Interval Estimation In Mixed Dose-Response Curves With Selective Component Labeling, Darl D. Flake Ii May 2016

Separation Of Points And Interval Estimation In Mixed Dose-Response Curves With Selective Component Labeling, Darl D. Flake Ii

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Dose-response experiments are those that involve giving subjects different amounts of a treatment and observing the outcome. For example, plants may be given fertilizer and their growth could be measured or cancer patients could be given different doses of chemotherapy and their response could be monitored. These experiments are used to understand the relationship between the amount of, and response to, the treatment. Logistic regression models are often used to summarize data from these types of experiments. The dose-response experiment that motivated this dissertation involved treating a grain-pest with a pesticide. Some of the beetles had genes that made them …


An Analysis Of Accuracy Using Logistic Regression And Time Series, Edwin Baidoo, Jennifer L. Priestley Jan 2016

An Analysis Of Accuracy Using Logistic Regression And Time Series, Edwin Baidoo, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

This paper analyzes the accuracy rates for logistic regression and time series models. It also examines a relatively new performance index that takes into consideration the business assumptions of credit markets. Although prior research has focused on evaluation metrics, such as AUC and Gini index, this new measure has a more intuitive interpretation for various managers and decision makers and can be applied to both Logistic and Time Series models.


Application Of Isotonic Regression In Predicting Business Risk Scores, Linh T. Le, Jennifer L. Priestley Jan 2016

Application Of Isotonic Regression In Predicting Business Risk Scores, Linh T. Le, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

An isotonic regression model fits an isotonic function of the explanatory variables to estimate the expectation of the response variable. In other words, as the function increases, the estimated expectation of the response must be non-decreasing. With this characteristic, isotonic regression could be a suitable option to analyze and predict business risk scores. A current challenge of isotonic regression is the decrease of performance when the model is fitted in a large data set e.g. more than four or five dimensions. This paper attempts to apply isotonic regression models into prediction of business risk scores using a large data set …


System-Wide Prediction Of General, All-Cause, Preventable Hospital Readmissions, Ken Musselman, Brandon Pope, Steve Witz, Zhiyi Tian, Lingsong Zhang, Linda Leon, Ann Davis Dec 2015

System-Wide Prediction Of General, All-Cause, Preventable Hospital Readmissions, Ken Musselman, Brandon Pope, Steve Witz, Zhiyi Tian, Lingsong Zhang, Linda Leon, Ann Davis

RCHE Publications

Existing studies of hospital readmissions typically focus on specific diagnoses, age groups, discharge dispositions, payer classes, or hospitals, and often use small samples. It is not clear how predictive models generated from such studies generalize across diseases, hospitals, or time periods. In this study, a logistic regression model of readmission risk within 30 days based on hospital administrative data was constructed and validated across hospitals and time periods. The hospitals included both general and specialty hospitals such as long-term care, women’s, and children’s hospitals. The administrative data included information on patient’s demographics, diagnoses, procedures, and discharge disposition. Derivation and validation …


Analysis Of Rheumatoid Arthritis Data Using Logistic Regression And Penalized Approach, Wei Chen Nov 2015

Analysis Of Rheumatoid Arthritis Data Using Logistic Regression And Penalized Approach, Wei Chen

USF Tampa Graduate Theses and Dissertations

In this paper, a rheumatoid arthritis (RA) medicine clinical dataset with an ordinal response is selected to study this new medicine. In the dataset, there are four features, sex, age,treatment, and preliminary. Sex is a binary categorical variable with 1 indicates male, and 0 indicates female. Age is the numerical age of the patients. And treatment is a binary categorical variable with 1 indicates has RA, and 0 indicates does not have RA. And preliminary is a five class categorical variable indicates the patient’s RA severity status before taking the medication. The response Y is 5 class ordinal variable shows …


Bird Keeping And Lung Cancer, Andrew Tackmann, Jonathan Hellman, Jamie Johnson Aug 2014

Bird Keeping And Lung Cancer, Andrew Tackmann, Jonathan Hellman, Jamie Johnson

Journal of Undergraduate Research at Minnesota State University, Mankato

Logistic regression is reviewed in estimating parameters and in making inferences about the parameters. A contingency table approach in computing goodness of fit in logistic regression is elaborated. An existing data on a sample of lung cancer patients and a control group is used to apply the procedures discussed. The data reveals that between the groups considered, the factors ‘bird keeping’ and ‘the number of years of smoking’ are significant as the causes for lung cancer.


A Note On The Control Function Approach With An Instrumental Variable And A Binary Outcome, Eric Tchetgen Tchetgen Jul 2014

A Note On The Control Function Approach With An Instrumental Variable And A Binary Outcome, Eric Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


The Association Of Calcium Intake And Other Risk Factors With Cardiovascular Disease Among Obese Adults In Usa, Yang Chen, Sheryl Strasser, Katie Callahan, David Blackley, Yan Cao, Liang Wang, Shimin Zheng Mar 2014

The Association Of Calcium Intake And Other Risk Factors With Cardiovascular Disease Among Obese Adults In Usa, Yang Chen, Sheryl Strasser, Katie Callahan, David Blackley, Yan Cao, Liang Wang, Shimin Zheng

ETSU Faculty Works

In this study, we used a cross-sectional study design to examine the relationship between the calcium intake and risk factors for CVD among obese adults by using continuous waves of National Health and Nutrition Examination Survey (NHANES) data 1999-2010. The association between calcium intake and risk factors of CVD (hypertension, total cholesterol, HDL, glycohemoglobin), CRP, albuminuria) is assessed among obese adults in USA. The incidence of Cardiovascular Disease (CVD) is high among obese people. The potential effects of inadequate calcium intake on CVD are receiving increased epidemiologic attention. Understanding the association between risk factors for CVD and calcium intake among …


Generation And Statistical Modeling Of Active Protein Chimeras: A Sequence Based Approach, Nicholas Fico Oct 2013

Generation And Statistical Modeling Of Active Protein Chimeras: A Sequence Based Approach, Nicholas Fico

Open Access Dissertations

Generation of active protein chimeras is a valuable tool to probe the functional space of proteins. Statistical modeling is the next logical step, allowing us to build a model of gene fragment replaceability between species. In this thesis I begin to develop the statistical tools that are needed to systematically describe combinatorial protein libraries. I present three sets of diverse chimeric protein libraries developed using sequence information. The statistical model of the human N-Ras and human K-Ras-4B genes reveal a set previously unidetifed surface residues on the N-Ras G-Domain that may be involved in cellular localization. Statistical modeling of a …


Bayesian Phase I Dose Finding In Cancer Trials, Lin Yang Aug 2011

Bayesian Phase I Dose Finding In Cancer Trials, Lin Yang

Dissertations and Theses (Open Access)

This dissertation explores phase I dose-finding designs in cancer trials from three perspectives: the alternative Bayesian dose-escalation rules, a design based on a time-to-dose-limiting toxicity (DLT) model, and a design based on a discrete-time multi-state (DTMS) model.

We list alternative Bayesian dose-escalation rules and perform a simulation study for the intra-rule and inter-rule comparisons based on two statistical models to identify the most appropriate rule under certain scenarios. We provide evidence that all the Bayesian rules outperform the traditional ``3+3'' design in the allocation of patients and selection of the maximum tolerated dose.

The design based on a time-to-DLT model …


Statistical Analysis Of Fatalities Due To Vehicle Accidents In Las Vegas, Nv, Annabelle Marie Mathis Aug 2011

Statistical Analysis Of Fatalities Due To Vehicle Accidents In Las Vegas, Nv, Annabelle Marie Mathis

UNLV Theses, Dissertations, Professional Papers, and Capstones

The goal of this thesis is to investigate factors that affect the odds of having a fatality in a vehicle collision. We will be looking at characteristics of the driver that caused the accident (age, gender, behavior, actions, influences, and seat belt worn), the characteristics of the vehicle the driver drove (type of vehicle, and air bag deployment), the characteristics of the environment in which the accident occurred (weather, road condition, lighting, time of day, the day of the week, and month of the year), the characteristics of the crash (direction of accident and how many vehicles were involved), and …


Logistic Regression Models For Higher Order Transition Probabilities Of Markov Chain For Analyzing The Occurrences Of Daily Rainfall Data, Narayan Chanra Sinha, M. Ataharul Islam, Kazi Saleh Ahamed May 2011

Logistic Regression Models For Higher Order Transition Probabilities Of Markov Chain For Analyzing The Occurrences Of Daily Rainfall Data, Narayan Chanra Sinha, M. Ataharul Islam, Kazi Saleh Ahamed

Journal of Modern Applied Statistical Methods

Logistic regression models for transition probabilities of higher order Markov models are developed for the sequence of chain dependent repeated observations. To identify the significance of these models and their parameters a test procedure for a likelihood ratio criterion is developed. A method of model selection is suggested on the basis of AIC and BIC procedures. The proposed models and test procedures are applied to analyze the occurrences of daily rainfall data for selected stations in Bangladesh. Based on results from these models, the transition probabilities of first order Markov model for temperature and humidity provided the most suitable option …


Bayesian Semiparametric Generalizations Of Linear Models Using Polya Trees, Angela Schoergendorfer Jan 2011

Bayesian Semiparametric Generalizations Of Linear Models Using Polya Trees, Angela Schoergendorfer

University of Kentucky Doctoral Dissertations

In a Bayesian framework, prior distributions on a space of nonparametric continuous distributions may be defined using Polya trees. This dissertation addresses statistical problems for which the Polya tree idea can be utilized to provide efficient and practical methodological solutions.

One problem considered is the estimation of risks, odds ratios, or other similar measures that are derived by specifying a threshold for an observed continuous variable. It has been previously shown that fitting a linear model to the continuous outcome under the assumption of a logistic error distribution leads to more efficient odds ratio estimates. We will show that deviations …


Robust Estimators In Logistic Regression: A Comparative Simulation Study, Sanizah Ahmad, Norazan Mohamed Ramli, Habshah Midi Nov 2010

Robust Estimators In Logistic Regression: A Comparative Simulation Study, Sanizah Ahmad, Norazan Mohamed Ramli, Habshah Midi

Journal of Modern Applied Statistical Methods

The maximum likelihood estimator (MLE) is commonly used to estimate the parameters of logistic regression models due to its efficiency under a parametric model. However, evidence has shown the MLE has an unduly effect on the parameter estimates in the presence of outliers. Robust methods are put forward to rectify this problem. This article examines the performance of the MLE and four existing robust estimators under different outlier patterns, which are investigated by real data sets and Monte Carlo simulation.


A Poisson-Like Model Of Sub-Clinical Signs From The Examination Of Healthy Aging Subjects, Stephen Merrill, Barbara Myklebust, Joel B. Myklebust, Norman Reynolds, Edmund Duthie Aug 2008

A Poisson-Like Model Of Sub-Clinical Signs From The Examination Of Healthy Aging Subjects, Stephen Merrill, Barbara Myklebust, Joel B. Myklebust, Norman Reynolds, Edmund Duthie

Mathematics, Statistics and Computer Science Faculty Research and Publications

Background and aims: Our studies of the standard neurological examination on 66 middle-aged (50–64 yrs) and elderly subjects (65–84 yrs) demonstrate that healthy elders have neurological deficits (or “signs”) that are not associated with specific known neurological disease. The purpose of the current study is to describe this loss of neurological function in healthy aging subjects as seen through accumulated subclinical neurological signs present.

Methods: Logistic regression is applied to the data on each of six signs. Parameters determined are used to describe the distribution of first occurrence times for each sign. The results are then used to construct a …


Perbandingan Analisis Regresi Logistik Dengan Analisis Propensity Score Matching Pada Studi Kasus Imunisasi Bayi, Waras Budi Utomo Jun 2008

Perbandingan Analisis Regresi Logistik Dengan Analisis Propensity Score Matching Pada Studi Kasus Imunisasi Bayi, Waras Budi Utomo

Kesmas

Analisis multivariat konvensioanal tidak selalu merupakan metode ideal untuk memprediksi efek pajanan pada studi-studi observasional. Ketika distribusi kovariat antara kelompok pajanan berbeda besar, penyesuaan dengan teknik multivariat konvensioanl tidak cukup menyeimbangkan kelompok tersebut. Bias yang tersisa dapat menghambat penarikan kesimpulan yang valid. Tujuan penelitian ini adalah membandingkan hasil analisis multivariat konvensional dengan analisis metoda propensity score matching pada studi kasus data sekunder imunisasi bayi ASUH KAP2 2003. Penelitian ini menemukan nilai OR metoda regresi logistik (0,99) berbeda dengan metoda propensity score matching (0,96). Metoda propensity score matching berhasil menjodohkan 574 subjek (68,27%). Untuk evaluasi pengaruh faktor risiko disarankan menggunakan model …


Estimation Of Risk For Developing Cardiac Problem In Patients Of Type 2 Diabetes As Obtained By The Technique Of Density Estimation, Ajit Mukherjee, Ajit Mathur, Rakesh Mittal May 2007

Estimation Of Risk For Developing Cardiac Problem In Patients Of Type 2 Diabetes As Obtained By The Technique Of Density Estimation, Ajit Mukherjee, Ajit Mathur, Rakesh Mittal

Journal of Modern Applied Statistical Methods

High levels of cholesterol and triglyceride are known to be strongly associated with development of cardiac problem in patients of type 2 diabetes. In a hospital-based study, patients showing ECG positive were compared with those who were not. The observations on cholesterol and triglyceride were considered for estimation of risk for developing the cardiac problem. The technique of density estimation employing Epanechnikov kernel was used for estimating bivariate probability density functions with respect to observations on cholesterol and triglyceride of the two groups. Using the odds form of Bayes’ rule, the estimates of posterior odds were computed.