Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (9)
- Biostatistics (5)
- Statistical Methodology (4)
- Data Science (3)
- Multivariate Analysis (3)
-
- Life Sciences (2)
- Analytical Chemistry (1)
- Behavior and Ethology (1)
- Business (1)
- Business Analytics (1)
- Business Intelligence (1)
- Chemistry (1)
- Clinical Trials (1)
- Design of Experiments and Sample Surveys (1)
- Disease Modeling (1)
- Diseases (1)
- Earth Sciences (1)
- Ecology and Evolutionary Biology (1)
- Environmental Sciences (1)
- Geomorphology (1)
- Laboratory and Basic Science Research (1)
- Medicine and Health Sciences (1)
- Natural Resources Management and Policy (1)
- Natural Resources and Conservation (1)
- Other Chemistry (1)
- Other Ecology and Evolutionary Biology (1)
- Population Biology (1)
- Institution
- Publication Year
- Publication
-
- Published and Grey Literature from PhD Candidates (2)
- SMU Data Science Review (2)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (1)
- Dissertations and Theses (Open Access) (1)
- Electronic Theses and Dissertations (1)
-
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Mathematics & Statistics Faculty Publications (1)
- Research Symposium (1)
- School of Mathematical & Statistical Sciences Faculty Publications (1)
- U.C. Berkeley Division of Biostatistics Working Paper Series (1)
- Virginia Journal of Science (1)
- Publication Type
Articles 1 - 13 of 13
Full-Text Articles in Statistical Models
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Research Symposium
Background: With a rapid development of data collection technology, high dimensional data, whose model dimension k may be growing or much larger than the sample size n, is becoming increasingly prevalent in different fields of study, such as ecology, genetics, among others. This data deluge is introducing new challenges to traditional statistical procedures and theories and is thus generating a renewed interest in the problems of variable selection and classification in high dimensional regression models. In large k, small n settings, variable selection is usually the first step for dimension reduction to uncover significant covariates, which contribute to …
Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang
Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang
School of Mathematical & Statistical Sciences Faculty Publications
In this paper, we propose a sparse Bayesian procedure with global and local(GL) shrinkage priors for the problems of variable selection and classification in high-dimensional logistic regression models. In particular, we consider two types of GL shrinkage priors for the regression coefficients, the horseshoe (HS)prior and the normal-gamma (NG) prior, and then specify a correlated prior for the binary vector to distinguish models with the same size. The GL priors are then combined with mixture representations of logistic distribution to construct a hierarchical Bayes model that allows efficient implementation of a Markov chain Monte Carlo (MCMC) to generate samples from …
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
SMU Data Science Review
Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …
Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair
Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair
Graduate Student Theses, Dissertations, & Professional Papers
Landslides are a globally pervasive problem with the potential to cause significant fatalities and economic losses. Although landslides are widespread, many at-risk regions may not have the high-quality data or resources used in most landslide susceptibility analyses. This study aims to develop regional susceptibility relationships that are versatile and use publicly available data and open-sourced software. Logistic Regression and Frequency Ratio susceptibility relationships were developed in 23 regions in Washington, Utah, North Carolina, and Kentucky, with a region referring to a unique area and data combination. Regions were diverse in their geology, morphology, climate, and nature and quality of their …
An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone
An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone
Published and Grey Literature from PhD Candidates
Data mining techniques have numerous applications in bankcard response modeling. Logistic regression has been used as the standard modeling tool in the financial industry because of its almost always desirable performance and its interpretability. In this paper, we propose a hybrid bankcard response model, which integrates decision tree-based chi-square automatic interaction detection (CHAID) into logistic regression. In the first stage of the hybrid model, CHAID analysis is used to detect the possible potential variable interactions. Then in the second stage, these potential interactions are served as the additional input variables in logistic regression. The motivation of the proposed hybrid model …
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
SMU Data Science Review
In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …
Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell
Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
Movement and habitat selection by Greater Sage-grouse (Centrocercus uropasianus) is of great interest to wildlife managers tasked with applying conservation measures for this iconic western species. Current technology has created small and lightweight GPS (Global Positioning Systems) transmitters that can be attached to sage-grouse. Using GIS software and statistical programs such as Program R, land managers can analyze GPS location data to assess how sage-grouse are geospatially interacting with their habitats. Within the Panguitch Sage-Grouse Management Area (SGMA) thousands of acres of land have been restored or manipulated to enhance sage-grouse habitat; this usually involves removal of pinyon pine …
Supervised Classification Using Finite Mixture Copula, Sumen Sen, Norou Diawara
Supervised Classification Using Finite Mixture Copula, Sumen Sen, Norou Diawara
Mathematics & Statistics Faculty Publications
Use of copula for statistical classification is recent and gaining popularity. For example, statistical classification using copula has been proposed for automatic character recognition, medical diagnostic and most recently in data mining. Classical discrimination rules assume normality. But in this data age time, this assumption is often questionable. In fact features of data could be a mixture of discrete and continues random variables. In this paper, mixture copula densities are used to model class conditional distributions. Such types of densities are useful when the marginal densities of the vector of features are not normally distributed and are of a mixed …
Application Of Support Vector Machine Modeling And Graph Theory Metrics For Disease Classification, Jessica M. Rudd
Application Of Support Vector Machine Modeling And Graph Theory Metrics For Disease Classification, Jessica M. Rudd
Published and Grey Literature from PhD Candidates
Disease classification is a crucial element of biomedical research. Recent studies have demonstrated that machine learning techniques, such as Support Vector Machine (SVM) modeling, produce similar or improved predictive capabilities in comparison to the traditional method of Logistic Regression. In addition, it has been found that social network metrics can provide useful predictive information for disease modeling. In this study, we combine simulated social network metrics with SVM to predict diabetes in a sample of data from the Behavioral Risk Factor Surveillance System. In this dataset, Logistic Regression outperformed SVM with ROC index of 81.8 and 81.7 for models with …
A Multi-Indexed Logistic Model For Time Series, Xiang Liu
A Multi-Indexed Logistic Model For Time Series, Xiang Liu
Electronic Theses and Dissertations
In this thesis, we explore a multi-indexed logistic regression (MILR) model, with particular emphasis given to its application to time series. MILR includes simple logistic regression (SLR) as a special case, and the hope is that it will in some instances also produce significantly better results. To motivate the development of MILR, we consider its application to the analysis of both simulated sine wave data and stock data. We looked at well-studied SLR and its application in the analysis of time series data. Using a more sophisticated representation of sequential data, we then detail the implementation of MILR. We compare …
Exploring New Models For Seatbelt Use In Survey Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter
Exploring New Models For Seatbelt Use In Survey Data, Mark K. Ledbetter, Norou Diawara, Bryan E. Porter
Virginia Journal of Science
Problem: Several approaches to analyze seatbelt use have been proposed in the literature. Two methods that have not been explored are the use of unweighted and weighted logistic regression models and the use of item response theory (IRT) or the Rasch model. Since accurate methods to predict seatbelt use behavior based upon observed data must include a built-in design method and model and overcome computation challenges, weighted and IRT methods deem to be other options for an observational survey of seatbelt use in the state of Virginia.
Method: The data observed from 136 sites within the Commonwealth of …
Bayesian Phase I Dose Finding In Cancer Trials, Lin Yang
Bayesian Phase I Dose Finding In Cancer Trials, Lin Yang
Dissertations and Theses (Open Access)
This dissertation explores phase I dose-finding designs in cancer trials from three perspectives: the alternative Bayesian dose-escalation rules, a design based on a time-to-dose-limiting toxicity (DLT) model, and a design based on a discrete-time multi-state (DTMS) model.
We list alternative Bayesian dose-escalation rules and perform a simulation study for the intra-rule and inter-rule comparisons based on two statistical models to identify the most appropriate rule under certain scenarios. We provide evidence that all the Bayesian rules outperform the traditional ``3+3'' design in the allocation of patients and selection of the maximum tolerated dose.
The design based on a time-to-DLT model …
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …