Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (3)
- Statistical Models (3)
- Life Sciences (2)
- Statistical Methodology (2)
- Analytical Chemistry (1)
-
- Bioinformatics (1)
- Biostatistics (1)
- Business (1)
- Business Analytics (1)
- Business Intelligence (1)
- Chemistry (1)
- Computational Biology (1)
- Data Science (1)
- Demography, Population, and Ecology (1)
- Epidemiology (1)
- Genetics (1)
- Genetics and Genomics (1)
- Laboratory and Basic Science Research (1)
- Medicine and Health Sciences (1)
- Microarrays (1)
- Other Chemistry (1)
- Public Affairs, Public Policy and Public Administration (1)
- Public Health (1)
- Social and Behavioral Sciences (1)
- Sociology (1)
- Statistical Theory (1)
- Transportation (1)
- Institution
- Publication
- Publication Type
Articles 1 - 6 of 6
Full-Text Articles in Multivariate Analysis
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
SMU Data Science Review
Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …
A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani
A Comparison Of Logistic, Ridge, And Lasso Regression With Heart Failure Risk Data: Effects Of Sample Size, Predictor Correlation, And Predictor Weight On Outcome Accuracy, Mahmoud M. Aljuhani
Electronic Theses and Dissertations
Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.
Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance …
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
SMU Data Science Review
In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …
Statistical Analysis Of Fatalities Due To Vehicle Accidents In Las Vegas, Nv, Annabelle Marie Mathis
Statistical Analysis Of Fatalities Due To Vehicle Accidents In Las Vegas, Nv, Annabelle Marie Mathis
UNLV Theses, Dissertations, Professional Papers, and Capstones
The goal of this thesis is to investigate factors that affect the odds of having a fatality in a vehicle collision. We will be looking at characteristics of the driver that caused the accident (age, gender, behavior, actions, influences, and seat belt worn), the characteristics of the vehicle the driver drove (type of vehicle, and air bag deployment), the characteristics of the environment in which the accident occurred (weather, road condition, lighting, time of day, the day of the week, and month of the year), the characteristics of the crash (direction of accident and how many vehicles were involved), and …
Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond
Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond
UW Biostatistics Working Paper Series
Whole-genome studies are becoming a mainstay of biomedical research. Examples include expression array experiments, comparative genomic hybridization analyses and large case-control studies for detecting polymorphism/disease associations. The tactic of applying a regression model to every locus to obtain test statistics is useful in such studies. However, this approach ignores potential correlation structure in the data that could be used to gain power, particularly when a Bonferroni correction is applied to adjust for multiple testing. In this article, we propose using regression techniques for misspecified multivariate outcomes to increase statistical power over independence-based modeling at each locus. Even when the outcome …
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …