Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Utah State University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 91 - 120 of 317

Full-Text Articles in Statistics and Probability

Using A Discrete Choice Experiment To Estimate Willingness To Pay For Location Based Housing Attributes, Kristopher C. Toll Dec 2019

Using A Discrete Choice Experiment To Estimate Willingness To Pay For Location Based Housing Attributes, Kristopher C. Toll

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

In 1993, a travel study was conducted along the Wasatch front in Utah (Research Systems Group INC, 2013). The main purpose of this study was to assess travel behavior to understand the needs for future growth in Utah. Since then, the Research Service Group (RSG), conducted a new study in 2012 to understand current travel preferences in Utah. This survey, called the Residential Choice Stated Preference survey, asked respondents to make ten choice comparisons between two hypothetical homes. Each home in the choice comparison was described by different attributes, those attributes that were used are, type of neighborhood, distance from …


Student Insights Report, Fall 2019, The Center For Student Analytics Sep 2019

Student Insights Report, Fall 2019, The Center For Student Analytics

Publications

For the past three years, the staff of the Center for Student Analytics have worked to discover and expose meaningful, data-informed insights into what helps students succeed at Utah State University. The following pages highlight 20 of the most useful insights we found provided here in small sets that will be useful to students, faculty, staff, university leadership, parents, and even prospective students. As you explore this report, we encourage you to see the student data as a window into USU itself. While big data helps us understand how individual students are performing, it tells us a great deal more …


Tuning Hyperparameters In Supervised Learning Models And Applications Of Statistical Learning In Genome-Wide Association Studies With Emphasis On Heritability, Jill F. Lundell Aug 2019

Tuning Hyperparameters In Supervised Learning Models And Applications Of Statistical Learning In Genome-Wide Association Studies With Emphasis On Heritability, Jill F. Lundell

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Machine learning is a buzz word that has inundated popular culture in the last few years. This is a term for a computer method that can automatically learn and improve from data instead of being explicitly programmed at every step. Investigations regarding the best way to create and use these methods are prevalent in research. Machine learning models can be difficult to create because models need to be tuned. This dissertation explores the characteristics of tuning three popular machine learning models and finds a way to automatically select a set of tuning parameters. This information was used to create an …


Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark May 2019

Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Filtered historical simulation with an underlying GARCH process can be used as a valuable tool in VaR analysis, as it derives risk estimates that are sensitive to the distributional properties of the historical data of the produced predictive density. I examine the applications to risk analysis that filtered historical simulation can provide, as well as an interpretation of the predictive density as a poor man’s Bayesian posterior distribution. The predictive density allows us to make associated probabilistic statements regarding the results for VaR analysis, giving greater measurement of risk and the ability to maintain the optimal level of risk per …


Feasibility Of Multi-Year Forecast For The Colorado River Water Supply: Time Series Modeling, Brian Plucinski May 2019

Feasibility Of Multi-Year Forecast For The Colorado River Water Supply: Time Series Modeling, Brian Plucinski

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

The Colorado River is one of the largest resources for water in the United States, as well as being an important asset to the economy. Previous studies have shown a connection between the Great Salt Lake and the Colorado River. This study used time series analysis to build models to predict the water supply of the Colorado River ten years out. These models used data from the Colorado River in addition to Great Salt Lake water elevation. Several models suggest a decline in water supply from 2013 – 2020, before starting to increase. These predictions differ from predictions published by …


Drones And “Ghost Guns”: Unregulated Legal Space, Tori Bodine Mar 2019

Drones And “Ghost Guns”: Unregulated Legal Space, Tori Bodine

Research on Capitol Hill

Law enforcement agencies are fighting a two - pronged battle when it comes to emerging technologies: keeping up with new ways criminals are using technology and developing effective ways to combat these innovations, while balancing these challenges against preserving the individual liberties of law - abiding citizens. This conflict is especially apparent with regard to criminal use of commercial drones and the developing fringe market surrounding homemade untraceable firearms (“ghost guns”).


Power In Pairs: Assessing The Statistical Value Of Paired Samples In Tests For Differential Expression, John R. Stevens, Jennifer S. Herrick, Roger K. Wolff, Martha L. Slattery Dec 2018

Power In Pairs: Assessing The Statistical Value Of Paired Samples In Tests For Differential Expression, John R. Stevens, Jennifer S. Herrick, Roger K. Wolff, Martha L. Slattery

Mathematics and Statistics Faculty Publications

Background: When genomics researchers design a high-throughput study to test for differential expression, some biological systems and research questions provide opportunities to use paired samples from subjects, and researchers can plan for a certain proportion of subjects to have paired samples. We consider the effect of this paired samples proportion on the statistical power of the study, using characteristics of both count (RNA-Seq) and continuous (microarray) expression data from a colorectal cancer study.

Results: We demonstrate that a higher proportion of subjects with paired samples yields higher statistical power, for various total numbers of samples, and for various strengths of …


Rfviz: An Interactive Visualization Package For Random Forests In R, Christopher Beckett Dec 2018

Rfviz: An Interactive Visualization Package For Random Forests In R, Christopher Beckett

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Random forests are very popular tools for predictive analysis and data science. They work for both classification (where there is a categorical response variable) and regression (where the response is continuous). Random forests provide proximities, and both local and global measures of variable importance. However, these quantities require special tools to be effectively used to interpret the forest. Rfviz is a sophisticated interactive visualization package and toolkit in R, specially designed for interpreting the results of a random forest in a user-friendly way. Rfviz uses a recently developed R package (loon) from the Comprehensive R Archive Network (CRAN) to create …


Comparing Performance Of Gene Set Test Methods Using Biologically Relevant Simulated Data, Richard M. Lambert Dec 2018

Comparing Performance Of Gene Set Test Methods Using Biologically Relevant Simulated Data, Richard M. Lambert

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Today we know that there are many genetically driven diseases and health conditions. These problems often manifest only when a set of genes are either active or inactive. Recent technology allows us to measure the activity level of genes in cells, which we call gene expression. It is of great interest to society to be able to statistically compare the gene expression of a large number of genes between two or more groups. For example, we may want to compare the gene expression of a group of cancer patients with a group of non-cancer patients to better understand the genetic …


Statistical Methods To Account For Gene-Level Covariates In Normalization Of High-Dimensional Read-Count Data, Lauren Holt Lenz Dec 2018

Statistical Methods To Account For Gene-Level Covariates In Normalization Of High-Dimensional Read-Count Data, Lauren Holt Lenz

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

The goal of genetic-based cancer research is often to identify which genes behave differently in cancerous and healthy tissue. This difference in behavior, referred to as differential expression, may lead researchers to more targeted preventative care and treatment. One way to measure the expression of genes is though a process called RNA-Seq, that takes physical tissue samples and maps gene products and fragments in the sample back to the gene that created it, resulting in a large read-count matrix with genes in the rows and a column for each sample. The read-counts for tumor and normal samples are then compared …


The Power Law Distribution Of Agricultural Land Size, Lauren Chamberlain Dec 2018

The Power Law Distribution Of Agricultural Land Size, Lauren Chamberlain

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

This paper demonstrates that the distribution of county level agricultural land size in the United States is best described by a power-law distribution, a distribution that displays extremely heavy tails. This indicates that the majority of farmland exists in the upper tail. Our analysis indicates that the top 5% of agricultural counties account for about 25% of agricultural land between 1997-2012. The power-law distribution of farm size has important implications for the design of more efficient regional and national agricultural policies as counties close to the mean account for little of the cumulative distribution of total agricultural land. This has …


Surviving A Civil War: Expanding The Scope Of Survival Analysis In Political Science, Andrew B. Whetten Dec 2018

Surviving A Civil War: Expanding The Scope Of Survival Analysis In Political Science, Andrew B. Whetten

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Survival Analysis in the context of Political Science is frequently used to study the duration of agreements, political party influence, wars, senator term lengths, etc. This paper surveys a collection of methods implemented on a modified version of the Power-Sharing Event Dataset (which documents civil war peace agreement durations in the Post-Cold War era) in order to identify the research questions that are optimally addressed by each method. A primary comparison will be made between a Cox Proportional Hazards Model using some advanced capabilities in the glmnet package, a Survival Random Forest Model, and a Survival SVM. En route to …


Feature Screening Of Ultrahigh Dimensional Feature Spaces With Applications In Interaction Screening, Randall D. Reese Aug 2018

Feature Screening Of Ultrahigh Dimensional Feature Spaces With Applications In Interaction Screening, Randall D. Reese

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Data for which the number of predictors exponentially exceeds the number of observations is becoming increasingly prevalent in fields such as bioinformatics, medical imaging, computer vision, And social network analysis. One of the leading questions statisticians must answer when confronted with such “big data” is how to reduce a set of exponentially many predictors down to a set of a mere few predictors which have a truly causative effect on the response being modelled. This process is often referred to as feature screening. In this work we propose three new methods for feature screening. The first method we propose (TC-SIS) …


A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua Aug 2018

A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

The Random Forest method is a useful machine learning tool developed by Leo Breiman. There are many existing implementations across different programming languages; the most popular of which exist in R, SAS, and Python. In this paper, we conduct a comprehensive comparison of these implementations with regards to the accuracy, variable importance measurements, and timing. This comparison was done on a variety of real and simulated data with different classification difficulty levels, number of predictors, and sample sizes. The comparison shows unexpectedly different results between the three implementations.


Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen Aug 2018

Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Integrating real data into a classroom is one of the recommendations in the Guidelines for Assessment and Instruction in Statistics Education (GAISE) college report which lays out guidelines for an introductory statistics course (Committee, GAISE College Report ASA Revision, 2016). In order to assess the effect of using real data in a classroom, the students received physical activity trackers to wear during an undergraduate introductory statistics course taught in the summer. This tracker, a Fitbit, enabled students to monitor and record their steps, calories, and active time throughout the class. Collecting personal activity data (PAD) creates a large database which …


Pooling Of Variances: The Skeleton In The Mixed Model Closet?, Philip M. Dixon May 2018

Pooling Of Variances: The Skeleton In The Mixed Model Closet?, Philip M. Dixon

Conference on Applied Statistics in Agriculture and Natural Resources

I explore three related issues concerning pooling of error variances: when is it appropriate (or not) to pool, how best to evaluate equality of variances, and whether there is a cost to never pooling. I focus on pooling decisions in a combined analysis of a multi-site experiment. A-priori, sites should have different error variances. My primary question is whether an analysis that ignores unequal variances is wrong.

I find that ignoring heteroscedasticity between sites maintains, or provides slightly conservative, tests of average treatment effects and treatment-by-site interactions. Models with site-specific variances do provide more powerful tests when variances are different. …


Examining Quadratic Relationships Between Traits And Methods In Two Multitrait-Multimethod Models, Fredric A. Hintz May 2018

Examining Quadratic Relationships Between Traits And Methods In Two Multitrait-Multimethod Models, Fredric A. Hintz

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Psychological researchers are interested in the validity of the measures they use, and the multitrait-multimethod design is one of the most frequently employed methods to examine validity. Confirmatory factor analysis is now a commonly used analytic tool for examining multitrait-multimethod data, where an underlying mathematical model is fit to data and the amount of variance due to the trait and method factors is estimated. While most contemporary confirmatory factor analysis methods for examining multi-trait multi-method data do not allow relationships between the trait and method factors, a few recently proposed models allow for the examination of linear relationships between traits …


Mindset, Attitudes, And Success In Statistics, Matthew Isaac May 2018

Mindset, Attitudes, And Success In Statistics, Matthew Isaac

Undergraduate Honors Capstone Projects

Students in many disciplines are required to take an introductory statistics course while pursuing a college education. Despite the utility of statistical methods in future research and career pursuits, many students have negative views of statistics. We are interested in how students' mindsets and attitudes towards statistics impact their performance in an undergraduate statistics course. We administered a survey to students in several undergraduate statistics courses at Utah State University. This survey included questions addressing mathematics experience, attitudes towards statistics, mindset, and course performance. We observed that the majority of students indicated the presence of a growth mindset and positive …


Partitioning The Effects Of Eco-Evolutionary Feedbacks On Community Stability, Swati Patel, Michael H. Cortez, Sebastian J. Schreiber Mar 2018

Partitioning The Effects Of Eco-Evolutionary Feedbacks On Community Stability, Swati Patel, Michael H. Cortez, Sebastian J. Schreiber

Mathematics and Statistics Faculty Publications

A fundamental challenge in ecology continues to be identifying mechanisms that stabilize community dynamics. By altering the interactions within a community, eco-evolutionary feedbacks may play a role in community stability. Indeed, recent empirical and theoretical studies demonstrate that these feedbacks can stabilize or destabilize communities and, moreover, that this sometimes depends on the relative rate of ecological to evolutionary processes. So far, theory on how eco-evolutionary feedbacks impact stability exists only for a few special cases. In our work, we develop a general theory for determining the effects of eco-evolutionary feedbacks on stability in communities with an arbitrary number of …


Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell Dec 2017

Seasonal Resource Selection And Habitat Treatment Use By A Fringe Population Of Greater Sage-Grouse, Rhett Boswell

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Movement and habitat selection by Greater Sage-grouse (Centrocercus uropasianus) is of great interest to wildlife managers tasked with applying conservation measures for this iconic western species. Current technology has created small and lightweight GPS (Global Positioning Systems) transmitters that can be attached to sage-grouse. Using GIS software and statistical programs such as Program R, land managers can analyze GPS location data to assess how sage-grouse are geospatially interacting with their habitats. Within the Panguitch Sage-Grouse Management Area (SGMA) thousands of acres of land have been restored or manipulated to enhance sage-grouse habitat; this usually involves removal of pinyon pine …


Novel Statistical Models For Quantitative Shape-Gene Association Selection, Xiaotian Dai Dec 2017

Novel Statistical Models For Quantitative Shape-Gene Association Selection, Xiaotian Dai

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Other research reported that genetic mechanism plays a major role in the development process of biological shapes. The primary goal of this dissertation is to develop novel statistical models to investigate the quantitative relationships between biological shapes and genetic variants. However, these problems can be extremely challenging to traditional statistical models for a number of reasons: 1) the biological phenotypes cannot be effectively represented by single-valued traits, while traditional regression only handles one dependent variable; 2) in real-life genetic data, the number of candidate genes to be investigated is extremely large, and the signal-to-noise ratio of candidate genes is expected …


Extracting And Visualizing Data From Mobile And Static Eye Trackers In R And Matlab, Chunyang Li Dec 2017

Extracting And Visualizing Data From Mobile And Static Eye Trackers In R And Matlab, Chunyang Li

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Eye tracking is the process of measuring where people are looking at with an eye tracker device. Eye tracking has been used in many scientific fields, such as education, usability research, sports, psychology, and marketing. Eye tracking data are often obtained from a static eye tracker or are manually extracted from a mobile eye tracker. Visualization usually plays an important role in the analysis of eye tracking data. So far, there existed no software package that contains a whole collection of eye tracking data processing and visualization tools. In this dissertation, we review the eye tracking technology, the eye tracking …


Exact Approaches For Bias Detection And Avoidance With Small, Sparse, Or Correlated Categorical Data, Sarah E. Schwartz Dec 2017

Exact Approaches For Bias Detection And Avoidance With Small, Sparse, Or Correlated Categorical Data, Sarah E. Schwartz

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Every day, traditional statistical methodology are used world wide to study a variety of topics and provides insight regarding countless subjects. Each technique is based on a distinct set of assumptions to ensure valid results. Additionally, many statistical approaches rely on large sample behavior and may collapse or degenerate in the presence of small, spare, or correlated data. This dissertation details several advancements to detect these conditions, avoid their consequences, and analyze data in a different way to yield trustworthy results.

One of the most commonly used modeling techniques for outcomes with only two possible categorical values (eg. live/die, pass/fail, …


Application Of Machine Learning And Statistical Learning Methods For Prediction In A Large-Scale Vegetation Map, Carla M. Brookey Dec 2017

Application Of Machine Learning And Statistical Learning Methods For Prediction In A Large-Scale Vegetation Map, Carla M. Brookey

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Original analyses of a large vegetation cover dataset from Roosevelt National Forest in northern Colorado were carried out by Blackard (1998) and Blackard and Dean (1998; 2000). They compared the classification accuracies of linear and quadratic discriminant analysis (LDA and QDA) with artificial neural networks (ANN) and obtained an overall classification accuracy of 70.58% for a tuned ANN compared to 58.38% for LDA and 52.76% for QDA.

Because there has been tremendous development of machine learning classification methods over the last 35 years in both computer science and statistics, as well as substantial improvements in the speed of computer hardware, …


Using Data To Improve Services For Infants With Hearing Loss: Linking Newborn Hearing Screening Records With Early Intervention Records, Maria Gonzalez, Lori Iarossi, Yan Wu, Ying Huang, Kirsten Siegenthaler Nov 2017

Using Data To Improve Services For Infants With Hearing Loss: Linking Newborn Hearing Screening Records With Early Intervention Records, Maria Gonzalez, Lori Iarossi, Yan Wu, Ying Huang, Kirsten Siegenthaler

Journal of Early Hearing Detection and Intervention

The purpose of this study was to match records of infants with permanent hearing loss from the New York Early Hearing Detection and Intervention Information System (NYEHDI-IS) to records of infants with permanent hearing loss receiving early intervention services from the New York State Early Intervention Program (NYSEIP) to identify areas in the state where hearing screening, diagnostic evaluations and referrals to the NYSEIP were not being made or documented in a timely manner. Data from 2014-2016 NYEHDI-IS and NYEIS information systems were matched using The Link King. There were 274 infants documented in NYEIS Information System as receiving early …


A Bivariate Hypothesis Testing Approach For Mapping The Trait-Influential Gene, Garrett Saunders, Matthew D. Meng, John R. Stevens Oct 2017

A Bivariate Hypothesis Testing Approach For Mapping The Trait-Influential Gene, Garrett Saunders, Matthew D. Meng, John R. Stevens

Mathematics and Statistics Faculty Publications

The linkage disequilibrium (LD) based quantitative trait loci (QTL) model involves two indispensable hypothesis tests: the test of whether or not a QTL exists, and the test of the LD strength between the QTaL and the observed marker. The advantage of this two-test framework is to test whether there is an influential QTL around the observed marker instead of just having a QTL by random chance. There exist unsolved, open statistical questions about the inaccurate asymptotic distributions of the test statistics. We propose a bivariate null kernel (BNK) hypothesis testing method, which characterizes the joint distribution of the two test …


Prediction Of Stress Increase In Unbonded Tendons Using Sparse Principal Component Analysis, Eric Mckinney Aug 2017

Prediction Of Stress Increase In Unbonded Tendons Using Sparse Principal Component Analysis, Eric Mckinney

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

While internal and external unbonded tendons are widely utilized in concrete structures, the analytic solution for the increase in unbonded tendon stress, Δ���, is challenging due to the lack of bond between strand and concrete. Moreover, most analysis methods do not provide high correlation due to the limited available test data. In this thesis, Principal Component Analysis (PCA), and Sparse Principal Component Analysis (SPCA) are employed on different sets of candidate variables, amongst the material and sectional properties from the database compiled by Maguire et al. [18]. Predictions of Δ��� are made via Principal Component Regression models, and the method …


Imputation For Random Forests, Joshua Young Aug 2017

Imputation For Random Forests, Joshua Young

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

This project introduces two new methods for imputation of missing data in random forests. The new methods are compared against other frequently used imputation methods, including those used in the randomForest package in R. To test the effectiveness of these methods, missing data are imputed into datasets that contain two missing data mechanisms including missing at random and missing completely at random. After imputation, random forests are run on the data and accuracies for the predictions are obtained. Speed is an important aspect in computing; the speeds for all the tested methods are also compared.

One of the new methods …


Tree-Based Regression For Interval-Valued Data, Chih-Ching Yeh Aug 2017

Tree-Based Regression For Interval-Valued Data, Chih-Ching Yeh

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Regression methods for interval-valued data have been increasingly studied in recent years. As most of the existing works focus on linear models, it is important to note that many problems in practice are nonlinear in nature and therefore development of nonlinear regression tools for intervalvalued data is crucial. In this project, we propose a tree-based regression method for interval-valued data, which is well applicable to both linear and nonlinear problems. Unlike linear regression models that usually require additional constraints to ensure positivity of the predicted interval length, the proposed method estimates the regression function in a nonparametric way, so the …


A Comparison Of Five Statistical Methods For Predicting Stream Temperature Across Stream Networks, Maike F. Holthuijzen Aug 2017

A Comparison Of Five Statistical Methods For Predicting Stream Temperature Across Stream Networks, Maike F. Holthuijzen

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

The health of freshwater aquatic systems, particularly stream networks, is mainly influenced by water temperature, which controls biological processes and influences species distributions and aquatic biodiversity. Thermal regimes of rivers are likely to change in the future, due to climate change and other anthropogenic impacts, and our ability to predict stream temperatures will be critical in understanding distribution shifts of aquatic biota. Spatial statistical network models take into account spatial relationships but have drawbacks, including high computation times and data pre-processing requirements. Machine learning techniques and generalized additive models (GAM) are promising alternatives to the SSN model. Two machine learning …