Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 301 - 330 of 618

Full-Text Articles in Statistics and Probability

Mentoring Undergraduate Research In Statistics: Reaping The Benefits And Overcoming The Barriers, Joseph R. Nolan, Kelly S. Mcconville, Vittorio Addona, Nathan L. Tintle, Dennis K. Pearl May 2020

Mentoring Undergraduate Research In Statistics: Reaping The Benefits And Overcoming The Barriers, Joseph R. Nolan, Kelly S. Mcconville, Vittorio Addona, Nathan L. Tintle, Dennis K. Pearl

Faculty Work Comprehensive List

Undergraduate research experiences (UREs), whether within the context of a mentor-mentee experience or a classroom framework, represent an excellent opportunity to expose students to the independent scholarship model. The high impact of undergraduate research has received recent attention in the context of STEM disciplines. Reflecting a 2017 survey of statistics faculty, this article examines the perceived benefits of UREs, as well as barriers to the incorporation of UREs, specifically within the field of statistics. Viewpoints of students, faculty mentors, and institutions are investigated. Further, the article offers several strategies for leveraging characteristics unique to the field of statistics to overcome …


Forecasting Daily Stock Market Return With Multiple Linear Regression, Shengxuan Chen May 2020

Forecasting Daily Stock Market Return With Multiple Linear Regression, Shengxuan Chen

Mathematics Senior Capstone Papers

The purpose of this project is to use data mining and big data analytic techniques to forecast daily stock market return with multiple linear regression. Using mathematical and statistical models to analyze the stock market is important and challenging. The accuracy of the final results relies on the quality of the input data and the validity of the methodology. In the report, within 5-year period, the data regarding eleven financial and economical features are observed and recorded on each trading day. After preprocessing the raw data with statistical method, we use the multiple linear regression to predict the daily return …


Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang May 2020

Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang

LSU Doctoral Dissertations

Large volumes of temporal event data, such as online check-ins and electronic records of hospital admissions, are becoming increasingly available in a wide variety of applications including healthcare analytics, smart cities, and social network analysis. Those temporal events are often asynchronous, interdependent, and exhibiting self-exciting properties. For example, in the patient's diagnosis events, the elevated risk exists for a patient that has been recently at risk. Machine learning that leverages event sequence data can improve the prediction accuracy of future events and provide valuable services. For example, in e-commerce and network traffic diagnosis, the analysis of user activities can be …


A Statistical Analysis Of The Unm Facets Design Identity & Beliefs Survey Data, Clarissa A. Sorensen-Unruh May 2020

A Statistical Analysis Of The Unm Facets Design Identity & Beliefs Survey Data, Clarissa A. Sorensen-Unruh

Mathematics & Statistics ETDs

The NSF-funded FACETS (Formation of Accomplished Chemical Engineers for Transforming Society, NSF Award 1623105) grant aims to transform the undergraduate engineering experience in the Department of Chemical and Biological Engineering at the University of New Mexico to address attrition within engineering majors, especially among underserved populations (Brainard & Carlin, 1998). The UNM FACETS Design Identity & Beliefs survey, an assessment tool used as part of the research of the grant, generated the dataset used in this study. I performed several different statistical analyses on the dataset, including confirmatory factor analysis (CFA), principal component analysis (PCA), and cluster analysis. The …


Decision Tree For Predicting The Party Of Legislators, Afsana Mimi May 2020

Decision Tree For Predicting The Party Of Legislators, Afsana Mimi

Publications and Research

The motivation of the project is to identify the legislators who voted frequently against their party in terms of their roll call votes using Office of Clerk U.S. House of Representatives Data Sets collected in 2018 and 2019. We construct a model to predict the parties of legislators based on their votes. The method we used is Decision Tree from Data Mining. Python was used to collect raw data from internet, SAS was used to clean data, and all other calculations and graphical presentations are performed using the R software.


Statistical Analysis Of Land Cover Conversion Trends In Northwest Ohio, Chaska Mcgowan May 2020

Statistical Analysis Of Land Cover Conversion Trends In Northwest Ohio, Chaska Mcgowan

Honors Projects

Agricultural land in the U.S. is abundant but not infinite. Change in cropland impacts national and local economies and the natural environment. The Black Swamp Conservancy (BSC), a non-profit land trust in Perrysburg, Ohio, is committed to preserving agricultural and natural lands in the Northwest Ohio region for future generations. This project was designed in collaboration with the BSC to illustrate the spatial distribution of land cover change within their sixteen-county service area in Northwest Ohio and to find a list of factors associated with land cover change in the region. The primary data source was the National Land Cover …


Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder May 2020

Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder

Dissertations, 2020-current

The Rasch model is commonly used to calibrate multiple choice items. However, the sample sizes needed to estimate the Rasch model can be difficult to attain (e.g., consider a small testing company trying to pretest new items). With small sample sizes, auxiliary information besides the item responses may improve estimation of the item parameters. The purpose of this study was to determine if incorporating item property information (i.e., characteristics of the items related to item difficulty) in a random effects linear logistic test model (RE-LLTM) would improve estimation of item difficulty. A simulation study was conducted that varied sample size, …


Modeling Species Distribution And Habitat Suitability Of American Ginseng (Panax Quinquefolius) In Virginia, Jacob D. J. Peters May 2020

Modeling Species Distribution And Habitat Suitability Of American Ginseng (Panax Quinquefolius) In Virginia, Jacob D. J. Peters

Masters Theses, 2020-current

American ginseng (Panax quinquefolius) is a well-known and sought-after medicinal plant native to North America that is facing increased threat of extinction due to overharvesting, herbivory, and habitat loss. Species distribution and habitat suitability models may be valuable to landowners interested in sustainable harvest or to institutions interested in the conservation and restoration of the species. With unequal sampling efforts across a region of interest, it is likely that some locations with appropriate habitat may be misrepresented in model predictions. This study refined a state-derived species distribution model for ginseng through increased sampling effort across the Cumberland Plateau …


Splitting Up A Complex Mess: The Effectiveness Of Statistical Analysis On Delimiting Species Complexes, Sara N. Schoen May 2020

Splitting Up A Complex Mess: The Effectiveness Of Statistical Analysis On Delimiting Species Complexes, Sara N. Schoen

Masters Theses, 2020-current

Recent studies have highlighted a need for more refined tools in species delimitation. This is especially true when considering diversity within species complexes, where members are morphologically similar and traditional tools have thus far failed to provide clearly defined boundaries between species. This project seeks to refine our traditional tools of species delimitation and apply new tools to the challenges created by species complexes. The focus organisms of this study are the anurans of the Limnonectes kuhlii complex. This species complex comprises more than 25 species of stream frogs from Southeast Asia. Traditionally, morphometrics (particularly linear measures) has been the …


Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig May 2020

Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig

Masters Theses, 2020-current

In the absence of random assignment, researchers must consider the impact of selection bias – pre-existing covariate differences between groups due to differences among those entering into treatment and those otherwise unable to participate. Propensity score matching (PSM) and generalized boosted modeling (GBM) are two quasi-experimental pre-processing methods that strive to reduce the impact of selection bias before analyzing a treatment effect. PSM and GBM both examine a treatment and comparison group and either match or weight members of those groups to create new, balanced groups. The new, balanced groups theoretically can then be used as a proxy for the …


Attack And Defense In Security Analytics, Yiyun Zhou May 2020

Attack And Defense In Security Analytics, Yiyun Zhou

Doctor of Data Science and Analytics Dissertations

The security problem has gained increasing awareness due to the various kinds of global threats. Security analytics is the process of using streaming data acquisition, collection, and artificial intelligence algorithms for security monitoring and threat disclosure. In this dissertation work, we utilize practical data-driven security analytics to identify the potential threat and explore the robustness of the machine learning model. We focus on two aspects: (1) Security Analytics: utilize machine learning and statistical analytics tools to identify and resolve the threat in real life, such as cybersecurity, abnormal activities. (2) Analytic Security: Explore the security issues of the machine learning …


Age At Migration And The Risk Of Psychotic Disorders: A Systematic Review And Meta-Analysis., Kelly K. Anderson, Jordan Edwards May 2020

Age At Migration And The Risk Of Psychotic Disorders: A Systematic Review And Meta-Analysis., Kelly K. Anderson, Jordan Edwards

Epidemiology and Biostatistics Publications

OBJECTIVE: To conduct a systematic review and meta-analysis of the existing evidence on the association between age at migration and the risk of psychotic disorders.

METHODS: Observational studies were eligible for inclusion if they presented data on the association between age at migration and the risk of psychotic disorders among first-generation migrant groups. We used two random effects meta-analyses to pool effect estimates for each stratum of age at migration relative to (i) a native-born reference category and (ii) the youngest age stratum (0 to 2 years).

RESULTS: Ten studies met inclusion criteria, and five were included in the meta-analysis. …


Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield May 2020

Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield

Electronic Theses and Dissertations

The prevalence of conducting statistical inference for the mean of the beta distribution has been rising in various fields of academic research, such as in immunology that analyzes proportions of rare cell population subsets. For our purposes, we will address this statistical inference problem by using likelihood-based applications to hypothesis testing, along with a relatively new statistical method called saddlepoint approximations. Through simulation work, we will compare the performance of these statistical procedures and provide both the statistical and scientific communities with recommendations on best practices.


An Analysis Of Dredge Efficiency For Surfclam And Ocean Quahog Commercial Dredges, Leanne Poussard May 2020

An Analysis Of Dredge Efficiency For Surfclam And Ocean Quahog Commercial Dredges, Leanne Poussard

Master's Theses

Between 1997 and 2011, The National Marine Fisheries Service conducted 50 depletion experiments to estimate survey gear efficiency and stock density for Atlantic surfclam (Spisula solidissima) and ocean quahog (Arctica islandica) populations using commercial hydraulic dredges. The Patch Model was formulated to estimate gear efficiency and organism density from the data. The range of efficiencies estimated is substantial, leading to uncertainty in the application of these estimates in stock assessment. Analysis of depletion experiment simulations showed that uncertainty in the estimates of gear efficiency from depletion experiments was reduced by higher numbers of dredge tows per experiment, more tow overlap …


Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula May 2020

Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula

Electronic Theses and Dissertations

With single-cell RNA sequencing (scRNA-seq) technology, researchers are able to gain a better understanding of health and disease through the analysis of gene expression data at the cellular-level; however, scRNA-seq data tend to have high proportions of zero values, increased cell-to-cell variability, and overdispersion due to abnormally large expression counts, which create new statistical problems that need to be addressed. This dissertation includes three research projects that propose Bayesian methodology suitable for scRNA-seq analysis. In the first project, a hurdle model for identifying differentially expressed genes across cell types in scRNA-seq data is presented. This model incorporates a correlated random …


Using Stability To Select A Shrinkage Method, Dean Dustin May 2020

Using Stability To Select A Shrinkage Method, Dean Dustin

Department of Statistics: Dissertations, Theses, and Student Research

Shrinkage methods are estimation techniques based on optimizing expressions to find which variables to include in an analysis, typically a linear regression. The general form of these expressions is the sum of an empirical risk plus a complexity penalty based on the number of parameters. Many shrinkage methods are known to satisfy an ‘oracle’ property meaning that asymptotically they select the correct variables and estimate their coefficients efficiently. In Section 1.2, we show oracle properties in two general settings. The first uses a log likelihood in place of the empirical risk and allows a general class of penalties. The second …


Predicting Disease Progression Using Deep Recurrent Neural Networks And Longitudinal Electronic Health Record Data, Seunghwan Kim May 2020

Predicting Disease Progression Using Deep Recurrent Neural Networks And Longitudinal Electronic Health Record Data, Seunghwan Kim

McKelvey School of Engineering Graduate Student Theses & Dissertations

Electronic Health Records (EHR) are widely adopted and used throughout healthcare systems and are able to collect and store longitudinal information data that can be used to describe patient phenotypes. From the underlying data structures used in the EHR, discrete data can be extracted and analyzed to improve patient care and outcomes via tasks such as risk stratification and prospective disease management. Temporality in EHR is innately present given the nature of these data, however, and traditional classification models are limited in this context by the cross- sectional nature of training and prediction processes. Finding temporal patterns in EHR is …


Analyzing Competitive Balance In Professional Sport, Kevin Alwell May 2020

Analyzing Competitive Balance In Professional Sport, Kevin Alwell

Honors Scholar Theses

In this paper we review several measures to statistically analyze competitive balance and report which leagues have a wider variance of performance amongst its competitors. Each league seeks to maintain high levels of parity, making matches and overall season more unpredictable and appealing to the general audience. Here we quantify competitive advantage across major sports leagues in numbers using several statistical methods in order for leagues to optimize their revenue.


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya May 2020

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


The Effects Of Zoledronate And Sleep Deprivation On The Distal Femur Trabecular Thickness Of Ovariectomized Rats: Application Of Different Statistical Methods, Erin Nolte May 2020

The Effects Of Zoledronate And Sleep Deprivation On The Distal Femur Trabecular Thickness Of Ovariectomized Rats: Application Of Different Statistical Methods, Erin Nolte

Student Scholar Symposium Abstracts and Posters

Osteoporosis is a disease that causes the degradation of bone, leading to an increased risk of fracture. 1 in 3 women over the age of 50 will be affected by Osteoporosis. This study aims to understand how bone is affected by sleep deprivation in estrogen-deficient rats, and how Zoledronate might negate the inimical effects of sleep deprivation on bone. As bone mineral density (BMD) is a crude evaluation of the architectural changes seen in Osteoporosis, trabecular thickness may serve as a better single evaluation of bone health. 31 Wistar female rats were ovariectomized and separated into 4 random groups. The …


Simplicity As A New Environmental Virtue, Justin Wheeler May 2020

Simplicity As A New Environmental Virtue, Justin Wheeler

Undergraduate Honors Capstone Projects

This paper argues for the addition of a new environmentally focused virtue, simplicity, to the virtue ethical framework developed by Aristotle. First, relevant background from Aristotle’s virtue ethics are developed including the crucial, “doctrine of the mean”, a balance between excess and deficiency of a specified character trait. The tenets of the new virtue simplicity are developed with practical examples based on Aristotle’s method of developing a virtue of character. Simplicity is proposed as a desire to take the appropriate amount from the natural world and an acceptance of one’s circumstances. Those possessing simplicity will not fall victim to the …


Analysis Of Sat And Isat Scores For Madison School District In Rexburg, Idaho, Holly Dawn Palmer May 2020

Analysis Of Sat And Isat Scores For Madison School District In Rexburg, Idaho, Holly Dawn Palmer

Undergraduate Honors Capstone Projects

Testing is an integral part of measuring education. If used properly SAT scores can be compared across the nation, and statewide tests can compare different school districts to each other if done properly to avoid certain pitfalls (Fetler, 1991). However, if tests do not have a significant impact on a student, their motivation to take the test will be low and test quality cannot be assumed. When the state funds two separate tests for their students but only one has a significant impact on the student, how should the scores for each test be used, and is it okay to …


The Two Types Of Society: Computationally Revealing Recurrent Social Formations And Their Evolutionary Trajectories, Lux Miranda May 2020

The Two Types Of Society: Computationally Revealing Recurrent Social Formations And Their Evolutionary Trajectories, Lux Miranda

Undergraduate Honors Capstone Projects

Comparative social science has a long history of attempts to classify societies and cultures in terms of shared characteristics. However, only recently has it become feasible to conduct quantitative analysis of large historical datasets to mathematically approach the study of social complexity and classify shared societal characteristics. Such methods have the potential to identify recurrent social formations in human societies and contribute to social evolutionary theory. However, in order to achieve this potential, repeated studies are needed to assess the robustness of results to changing methods and data sets. Using an improved derivative of the Seshat: Global History Databank, we …


Demystification Of Graph And Information Entropy, Bryce Frederickson May 2020

Demystification Of Graph And Information Entropy, Bryce Frederickson

Undergraduate Honors Capstone Projects

Shannon entropy is an information-theoretic measure of unpredictability in probabilistic models. Recently, it has been used to form a tool, called the von Neumann entropy, to study quantum mechanics and network flows by appealing to algebraic properties of graph matrices. But still, little is known about what the von Neumann entropy says about the combinatorial structure of the graphs themselves. This paper gives a new formulation of the von Neumann entropy that describes it as a rate at which random movement settles down in a graph. At the same time, this new perspective gives rise to a generalization of von …


Equivalency Testing For Two Formulations Of A Clinical Laboratory Control Material, Jessica M. Hart May 2020

Equivalency Testing For Two Formulations Of A Clinical Laboratory Control Material, Jessica M. Hart

Capstone Experience: Master of Public Health

Clinical laboratory control materials are an integral part of legally-mandated and highly regulated quality control protocols in all clinical laboratories. These controls ensure accurate performance of the laboratory testing and instrumentation used to produce medical test results for millions of patients. It is of clinical and public health interest to ensure the diagnostic test results which affect so many people are regulated by the most accurate and precise controls.

Formulation changes in control materials have the potential to impact laboratory quality control. In this study, data from two formulations of a hematology control were compared to assess equivalency of the …


Predicting The Federal Funds Rate, Danielle Herzberg May 2020

Predicting The Federal Funds Rate, Danielle Herzberg

Undergraduate Theses and Capstone Projects

This thesis examines various economic indicators to select those that are the most significant in a predictive model of the Effective Federal Funds Rate. Three different statistical models were built to show how monetary policy changed over time. These three models frame the last economic downturns in the United States; the tech bubble, the housing bubble, and the Great Recession. Many iterations of statistical regressions were conducted in order to achieve the final three models that highlight variables with the highest levels of significance. It is important to note the economic data has high levels of autocorrelation, and that these …


Applications Of Machine Learning In High-Frequency Trade Direction Classification, Jared E. Hansen May 2020

Applications Of Machine Learning In High-Frequency Trade Direction Classification, Jared E. Hansen

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

The correct assignment of trades as buyer-initiated or seller-initiated is paramount in many quantitative finance studies. Simple decision rule methods have been used for signing trades since many data sets available to researchers do not include the sign of each trade executed. By utilizing these decision rule methods, as well as engineering new variables from available data, we have demonstrated that machine learning models outperform prior methods for accurately signing trades as buys and sells, achieving state-of-the-art results. The best model developed was 4.5 percentage points more accurate than older methods when predicting onto unseen data. Since finance and economics …


On Arnold–Villasenor Conjectures For Characterizaing Exponential Distribution Based On Sample Of Size Three, George Yanev May 2020

On Arnold–Villasenor Conjectures For Characterizaing Exponential Distribution Based On Sample Of Size Three, George Yanev

School of Mathematical & Statistical Sciences Faculty Publications

Arnold and Villasenor [4] obtain a series of characterizations of the exponential distribution based on random samples of size two. These results were already applied in constructing goodness-of-fit tests. Extending the techniques from [4], we prove some of Arnold and Villasenor’s conjectures for samples of size three. An example with simulated data is discussed.


Physical Therapy Nontreatment Events With Primary Physical Therapist, Stephen Johnson May 2020

Physical Therapy Nontreatment Events With Primary Physical Therapist, Stephen Johnson

UNLV Theses, Dissertations, Professional Papers, and Capstones

Background: Physical therapy improves prognosis reduces stay and is generally helpful in aiding recovery from a wide range of ailments. Nontreatment rates occur for multiple reasons and are also related to the personalities of physical therapists.

Methods: We used data from a research project involving physical therapy at an acute care facility in our community. Our study focused on the retrospectively determined primary physical therapist for each patient. We used the chi-squared tests to compare nontreatment rates between days of the week and disease type and the reasons for nontreatment events. Repeated-measure models were used to evaluate the effect of …


Life And Death: Quantifying The Risk Of Heart Disease With Machine Learning, Jack Scott Glienke May 2020

Life And Death: Quantifying The Risk Of Heart Disease With Machine Learning, Jack Scott Glienke

Honors Program Theses

Coronary heart disease has long been a key area of focus in the discussion of public health. As such, numerous studies have been conducted throughout history with the sole intention of identifying risk factors leading to the onset of cardiovascular conditions. A plethora of statistical procedures can be used to identify an individual’s risk of developing heart disease, yet regression models tend to be the default tool used by researchers. Using the data obtained from the most influential cardiovascular study to date, the Framingham Heart Study, this analysis uses machine learning techniques to generate and test the predictive power of …