Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,821 Full-Text Articles 23,918 Authors 9,922,835 Downloads 283 Institutions

All Articles in Statistics and Probability

Faceted Search

12,821 full-text articles. Page 212 of 487.

Mentoring Undergraduate Research In Statistics: Reaping The Benefits And Overcoming The Barriers, Joseph R. Nolan, Kelly S. McConville, Vittorio Addona, Nathan L. Tintle, Dennis K. Pearl 2020 Northern Kentucky University

Mentoring Undergraduate Research In Statistics: Reaping The Benefits And Overcoming The Barriers, Joseph R. Nolan, Kelly S. Mcconville, Vittorio Addona, Nathan L. Tintle, Dennis K. Pearl

Faculty Work Comprehensive List

Undergraduate research experiences (UREs), whether within the context of a mentor-mentee experience or a classroom framework, represent an excellent opportunity to expose students to the independent scholarship model. The high impact of undergraduate research has received recent attention in the context of STEM disciplines. Reflecting a 2017 survey of statistics faculty, this article examines the perceived benefits of UREs, as well as barriers to the incorporation of UREs, specifically within the field of statistics. Viewpoints of students, faculty mentors, and institutions are investigated. Further, the article offers several strategies for leveraging characteristics unique to the field of statistics to overcome …


Forecasting Daily Stock Market Return With Multiple Linear Regression, Shengxuan Chen 2020 Louisiana Tech University

Forecasting Daily Stock Market Return With Multiple Linear Regression, Shengxuan Chen

Mathematics Senior Capstone Papers

The purpose of this project is to use data mining and big data analytic techniques to forecast daily stock market return with multiple linear regression. Using mathematical and statistical models to analyze the stock market is important and challenging. The accuracy of the final results relies on the quality of the input data and the validity of the methodology. In the report, within 5-year period, the data regarding eleven financial and economical features are observed and recorded on each trading day. After preprocessing the raw data with statistical method, we use the multiple linear regression to predict the daily return …


Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang 2020 Louisiana State University

Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang

LSU Doctoral Dissertations

Large volumes of temporal event data, such as online check-ins and electronic records of hospital admissions, are becoming increasingly available in a wide variety of applications including healthcare analytics, smart cities, and social network analysis. Those temporal events are often asynchronous, interdependent, and exhibiting self-exciting properties. For example, in the patient's diagnosis events, the elevated risk exists for a patient that has been recently at risk. Machine learning that leverages event sequence data can improve the prediction accuracy of future events and provide valuable services. For example, in e-commerce and network traffic diagnosis, the analysis of user activities can be …


Decision Tree For Predicting The Party Of Legislators, Afsana Mimi 2020 CUNY New York City College of Technology

Decision Tree For Predicting The Party Of Legislators, Afsana Mimi

Publications and Research

The motivation of the project is to identify the legislators who voted frequently against their party in terms of their roll call votes using Office of Clerk U.S. House of Representatives Data Sets collected in 2018 and 2019. We construct a model to predict the parties of legislators based on their votes. The method we used is Decision Tree from Data Mining. Python was used to collect raw data from internet, SAS was used to clean data, and all other calculations and graphical presentations are performed using the R software.


A Statistical Analysis Of The Unm Facets Design Identity & Beliefs Survey Data, Clarissa A. Sorensen-Unruh 2020 University of New Mexico - Main Campus

A Statistical Analysis Of The Unm Facets Design Identity & Beliefs Survey Data, Clarissa A. Sorensen-Unruh

Mathematics & Statistics ETDs

The NSF-funded FACETS (Formation of Accomplished Chemical Engineers for Transforming Society, NSF Award 1623105) grant aims to transform the undergraduate engineering experience in the Department of Chemical and Biological Engineering at the University of New Mexico to address attrition within engineering majors, especially among underserved populations (Brainard & Carlin, 1998). The UNM FACETS Design Identity & Beliefs survey, an assessment tool used as part of the research of the grant, generated the dataset used in this study. I performed several different statistical analyses on the dataset, including confirmatory factor analysis (CFA), principal component analysis (PCA), and cluster analysis. The …


Statistical Analysis Of Land Cover Conversion Trends In Northwest Ohio, Chaska McGowan 2020 Bowling Green State University

Statistical Analysis Of Land Cover Conversion Trends In Northwest Ohio, Chaska Mcgowan

Honors Projects

Agricultural land in the U.S. is abundant but not infinite. Change in cropland impacts national and local economies and the natural environment. The Black Swamp Conservancy (BSC), a non-profit land trust in Perrysburg, Ohio, is committed to preserving agricultural and natural lands in the Northwest Ohio region for future generations. This project was designed in collaboration with the BSC to illustrate the spatial distribution of land cover change within their sixteen-county service area in Northwest Ohio and to find a list of factors associated with land cover change in the region. The primary data source was the National Land Cover …


Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder 2020 James Madison University

Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder

Dissertations, 2020-current

The Rasch model is commonly used to calibrate multiple choice items. However, the sample sizes needed to estimate the Rasch model can be difficult to attain (e.g., consider a small testing company trying to pretest new items). With small sample sizes, auxiliary information besides the item responses may improve estimation of the item parameters. The purpose of this study was to determine if incorporating item property information (i.e., characteristics of the items related to item difficulty) in a random effects linear logistic test model (RE-LLTM) would improve estimation of item difficulty. A simulation study was conducted that varied sample size, …


Modeling Species Distribution And Habitat Suitability Of American Ginseng (Panax Quinquefolius) In Virginia, Jacob D. J. Peters 2020 James Madison University

Modeling Species Distribution And Habitat Suitability Of American Ginseng (Panax Quinquefolius) In Virginia, Jacob D. J. Peters

Masters Theses, 2020-current

American ginseng (Panax quinquefolius) is a well-known and sought-after medicinal plant native to North America that is facing increased threat of extinction due to overharvesting, herbivory, and habitat loss. Species distribution and habitat suitability models may be valuable to landowners interested in sustainable harvest or to institutions interested in the conservation and restoration of the species. With unequal sampling efforts across a region of interest, it is likely that some locations with appropriate habitat may be misrepresented in model predictions. This study refined a state-derived species distribution model for ginseng through increased sampling effort across the Cumberland Plateau …


Splitting Up A Complex Mess: The Effectiveness Of Statistical Analysis On Delimiting Species Complexes, Sara N. Schoen 2020 James Madison University

Splitting Up A Complex Mess: The Effectiveness Of Statistical Analysis On Delimiting Species Complexes, Sara N. Schoen

Masters Theses, 2020-current

Recent studies have highlighted a need for more refined tools in species delimitation. This is especially true when considering diversity within species complexes, where members are morphologically similar and traditional tools have thus far failed to provide clearly defined boundaries between species. This project seeks to refine our traditional tools of species delimitation and apply new tools to the challenges created by species complexes. The focus organisms of this study are the anurans of the Limnonectes kuhlii complex. This species complex comprises more than 25 species of stream frogs from Southeast Asia. Traditionally, morphometrics (particularly linear measures) has been the …


Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig 2020 James Madison University

Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig

Masters Theses, 2020-current

In the absence of random assignment, researchers must consider the impact of selection bias – pre-existing covariate differences between groups due to differences among those entering into treatment and those otherwise unable to participate. Propensity score matching (PSM) and generalized boosted modeling (GBM) are two quasi-experimental pre-processing methods that strive to reduce the impact of selection bias before analyzing a treatment effect. PSM and GBM both examine a treatment and comparison group and either match or weight members of those groups to create new, balanced groups. The new, balanced groups theoretically can then be used as a proxy for the …


Attack And Defense In Security Analytics, Yiyun Zhou 2020 Kennesaw State University

Attack And Defense In Security Analytics, Yiyun Zhou

Doctor of Data Science and Analytics Dissertations

The security problem has gained increasing awareness due to the various kinds of global threats. Security analytics is the process of using streaming data acquisition, collection, and artificial intelligence algorithms for security monitoring and threat disclosure. In this dissertation work, we utilize practical data-driven security analytics to identify the potential threat and explore the robustness of the machine learning model. We focus on two aspects: (1) Security Analytics: utilize machine learning and statistical analytics tools to identify and resolve the threat in real life, such as cybersecurity, abnormal activities. (2) Analytic Security: Explore the security issues of the machine learning …


Age At Migration And The Risk Of Psychotic Disorders: A Systematic Review And Meta-Analysis., Kelly K. Anderson, Jordan Edwards 2020 Western University

Age At Migration And The Risk Of Psychotic Disorders: A Systematic Review And Meta-Analysis., Kelly K. Anderson, Jordan Edwards

Epidemiology and Biostatistics Publications

OBJECTIVE: To conduct a systematic review and meta-analysis of the existing evidence on the association between age at migration and the risk of psychotic disorders.

METHODS: Observational studies were eligible for inclusion if they presented data on the association between age at migration and the risk of psychotic disorders among first-generation migrant groups. We used two random effects meta-analyses to pool effect estimates for each stratum of age at migration relative to (i) a native-born reference category and (ii) the youngest age stratum (0 to 2 years).

RESULTS: Ten studies met inclusion criteria, and five were included in the meta-analysis. …


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya 2020 University of Louisville

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield 2020 Stephen F. Austin State University

Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield

Electronic Theses and Dissertations

The prevalence of conducting statistical inference for the mean of the beta distribution has been rising in various fields of academic research, such as in immunology that analyzes proportions of rare cell population subsets. For our purposes, we will address this statistical inference problem by using likelihood-based applications to hypothesis testing, along with a relatively new statistical method called saddlepoint approximations. Through simulation work, we will compare the performance of these statistical procedures and provide both the statistical and scientific communities with recommendations on best practices.


Predicting Disease Progression Using Deep Recurrent Neural Networks And Longitudinal Electronic Health Record Data, Seunghwan Kim 2020 Washington University in St. Louis

Predicting Disease Progression Using Deep Recurrent Neural Networks And Longitudinal Electronic Health Record Data, Seunghwan Kim

McKelvey School of Engineering Graduate Student Theses & Dissertations

Electronic Health Records (EHR) are widely adopted and used throughout healthcare systems and are able to collect and store longitudinal information data that can be used to describe patient phenotypes. From the underlying data structures used in the EHR, discrete data can be extracted and analyzed to improve patient care and outcomes via tasks such as risk stratification and prospective disease management. Temporality in EHR is innately present given the nature of these data, however, and traditional classification models are limited in this context by the cross- sectional nature of training and prediction processes. Finding temporal patterns in EHR is …


Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula 2020 University of Louisville

Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula

Electronic Theses and Dissertations

With single-cell RNA sequencing (scRNA-seq) technology, researchers are able to gain a better understanding of health and disease through the analysis of gene expression data at the cellular-level; however, scRNA-seq data tend to have high proportions of zero values, increased cell-to-cell variability, and overdispersion due to abnormally large expression counts, which create new statistical problems that need to be addressed. This dissertation includes three research projects that propose Bayesian methodology suitable for scRNA-seq analysis. In the first project, a hurdle model for identifying differentially expressed genes across cell types in scRNA-seq data is presented. This model incorporates a correlated random …


Using Stability To Select A Shrinkage Method, Dean Dustin 2020 University of Nebraska-Lincoln

Using Stability To Select A Shrinkage Method, Dean Dustin

Department of Statistics: Dissertations, Theses, and Student Research

Shrinkage methods are estimation techniques based on optimizing expressions to find which variables to include in an analysis, typically a linear regression. The general form of these expressions is the sum of an empirical risk plus a complexity penalty based on the number of parameters. Many shrinkage methods are known to satisfy an ‘oracle’ property meaning that asymptotically they select the correct variables and estimate their coefficients efficiently. In Section 1.2, we show oracle properties in two general settings. The first uses a log likelihood in place of the empirical risk and allows a general class of penalties. The second …


Physical Therapy Nontreatment Events With Primary Physical Therapist, Stephen Johnson 2020 University of Nevada, Las Vegas

Physical Therapy Nontreatment Events With Primary Physical Therapist, Stephen Johnson

UNLV Theses, Dissertations, Professional Papers, and Capstones

Background: Physical therapy improves prognosis reduces stay and is generally helpful in aiding recovery from a wide range of ailments. Nontreatment rates occur for multiple reasons and are also related to the personalities of physical therapists.

Methods: We used data from a research project involving physical therapy at an acute care facility in our community. Our study focused on the retrospectively determined primary physical therapist for each patient. We used the chi-squared tests to compare nontreatment rates between days of the week and disease type and the reasons for nontreatment events. Repeated-measure models were used to evaluate the effect of …


Life And Death: Quantifying The Risk Of Heart Disease With Machine Learning, Jack Scott Glienke 2020 University of Northern Iowa

Life And Death: Quantifying The Risk Of Heart Disease With Machine Learning, Jack Scott Glienke

Honors Program Theses

Coronary heart disease has long been a key area of focus in the discussion of public health. As such, numerous studies have been conducted throughout history with the sole intention of identifying risk factors leading to the onset of cardiovascular conditions. A plethora of statistical procedures can be used to identify an individual’s risk of developing heart disease, yet regression models tend to be the default tool used by researchers. Using the data obtained from the most influential cardiovascular study to date, the Framingham Heart Study, this analysis uses machine learning techniques to generate and test the predictive power of …


Analyzing Competitive Balance In Professional Sport, Kevin Alwell 2020 University of Connecticut - Storrs

Analyzing Competitive Balance In Professional Sport, Kevin Alwell

Honors Scholar Theses

In this paper we review several measures to statistically analyze competitive balance and report which leagues have a wider variance of performance amongst its competitors. Each league seeks to maintain high levels of parity, making matches and overall season more unpredictable and appealing to the general audience. Here we quantify competitive advantage across major sports leagues in numbers using several statistical methods in order for leagues to optimize their revenue.


Digital Commons powered by bepress