Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,818 Full-Text Articles 23,898 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,818 full-text articles. Page 181 of 486.

Joint Spacing In The Caples Lake Granodiorite Of The Sierra Nevada Batholith In Eldorado National Forest, California: A Comparative Analysis Of Joint Sets And Data Resolution, Jimmy Wood 2021 University of Nebraska at Omaha

Joint Spacing In The Caples Lake Granodiorite Of The Sierra Nevada Batholith In Eldorado National Forest, California: A Comparative Analysis Of Joint Sets And Data Resolution, Jimmy Wood

Theses/Capstones/Creative Projects

Joints are the most common deformation structure in the Earth’s upper crust and exert a significant influence on structural stability, landscape morphology, and fluid flow . Therefore, a greater understanding of fracture parameters (e.g., length, aperture, etc.) allows us to more accurately predict their presence, persistence, and prevalence, in the subsurface . We study the fracture spacing of two sub-orthogonal joint sets—66 NE-246 SW and 330 NW-150 SE—in the Caples Lake granodiorite of the Sierra Nevada Batholith, California. Specifically, we investigate 1) their spacing distributions with a keen interest in power-law (fractal) spacing, 2) distribution comparisons between master and cross …


Biases And Blind-Spots In Genome-Wide Crispr-Cas9 Knockout Screens, Merve Dede 2021 The University of Texas MD Anderson Cancer Center UTHealth Graduate School of Biomedical Sciences

Biases And Blind-Spots In Genome-Wide Crispr-Cas9 Knockout Screens, Merve Dede

Dissertations and Theses (Open Access)

Adaptation of the bacterial CRISPR-Cas9 system to mammalian cells revolutionized the field of functional genomics, enabling genome-scale genetic perturbations to study essential genes, whose loss of function results in a severe fitness defect. There are two types of essential genes in a cell. Core essential genes are absolutely required for growth and proliferation in every cell type. On the other hand, context-dependent essential genes become essential in an environmental or genetic context. The concept of context-dependent gene essentiality is particularly important in cancer, since killing cancer cells selectively without harming surrounding healthy tissue remains a major challenge. The toxicity of …


High-Dimensional Random Forests, Roland Fiagbe 2021 University of Texas at El Paso

High-Dimensional Random Forests, Roland Fiagbe

Open Access Theses & Dissertations

The significant advances in technology have enabled easy collection and management of high-dimensional data in many fields, however, the process of modeling these data imposes a huge problem in the field of data science. Dealing with high-dimensional data is one of the significant challenges that degenerate the performance and precision of most classification and regression algorithms, e.g., random forests. Random Forest (RF) is among the few methods that can be extended to model high-dimensional data; nevertheless, its performance and precision, like others, are highly affected by high dimensions, especially when the dataset contains a huge number of noise or noninformative …


Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil 2021 University of Texas at El Paso

Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil

Open Access Theses & Dissertations

With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …


Making Valid Inferences With Decision Tree, George Ekow Quaye 2021 University of Texas at El Paso

Making Valid Inferences With Decision Tree, George Ekow Quaye

Open Access Theses & Dissertations

HypoThesis testing and Confidence Interval (CI) estimates are key statistics in predicting future values in data analysis. Most often, CI estimates are directly obtained from the summary statistics of a particular statistical methodology output. However, when it comes to the summary of decision tree outputs, these CI estimates are not directly obtained. So a na\"{i}ve way of making node-level inference is to construct a $(1-\alpha) \times 100\%$ confidence interval for a node mean $\bar{y}_t$ using the relation: $\bar{y}_t \, \pm \, z_{1-\alpha/2} \, \frac{s_t}{\sqrt{n_t}}$, where $\bar{y}_t$ is the node mean and $s_t$ is the standard deviation estimates from the decision …


Refined Moderation Analysis With Binary Outcomes, Eric Anto 2021 University of Texas at El Paso

Refined Moderation Analysis With Binary Outcomes, Eric Anto

Open Access Theses & Dissertations

With the growing interest in personalized or precision medicine, it is indispensable thatmoderation analysis which is primarily related to the study of differential treatment effects among patients with different characteristics, also serves as the bedrock for precision medicine is taken more seriously. Concerning moderation analysis with binary outcomes, we start with an interesting observation, which shows that heterogeneous treatment effects could be equivalently estimated via a role exchange between the outcome and the treatment variable. The result holds for both experimental data and observational data, yet with an important difference in interpretation. Two estimators of moderating effects corresponding to two …


The Effects Of The Nba Covid Bubble On The Nba Playoffs: A Case Study For Home-Court Advantage, Michael Price 2021 University of Connecticut

The Effects Of The Nba Covid Bubble On The Nba Playoffs: A Case Study For Home-Court Advantage, Michael Price

Honors Scholar Theses

The 2020 NBA playoffs were played inside of a bubble in Disney World because of the COVID-19 pandemic. This meant that there were no fans in attendance, games played on neutral courts and no traveling for teams, which in theory removes home-court advantage from the games. This setting has attracted much discussion as analysts and fans debated the possible effects it may have on the outcome of games. Home-court advantage has historically played an influential role in NBA playoff series outcomes. The 2020 playoff provided a unique opportunity to study the effects of the bubble and home-court advantage by comparing …


Association Between Dietary Inflammatory Index, Dietary Patterns, Plant-Based Dietary Index And The Risk Of Obesity, Yoko B. Wang, Nitin Shivappa, James R. Hébert ScD, Amanda J. Page, Yohannes Adama Melaku 2021 University of South Carolina

Association Between Dietary Inflammatory Index, Dietary Patterns, Plant-Based Dietary Index And The Risk Of Obesity, Yoko B. Wang, Nitin Shivappa, James R. Hébert Scd, Amanda J. Page, Yohannes Adama Melaku

Faculty Publications

Evidence on the association between various dietary constructs and obesity risk is limited. This study aims to investigate the longitudinal relationship between different diet indices and dietary patterns with the risk of obesity. Non-obese participants (n = 787) in the North West Adelaide Health Study were followed from 2010 to 2015. The dietary inflammatory index (DII®), plant-based dietary index (PDI) and factor-derived dietary pattern scores were computed based on food frequency questionnaire data. We found the incidence of obesity was 7.62% at the 5-year follow up. In the adjusted model, results from multivariable log-binomial logistic regression showed that a prudent …


Diet Quality And Risk Of Lung Cancer In The Multiethnic Cohort Study, Song-Yi Park, Carol J. Boushey, Yurii B. Shvetsov, Michael David Wirth MSPH,Ph.D., Nitin Shivappa Ph.D., James R. Hébert Sc.D., Christopher A. Haiman, Lynee R. Wilkens, Loic Le Marchand 2021 University of South Carolina

Diet Quality And Risk Of Lung Cancer In The Multiethnic Cohort Study, Song-Yi Park, Carol J. Boushey, Yurii B. Shvetsov, Michael David Wirth Msph,Ph.D., Nitin Shivappa Ph.D., James R. Hébert Sc.D., Christopher A. Haiman, Lynee R. Wilkens, Loic Le Marchand

Faculty Publications

Diet quality, assessed by the Healthy Eating Index-2015 (HEI-2015), the Alternative Healthy Eating Index-2010 (AHEI-2010), the alternate Mediterranean Diet (aMED) score, the Dietary Approaches to Stop Hypertension (DASH) score, and the Dietary Inflammatory Index (DII®), was examined in relation to risk of lung cancer in the Multiethnic Cohort Study. The analysis included 179,318 African Americans, Native Hawaiians, Japanese Americans, Latinos, and Whites aged 45–75 years, with 5350 incident lung cancer cases during an average follow-up of 17.5 ± 5.4 years. In multivariable Cox models comprehensively adjusted for cigarette smoking, the hazard ratios (95% confidence intervals) for the highest vs. lowest …


Association Between Appendicular Skeletal Muscle Index And Leukocyte Telomere Length In Adults: A Study From National Health And Nutrition Examination Survey (Nhanes) 1999-2002, Lingzhi Chen, Nitin Shivappa MBBS, MPH, Ph.D., Xiuxun Dong, Jinjing Ming 2021 University of South Carolina

Association Between Appendicular Skeletal Muscle Index And Leukocyte Telomere Length In Adults: A Study From National Health And Nutrition Examination Survey (Nhanes) 1999-2002, Lingzhi Chen, Nitin Shivappa Mbbs, Mph, Ph.D., Xiuxun Dong, Jinjing Ming

Faculty Publications

Background

A higher body mass index (BMI) is associated with shorter telomeres. The loss of muscle mass with aging is associated with adverse outcomes. The appendicular skeletal muscle index (ASMI) is currently used to quantify muscle mass.

Objective

We investigated the association of the ASMI with leukocyte telomere length in adult Americans.

Methods

This cross-sectional study used the National Health and Nutrition Examination Survey (NHANES) 1999–2002 dataset. Body composition was measured by dual-energy X-ray absorptiometry. Low muscle mass was defined using sex-specific thresholds of the appendicular skeletal muscle mass index (ASMI). The telomere-to-single-copy gene ratio (T/S ratio) was converted to …


Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang 2021 The University of Texas MD Anderson Cancer Center UTHealth Graduate School of Biomedical Sciences

Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang

Dissertations and Theses (Open Access)

Integrative genomic data analysis is a powerful tool to study the complex biological processes behind a disease. Statistical methods can model the interrelationships of the involved gene activities through jointly analyzing multiple types of genomic data from different platforms (vertical integration), or improve the power of a study through aggregating the same type of genomic data across studies (horizontal integration). In this dissertation, we propose statistical methods and strategies for integrative multi-omics data in association analysis of disease phenotypes, with an emphasis on cancer applications.

We develop a new strategy based on horizontal integration by leveraging publicly available datasets into …


Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame 2021 East Tennessee State University

Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame

Electronic Theses and Dissertations

Missing data continues to be one of the main problems in data analysis as it reduces sample representativeness and consequently, causes biased estimates. Multiple imputation methods have been established as an effective method of handling missing data. In this study, we examined multiple imputation methods for quantitative variables on twelve data sets with varied sizes and variability that were pseudo generated from an original data. The multiple imputation methods examined are the predictive mean matching, Bayesian linear regression and linear regression, non-Bayesian in the MICE (Multiple Imputation Chain Equation) package in the statistical software, R. The parameter estimates generated from …


Zeta Function Regularization And Its Relationship To Number Theory, Stephen Wang 2021 East Tennessee State University

Zeta Function Regularization And Its Relationship To Number Theory, Stephen Wang

Electronic Theses and Dissertations

While the "path integral" formulation of quantum mechanics is both highly intuitive and far reaching, the path integrals themselves often fail to converge in the usual sense. Richard Feynman developed regularization as a solution, such that regularized path integrals could be calculated and analyzed within a strictly physics context. Over the past 50 years, mathematicians and physicists have retroactively introduced schemes for achieving mathematical rigor in the study and application of regularized path integrals. One such scheme was introduced in 2007 by the mathematicians Klaus Kirsten and Paul Loya. In this thesis, we reproduce the Kirsten and Loya approach to …


Use Of Linear Discriminant Analysis In Song Classification: Modeling Based On Wilco Albums, Caroline Pollard 2021 University of Mississippi

Use Of Linear Discriminant Analysis In Song Classification: Modeling Based On Wilco Albums, Caroline Pollard

Honors Theses

The study of music recommender algorithms is a relatively new area of study. Although these algorithms serve a variety of functions, they primarily help advertise and suggest music to users on music streaming services. This thesis explores the use of linear discriminant analysis in music categorization for the purpose of serving as a cheaper and simpler content-based recommender algorithm. The use of linear discriminant analysis was tested by creating lineardiscriminant functions that classify Wilco’s songs into their respective albums, specifically A.M., Yankee Hotel Foxtrot, and Sky Blue Sky. 4 sample songs were chosen from each album, and song data was …


The Role And Challenges Of Cluster Randomised Trials For Global Health, Louis Dron, Monica Taljaard, Yin Bun Cheung, Rebecca Grais, Nathan Ford, Kristian Thorlund, Fyezah Jehan, Etheldreda Nakimuli-Mpungu, Denis Xavier, Zulfiqar Ahmed Bhutta, Jay J H. Park, Edward J. Mills 2021 McMaster University, Hamilton, Canada

The Role And Challenges Of Cluster Randomised Trials For Global Health, Louis Dron, Monica Taljaard, Yin Bun Cheung, Rebecca Grais, Nathan Ford, Kristian Thorlund, Fyezah Jehan, Etheldreda Nakimuli-Mpungu, Denis Xavier, Zulfiqar Ahmed Bhutta, Jay J H. Park, Edward J. Mills

Department of Paediatrics and Child Health

Evaluating whether an intervention works when trialled in groups of individuals can pose complex challenges for clinical research. Cluster randomised controlled trials involve the random allocation of groups or clusters of individuals to receive an intervention, and they are commonly used in global health research. In this paper, we describe the potential reasons for the increasing popularity of cluster trials in low-income and middle-income countries. We also draw on key areas of global health research for an assessment of common trial planning practices, and we address their methodological shortcomings and pitfalls. Lastly, we discuss alternative approaches for population-level intervention trials …


Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell 2021 University of Lynchburg

Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell

Undergraduate Theses and Capstone Projects

This thesis analyzes the correlation between a team’s statistics and the success of their performances, and develops a predictive model that can be used to forecast final season results for that team. Data from the 2017-2018 Premier League season is to be gathered and broken down within R to highlight what factors and variables are largely contributing to the success or downfall of a team. A multiple linear regression model and stepwise selection process is then used to include any factors that are significant in predicting in match results.

The predictions about the 17-18 season results based on the model …


Adaptive Optimal Market Making Strategies With Inventory Liquidation Cost, YI ZHANG 2021 Washington University in St. Louis

Adaptive Optimal Market Making Strategies With Inventory Liquidation Cost, Yi Zhang

Arts & Sciences Graduate Student Theses and Dissertations

Along the lines of the paper \cite{zoe}, we find a general form of the optimal market making strategy for a high-frequency market maker (HFM) in a discrete-time Limit Order Book (LOB) model. Unlike \cite{zoe}, the optimal market making strategy is adaptive depending on the arrival of Market Order (MO) in the previous time intervals. We provide a method to make each placement of Limit Orders (LO) dependent on previous information in the same trading day and prove the admissibility of the optimal market making strategy under some general assumptions. Empirical study shows the adaptive optimal strategies outperform the non-adaptive strategy …


Identifying Armed Group Presence Using Hidden Markov Models, MAURICIO VELA BARON 2021 Washington University in St. Louis

Identifying Armed Group Presence Using Hidden Markov Models, Mauricio Vela Baron

Arts & Sciences Graduate Student Theses and Dissertations

Identifying armed group presence is often helpful for conflict studies to examine patterns of conflict. Armed group presence is often used as the main variable of interest in several studies, and in some cases, this variable is ignored. Many of these studies use expert data or proxy variables to analyze armed group presence. This paper proposes Hidden Markov Models (HMMs) as a method to identify armed group presence. HMMs permit identifying armed group presence at a sub-national level and in long panel data sets. A HMM is used in this paper to identify paramilitary and FARC presence in Colombia. The …


Does Social Support Moderate The Association Between Income And Food Security Status Among Seniors Living In Southern Nevada?, Adugna Teka Siweya 2021 University of Nevada, Las Vegas

Does Social Support Moderate The Association Between Income And Food Security Status Among Seniors Living In Southern Nevada?, Adugna Teka Siweya

UNLV Theses, Dissertations, Professional Papers, and Capstones

Background: Income is the strongest predictor of food insecurity among seniors, and social support also an essential factor to help mitigate the effects of food insecurity. However, little is known about the potential role that social support may play as a moderator of the association between income and food insecurity. Thus, we aim to examine social support as a moderator for the relationship between income and food insecurity among seniors. Methods: Logistic regression models were used to analyze data collected in 2019 from seniors residing in Southern Nevada. Predictors of food insecurity, sociodemographic factors, social support variables, and income and …


The Effect Of High Elevation Weather Stations On The Usda's Pasture, Rangeland, And Forage Insurance Program, Wyatt Matthew Feuz 2021 Utah State University

The Effect Of High Elevation Weather Stations On The Usda's Pasture, Rangeland, And Forage Insurance Program, Wyatt Matthew Feuz

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

This paper examines the effect of high elevation weather stations on the rainfall index used by the Pasture, Rangeland, and Forage insurance program. Weather station data for the state of Utah is used to identify high elevation weather stations and their location. Utilizing the corresponding rainfall index data, the effect of the high elevation weather stations is determined. This paper finds when high elevation weather stations begin reporting there is a jump up of 19.01–27.88 percentage points on average in the rainfall index for the corresponding grid locations. This indicates the rainfall index may not accurately represent actual precipitation amounts …


Digital Commons powered by bepress