Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2021

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 331 - 360 of 662

Full-Text Articles in Statistics and Probability

Confidence Intervals Of Covid-19 Vaccine Efficacy Rates, Frank Wang May 2021

Confidence Intervals Of Covid-19 Vaccine Efficacy Rates, Frank Wang

Numeracy

This tutorial uses publicly available data from drug makers and the Food and Drug Administration to guide learners to estimate the confidence intervals of COVID-19 vaccine efficacy rates with a Bayesian framework. Under the classical approach, there is no probability associated with a parameter, and the meaning of confidence intervals can be misconstrued by inexperienced students. With Bayesian statistics, one can find the posterior probability distribution of an unknown parameter, and state the probability of vaccine efficacy rate, which makes the communication of uncertainty more flexible. We use a hypothetical example and a real baseball example to guide readers to …


Species In Vernal Pools: Anova, Lisa Manne May 2021

Species In Vernal Pools: Anova, Lisa Manne

Open Educational Resources

A one-way analysis of variance exercise using data on species diversities from vernal pools.Data are from vernal pools in Willowbrook Park (adjacent to College of Staten Island's campus) in spring.

The typical ANOVA gives a straightforward result (significant anova, easily-interpreted Tukey-Kramer analysis). This data set requires more nuanced interpretation, as the ANOVA is marginally significant, and Tukey-Kramer yields one significant pairwise comparison between groups. Relative lack of variation within groups explains this apparent enigma.


Improving Bayesian Graph Convolutional Networks Using Markov Chain Monte Carlo Graph Sampling, Aneesh Komanduri May 2021

Improving Bayesian Graph Convolutional Networks Using Markov Chain Monte Carlo Graph Sampling, Aneesh Komanduri

Computer Science and Computer Engineering Undergraduate Honors Theses

In the modern age of social media and networks, graph representations of real-world phenomena have become incredibly crucial. Often, we are interested in understanding how entities in a graph are interconnected. Graph Neural Networks (GNNs) have proven to be a very useful tool in a variety of graph learning tasks including node classification, link prediction, and edge classification. However, in most of these tasks, the graph data we are working with may be noisy and may contain spurious edges. That is, there is a lot of uncertainty associated with the underlying graph structure. Recent approaches to modeling uncertainty have been …


Applying Emotional Analysis For Automated Content Moderation, John Shelnutt May 2021

Applying Emotional Analysis For Automated Content Moderation, John Shelnutt

Computer Science and Computer Engineering Undergraduate Honors Theses

The purpose of this project is to explore the effectiveness of emotional analysis as a means to automatically moderate content or flag content for manual moderation in order to reduce the workload of human moderators in moderating toxic content online. In this context, toxic content is defined as content that features excessive negativity, rudeness, or malice. This often features offensive language or slurs. The work involved in this project included creating a simple website that imitates a social media or forum with a feed of user submitted text posts, implementing an emotional analysis algorithm from a word emotions dataset, designing …


Retail Trading And Stock Volatility: The Case Of Robinhood, Cooper Jones May 2021

Retail Trading And Stock Volatility: The Case Of Robinhood, Cooper Jones

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

We examine the relation between Robinhood usership and stock market volatility. We show that daily fluctuations in Robinhood usership, which is used to proxy retail trading, significantly influence various measures of volatility. These results might suggest that Robinhood users contribute to noise trading as they are generally individuals trading on name recognition, media coverage, popularity, and familiarity of products, rather than on fundamental values. In our empirical approach, we find that the percentage increase in Robinhood usership Granger causes increases in daily stock volatility.


Randomised Trials At The Level Of The Individual, Jay J H. Park, Nathan Ford, Denis Xavier, Per Ashorn, Rebecca F. Grais, Zulfiqar Ahmed Bhutta, Herman Goossens, Kristian Thorlund, Maria Eugenia Socias, Edward J. Mills May 2021

Randomised Trials At The Level Of The Individual, Jay J H. Park, Nathan Ford, Denis Xavier, Per Ashorn, Rebecca F. Grais, Zulfiqar Ahmed Bhutta, Herman Goossens, Kristian Thorlund, Maria Eugenia Socias, Edward J. Mills

Centre of Excellence in Women and Child Health

In global health research, short-term, small-scale clinical trials with fixed, two-arm trial designs that generally do not allow for major changes throughout the trial are the most common study design. Building on the introductory paper of this Series, this paper discusses data-driven approaches to clinical trial research across several adaptive trial designs, as well as the master protocol framework that can help to harmonise clinical trial research efforts in global health research. We provide a general framework for more efficient trial research, and we discuss the importance of considering different study designs in the planning stage with statistical simulations. We …


Improving Access And Health Outcomes Through Carepartners And Medaccess Programs, Elora Way, Becky Wurwarg May 2021

Improving Access And Health Outcomes Through Carepartners And Medaccess Programs, Elora Way, Becky Wurwarg

Publications

Access to Care, a division of MaineHealth, works to ensure that Maine residents have access to comprehensive and affordable healthcare that improves community wellbeing. In 2020, as part of its ongoing commitment to evaluating the effectiveness of its programs, Access to Care partnered with the University of Southern Maine’s Data Innovation Project to conduct a multi-year retrospective evaluation of its two longest-running initiatives, CarePartners and MedAccess. Established in 2001, CarePartners coordinates donated healthcare services for low-income, uninsured residents across six Maine counties by connecting participants with case managers, primary care providers, and pharmacy benefits. Between 2016 and 2019, 4,426 individuals …


Applications Of Evidence Theory To High-Consequence Systems Safety, Christina Marie Deffenbaugh May 2021

Applications Of Evidence Theory To High-Consequence Systems Safety, Christina Marie Deffenbaugh

Mathematics & Statistics ETDs

Issues linked to abnormal environments (like high-consequence systems safety, e.g., nuclear weapon components, bridges, apartment buildings, etc.) may have insufficient information to use either classical statistical methods or Bayesian approaches for calculating associated probabilistic risks, so there is often a requirement for another method that can deal with a low-information situation to obtain a risk assessment. Belief/plausibility measures of uncertainty from A. P. Dempster and G. Shafer’s Evidence Theory is one such method. This thesis has two goals. First, a brief discussion on belief/plausibility measures as an application of Evidence Theory will familiarize the audience with its history and how …


Cointegration And Statistical Arbitrage Of Precious Metals, Judge Van Horn May 2021

Cointegration And Statistical Arbitrage Of Precious Metals, Judge Van Horn

Finance Undergraduate Honors Theses

When talking about financial instruments correlation is often thrown around as a measure of the relation between two securities. An often more useful or tradeable measure is cointegration. Cointegration is the measure of two securities tendency to revert to an average price over time. In other words, cointegration ignores directionality and only cares about the distance between two securities. For a mean reversion strategy such as statistical arbitrage cointegration proves to be a far more reliable statistical measure of mean reversion, and while it is more reliable than correlation it still has its own problems. One thing to consider is …


The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun May 2021

The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun

Open Access Theses & Dissertations

Fast and accurate calculations of the electrostatic features for highly charged biomolecules such as DNA, RNA, highly charged proteins, are crucial but challenging tasks. Traditional implicit solvent methods calculate the electrostatic features fast, but they are not able to balance the high net charges in the biomolecules effectively. Explicit solvent methods add unbalanced ions to neutralize the highly charged biomolecules in molecular dynamic simulations, which require more expensive computing resources. Here we developed a novel method, the Hybridizing Ions Treatment (HIT) method, which hybridizes the implicit solvent method with the explicit method to realistically calculate the electrostatic potential for highly …


Robust Variable Selection In Multiple Linear Regression Via Penalized Least Trimmed Squares., Reagan Kesseku May 2021

Robust Variable Selection In Multiple Linear Regression Via Penalized Least Trimmed Squares., Reagan Kesseku

Open Access Theses & Dissertations

Variable selection has been studied using different approaches. Its growing importance lies in numerous applications to high-dimensional data from experiments and natural phenomena. Often, models are to be constructed from such data based on significant variables for estimation or prediction purposes. This demands not just any variable selectionmethod, but one that is robust, computationally efficient and with other desirable statistical properties. Besides the high-dimensionality of such data, the presence of outliers is common due to heterogeneous sources. Though outliers often contain useful information, they can unduly influence non-robust estimators to produce misleading results. This is the case for ordinary least …


A Data Adaptive Model For Retail Sales Of Electricity, Johanna Marcelia May 2021

A Data Adaptive Model For Retail Sales Of Electricity, Johanna Marcelia

Boise State University Theses and Dissertations

When fitting a model to a data set, the goal is to create a model that captures the trends present in the data. However, data often contains regions where the underlying model changes or exhibits shifts in certain parameters due to economic events. These locations in the data are known as changepoints, and ignoring them can result in high error and incorrect forecasts. By developing a specific cost function and optimizing using the genetic algorithm, we are able to locate and account for the changepoints in a given data set. We specifically apply this process to the retail sales of …


Joint Spacing In The Caples Lake Granodiorite Of The Sierra Nevada Batholith In Eldorado National Forest, California: A Comparative Analysis Of Joint Sets And Data Resolution, Jimmy Wood May 2021

Joint Spacing In The Caples Lake Granodiorite Of The Sierra Nevada Batholith In Eldorado National Forest, California: A Comparative Analysis Of Joint Sets And Data Resolution, Jimmy Wood

Theses/Capstones/Creative Projects

Joints are the most common deformation structure in the Earth’s upper crust and exert a significant influence on structural stability, landscape morphology, and fluid flow . Therefore, a greater understanding of fracture parameters (e.g., length, aperture, etc.) allows us to more accurately predict their presence, persistence, and prevalence, in the subsurface . We study the fracture spacing of two sub-orthogonal joint sets—66 NE-246 SW and 330 NW-150 SE—in the Caples Lake granodiorite of the Sierra Nevada Batholith, California. Specifically, we investigate 1) their spacing distributions with a keen interest in power-law (fractal) spacing, 2) distribution comparisons between master and cross …


Adaptive Optimal Market Making Strategies With Inventory Liquidation Cost, Yi Zhang May 2021

Adaptive Optimal Market Making Strategies With Inventory Liquidation Cost, Yi Zhang

Arts & Sciences Graduate Student Theses and Dissertations

Along the lines of the paper \cite{zoe}, we find a general form of the optimal market making strategy for a high-frequency market maker (HFM) in a discrete-time Limit Order Book (LOB) model. Unlike \cite{zoe}, the optimal market making strategy is adaptive depending on the arrival of Market Order (MO) in the previous time intervals. We provide a method to make each placement of Limit Orders (LO) dependent on previous information in the same trading day and prove the admissibility of the optimal market making strategy under some general assumptions. Empirical study shows the adaptive optimal strategies outperform the non-adaptive strategy …


Biases And Blind-Spots In Genome-Wide Crispr-Cas9 Knockout Screens, Merve Dede May 2021

Biases And Blind-Spots In Genome-Wide Crispr-Cas9 Knockout Screens, Merve Dede

Dissertations and Theses (Open Access)

Adaptation of the bacterial CRISPR-Cas9 system to mammalian cells revolutionized the field of functional genomics, enabling genome-scale genetic perturbations to study essential genes, whose loss of function results in a severe fitness defect. There are two types of essential genes in a cell. Core essential genes are absolutely required for growth and proliferation in every cell type. On the other hand, context-dependent essential genes become essential in an environmental or genetic context. The concept of context-dependent gene essentiality is particularly important in cancer, since killing cancer cells selectively without harming surrounding healthy tissue remains a major challenge. The toxicity of …


Estimating Cumulative Incidence Rate On Interval Censored Data In An Illness-Death Model., Chen Qian May 2021

Estimating Cumulative Incidence Rate On Interval Censored Data In An Illness-Death Model., Chen Qian

Electronic Theses and Dissertations

Phase IV clinical trials are designed to monitor long-term side effects caused overtime by the medical treatment. For instance, in advanced primary cancer treatment, childhood cancer survivors are often at risk of developing undesired events, such as cardiotoxicity, during their adulthood. Such problems could be due to their cancer or the treatment they received for their cancer such as radiation or intensive chemotherapy. Cardiotoxicity can be diagnosed with electrophysiology with measurements of fraction shortening, afterload, etc. Often the primary focus of a study could be on estimating the cumulative incidence of a particular outcome of interest such as cardiotoxicity. However, …


Observational Studies In Group Testing And Potential Applications., Alexander Christopher Noll May 2021

Observational Studies In Group Testing And Potential Applications., Alexander Christopher Noll

Electronic Theses and Dissertations

The use of group testing to identify individuals with targeted outcomes in a population can greatly improve the efficiency, speed, and cost effectiveness of testing a population for an outcome, or at least for identifying the prevalence of an outcome in a population. The implementation of causal inference techniques can provide the basis for an observational study that would allow an investigator to gather estimates for treatment effectiveness if group testing was conducted on the population in a certain way. This thesis examines a simulation of the above outlined principles in order to demonstrate a potential application for determining treatment …


High-Dimensional Random Forests, Roland Fiagbe May 2021

High-Dimensional Random Forests, Roland Fiagbe

Open Access Theses & Dissertations

The significant advances in technology have enabled easy collection and management of high-dimensional data in many fields, however, the process of modeling these data imposes a huge problem in the field of data science. Dealing with high-dimensional data is one of the significant challenges that degenerate the performance and precision of most classification and regression algorithms, e.g., random forests. Random Forest (RF) is among the few methods that can be extended to model high-dimensional data; nevertheless, its performance and precision, like others, are highly affected by high dimensions, especially when the dataset contains a huge number of noise or noninformative …


Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil May 2021

Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil

Open Access Theses & Dissertations

With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …


Making Valid Inferences With Decision Tree, George Ekow Quaye May 2021

Making Valid Inferences With Decision Tree, George Ekow Quaye

Open Access Theses & Dissertations

HypoThesis testing and Confidence Interval (CI) estimates are key statistics in predicting future values in data analysis. Most often, CI estimates are directly obtained from the summary statistics of a particular statistical methodology output. However, when it comes to the summary of decision tree outputs, these CI estimates are not directly obtained. So a na\"{i}ve way of making node-level inference is to construct a $(1-\alpha) \times 100\%$ confidence interval for a node mean $\bar{y}_t$ using the relation: $\bar{y}_t \, \pm \, z_{1-\alpha/2} \, \frac{s_t}{\sqrt{n_t}}$, where $\bar{y}_t$ is the node mean and $s_t$ is the standard deviation estimates from the decision …


Refined Moderation Analysis With Binary Outcomes, Eric Anto May 2021

Refined Moderation Analysis With Binary Outcomes, Eric Anto

Open Access Theses & Dissertations

With the growing interest in personalized or precision medicine, it is indispensable thatmoderation analysis which is primarily related to the study of differential treatment effects among patients with different characteristics, also serves as the bedrock for precision medicine is taken more seriously. Concerning moderation analysis with binary outcomes, we start with an interesting observation, which shows that heterogeneous treatment effects could be equivalently estimated via a role exchange between the outcome and the treatment variable. The result holds for both experimental data and observational data, yet with an important difference in interpretation. Two estimators of moderating effects corresponding to two …


The Effects Of The Nba Covid Bubble On The Nba Playoffs: A Case Study For Home-Court Advantage, Michael Price May 2021

The Effects Of The Nba Covid Bubble On The Nba Playoffs: A Case Study For Home-Court Advantage, Michael Price

Honors Scholar Theses

The 2020 NBA playoffs were played inside of a bubble in Disney World because of the COVID-19 pandemic. This meant that there were no fans in attendance, games played on neutral courts and no traveling for teams, which in theory removes home-court advantage from the games. This setting has attracted much discussion as analysts and fans debated the possible effects it may have on the outcome of games. Home-court advantage has historically played an influential role in NBA playoff series outcomes. The 2020 playoff provided a unique opportunity to study the effects of the bubble and home-court advantage by comparing …


Association Between Dietary Inflammatory Index, Dietary Patterns, Plant-Based Dietary Index And The Risk Of Obesity, Yoko B. Wang, Nitin Shivappa, James R. Hébert Scd, Amanda J. Page, Yohannes Adama Melaku May 2021

Association Between Dietary Inflammatory Index, Dietary Patterns, Plant-Based Dietary Index And The Risk Of Obesity, Yoko B. Wang, Nitin Shivappa, James R. Hébert Scd, Amanda J. Page, Yohannes Adama Melaku

Faculty Publications

Evidence on the association between various dietary constructs and obesity risk is limited. This study aims to investigate the longitudinal relationship between different diet indices and dietary patterns with the risk of obesity. Non-obese participants (n = 787) in the North West Adelaide Health Study were followed from 2010 to 2015. The dietary inflammatory index (DII®), plant-based dietary index (PDI) and factor-derived dietary pattern scores were computed based on food frequency questionnaire data. We found the incidence of obesity was 7.62% at the 5-year follow up. In the adjusted model, results from multivariable log-binomial logistic regression showed that a prudent …


Diet Quality And Risk Of Lung Cancer In The Multiethnic Cohort Study, Song-Yi Park, Carol J. Boushey, Yurii B. Shvetsov, Michael David Wirth Msph,Ph.D., Nitin Shivappa Ph.D., James R. Hébert Sc.D., Christopher A. Haiman, Lynee R. Wilkens, Loic Le Marchand May 2021

Diet Quality And Risk Of Lung Cancer In The Multiethnic Cohort Study, Song-Yi Park, Carol J. Boushey, Yurii B. Shvetsov, Michael David Wirth Msph,Ph.D., Nitin Shivappa Ph.D., James R. Hébert Sc.D., Christopher A. Haiman, Lynee R. Wilkens, Loic Le Marchand

Faculty Publications

Diet quality, assessed by the Healthy Eating Index-2015 (HEI-2015), the Alternative Healthy Eating Index-2010 (AHEI-2010), the alternate Mediterranean Diet (aMED) score, the Dietary Approaches to Stop Hypertension (DASH) score, and the Dietary Inflammatory Index (DII®), was examined in relation to risk of lung cancer in the Multiethnic Cohort Study. The analysis included 179,318 African Americans, Native Hawaiians, Japanese Americans, Latinos, and Whites aged 45–75 years, with 5350 incident lung cancer cases during an average follow-up of 17.5 ± 5.4 years. In multivariable Cox models comprehensively adjusted for cigarette smoking, the hazard ratios (95% confidence intervals) for the highest vs. lowest …


Association Between Appendicular Skeletal Muscle Index And Leukocyte Telomere Length In Adults: A Study From National Health And Nutrition Examination Survey (Nhanes) 1999-2002, Lingzhi Chen, Nitin Shivappa Mbbs, Mph, Ph.D., Xiuxun Dong, Jinjing Ming May 2021

Association Between Appendicular Skeletal Muscle Index And Leukocyte Telomere Length In Adults: A Study From National Health And Nutrition Examination Survey (Nhanes) 1999-2002, Lingzhi Chen, Nitin Shivappa Mbbs, Mph, Ph.D., Xiuxun Dong, Jinjing Ming

Faculty Publications

Background

A higher body mass index (BMI) is associated with shorter telomeres. The loss of muscle mass with aging is associated with adverse outcomes. The appendicular skeletal muscle index (ASMI) is currently used to quantify muscle mass.

Objective

We investigated the association of the ASMI with leukocyte telomere length in adult Americans.

Methods

This cross-sectional study used the National Health and Nutrition Examination Survey (NHANES) 1999–2002 dataset. Body composition was measured by dual-energy X-ray absorptiometry. Low muscle mass was defined using sex-specific thresholds of the appendicular skeletal muscle mass index (ASMI). The telomere-to-single-copy gene ratio (T/S ratio) was converted to …


Use Of Linear Discriminant Analysis In Song Classification: Modeling Based On Wilco Albums, Caroline Pollard May 2021

Use Of Linear Discriminant Analysis In Song Classification: Modeling Based On Wilco Albums, Caroline Pollard

Honors Theses

The study of music recommender algorithms is a relatively new area of study. Although these algorithms serve a variety of functions, they primarily help advertise and suggest music to users on music streaming services. This thesis explores the use of linear discriminant analysis in music categorization for the purpose of serving as a cheaper and simpler content-based recommender algorithm. The use of linear discriminant analysis was tested by creating lineardiscriminant functions that classify Wilco’s songs into their respective albums, specifically A.M., Yankee Hotel Foxtrot, and Sky Blue Sky. 4 sample songs were chosen from each album, and song data was …


Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang May 2021

Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang

Dissertations and Theses (Open Access)

Integrative genomic data analysis is a powerful tool to study the complex biological processes behind a disease. Statistical methods can model the interrelationships of the involved gene activities through jointly analyzing multiple types of genomic data from different platforms (vertical integration), or improve the power of a study through aggregating the same type of genomic data across studies (horizontal integration). In this dissertation, we propose statistical methods and strategies for integrative multi-omics data in association analysis of disease phenotypes, with an emphasis on cancer applications.

We develop a new strategy based on horizontal integration by leveraging publicly available datasets into …


Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame May 2021

Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame

Electronic Theses and Dissertations

Missing data continues to be one of the main problems in data analysis as it reduces sample representativeness and consequently, causes biased estimates. Multiple imputation methods have been established as an effective method of handling missing data. In this study, we examined multiple imputation methods for quantitative variables on twelve data sets with varied sizes and variability that were pseudo generated from an original data. The multiple imputation methods examined are the predictive mean matching, Bayesian linear regression and linear regression, non-Bayesian in the MICE (Multiple Imputation Chain Equation) package in the statistical software, R. The parameter estimates generated from …


Zeta Function Regularization And Its Relationship To Number Theory, Stephen Wang May 2021

Zeta Function Regularization And Its Relationship To Number Theory, Stephen Wang

Electronic Theses and Dissertations

While the "path integral" formulation of quantum mechanics is both highly intuitive and far reaching, the path integrals themselves often fail to converge in the usual sense. Richard Feynman developed regularization as a solution, such that regularized path integrals could be calculated and analyzed within a strictly physics context. Over the past 50 years, mathematicians and physicists have retroactively introduced schemes for achieving mathematical rigor in the study and application of regularized path integrals. One such scheme was introduced in 2007 by the mathematicians Klaus Kirsten and Paul Loya. In this thesis, we reproduce the Kirsten and Loya approach to …


Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell May 2021

Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell

Undergraduate Theses and Capstone Projects

This thesis analyzes the correlation between a team’s statistics and the success of their performances, and develops a predictive model that can be used to forecast final season results for that team. Data from the 2017-2018 Premier League season is to be gathered and broken down within R to highlight what factors and variables are largely contributing to the success or downfall of a team. A multiple linear regression model and stepwise selection process is then used to include any factors that are significant in predicting in match results.

The predictions about the 17-18 season results based on the model …