Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,820 Full-Text Articles 23,917 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,820 full-text articles. Page 194 of 487.

A Review Of Sample Size And Design Efficacy In Crossover Design In Peer-Reviewed Psychology Research, Kyle Moxley 2021 Wayne State University

A Review Of Sample Size And Design Efficacy In Crossover Design In Peer-Reviewed Psychology Research, Kyle Moxley

Wayne State University Dissertations

A REVIEW OF SAMPLE SIZE AND DESIGN EFFICACY IN CROSSOVER DESIGN IN PEER-REVIEWED PSYCHOLOGY RESEARCHby KYLE C. MOXLEY November 2021 Advisor: Dr. Shlomo S. Sawilowsky Major: Education Evaluation and Research Degree: Doctor of Philosophy The present study seeks to investigate the efficacy of crossover research designs, and the application of crossover designs, in the field of behavioral sciences. Under ideal conditions, crossover designs are assumed to be more efficacious than parallel studies in that participants are given both treatments. However, the presence of carryover effects from treatments may influence outcomes (Jones & Kenward, 2014). To prevent carryover effects, researchers frequently …


The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr 2021 Amherst College

The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr

Statistical and Data Sciences: Faculty Publications

We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.


Toward Uncharted Territory Of Cellular Heterogeneity: Advances And Applications Of Single-Cell Rna-Seq, Brandon Lieberman, Meena Kusi, Chia Nung Hung, Chih Wei Chou, Ning He, Yen Yi Ho, Josephine A. Taverna, Tim H.M. Huang, Chun Liang Chen 2021 University of South Carolina

Toward Uncharted Territory Of Cellular Heterogeneity: Advances And Applications Of Single-Cell Rna-Seq, Brandon Lieberman, Meena Kusi, Chia Nung Hung, Chih Wei Chou, Ning He, Yen Yi Ho, Josephine A. Taverna, Tim H.M. Huang, Chun Liang Chen

Faculty Publications

Among single-cell analysis technologies, single-cell RNA-seq (scRNA-seq) has been one of the front runners in technical inventions. Since its induction, scRNA-seq has been well received and undergone many fast-paced technical improvements in cDNA synthesis and amplification, processing and alignment of next generation sequencing reads, differentially expressed gene calling, cell clustering, subpopulation identification, and developmental trajectory prediction. scRNA-seq has been exponentially applied to study global transcriptional profiles in all cell types in humans and animal models, healthy or with diseases, including cancer. Accumulative novel subtypes and rare subpopulations have been discovered as potential underlying mechanisms of stochasticity, differentiation, proliferation, tumorigenesis, and …


Nutritional Approach For Increasing Public Health During Pandemic Of Covid-19: A Comprehensive Review Of Antiviral Nutrients And Nutraceuticals, Vahideh Ebrahimzadeh-Attari, Ghodratollah Panahi, James R. Hébert ScD, Alireza Ostadrahimi, Maryam Saghafi-Asl, Neda Lotfi-Yaghin, Behzad Baradaran 2021 University of South Carolina

Nutritional Approach For Increasing Public Health During Pandemic Of Covid-19: A Comprehensive Review Of Antiviral Nutrients And Nutraceuticals, Vahideh Ebrahimzadeh-Attari, Ghodratollah Panahi, James R. Hébert Scd, Alireza Ostadrahimi, Maryam Saghafi-Asl, Neda Lotfi-Yaghin, Behzad Baradaran

Faculty Publications

Background: The novel coronavirus (COVID-19) is considered as the most life-threatening pandemic disease during the last decade. The individual nutritional status, though usually ignored in the management of COVID-19, plays a critical role in the immune function and pathogenesis of infection. Accordingly, the present review article aimed to report the effects of nutrients and nutraceuticals on respiratory viral infections including COVID-19, with a focus on their mechanisms of action.

Methods: Studies were identified via systematic searches of the databases including PubMed/ MEDLINE, ScienceDirect, Scopus, and Google Scholar from 2000 until April 2020, using keywords. All relevant clinical and experimental studies …


The Need To Incorporate Communities In Compartmental Models, Michael J. Kane, Owais Gilani 2021 Yale University

The Need To Incorporate Communities In Compartmental Models, Michael J. Kane, Owais Gilani

Faculty Journal Articles

Tian et al. provide a framework for assessing population- level interventions of disease outbreaks through the construction of counterfactuals in a large-scale, natural experiment assessing the efficacy of mild, but early interventions compared to delayed interventions. The technique is applied to the recent SARS-CoV-2 outbreak with the population of Shenzhen, China acting as the mild-but-early treatment group and a combination of several US counties resembling Shenzhen but enacting a delayed intervention acting as the control. To help further the development of this framework and identify an avenue for further enhancement, we focus on the use and potential limitations of compartmental …


การเปรียบเทียบประสิทธิภาพของการประมาณค่าของพารามิเตอร์ด้วยวิธีลาสโซและวิธีการคัดเลือกชุดข้อมูลย่อยที่ดีที่สุดในการวิเคราะห์การถดถอยเชิงเส้นสำหรับข้อมูลที่มีมิติสูง, วรัญญา บุตรบุรี 2021 คณะพาณิชยศาสตร์และการบัญชี

การเปรียบเทียบประสิทธิภาพของการประมาณค่าของพารามิเตอร์ด้วยวิธีลาสโซและวิธีการคัดเลือกชุดข้อมูลย่อยที่ดีที่สุดในการวิเคราะห์การถดถอยเชิงเส้นสำหรับข้อมูลที่มีมิติสูง, วรัญญา บุตรบุรี

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยครั้งนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพของวิธีการประมาณค่าพารามิเตอร์สำหรับข้อมูลที่มีมิติสูงด้วยทั้งหมด 5 วิธี ได้แก่ วิธี L0Learn, L0L2Learn, L1, A-L1 และวิธี A-L1L2 โดยการเปรียบเทียบประสิทธิภาพจะเปรียบเทียบใน 2 ด้าน คือ 1) เปรียบเทียบประสิทธิภาพด้านการพยากรณ์ ซึ่งวัดจากค่าคลาดเคลื่อนการทำนาย (MSE) และ 2) ความถูกต้องในการคัดเลือกตัวแปรอิสระเข้าสู่ตัวแบบ ซึ่งพิจารณาจากของค่า Precision Recall และค่า AUC ข้อมูลที่มีมิติสูงที่ใช้ในการศึกษาครั้งนี้ได้จากการจำลอง โดยกำหนดให้ในแต่ละชุดข้อมูลประกอบด้วยจำนวนค่าสังเกต 100 ค่าสังเกต (n = 100) และมีตัวแปรอิสระจำนวน 100 ตัว (p = 1000) โดยตัวแปรอิสระมีการแจกแจงแบบปรกติหลายตัวแปรซึ่งมีความสัมพันธ์กันแบบยกกำลัง (Exponential Correlation) 3 ระดับคือ 0, 0.5 และ 0.9 ค่าความคลาดเคลื่อนสุ่มขึ้นอยู่กับอัตราส่วนสัญญาณต่อสัญญาณรบกวน (SNR) ซึ่งมี 6 ระดับคือ 0.1, 0.5, 1, 5, 10, และ 20 โดยจำลองข้อมูลจำนวน 100 ชุดในแต่ละสถานการณ์ จากการวัดประสิทธิภาพจากค่าเฉลี่ยของข้อมูลทั้ง 100 ชุด ผลการเปรียบเทียบประสิทธิภาพด้านการพยากรณ์พบว่า เมื่อข้อมูลมีค่า SNR ต่ำและตัวแปรอิสระมีความสัมพันธ์กันน้อยถึงปานกลาง วิธี L1 จะมีประสิทธิภาพสูงที่สุด ตามด้วยวิธี L0L2Leran วิธี L0Learn วิธี A-L1L2 และวิธี A-L1 ตามลำดับ แต่เมื่อข้อมูลมีค่า SNR เพิ่มสูงขึ้นและในขณะเดียวกันตัวแปรอิสระมีความสัมพันธ์กันมากขึ้นวิธี A-L1 และวิธี A-L1L2 จะมีประสิทธิภาพสูงที่สุด ตามด้วยวิธี L1 วิธี L0L2Leran วิธี L0Learn ตามลำดับ ส่วนผลการเปรียบเทียบประสิทธิภาพด้านการคัดเลือกตัวแปรเข้าสู่ตัวแบบ เมื่อพิจารณาจากค่าเฉลี่ยของค่า Precision …


Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee 2021 Claremont McKenna College

Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee

CMC Senior Theses

This paper attempts to quantify predictive power of social media sentiment and financial data in stock prediction by utilizing a comprehensive set of stock-related fundamental and technical variables and social media sentiments. For conducting sentiment analysis, this study employs a pretrained finBERT model that provides three different sentiment classifications and respective softmax scores. Hence, the significance of these variables is evaluated with XGBoost regression and Shapley Additive exPlanations (SHAP) frameworks. Through investigating feature importance, this study finds that statistical properties of sentiment variables provide a stronger predictive power than a weighted sentiment score and that it is possible to quantify …


Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard 2021 Claremont Colleges

Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard

CMC Senior Theses

Using a Twitter API, I gather and analyze tweets by performing sentiment analysis to solve the GOAT debate among professional athletes with the primary focus on comparing Michael Jordan and LeBron James. Athletes from the National Football League (NFL), the National Basketball Association (NBA), Major League Baseball (MLB), and the National Collegiate Athletic Association (NCAA) Division 1 Men's and Women's Basketball were selected to compare how sentiment polarity varies across sports. Sentiment polarity is measured by labeling text as "positive", "neutral", or "negative" which allows us to determine which athlete/sport is highly favored among the Twitter community when it comes …


An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein 2021 Claremont Colleges

An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein

CMC Senior Theses

Regression splines have an established value for producing quality fit at a relatively low-degree polynomial. This paper explores the implications of adopting new methods for knot selection in tandem with established methodology from the current literature. Structural features of generated datasets, as well as residuals collected from sequential iterative models are used to augment the equidistant knot selection process. From analyzing a simulated dataset and an application onto the Racial Animus dataset, I find that a B-spline basis paired with equally-spaced knots remains the best choice when data are evenly distributed, even when structural features of a dataset are known …


Bayesian Tail Probability Estimation And Model Selection, Nan Shen 2021 Northern Illinois University

Bayesian Tail Probability Estimation And Model Selection, Nan Shen

Graduate Research Theses & Dissertations

Bayesian statistics is a prevalent and important field in statistics that assigns Bayesian probabilities, which represent a state of knowledge, to unknown quantities. We study Bayesian statistics with its applications through two projects in this report.

In the first project, we investigate the reasons that the Bayesian estimator of the tail probability is always higher than the frequentist estimator. Sufficient conditions for this phenomenon are established by looking at Taylor series approximations about the tail and by using Jensen's Inequality, both of which point to the convexity of the distribution function.

The second project is about redefining the Bayesian information …


Gene-Based Disease Classification Using Bayesian Self-Organizing Map Neural Networks, Guangting Zhou 2021 Northern Illinois University

Gene-Based Disease Classification Using Bayesian Self-Organizing Map Neural Networks, Guangting Zhou

Graduate Research Theses & Dissertations

Genes perform vital roles in living beings. By taking charges of protein synthesis, genes are able to take control of the expression of living traits. There are a lot of diseases associated closely to our genes. By analyzing genetic information, we are able to detect or classify gene based diseases. Among genetic disease information technologies, microarray can be one of the widely used ones. Usually, microarray data records thousands of gene expression features from a small number of samples including both normal and abnormal expressed tissues. It provides standardized comparison information between normal and diseased tissues, so as to provide …


Robust Determinants Of Happiness: High-Dimensional Bayesian Treatment Of Model Uncertainty, Milivoje Davidovic 2021 Northern Illinois University

Robust Determinants Of Happiness: High-Dimensional Bayesian Treatment Of Model Uncertainty, Milivoje Davidovic

Graduate Research Theses & Dissertations

The thesis investigates the most relevant economic and institutional determinants of happiness in some 93 countries worldwide, covering the period 2006-2019. We employ the Bayesian Model Averaging (BMA) fixed effect model (country demeaned and time demeaned) using a working panel data set with 651 observations. Our initial goal is to address the problem of model uncertainty in panel data models of happiness, aiming at selecting a set variables that are likely to be included in as "true" model of happiness. In addition, we aim to investigate the causal relationship running from selected economic and institutional variable to index of happiness. …


Frequentist Methods In Handling Misrepresentation Risk, Rexford Mawunyegah Akakpo 2021 Pentecost University College

Frequentist Methods In Handling Misrepresentation Risk, Rexford Mawunyegah Akakpo

Graduate Research Theses & Dissertations

A commonly encountered risk in insurance business is misrepresentation risk. Misrepresentation is a type of insurance fraud where a policyholder or a policy applicant falsifies his or her risk status in order to pay cheaper premiums for more expensive future risks. It is difficult and expensive for insurance companies to detect this kind of risk. With high cost of sophisticated underwriting, it becomes a norm for insurance companies to regularly rely on the policy applicant to self-report most of their risk statuses. We employ a frequentist approach by using expectation-maximization (EM) algorithm to carry out maximum likelihood estimation of the …


Addressing The Ecological Fallacy With Lagrangian Inference, Michael Schwob 2021 University of Nevada, Las Vegas

Addressing The Ecological Fallacy With Lagrangian Inference, Michael Schwob

Calvert Undergraduate Research Awards

Most epidemiologists elect to use statistical models that use population-level data to make inference on the spread of some virus or disease. This has become commonplace in the fields of epidemiology and biostatistics since most data used to construct and verify epidemic models are recorded at the population-level. Obtaining inference from a population-level model may be beneficial in studying the spread of disease in a homogeneous population, but the use of such models to describe a heterogeneous population results in inadequate inference. The inaccuracy of these models is further amplified when one tries to make individual-level inference from these population-level …


Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis 2021 Virginia Commonwealth University

Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis

Theses and Dissertations

High-throughput chromosome conformation capture technology (Hi-C) has revealed extensive DNA looping and folding into discrete 3D domains. These include Topologically Associating Domains (TADs) and chromatin loops, the 3D domains critical for cellular processes like gene regulation and cell differentiation. The relatively low resolution of Hi-C data (regions of several kilobases in size) prevents precise mapping of domain boundaries by conventional TAD/loop-callers. However, high resolution genomic annotations associated with boundaries, such as CTCF and members of cohesin complex, suggest a computational approach for precise location of domain boundaries.

We developed preciseTAD, an optimized machine learning framework that leverages a random …


Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu 2021 Virginia Commonwealth University

Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu

Theses and Dissertations

For analyzing large electronic health records (EHR) with time-to-event endpoints, such as in kidney transplantation, a major challenge is to provide an accurate risk analyses, while accounting for a multitude of epidemiological and statistical complexities. Motivated by a right-censored kidney transplantation EHR dataset derived from the United Network of Organ Sharing (UNOS), this dissertation, through a culmination of two interrelated yet distinctly different projects, focuses on developments of novel statistical procedures and methodologies to address some pressing issues arising in EHR-based research. In the first project, we aim to decouple the causal effects of treatments (here, studying subgroups, such as …


The Relevance Of Credit Risk In The Determination Of Commercial Banks’ Profitability: Evidence From Ghana, Godwin Kwabla Ekpe 2021 Northern Illinois University

The Relevance Of Credit Risk In The Determination Of Commercial Banks’ Profitability: Evidence From Ghana, Godwin Kwabla Ekpe

Graduate Research Theses & Dissertations

Existing empirical literature on the relationship between credit risk and bank’s profitability is replete with mixed results. This research investigates the probable effect of credit risk on banks’ profitability by examining the nature of the relationship between two measures of credit risk (Loss provisioning rate and Actual provisioning charge rate) and two measures ofprofitability (Return on assets and Return on Equity). The investigation is conducted using data on the Ghanaian banking industry. Various modeling techniques are used to fit the data, including frequentist beta regression and Bayesian beta regression models. The results across all models suggest negative linear relationship between …


Traffic Fatality Rate Prediction Based On Deep Neural Network And Bayesian Neural Network, Yiqun Hu 2021 Northern Illinois University

Traffic Fatality Rate Prediction Based On Deep Neural Network And Bayesian Neural Network, Yiqun Hu

Graduate Research Theses & Dissertations

There have been numerous studies on traffic accidents and their fatality rate. For this challenging machine learning regression problem, Neural Networks (NNs) have produced state-of-the-art data. Despite their success, they are often used in a fre- quentist scheme, which means they cannot account for uncertainty in their forecasts. BNNs are comprised of a Probabilistic Model and a Neural Network. The aim of such a design is to bring together the benefits of Neural Networks and stochastic modeling. Neural networks have the ability to approximate continuous functions uni- versally. Statistical models allow for the direct definition of a model with known …


A Time Series Analysis Approach To Forecasting Covid-19 Cases And Deaths: An Analysis Of Covid-19 Data In Colombia, Andrea Jackson-Sagredo 2021 Northern Illinois University

A Time Series Analysis Approach To Forecasting Covid-19 Cases And Deaths: An Analysis Of Covid-19 Data In Colombia, Andrea Jackson-Sagredo

Graduate Research Theses & Dissertations

The novel Coronavirus, known as COVID-19 is a highly contagious and transmissible infectious disease that has taken a toll throughout the entire world for over a year. The inner workings and long term effects of COVID-19 continue to be misunderstood. While COVID-19 has impacted all countries tremendously, Latin American countries and specifically Colombia have been impacted significantly by the virus. This thesis investigates the potential to forecast COVID-19 cases and deaths using Time Series Analysis methods and models for the South American country of Colombia. Time series analysis on Colombian COVID-19 data begins with data processing on a data set …


Stochastic Infection In Network Models With Applications To Pollution Analysis, Alexander Thor Wold 2021 Northern Illinois University

Stochastic Infection In Network Models With Applications To Pollution Analysis, Alexander Thor Wold

Graduate Research Theses & Dissertations

The continued adoption and escalation of commercial surface extraction techniques threatens to contaminate adjacent river networks across the coal mining landscape. We look to simulate the movement of this pollution across connected graph structures using stochastic block model methodologies, fueled from work and theory derived from exponential random graph models. We begin our study by applying our virtual experiment to the motivating material, later offering an exhaustive walk through of the simulation itself. Afterwards, we present our findings and their implications to water pollution analysis, emphasizing a need for this research and expanding on some inferential statistics left for later …


Digital Commons powered by bepress