Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,812 Full-Text Articles 23,893 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,812 full-text articles. Page 327 of 486.

A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew DeJongh, Nathan Tintle 2016 Dordt College

A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew Dejongh, Nathan Tintle

Statistical and Data Sciences: Faculty Publications

Numerous methods for classifying gene activity states based on gene expression data have been proposed for use in downstream applications, such as incorporating transcriptomics data into metabolic models in order to improve resulting flux predictions. These methods often attempt to classify gene activity for each gene in each experimental condition as belonging to one of two states: active (the gene product is part of an active cellular mechanism) or inactive (the cellular mechanism is not active). These existing methods of classifying gene activity states suffer from multiple limitations, including enforcing unrealistic constraints on the overall proportions of active and inactive …


Passive Visual Analytics Of Social Media Data For Detection Of Unusual Events, Kush Rustagi, Junghoon Chae 2016 Purdue University

Passive Visual Analytics Of Social Media Data For Detection Of Unusual Events, Kush Rustagi, Junghoon Chae

The Summer Undergraduate Research Fellowship (SURF) Symposium

Now that social media sites have gained substantial traction, huge amounts of un-analyzed valuable data are being generated. Posts containing images and text have spatiotemporal data attached as well, having immense value for increasing situational awareness of local events, providing insights for investigations and understanding the extent of incidents, their severity, and consequences, as well as their time-evolving nature. However, the large volume of unstructured social media data hinders exploration and examination. To analyze such social media data, the S.M.A.R.T system provides the analyst with an interactive visual spatiotemporal analysis and spatial decision support environment that assists in evacuation planning …


Design Optimization Of A Stochastic Multi-Objective Problem: Gaussian Process Regressions For Objective Surrogates, Juan Sebastian Martinez, Piyush Pandita, Rohit K. Tripathy, Ilias Bilionis 2016 Universidad de Los Andes - Colombia

Design Optimization Of A Stochastic Multi-Objective Problem: Gaussian Process Regressions For Objective Surrogates, Juan Sebastian Martinez, Piyush Pandita, Rohit K. Tripathy, Ilias Bilionis

The Summer Undergraduate Research Fellowship (SURF) Symposium

Multi-objective optimization (MOO) problems arise frequently in science and engineering situations. In an optimization problem, we want to find the set of input parameters that generate the set of optimal outputs, mathematically known as the Pareto frontier (PF). Solving the MOO problem is a challenge since expensive experiments can be performed only a constrained number of times and there is a limited set of data to work with, e.g. a roll-to-roll microwave plasma chemical vapor deposition (MPCVD) reactor for manufacturing high quality graphene. State-of-the-art techniques, e.g. evolutionary algorithms; particle swarm optimization, require a large amount of observations and do not …


Mediation Analysis For A Survival Outcome With Time-Varying Exposures, Mediators, And Confounders, Sheng-Hsuan Lin, Jessica G. Young, Roger Logan, Tyler J. VanderWeele 2016 Department of Biostatistics, Columbia Mailman School of Public Health

Mediation Analysis For A Survival Outcome With Time-Varying Exposures, Mediators, And Confounders, Sheng-Hsuan Lin, Jessica G. Young, Roger Logan, Tyler J. Vanderweele

Harvard University Biostatistics Working Paper Series

We propose an approach to conduct mediation analysis for survival data with time-varying exposures, mediators, and confounders. We identify certain interventional direct and indirect effects through a survival mediational g-formula and describe the required assumptions. We also provide a feasible parametric approach along with an algorithm and software to estimate these effects. We apply this method to analyze the Framingham Heart Study data to investigate the causal mechanism of smoking on mortality through coronary artery disease. The risk ratio of smoking 30 cigarettes per day for ten years compared with no smoking on mortality is 2.34 (95 % CI = …


Assessing The Association Between Quantitative Maturity And Student Performance In Simulation-Based And Non-Simulation Based Introductory Statistics, Nathan L. Tintle 2016 Dordt College

Assessing The Association Between Quantitative Maturity And Student Performance In Simulation-Based And Non-Simulation Based Introductory Statistics, Nathan L. Tintle

Faculty Work Comprehensive List

The recent simulation-based inference movement in algebra-based introductory statistics courses has provided preliminary evidence of improved student conceptual understanding and retention of key statistical concepts. However, little is known about whether these positive effects in courses using simulation-based inference are preferentially distributed across different types of students. Recent studies investigating predictors of student performance in traditional, algebra-based introductory statistics courses (Stat 101) have focused primarily on mathematical achievement or competencies in high school and early college. Little consideration has been given to how prior experience and competency with statistical thinking may be associated with student performance in college-level courses. In …


Oscillation Criteria For Third-Order Functional Differential Equations With Damping, Martin Bohner, Said R. Grace, Irena Jadlovska 2016 Missouri University of Science and Technology

Oscillation Criteria For Third-Order Functional Differential Equations With Damping, Martin Bohner, Said R. Grace, Irena Jadlovska

Mathematics and Statistics Faculty Research & Creative Works

This paper is a continuation of the recent study by Bohner et al [9] on oscillation properties of nonlinear third order functional differential equation under the assumption that the second order differential equation is nonoscillatory. We consider both the delayed and advanced case of the studied equation. The presented results correct and extend earlier ones. Several illustrative examples are included.


Sensitivity Of Trial Performance To Delay Outcomes, Accrual Rates, And Prognostic Variables Based On A Simulated Randomized Trial With Adaptive Enrichment, Tiachen Qian, Elizabeth Colantuoni, Aaron Fisher, Michael Rosenblum 2016 Johns Hopkins Bloomberg School of Public Health, Department of Biostatistics

Sensitivity Of Trial Performance To Delay Outcomes, Accrual Rates, And Prognostic Variables Based On A Simulated Randomized Trial With Adaptive Enrichment, Tiachen Qian, Elizabeth Colantuoni, Aaron Fisher, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

Adaptive enrichment designs involve rules for restricting enrollment to a subset of the population during the course of an ongoing trial. This can be used to target those who benefit from the experimental treatment. To leverage prognostic information in baseline variables and short-term outcomes, we use a semiparametric, locally efficient estimator, and investigate its strengths and limitations compared to standard estimators. Through simulation studies, we assess how sensitive the trial performance (Type I error, power, expected sample size, trial duration) is to different design characteristics. Our simulation distributions mimic features of data from the Alzheimer’s Disease Neuroimaging Initiative, and involve …


The Impact Of Patient Navigation On The Delivery Of Diagnostic Breast Cancer Care In The National Patient Navigation Research Program: A Prospective Meta-Analysis., Tracy A Battaglia, Julie S Darnell, Naomi Ko, Fred Snyder, Electra D Paskett, Kristen J Wells, Elizabeth M Whitley, Jennifer J Griggs, Anand Karnad, Heather Young, Victoria Warren-Mears, Melissa A Simon, Elizabeth Calhoun 2016 George Washington University

The Impact Of Patient Navigation On The Delivery Of Diagnostic Breast Cancer Care In The National Patient Navigation Research Program: A Prospective Meta-Analysis., Tracy A Battaglia, Julie S Darnell, Naomi Ko, Fred Snyder, Electra D Paskett, Kristen J Wells, Elizabeth M Whitley, Jennifer J Griggs, Anand Karnad, Heather Young, Victoria Warren-Mears, Melissa A Simon, Elizabeth Calhoun

Epidemiology Faculty Publications

Patient navigation is emerging as a standard in breast cancer care delivery, yet multi-site data on the impact of navigation at reducing delays along the continuum of care are lacking. The purpose of this study was to determine the effect of navigation on reaching diagnostic resolution at specific time points after an abnormal breast cancer screening test among a national sample. A prospective meta-analysis estimated the adjusted odds of achieving timely diagnostic resolution at 60, 180, and 365 days. Exploratory analyses were conducted on the pooled sample to identify which groups had the most benefit from navigation. Clinics from six …


Modeling Internet Traffic Generations Based On Users And Activities For Telecommunication Applications, Sara Stoudt, Pamela Badian-Pessot, Blanche Ngo Mahop, Erika Earley, Jordan Menter, Yadira Flores, Danielle Williams, Weijia Zhang, Liza Maharjan, Yixin Bao, Laura Rosenbauer, Van Nguyen, Veena Mendiratta, Nessy Tania 2016 Smith College

Modeling Internet Traffic Generations Based On Users And Activities For Telecommunication Applications, Sara Stoudt, Pamela Badian-Pessot, Blanche Ngo Mahop, Erika Earley, Jordan Menter, Yadira Flores, Danielle Williams, Weijia Zhang, Liza Maharjan, Yixin Bao, Laura Rosenbauer, Van Nguyen, Veena Mendiratta, Nessy Tania

Mathematics Sciences: Faculty Publications

A traffic generation model is a stochastic model of the data flow in a communication network. These models are useful during the development of telecommunication technologies and for analyzing the performance and capacity of various protocols, algorithms, and network topologies. We present here two modeling approaches for simulating internet traffic. In our models, we simulate the length and interarrival times of individual packets, the discrete unit of data transfer over the internet. Our first modeling approach is based on fitting data to known theoretical distributions. The second method utilizes empirical copulae and is completely data driven. Our models were based …


Propensity Score Based Methods For Estimating The Treatment Effects Based On Observational Studies., Younathan Abdia 2016 University of Louisville

Propensity Score Based Methods For Estimating The Treatment Effects Based On Observational Studies., Younathan Abdia

Electronic Theses and Dissertations

This dissertation consists of two interconnected research projects. The first project was a study of propensity scores based statistical methods for estimating the average treatment effect (ATE) and the average treatment effect among treated (ATT) when there are two treatment groups. The ATE is defined as the mean of the individual causal effects in the whole population, while ATT is defined as the treatment effect for the treated population. Propensity score based statistical methods, such as matching, regression, stratification, inverse probability weighting (IPW), and doubly robust (DR) methods were used to estimate the ATE and ATT. Simulation studies and case …


A Two-Strain Tb Model With Multiple Latent Stages, Azizeh Jabbari, Carlos Castillo-Chavez, Fereshteh Nazari, Baojun Song, Hossein Kheiri 2016 University of Tabriz

A Two-Strain Tb Model With Multiple Latent Stages, Azizeh Jabbari, Carlos Castillo-Chavez, Fereshteh Nazari, Baojun Song, Hossein Kheiri

Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works

A two-strain tuberculosis (TB) transmission model incorporating antibiotic-generated TB resistant strains and long and variable waiting periods within the latently infected class is introduced. The mathematical analysis is carried out when the waiting periods are modeled via parametrically friendly gamma distributions, a reasonable alternative to the use of exponential distributed waiting periods or to integral equations involving "arbitrary" distributions. The model supports a globally-asymptotically stable disease-free equilibrium when the reproduction number is less than one and an endemic equilibriums, shown to be locally asymptotically stable, or l.a.s., whenever the basic reproduction number is greater than one. Conditions for the existence …


Is There A Symmetric Version Of Hindman's Theorem?, Ethan Akin, Eli Glasner 2016 Missouri University of Science and Technology

Is There A Symmetric Version Of Hindman's Theorem?, Ethan Akin, Eli Glasner

Mathematics and Statistics Faculty Research & Creative Works

We show that there does not exist a symmetric version of Hindman's Theorem, or more explicitly, that the property of containing a symmetric IP-set is not divisible. We consider several related dynamics questions.


Model-Free Variable Screening, Sparse Regression Analysis And Other Applications With Optimal Transformations, Qiming Huang 2016 Purdue University

Model-Free Variable Screening, Sparse Regression Analysis And Other Applications With Optimal Transformations, Qiming Huang

Open Access Dissertations

Variable screening and variable selection methods play important roles in modeling high dimensional data. Variable screening is the process of filtering out irrelevant variables, with the aim to reduce the dimensionality from ultrahigh to high while retaining all important variables. Variable selection is the process of selecting a subset of relevant variables for use in model construction. The main theme of this thesis is to develop variable screening and variable selection methods for high dimensional data analysis. In particular, we will present two relevant methods for variable screening and selection under a unified framework based on optimal transformations.

In the …


Maximum Empirical Likelihood Estimation In U-Statistics Based General Estimating Equations, Lingnan Li 2016 Purdue University

Maximum Empirical Likelihood Estimation In U-Statistics Based General Estimating Equations, Lingnan Li

Open Access Dissertations

In the first part of this thesis, we study maximum empirical likelihood estimates (MELE's) in U-statistics based general estimating equations (UGEE's). Our technical maneuver is the jackknife empirical likelihood (JEL) approach. We give the local uniform asymptotic normality condition for the log-JEL for UGEE's. We derive the estimating equations for finding MELE's and provide their asymptotic normality. We obtain easy MELE's which have less computational burden than the usual MELE's and can be easily implemented using existing software. We investigate the use of side information of the data to improve efficiency. We exhibit that the MELE's are fully efficient, and …


Controlling For Confounding Network Properties In Hypothesis Testing And Anomaly Detection, Timothy La Fond 2016 Purdue University

Controlling For Confounding Network Properties In Hypothesis Testing And Anomaly Detection, Timothy La Fond

Open Access Dissertations

An important task in network analysis is the detection of anomalous events in a network time series. These events could merely be times of interest in the network timeline or they could be examples of malicious activity or network malfunction. Hypothesis testing using network statistics to summarize the behavior of the network provides a robust framework for the anomaly detection decision process. Unfortunately, choosing network statistics that are dependent on confounding factors like the total number of nodes or edges can lead to incorrect conclusions (e.g., false positives and false negatives). In this dissertation we describe the challenges that face …


Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier 2016 Purdue University

Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier

Open Access Dissertations

Increasingly, new sources of data are being incorporated into plant breeding pipelines. Enormous amounts of data from field phenomics and genotyping technologies places data mining and analysis into a completely different level that is challenging from practical and theoretical standpoints. Intelligent decision-making relies on our capability of extracting from data useful information that may help us to achieve our goals more efficiently. Many plant breeders, agronomists and geneticists perform analyses without knowing relevant underlying assumptions, strengths or pitfalls of the employed methods. The study endeavors to assess statistical learning properties and plant breeding applications of supervised and unsupervised machine learning …


Extreme-Strike And Small-Time Asymptotics For Gaussian Stochastic Volatility Models, Xin Zhang 2016 Purdue University

Extreme-Strike And Small-Time Asymptotics For Gaussian Stochastic Volatility Models, Xin Zhang

Open Access Dissertations

Asymptotic behavior of implied volatility is of our interest in this dissertation. For extreme strike, we consider a stochastic volatility asset price model in which the volatility is the absolute value of a continuous Gaussian process with arbitrary prescribed mean and covariance. By exhibiting a Karhunen-Loève expansion for the integrated variance, and using sharp estimates of the density of a general second-chaos variable, we derive asymptotics for the asset price density for large or small values of the variable, and study the wing behavior of the implied volatility in these models. Our main result provides explicit expressions for the first …


The Design And Statistical Analysis Of Single-Cell Rna-Sequencing Experiments, Faye H. Zheng 2016 Purdue University

The Design And Statistical Analysis Of Single-Cell Rna-Sequencing Experiments, Faye H. Zheng

Open Access Dissertations

Next-generation DNA- and RNA-sequencing (RNA-seq) technologies have expanded rapidly in both throughput and accuracy within the last decade. The momentum continues as emerging techniques become increasingly capable of profiling molecular content at the level of individual cells. One goal of this research is to put forward best practices in the design of single-cell RNA-sequencing (scRNA-seq) experiments, specifically as it relates to choices regarding the trade-off between sequencing depth and sample size. In addition to general guidelines, an interactive tool is presented to aid researchers in making experiment-specific decisions that are informed by real data and practical constraints. Further, a new …


Some Nonparametric Ordered Restricted Inference Problems In The Context Of A Statistical Education Study, Bradford M. Dykes 2016 Western Michigan University

Some Nonparametric Ordered Restricted Inference Problems In The Context Of A Statistical Education Study, Bradford M. Dykes

Dissertations

Over the past 10 years, the Department of Statistics at Western Michigan University has developed a question generating system that can be used for creating multiple forms of exams, quizzes and homework for online and face-to-face use. This system can also be used to provide students with a form of instantaneous feedback. With the goal of analyzing how different levels of feedback in an online learning environment impacts students' performance on assignments, this study presents data collected on two semesters of students enrolled in three different meeting types (strictly online, typical face-to-face, and honors face-to-face) of an introductory Statistics course. …


Diabetes Is Associated With Cerebrovascular But Not Alzheimer's Disease Neuropathology, Erin L. Abner, Peter T. Nelson, Richard J. Kryscio, Frederick A. Schmitt, David W. Fardo, Randall L. Woltjer, Nigel J. Cairns, Lei Yu, Hiroko H. Dodge, Chengjie Xiong, Kamal Masaki, Suzanne L. Tyas, David A. Bennett, Julie A. Schneider, Zoe Arvanitakis 2016 University of Kentucky

Diabetes Is Associated With Cerebrovascular But Not Alzheimer's Disease Neuropathology, Erin L. Abner, Peter T. Nelson, Richard J. Kryscio, Frederick A. Schmitt, David W. Fardo, Randall L. Woltjer, Nigel J. Cairns, Lei Yu, Hiroko H. Dodge, Chengjie Xiong, Kamal Masaki, Suzanne L. Tyas, David A. Bennett, Julie A. Schneider, Zoe Arvanitakis

Sanders-Brown Center on Aging Faculty Publications

INTRODUCTION: The relationship of diabetes to specific neuropathologic causes of dementia is incompletely understood.

METHODS: We used logistic regression to evaluate the association between diabetes and infarcts, Braak neurofibrillary tangle stage, and neuritic plaque score in 2365 autopsied persons. In a subset of >1300 persons with available cognitive data, we examined the association between diabetes and cognition using Poisson regression.

RESULTS: Diabetes increased odds of brain infarcts (odds ratio [OR] = 1.57, P < .0001), specifically lacunes (OR = 1.71, P < .0001), but not Alzheimer's disease neuropathology. Diabetes plus infarcts was associated with lower cognitive scores at end of life than infarcts or diabetes alone, and diabetes plus high level of Alzheimer's neuropathologic changes was associated with lower mini-mental state examination scores than the pathology alone.

DISCUSSION: This study supports the conclusions that diabetes increases the risk of cerebrovascular but not Alzheimer's disease pathology, and at least some of diabetes' relationship to …


Digital Commons powered by bepress