Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Statistical Methodology

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 931 - 960 of 1562

Full-Text Articles in Statistics and Probability

Enhancing Models And Measurements Of Traffic-Related Air Pollutants For Health Studies Using Dispersion Modeling And Bayesian Data Fusion, Stuart A. Batterman, Veronica J. Berrocal, Chad Milando, Owais Gilani, Saravanan Arunachalam, K. Max Zhang Jan 2020

Enhancing Models And Measurements Of Traffic-Related Air Pollutants For Health Studies Using Dispersion Modeling And Bayesian Data Fusion, Stuart A. Batterman, Veronica J. Berrocal, Chad Milando, Owais Gilani, Saravanan Arunachalam, K. Max Zhang

Faculty Journal Articles

Research Report 202 describes a study led by Dr. Stuart Batterman at the University of Michigan, Ann Arbor and colleagues. The investigators evaluated the ability to predict traffic-related air pollution using a variety of methods and models, including a line source air pollution dispersion model and sophisticated spatiotemporal Bayesian data fusion methods. Exposure assessment for traffic-related air pollution is challenging because the pollutants are a complex mixture and vary greatly over space and time. Because extensive direct monitoring is difficult and expensive, a number of modeling approaches have been developed, but each model has its own limitations and errors.

Dr. …


Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals, Sihai Dave Zhao, Yet Tian Nguyen Jan 2020

Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals, Sihai Dave Zhao, Yet Tian Nguyen

Mathematics & Statistics Faculty Publications

It is frequently of interest to identify simultaneous signals, defined as features that exhibit statistical significance across each of several independent experiments. For example, genes that are consistently differentially expressed across experiments in different animal species can reveal evolutionarily conserved biological mechanisms. However, in some problems the test statistics corresponding to these features can have complicated or unknown null distributions. This paper proposes a novel nonparametric false discovery rate control procedure that can identify simultaneous signals even without knowing these null distributions. The method is shown, theoretically and in simulations, to asymptotically control the false discovery rate. It was also …


Generalization Of Kullback-Leibler Divergence For Multi-Stage Diseases: Application To Diagnostic Test Accuracy And Optimal Cut-Points Selection Criterion, Chen Mo Jan 2020

Generalization Of Kullback-Leibler Divergence For Multi-Stage Diseases: Application To Diagnostic Test Accuracy And Optimal Cut-Points Selection Criterion, Chen Mo

College of Graduate Studies: Theses & Dissertations

The Kullback-Leibler divergence (KL), which captures the disparity between two distributions, has been considered as a measure for determining the diagnostic performance of an ordinal diagnostic test. This study applies KL and further generalizes it to comprehensively measure the diagnostic accuracy test for multi-stage (K > 2) diseases, named generalized total Kullback-Leibler divergence (GTKL). Also, GTKL is proposed as an optimal cut-points selection criterion for discriminating subjects among different disease stages. Moreover, the study investigates a variety of applications of GTKL on measuring the rule-in/out potentials in the single-stage and multi-stage levels. Intensive simulation studies are conducted to compare the performance …


Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten Jan 2020

Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten

Theses and Dissertations

Humans are exposed to multiple chemicals every day. Epidemiological studies have shown that chemical mixtures are associated with cancers, allergies, neurodevelopmental disorders, and other adverse health effects. To assess these associations, investigators are increasingly using chemical mixture approaches like weighted quantile sum (WQS) regression. In these studies, the research objectives are to determine whether a mixture of correlated chemicals is associated with an adverse health outcome and to identify the important chemicals. However, as experimental equipment measures each exposure to a chemical-specific detection limit, the exposures are unknown between zero and the detection limit. Indeed, the number of exposures below …


The Effect Of Time And Temperature On The Quality Of Latent Fingerprints On Incandescent Lightbulbs, Varying Donors Age And Sex, Kinaysha M. Collazo Maldonado Jan 2020

The Effect Of Time And Temperature On The Quality Of Latent Fingerprints On Incandescent Lightbulbs, Varying Donors Age And Sex, Kinaysha M. Collazo Maldonado

Master of Science in Forensic Science Directed Research Projects

Fingerprints are used as a means of identification, but there are no established methodologies to determine time since deposition of latent fingerprints by visual means alone. This research considered the influence of age and sex on the quality of recovered latent prints from lit and unlit lightbulbs from 1 to 10 days, using accumulated degree hours (ADH) to account for both heat and time simultaneously. Two male and two female donors (one of each aged40 years) were used. A thermal imaging camera was used to monitor the lightbulbs top and middle regions, which were significantly different (p≤0.05) for the experimental …


Distribution Of Human Exposure To Ozone During Commuting Hours In Connecticut Using The Cellular Device Network, Owais Gilani, Simon Urbanek, Michael J. Kane Jan 2020

Distribution Of Human Exposure To Ozone During Commuting Hours In Connecticut Using The Cellular Device Network, Owais Gilani, Simon Urbanek, Michael J. Kane

Faculty Journal Articles

Epidemiologic studies have established associations between various air pollutants and adverse health outcomes for adults and children. Due to high costs of monitoring air pollutant concentrations for subjects enrolled in a study, statisticians predict exposure concentrations from spatial models that are developed using concentrations monitored at a few sites. In the absence of detailed information on when and where subjects move during the study window, researchers typically assume that the subjects spend their entire day at home, school, or work. This assumption can potentially lead to large exposure assignment bias. In this study, we aim to determine the distribution of …


How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller Jan 2020

How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller

CMC Senior Theses

In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …


Process Based Analysis Of Fluvial Stratigraphic Record: Middle Pennsylvanian Allegheny Formation, North-Central Wv, Oluwasegun O. Abatan Jan 2020

Process Based Analysis Of Fluvial Stratigraphic Record: Middle Pennsylvanian Allegheny Formation, North-Central Wv, Oluwasegun O. Abatan

Graduate Theses, Dissertations, and Problem Reports (ETD)

Fluvial deposits represent some of the best hydrocarbon reservoirs, but the quality of fluvial reservoirs varies depending on the reservoir architecture, which is controlled by allogenic and autogenic processes. Allogenic controls, including paleoclimate, tectonics, and glacio-eustasy, have long been debated as dominant controls in the deposition of fluvial strata. However, recent research has questioned the validity of this cyclicity and may indicate major influence from autogenic controls. To further investigate allogenic controls on stratal order, I analyzed the facies architecture, geomorphology, paleohydrology, and the stratigraphic framework of the Middle Pennsylvanian Allegheny Formation (MPAF), a fluvial depositional system in the Appalachian …


Generalized Matrix Decomposition Regression: Estimation And Inference For Two-Way Structured Data, Yue Wang, Ali Shojaie, Tim Randolph, Jing Ma Dec 2019

Generalized Matrix Decomposition Regression: Estimation And Inference For Two-Way Structured Data, Yue Wang, Ali Shojaie, Tim Randolph, Jing Ma

UW Biostatistics Working Paper Series

Analysis of two-way structured data, i.e., data with structures among both variables and samples, is becoming increasingly common in ecology, biology and neuro-science. Classical dimension-reduction tools, such as the singular value decomposition (SVD), may perform poorly for two-way structured data. The generalized matrix decomposition (GMD, Allen et al., 2014) extends the SVD to two-way structured data and thus constructs singular vectors that account for both structures. While the GMD is a useful dimension-reduction tool for exploratory analysis of two-way structured data, it is unsupervised and cannot be used to assess the association between such data and an outcome of interest. …


Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang Dec 2019

Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang

Statistical Science Theses and Dissertations

This dissertation contains two topics: (1) A Comparative Study of Statistical Methods for Quantifying and Testing Between-study Heterogeneity in Meta-analysis with Focus on Rare Binary Events; (2) Estimation of Variances in Cluster Randomized Designs Using Ranked Set Sampling.

Meta-analysis, the statistical procedure for combining results from multiple studies, has been widely used in medical research to evaluate intervention efficacy and safety. In many practical situations, the variation of treatment effects among the collected studies, often measured by the heterogeneity parameter, may exist and can greatly affect the inference about effect sizes. Comparative studies have been done for only one or …


Statistical Inference For Networks Of High-Dimensional Point Processes, Xu Wang, Mladen Kolar, Ali Shojaie Dec 2019

Statistical Inference For Networks Of High-Dimensional Point Processes, Xu Wang, Mladen Kolar, Ali Shojaie

UW Biostatistics Working Paper Series

Fueled in part by recent applications in neuroscience, high-dimensional Hawkes process have become a popular tool for modeling the network of interactions among multivariate point process data. While evaluating the uncertainty of the network estimates is critical in scientific applications, existing methodological and theoretical work have only focused on estimation. To bridge this gap, this paper proposes a high-dimensional statistical inference procedure with theoretical guarantees for multivariate Hawkes process. Key to this inference procedure is a new concentration inequality on the first- and second-order statistics for integrated stochastic processes, which summarizes the entire history of the process. We apply this …


Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova Dec 2019

Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova

College of Education and Human Sciences: Dissertations, Theses, and Student Research

When assessing a certain characteristic or trait using a multiple item measure, quality of that measure can be assessed by examining the reliability. To avoid multiple time points, reliability can be represented by internal consistency, which is most commonly calculated using Cronbach’s coefficient alpha. Almost every time human participants are involved in research, there is missing data involved. Missing data means that even though complete data were expected to be collected, some data are missing. Missing data can follow different patterns as well as be the result of different mechanisms. One traditional way to deal with missing data is listwise …


Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden Dec 2019

Communications And Methodologies In Crime Geography: Contemporary Approaches To Disseminating Criminal Incidence And Research, Mitchell Ogden

Electronic Theses and Dissertations

Many tools exist to assist law enforcement agencies in mitigating criminal activity. For centuries, academics used statistics in the study of crime and criminals, and more recently, police departments make use of spatial statistics and geographic information systems in that pursuit. Clustering and hot spot methods of analysis are popular in this application for their relative simplicity of interpretation and ease of process. With recent advancements in geospatial technology, it is easier than ever to publicly share data through visual communication tools like web applications and dashboards. Sharing data and results of analyses boosts transparency and the public image of …


Optimal Design For A Causal Structure, Zaher Kmail Aug 2019

Optimal Design For A Causal Structure, Zaher Kmail

Department of Statistics: Dissertations, Theses, and Student Research

Linear models and mixed models are important statistical tools. But in many natural phenomena, there is more than one endogenous variable involved and these variables are related in a sophisticated way. Structural Equation Modeling (SEM) is often used to model the complex relationships between the endogenous and exogenous variables. It was first implemented in research to estimate the strength and direction of direct and indirect effects among variables and to measure the relative magnitude of each causal factor.

Historically, traditional optimal design theory focuses on univariate linear, nonlinear, and mixed models. There is no current literature on the subject of …


Effect Of Cross-Validation On The Output Of Multiple Testing Procedures, Josh Dallas Price Aug 2019

Effect Of Cross-Validation On The Output Of Multiple Testing Procedures, Josh Dallas Price

Graduate Theses and Dissertations

High dimensional data with sparsity is routinely observed in many scientific disciplines. Filtering out the signals embedded in noise is a canonical problem in such situations requiring multiple testing. The Benjamini--Hochberg procedure using False Discovery Rate control is the gold standard in large scale multiple testing. In Majumder et al. (2009) an internally cross-validated form of the procedure is used to avoid a costly replicate study and the complications that arise from population selection in such studies (i.e. extraneous variables). I implement this procedure and run extensive simulation studies under increasing levels of dependence among parameters and different data generating …


Taking Multiple Regression Analysis To Task: A Review Of Mindware: Tools For Smart Thinking, By Richard Nisbett (2015), Jason Makansi Jul 2019

Taking Multiple Regression Analysis To Task: A Review Of Mindware: Tools For Smart Thinking, By Richard Nisbett (2015), Jason Makansi

Numeracy

Richard Nisbett. 2015. Mindware: Tools for Smart Thinking.(New York, NY: Farrar, Strauss, and Giroux). 336 pp. ISBN: 9780374536244

Nisbett, a psychologist, may not achieve his stated goal of teaching readers to “effortlessly” extend their common sense when it comes to quantitative analysis applied to everyday issues, but his critique of multiple regression analysis (MRA) in the middle chapters of Mindware is worth attention from, and contemplation by, the QL/QR and Numeracy community. While in at least one other source, Nisbett’s critique has been called a “crusade” against MRA, what he really advocates is that it not be used as …


Using Meta-Analysis To Assess Affective Outcomes In A Multi-Course Qr Module Intervention, James Friedrich, Kelley D. Strawn Jul 2019

Using Meta-Analysis To Assess Affective Outcomes In A Multi-Course Qr Module Intervention, James Friedrich, Kelley D. Strawn

Numeracy

When quantitative reasoning(QR) interventions share a common hypothesis or goal, a promising approach for evaluation involves integrating separate analyses through the use of meta-analysis. This paper reports an assessment of a module-based QR intervention distributed across 20 courses at a single institution. Topics and participating courses were diverse, including arts & humanities, quantitative behavioral sciences, and natural sciences & mathematics groupings, but all addressed the shared affective goals of reducing student QR self-doubt and increasing appreciation for QR value and utility. With a local framework to guide module development, we assess these outcomes using reliable self-report measures in a pre-post …


Interpreting Patient Reported Outcomes In Orthopaedic Surgery: A Systematic Review, Shgufta Docter, Zina Fathalla, Michael Lukacs, Michaela Khan, Morgan Jennings, Shu-Hsuan Liu, Dong Zi, Dianne Bryant Jun 2019

Interpreting Patient Reported Outcomes In Orthopaedic Surgery: A Systematic Review, Shgufta Docter, Zina Fathalla, Michael Lukacs, Michaela Khan, Morgan Jennings, Shu-Hsuan Liu, Dong Zi, Dianne Bryant

Western Research Forum

Background: Reporting methods of patient reported outcome measures (PROMs) vary in orthopaedic surgery literature. While most studies report statistical significance, the interpretation of results would be improved if authors reported confidence intervals (CIs), the minimally clinically important difference (MCID), and number needed to treat (NNT).

Objective: To assess the quality and interpretability of reporting the results of PROMs. To evaluate reporting, we will assess the proportion of studies that reported (1) 95% CIs, (2) MCID, and (3) NNT. To evaluate interpretation, we will assess the proportion of studies that discussed results using the MCID or the effect sizes and how …


Modelling Weighted Signed Networks, Alberto Caimo, Isabella Gollini Jun 2019

Modelling Weighted Signed Networks, Alberto Caimo, Isabella Gollini

Conference papers

In this paper we introduce a new modelling approach to analyse weighted signed networks by assuming that their generative process consists of two models: the interaction model which describes the overall connectivity structure of the relations in the network without taking into account neither the weight nor the sign of the dyadic relations; and the conditional weighted signed network model describes how the edge signed weights form given the interaction structure. We then show how this modelling approach can facilitate the interpretation of the overall network process. Finally, we adopt a Bayesian inferential approach to illustrate the new methodology by …


Optimizing Electrospun Ceramic Nanofiber Strength Through Two-Step Sintering, Michael Ross Jun 2019

Optimizing Electrospun Ceramic Nanofiber Strength Through Two-Step Sintering, Michael Ross

Materials Engineering

Two-step sintering (TSS) consists of a high-temperature step and immediate cooling to a sintering temperature for an extended sintering time, where grain growth is suppressed by severe densification during the high-temperature step. TSS is adopted to enhance mechanical properties of electrospun ceramic nanofibers (CNFs), a class of porous ceramics used for environmental remediation, optoelectronics, and filtration. PVP and Ga(NO3)3 nanofiber mesh, provided by Lawrence Livermore National Laboratory, was shaped, oxidized, and two-step sintered to form a nanocrystalline β-Ga2O3 CNF tube using a high-temperature step of 1,000oC. Sintering temperatures and times varied from …


Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan May 2019

Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan

Dissertations

Spatial and temporal dependencies are ubiquitous properties of data in numerous domains. The popularity of spatial and temporal data mining has thus grown with the increasing prevalence of massive data. The presence of spatial and temporal attributes not only provides complementary useful perspectives, but also poses new challenges to the representation and integration into the learning procedure. In this dissertation, the involved spatial and temporal dependencies are explored with three genres: sample-wise, feature-wise, and target-wise. A family of novel methodologies is developed accordingly for the dependency representation in respective scenarios.

First, dependencies among discrete, continuous and repeated observations are studied …


Advances In Measurement Error Modeling, Linh Nghiem May 2019

Advances In Measurement Error Modeling, Linh Nghiem

Statistical Science Theses and Dissertations

Measurement error in observations is widely known to cause bias and a loss of power when fitting statistical models, particularly when studying distribution shape or the relationship between an outcome and a variable of interest. Most existing correction methods in the literature require strong assumptions about the distribution of the measurement error, or rely on ancillary data which is not always available. This limits the applicability of these methods in many situations. Furthermore, new correction approaches are also needed for high-dimensional settings, where the presence of measurement error in the covariates adds another level of complexity to the desirable structure …


Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams May 2019

Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams

Statistical Science Theses and Dissertations

Much recent research has focused on methods for combining a probability sample with a non-probability sample to improve estimation by making use of information from both sources. If units exist in both samples, it becomes necessary to link the information from the two samples for these units. Record linkage is a technique to link records from two lists that refer to the same unit but lack a unique identifier across both lists. Record linkage assigns a probability to each potential pair of records from the lists so that principled matching decisions can be made. Because record linkage is a probabilistic …


Measuring Clinical Weight Loss In Young Children With Severe Obesity: Comparison Of Outcomes Using Zbmi, Modified Zbmi, And Percent Of 95th Percentile, Carolyn Bates May 2019

Measuring Clinical Weight Loss In Young Children With Severe Obesity: Comparison Of Outcomes Using Zbmi, Modified Zbmi, And Percent Of 95th Percentile, Carolyn Bates

Research Days

No abstract provided.


What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren May 2019

What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren

The Qualitative Report

Statistical and research consulting is defined as the collaboration of a statistician or methodologist with another professional for devising solutions to research problems. An in-depth, interview qualitative approach was taken to answer the research question of what makes a good research consultant. The authors interviewed four faculty members in the field of statistics and research methods and two experienced graduate student consultants. In-depth, face-to-face interviews revealed common themes regarding consultancy skills, resourcefulness, communication and interpersonal skills. The participants discussed how to improve consulting sessions and deal with clients with different statistics levels and backgrounds. Participants felt there was no difference …


Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia May 2019

Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia

SMU Data Science Review

In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …


Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt May 2019

Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt

Electronic Theses and Dissertations

A statistician's job is to produce statistical models. When these models are precise and unbiased, we can relate them to new data appropriately. However, when data sets have missing values, assumptions to statistical methods are violated and produce biased results. The statistician's objective is to implement methods that produce unbiased and accurate results. Research in missing data is becoming popular as modern methods that produce unbiased and accurate results are emerging, such as MICE in R, a statistical software. Using real data, we compare four common imputation methods, in the MICE package in R, at different levels of missingness. The …


Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis, Samuel Anyaso-Samuel May 2019

Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis, Samuel Anyaso-Samuel

Boise State University Theses and Dissertations

The statistical process control (SPC) chart is an effective tool for the analysis, interpretation, and visualization of data from sequential processes. Commonly used SPC charts such as the Shewhart, CUSUM and EWMA charts are widely implemented in detecting distributional shifts in various processes. With recent scientific and technological advancements, massive amounts of data continue to be generated by production, medical, agricultural and many other industrial processes. Conventional SPC charts have significant drawbacks in monitoring such processes, specifically when the velocity of the data flow is greater than the run time of the monitoring procedure. In the literature, dynamic sampling control …


Vulnerability Of Industrial Facilities In The Lower Mississippi River Industrial Corridor To Relative Sea Level Rise And Tropical Cyclone Storm Surge, Joseph Blake Harris Mar 2019

Vulnerability Of Industrial Facilities In The Lower Mississippi River Industrial Corridor To Relative Sea Level Rise And Tropical Cyclone Storm Surge, Joseph Blake Harris

LSU Doctoral Dissertations

Relative sea level rise (RSLR) and tropical cyclone-induced storm surge are major threats to the Lower Mississippi River Industrial Corridor (LMRIC) which has approximately 120 industrial complexes located within the corridor. Spatial interpolation methods were applied to the 2004 National Oceanic and Atmospheric published Technical Report #50 subsidence dataset and cross-validation techniques were used to determine the accuracy of each method. Digital elevation models (DEMs) were created for the years 2025, 2050, and 2075, based on these predictive surface of subsidence rates. Future DEMs were utilized to model RSLR and determine the extent of storm surge on the LMRIC by …


Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan Mar 2019

Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan

COBRA Preprint Series

One of the major goals in large-scale genomic studies is to identify genes with a prognostic impact on time-to-event outcomes which provide insight into the disease's process. With rapid developments in high-throughput genomic technologies in the past two decades, the scientific community is able to monitor the expression levels of tens of thousands of genes and proteins resulting in enormous data sets where the number of genomic features is far greater than the number of subjects. Methods based on univariate Cox regression are often used to select genomic features related to survival outcome; however, the Cox model assumes proportional hazards …