Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2017

Discipline
Institution
Keyword
Publication
Publication Type

Articles 541 - 570 of 628

Full-Text Articles in Statistics and Probability

Socio-Demographic Determinants Of Racial Disparities In Stage At Diagnosis Of Prostate Cancer In New York State, Christophe Maxime Fokoua Dongmo Jan 2017

Socio-Demographic Determinants Of Racial Disparities In Stage At Diagnosis Of Prostate Cancer In New York State, Christophe Maxime Fokoua Dongmo

Legacy Theses & Dissertations (2009 - 2024)

ABSTRACT


Kz Spatial Wave Separation With Applications To Atmospheric Data, Ming Luo Jan 2017

Kz Spatial Wave Separation With Applications To Atmospheric Data, Ming Luo

Legacy Theses & Dissertations (2009 - 2024)

Unlike one-dimensional wave reconstruction, reconstruction 2D spatial wave via Fourier Transform doesn’t look like a non-parametric algorithm. In other words, we need the wave frequency and wave direction information to recover the spatial wave via Fourier Transform, especially when the stress of noise is present. The direct consequence is that accurate estimations of wave parameters are need for reconstructing of spatial waves. To this end, we propose to improve the accuracy of motion image scale detection and parameter estimations with optimization based on Kolmogorov-Zurbenko periodogram (KZP) information. Related methods and algorithms are denoted under the name of Kolmogorov-Zurbenko wave separations. …


Integrative Pathway Analysis Pipeline For Mirna And Mrna Data, Diana Mabel Diaz Herrera Jan 2017

Integrative Pathway Analysis Pipeline For Mirna And Mrna Data, Diana Mabel Diaz Herrera

Wayne State University Theses

The identification of pathways that are involved in a particular phenotype helps us understand the underlying biological processes. Traditional pathway analysis techniques aim to infer the impact on individual pathways using only mRNA levels. However, recent studies showed that gene expression alone is unable to capture the whole picture of biological phenomena. At the same time, MicroRNAs (miRNAs) are newly discovered gene regulators that have shown to play an important role in diagnosis, and prognosis for different types of diseases. Current pathway analysis techniques do not take miRNAs into consideration. In this project, we investigate the effect of integrating miRNA …


An Empirical Demonstration Of The Need For Exact Tests, Vance W. Berger Jan 2017

An Empirical Demonstration Of The Need For Exact Tests, Vance W. Berger

Journal of Modern Applied Statistical Methods

The robustness of parametric analyses is rarely questioned or qualified. Robustness, generally understood, means the exact and approximate p-values will lie on the same side of alpha for any reasonable data set; and 1) any data set would qualify as reasonable and 2) robustness holds universally, for all alpha levels and approximations. For this to be true, the approximation would need to be perfect all of the time. Any discrepancy between the approximation and the exact p-value, for any combination of alpha level and data set, would constitute a violation. Clearly, this is not true, and when confronted with this …


Teaching Size And Power Properties Of Hypothesis Tests Through Simulations, Suleyman Taspinar, Osman Dogan Jan 2017

Teaching Size And Power Properties Of Hypothesis Tests Through Simulations, Suleyman Taspinar, Osman Dogan

Publications and Research

In this study, we review the graphical methods suggested in Davidson and MacKinnon (Davidson, Russell, and James G. MacKinnon. 1998. “Graphical Methods for Investigating the Size and Power of Hypothesis Tests.” The Manchester School 66 (1): 1–26.) that can be used to investigate size and power properties of hypothesis tests for undergraduate and graduate econometrics courses. These methods can be used to assess finite sample properties of various hypothesis tests through simulation studies. In addition, these methods can be effectively used in classrooms to reinforce students’ understanding of basic hypothesis testing concepts such as Type I error, Type II error, …


06. Sas Program Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic Jan 2017

06. Sas Program Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic

Design and Analysis of Experiments

SAS program files for use with Design and Analysis of Experiments.


08. R Program Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic Jan 2017

08. R Program Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic

Design and Analysis of Experiments

R program files for use with Design and Analysis of Experiments.


07. Sas Data Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic Jan 2017

07. Sas Data Files For Design And Analysis Of Experiments, Angela Dean, Dan Voss, Danel Draguljic

Design and Analysis of Experiments

SAS Data files for use with Design and Analysis of Experiments.


Flexibility Of Projective-Planar Embeddings, John Maharry, Neil Robertson, Vaidy Sivaraman, Dan Slilaty Jan 2017

Flexibility Of Projective-Planar Embeddings, John Maharry, Neil Robertson, Vaidy Sivaraman, Dan Slilaty

Mathematics and Statistics Faculty Publications

Given two embeddings σ1 and σ2 of a labeled nonplanar graph in the projective plane, we give a collection of maneuvers on projective-planar embeddings that can be used to take σ1 to σ2


The Nonparametric Estimation Of Elliptical Distributions, Panfeng Liang Jan 2017

The Nonparametric Estimation Of Elliptical Distributions, Panfeng Liang

Open Access Theses & Dissertations

In practice, many multivariate datasets have identical marginal distributions. Elliptical distributions can be used to model many of those datasets. In this Thesis, we will propose a Bayesian method using Markov chain Monte Carlo (MCMC) methods to estimate the density function underlying multivariate datasets assuming it is an elliptical distribution.


Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett Jan 2017

Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett

Theses and Dissertations--Statistics

The extension of the CJS model to include individual, continuous, time-varying covariates relies on the estimation of covariate values on occasions on which individuals were not captured. Fitting this model in a Bayesian framework typically involves the implementation of a Markov chain Monte Carlo (MCMC) algorithm, such as a Gibbs sampler, to sample from the posterior distribution. For large data sets with many missing covariate values that must be estimated, this creates a computational issue, as each iteration of the MCMC algorithm requires sampling from the full conditional distributions of each missing covariate value. This dissertation examines two solutions to …


Novel Computational Methods For Censored Data And Regression, Yifan Yang Jan 2017

Novel Computational Methods For Censored Data And Regression, Yifan Yang

Theses and Dissertations--Statistics

This dissertation can be divided into three topics. In the first topic, we derived a recursive algorithm for the constrained Kaplan-Meier estimator, which promotes the computation speed up to fifty times compared to the current method that uses EM algorithm. We also showed how this leads to the vast improvement of empirical likelihood analysis with right censored data. After a brief review of regularized regressions, we investigated the computational problems in the parametric/non-parametric hybrid accelerated failure time models and its regularization in a high dimensional setting. We also illustrated that, when the number of pieces increases, the discussed models are …


Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu Jan 2017

Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu

Theses and Dissertations--Statistics

Firstly, we reviewed some popular nonparameteric regression methods during the past several decades. Then we extended the compound estimation (Charnigo and Srinivasan [2011]) to adapt random design points and heteroskedasticity and proposed a modified Cp criteria for tuning parameter selection. Moreover, we developed a DCp criteria for tuning paramter selection problem in general nonparametric derivative estimation. This extends GCp criteria in Charnigo, Hall and Srinivasan [2011] with random design points and heteroskedasticity. Next, we proposed a change point detection method via compound estimation for both fixed design and random design case, the adaptation of heteroskedasticity was considered for the method. …


A Predictive Probability Interim Design For Phase Ii Clinical Trials With Continuous Endpoints, Meng Liu Jan 2017

A Predictive Probability Interim Design For Phase Ii Clinical Trials With Continuous Endpoints, Meng Liu

Theses and Dissertations--Epidemiology and Biostatistics

Phase II clinical trials aim to potentially screen out ineffective and identify effective therapies to move forward to randomized phase III trials. Single-arm studies remain the most utilized design in phase II oncology trials, especially in scenarios where a randomized design is simply not practical. Due to concerns regarding excessive toxicity or ineffective new treatment strategies, interim analyses are typically incorporated in the trial, and the choice of statistical methods mainly depends on the type of primary endpoints. For oncology trials, the most common primary objectives in phase II trials include tumor response rate (binary endpoint) and progression disease-free survival …


An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert Jan 2017

An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert

Theses and Dissertations--Epidemiology and Biostatistics

It is estimated that Periodontal Diseases effects up to 90% of the adult population. Given the complexity of the host environment, many factors contribute to expression of the disease. Age, Gender, Socioeconomic Status, Smoking Status, and Race/Ethnicity are all known risk factors, as well as a handful of known comorbidities. Certain vitamins and minerals have been shown to be protective for the disease, while some toxins and chemicals have been associated with an increased prevalence. The role of toxins, chemicals, vitamins, and minerals in relation to disease is believed to be complex and potentially modified by known risk factors. A …


Statistical Analyses To Detect And Refine Genetic Associations With Neurodegenerative Diseases, Yuriko Katsumata Jan 2017

Statistical Analyses To Detect And Refine Genetic Associations With Neurodegenerative Diseases, Yuriko Katsumata

Theses and Dissertations--Epidemiology and Biostatistics

Dementia is a clinical state caused by neurodegeneration and characterized by a loss of function in cognitive domains and behavior. Alzheimer’s disease (AD) is the most common form of dementia. Although the amyloid β (Aβ) protein and hyperphosphorylated tau aggregates in the brain are considered to be the key pathological hallmarks of AD, the exact cause of AD is yet to be identified. In addition, clinical diagnoses of AD can be error prone. Many previous studies have compared the clinical diagnosis of AD against the gold standard of autopsy confirmation and shown substantial AD misdiagnosis Hippocampal sclerosis of aging (HS-Aging) …


Improved Simultaneous Estimation Of Location And System Reliability Via Shrinkage Ideas, Beidi Qiang Jan 2017

Improved Simultaneous Estimation Of Location And System Reliability Via Shrinkage Ideas, Beidi Qiang

Theses and Dissertations

In decision theory, when several parameters need to be estimated simultaneously, many standard estimators can be improved, in terms of a combined loss function. The problem of finding such estimators has been well studied in the literature, but mostly under parametric settings, which is inappropriate for heavy-tailed distributions. In the first part of this dissertation, a robust simultaneous estimator of location is proposed using the shrinkage idea. A nonparametric Bayesian estimator is also discussed as an alternative. The proposed estimators do not assume a specific parametric distribution and they do not require the existence of finite moments. The performance of …


A Comprehensive Analysis Of Team Streakiness In Major League Baseball: 1962-2016, Paul H. Kvam, Zezhong Chen Jan 2017

A Comprehensive Analysis Of Team Streakiness In Major League Baseball: 1962-2016, Paul H. Kvam, Zezhong Chen

Department of Math & Statistics Faculty Publications

A baseball team would be considered “streaky” if its record exhibits an unusually high number of consecutive wins or losses, compared to what might be expected if the team’s performance does not really depend on whether or not they won their previous game. If an average team in Major League Baseball (i.e., with a record of 81-81) is not streaky, we assume its win probability would be stable at around 50% for most games, outside of peculiar details of day-to-day outcomes, such as whether the game is home or away, who is the starting pitcher, and so on.

In this …


City Life - Inquiry And Problem Solving Exercise, Milena Cuellar Jan 2017

City Life - Inquiry And Problem Solving Exercise, Milena Cuellar

Open Educational Resources

No abstract provided.


Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan Jan 2017

Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan

Theses and Dissertations--Statistics

We introduce a new class of measures for testing independence between two random vectors, which uses expected difference of conditional and marginal characteristic functions. By choosing a particular weight function in the class, we propose a new index for measuring independence and study its property. Two empirical versions are developed, their properties, asymptotics, connection with existing measures and applications are discussed. Implementation and Monte Carlo results are also presented.

We propose a two-stage sufficient variable selections method based on the new index to deal with large p small n data. The method does not require model specification and especially focuses …


Inference Using Bhattacharyya Distance To Model Interaction Effects When The Number Of Predictors Far Exceeds The Sample Size, Sarah A. Janse Jan 2017

Inference Using Bhattacharyya Distance To Model Interaction Effects When The Number Of Predictors Far Exceeds The Sample Size, Sarah A. Janse

Theses and Dissertations--Statistics

In recent years, statistical analyses, algorithms, and modeling of big data have been constrained due to computational complexity. Further, the added complexity of relationships among response and explanatory variables, such as higher-order interaction effects, make identifying predictors using standard statistical techniques difficult. These difficulties are only exacerbated in the case of small sample sizes in some studies. Recent analyses have targeted the identification of interaction effects in big data, but the development of methods to identify higher-order interaction effects has been limited by computational concerns. One recently studied method is the Feasible Solutions Algorithm (FSA), a fast, flexible method that …


Family-Based Association Studies Of Autism In Boys Via Facial-Feature Clusters, Luke Andrew Settles Jan 2017

Family-Based Association Studies Of Autism In Boys Via Facial-Feature Clusters, Luke Andrew Settles

Masters Theses

"Autism spectrum disorder (ASD) refers to a set of developmental disorders with varied attributes. Due to its substantial heterogeneity in terms of behavioral and clinical phenotypes, it is challenging to discern the genetic biomarkers behind ASD, even though the disease is known to be genetic in nature. This serves as a motivation to detect relationships between single nucleotide polymorphisms (SNPs) and a causal autism disease susceptibility locus (DSL) within more homogeneous subgroups. Recently, clinically meaningful subclassifications of ASD have been discovered utilizing facial features of prepubescent boys. Therefore, through the employment of data from 44 prepubertal Caucasian boys with ASD …


Comparing Region Level Testing Methods For Differential Dna Methylation Analysis, Arnold Albert Harder Jan 2017

Comparing Region Level Testing Methods For Differential Dna Methylation Analysis, Arnold Albert Harder

Masters Theses

”Finding possible connections and solutions to help fight progression of diseases is a major area of research. Genomics is a primary path of research in disease research. Through the DNA sequence, possible connections to diseases have been found. However, most methods for fixing issues within a DNA sequence are still out of reach. One potential path is to investigate epigenetic modifications, such as DNA methylation. DNA methylation occurs when a methyl group attached to cytosines on the DNA sequence. Statistical methods can be used to identify sites or regions of significant differences in methylation levels between groups ( e. g. …


Development Of A Variogram Approach To Spatial Outlier Detection Using A Supplemental Digital Elevation Model Dataset, Zane Daniel Helwig Jan 2017

Development Of A Variogram Approach To Spatial Outlier Detection Using A Supplemental Digital Elevation Model Dataset, Zane Daniel Helwig

Masters Theses

"When developing a ground water model, the quality of the dataset should first be evaluated. Spatial outliers can lead to predictions which are not representative of actual conditions. In order to isolate misrepresentative points, a method is presented which examines the experimental variogram of a ground water elevation dataset. To define a threshold variance between pairs of ground water elevation measures, ground elevation values from a digital elevation model (DEM) are used to determine a maximum reasonable variance expected to occur on the experimental variogram. To determine appropriate DEM parameters, a separate study was also done which observed characteristic behavior …


Advance Care Planning As A Shared Endeavor: Completion Of Acp Documents In A Multidisciplinary Cancer Program, Melissa A. Clark, Miles Q. Ott, Michelle L. Rogers, Mary C. Politi, Susan C. Miller, Laura Moynihan, Katina Robison, Ashley Stuckey, Don Dizon Jan 2017

Advance Care Planning As A Shared Endeavor: Completion Of Acp Documents In A Multidisciplinary Cancer Program, Melissa A. Clark, Miles Q. Ott, Michelle L. Rogers, Mary C. Politi, Susan C. Miller, Laura Moynihan, Katina Robison, Ashley Stuckey, Don Dizon

Statistical and Data Sciences: Faculty Publications

Objective—We examined the roles of oncology providers in advance care planning (ACP) delivery in the context of a multidisciplinary cancer program.

Methods—Semi-structured interviews were conducted with 200 women with recurrent and/or metastatic breast or gynecologic cancer. Participants were asked to name providers they deemed important in their cancer care and whether they had discussed and/or completed ACP documentation. Evidence of ACP documentation was obtained from chart reviews.

Results—Fifty percent of participants self-reported completing an advance directive (AD) and 48.5% had named a healthcare power of attorney (HPA), 38.5% had completed both, and 39.0% had completed neither document. Among women who …


Alcohol Perceptions And Behavior In A Residential Peer Social Network, Shannon R. Kenney, Miles Q. Ott, Matthew Meisel, Nancy P. Barnett Jan 2017

Alcohol Perceptions And Behavior In A Residential Peer Social Network, Shannon R. Kenney, Miles Q. Ott, Matthew Meisel, Nancy P. Barnett

Statistical and Data Sciences: Faculty Publications

Personalized normative feedback is a recommended component of alcohol interventions targeting college students. However, normative data are commonly collected through campus-based surveys, not through actual participant-referent relationships. In the present investigation, we examined how misperceptions of residence hall peers, both overall using a global question and those designated as important peers using person-specific questions, were related to students’ personal drinking behaviors. Participants were 108 students (88% freshman, 54% White, 51% female) residing in a single campus residence hall. Participants completed an online baseline survey in which they reported their own alcohol use and perceptions of peer alcohol use using both …


Curriculum Guidelines For Undergraduate Programs In Data Science, Richard D. De Veaux, Mahesh Agarwal, Maia Averett, Benjamin Baumer, Andrew Bray, Thomas C. Bressoud, Lance Bryant, Lei Z. Cheng, Amanda Francis, Robert Gould, Albert Y. Kim, Matt Kretchmar, Qin Lu, Ann Moskol, Deborah Nolan, Roberto Pelayo, Sean Raleigh, Ricky J. Sethi, Mutiara Sondjaja, Neelesh Tiruviluamala, Paul X. Uhlig, Talitha M. Washington, Curtis L. Wesley, David White, Ping Ye Jan 2017

Curriculum Guidelines For Undergraduate Programs In Data Science, Richard D. De Veaux, Mahesh Agarwal, Maia Averett, Benjamin Baumer, Andrew Bray, Thomas C. Bressoud, Lance Bryant, Lei Z. Cheng, Amanda Francis, Robert Gould, Albert Y. Kim, Matt Kretchmar, Qin Lu, Ann Moskol, Deborah Nolan, Roberto Pelayo, Sean Raleigh, Ricky J. Sethi, Mutiara Sondjaja, Neelesh Tiruviluamala, Paul X. Uhlig, Talitha M. Washington, Curtis L. Wesley, David White, Ping Ye

Statistical and Data Sciences: Faculty Publications

The Park City Math Institute 2016 Summer Undergraduate Faculty Program met for the purpose of composing guidelines for undergraduate programs in data science. The group consisted of 25 undergraduate faculty from a variety of institutions in the United States, primarily from the disciplines of mathematics, statistics, and computer science. These guidelines are meant to provide some structure for institutions planning for or revising a major in data science.


Augmenting Bottom-Up Metamodels With Predicates, Ross J. Gore, Saikou Diallo, Christopher Lynch, Jose Padilla Jan 2017

Augmenting Bottom-Up Metamodels With Predicates, Ross J. Gore, Saikou Diallo, Christopher Lynch, Jose Padilla

VMASC Publications

Metamodeling refers to modeling a model. There are two metamodeling approaches for ABMs: (1) top-down and (2) bottom-up. The top down approach enables users to decompose high-level mental models into behaviors and interactions of agents. In contrast, the bottom-up approach constructs a relatively small, simple model that approximates the structure and outcomes of a dataset gathered fromthe runs of an ABM. The bottom-up metamodel makes behavior of the ABM comprehensible and exploratory analyses feasible. Formost users the construction of a bottom-up metamodel entails: (1) creating an experimental design, (2) running the simulation for all cases specified by the design, (3) …


The Use Of The Pivot Pairwise Relative Criteria Importance Assessment Method For Determining The Weights Of Criteria, Florentin Smarandache, Dragisa Stanujkic, Edmundas Kazimieras Zavadskas, Darjan Karabasevic, Zenonas Turskis Jan 2017

The Use Of The Pivot Pairwise Relative Criteria Importance Assessment Method For Determining The Weights Of Criteria, Florentin Smarandache, Dragisa Stanujkic, Edmundas Kazimieras Zavadskas, Darjan Karabasevic, Zenonas Turskis

Branch Mathematics and Statistics Faculty and Staff Publications

The weights of evaluation criteria could have a significant impact on the results obtained by applying multiple criteria decision-making methods. Therefore, the two extensions of the SWARA method that can be used in cases when it is not easy, or even is impossible to reach a consensus on the expected importance of the evaluation criteria are proposed in this paper. The primary objective of the proposed extensions is to provide an understandable and easy-to-use approach to the collecting of respondents’ real attitudes towards the significance of evaluation criteria and to also provide an approach to the checking of the reliability …


A Comparison Of Decision Tree With Logistic Regression Model For Prediction Of Worst Non-Financial Payment Status In Commercial Credit, Jessica M. Rudd Mph, Gstat, Jennifer L. Priestley Jan 2017

A Comparison Of Decision Tree With Logistic Regression Model For Prediction Of Worst Non-Financial Payment Status In Commercial Credit, Jessica M. Rudd Mph, Gstat, Jennifer L. Priestley

Published and Grey Literature from PhD Candidates

Credit risk prediction is an important problem in the financial services domain. While machine learning techniques such as Support Vector Machines and Neural Networks have been used for improved predictive modeling, the outcomes of such models are not readily explainable and, therefore, difficult to apply within financial regulations. In contrast, Decision Trees are easy to explain, and provide an easy to interpret visualization of model decisions. The aim of this paper is to predict worst non-financial payment status among businesses, and evaluate decision tree model performance against traditional Logistic Regression model for this task. The dataset for analysis is provided …