Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Statistical Methodology

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1111 - 1140 of 1562

Full-Text Articles in Statistics and Probability

Addressing The Zeros Problem: Regression Models For Outcomes With A Large Proportion Of Zeros, With An Application To Trial Outcomes, Theodore Eisenberg, Thomas Eisenberg, Martin T. Wells, Min Zhang Mar 2015

Addressing The Zeros Problem: Regression Models For Outcomes With A Large Proportion Of Zeros, With An Application To Trial Outcomes, Theodore Eisenberg, Thomas Eisenberg, Martin T. Wells, Min Zhang

Cornell Law Faculty Publications

In law‐related and other social science contexts, researchers need to account for data with an excess number of zeros. In addition, dollar damages in legal cases also often are skewed. This article reviews various strategies for dealing with this data type. Tobit models are often applied to deal with the excess number of zeros, but these are more appropriate in cases of true censoring (e.g., when all negative values are recorded as zeros) and less appropriate when zeros are in fact often observed as the amount awarded. Heckman selection models are another methodology that is applied in this setting, yet …


Best Practice Recommendations For Data Screening, Justin A. Desimone, Peter D. Harms, Alice J. Desimone Feb 2015

Best Practice Recommendations For Data Screening, Justin A. Desimone, Peter D. Harms, Alice J. Desimone

Department of Management: Faculty Publications

Survey respondents differ in their levels of attention and effort when responding to items. There are a number of methods researchers may use to identify respondents who fail to exert sufficient effort in order to increase the rigor of analysis and enhance the trustworthiness of study results. Screening techniques are organized into three general categories, which differ in impact on survey design and potential respondent awareness. Assumptions and considerations regarding appropriate use of screening techniques are discussed along with descriptions of each technique. The utility of each screening technique is a function of survey design and administration. Each technique has …


Review Of Naked Statistics: Stripping The Dread From Data By Charles Wheelan, Michael T. Catalano Jan 2015

Review Of Naked Statistics: Stripping The Dread From Data By Charles Wheelan, Michael T. Catalano

Numeracy

Wheelan, Charles. Naked Statistics: Stripping the Dread from Data (New York, NY, W. W. Norton & Company, 2014). 282 pp. ISBN 978-0-393-07195-5

In his review of What Numbers Say and The Numbers Game, Rob Root (Numeracy 3(1): 9) writes “Popular books on quantitative literacy need to be easy to read, reasonably comprehensive in scope, and include examples that are thought-provoking and memorable.” Wheelan’s book certainly meets this description, and should be of interest to both the general public and those with a professional interest in numeracy. A moderately diligent learner can get a decent understanding of basic statistics …


Empirical Likelihood Confidence Band, Shihong Zhu Jan 2015

Empirical Likelihood Confidence Band, Shihong Zhu

Theses and Dissertations--Statistics

The confidence band represents an important measure of uncertainty associated with a functional estimator and empirical likelihood method has been proved to be a viable approach to constructing confidence bands in many cases. Using the empirical likelihood ratio principle, this dissertation developed simultaneous confidence bands for many functions of fundamental importance in survival analysis, including the survival function, the difference and ratio of survival functions, the hazards ratio function, and other parameters involving residual lifetimes. Covariate adjustment was incorporated under the proportional hazards assumption. The proposed method can be very useful when, for example, an individualized survival function is desired …


Statistics In The Billera-Holmes-Vogtmann Treespace, Grady S. Weyenberg Jan 2015

Statistics In The Billera-Holmes-Vogtmann Treespace, Grady S. Weyenberg

Theses and Dissertations--Statistics

This dissertation is an effort to adapt two classical non-parametric statistical techniques, kernel density estimation (KDE) and principal components analysis (PCA), to the Billera-Holmes-Vogtmann (BHV) metric space for phylogenetic trees. This adaption gives a more general framework for developing and testing various hypotheses about apparent differences or similarities between sets of phylogenetic trees than currently exists.

For example, while the majority of gene histories found in a clade of organisms are expected to be generated by a common evolutionary process, numerous other coexisting processes (e.g. horizontal gene transfers, gene duplication and subsequent neofunctionalization) will cause some genes to exhibit a …


Bayesian Function-On-Function Regression For Multilevel Functional Data, Mark J. Meyer, Brent A. Coull, Francesco Versace, Paul Cinciripini, Jeffrey S. Morris Jan 2015

Bayesian Function-On-Function Regression For Multilevel Functional Data, Mark J. Meyer, Brent A. Coull, Francesco Versace, Paul Cinciripini, Jeffrey S. Morris

Faculty Journal Articles

Medical and public health research increasingly involves the collection of complex and high dimensional data. In particular, functional data—where the unit of observation is a curve or set of curves that are finely sampled over a grid—is frequently obtained. Moreover, researchers often sample multiple curves per person resulting in repeated functional measures. A common question is how to analyze the relationship between two functional variables. We propose a general function-on-function regression model for repeatedly sampled functional data on a fine grid, presenting a simple model as well as a more extensive mixed model framework, and introducing various functional Bayesian inferential …


Understanding Vulnerability In Alaska Fishing Communities: A Validation Methodology For Rapid Assessment Of Well-Being Indices, Conor M. Maguire Jan 2015

Understanding Vulnerability In Alaska Fishing Communities: A Validation Methodology For Rapid Assessment Of Well-Being Indices, Conor M. Maguire

All Master's Theses

Social well-being indices measure how fishing communities are likely to be affected by social-ecological perturbations, and are a significant tool to identify the primary issues influencing communities’ sustained participation in fishing activities. In an attempt to further our understanding of how communities are affected by such perturbations, we have developed a rapid assessment methodology to test the external validity of a set of well-being indices that measure community vulnerability. This methodology informs how well such indices reflect the communities they represent by measuring elements of well-being through field observations, and comparing them to corresponding index components created from secondary data …


Statistical Learning With Artificial Neural Network Applied To Health And Environmental Data, Taysseer Sharaf Jan 2015

Statistical Learning With Artificial Neural Network Applied To Health And Environmental Data, Taysseer Sharaf

USF Tampa Graduate Theses and Dissertations

The current study illustrates the utilization of artificial neural network in statistical methodology. More specifically in survival analysis and time series analysis, where both holds an important and wide use in many applications in our real life. We start our discussion by utilizing artificial neural network in survival analysis. In literature there exist two important methodology of utilizing artificial neural network in survival analysis based on discrete survival time method. We illustrate the idea of discrete survival time method and show how one can estimate the discrete model using artificial neural network. We present a comparison between the two methodology …


Comparing Group Means When Nonresponse Rates Differ, Gabriela M. Stegmann Jan 2015

Comparing Group Means When Nonresponse Rates Differ, Gabriela M. Stegmann

UNF Graduate Theses and Dissertations

Missing data bias results if adjustments are not made accordingly. This thesis addresses this issue by exploring a scenario where data is missing at random depending on a covariate x. Four methods for comparing groups while adjusting for missingness are explored by conducting simulations: independent samples t-test with predicted mean stratification, independent samples t-test with response propensity stratification, independent samples t-test with response propensity weighting, and an analysis of covariance. Results show that independent samples t-test with response propensity weighting and analysis of covariance can appropriately adjust for bias. ANCOVA is the stronger method when …


The Bootstrap Estimation In Time Series, Yun Liu Jan 2015

The Bootstrap Estimation In Time Series, Yun Liu

Dissertations, Master's Theses and Master's Reports

Time series, a special case in dependent data sequence, is widely used in many fields. In time series, linear process models are quite popularly used. General form of linear process indicates the time dependence property of time series, AR(p), MA(q) and ARMA(p;,q) models are all linear process models. In this report, simulations are based on the simplest models of these linear process models, such as AR(1), MA(1) and ARMA(1,1) models. AR(1)-SEASON, which is developed based on AR(1) model by changing the weight of residuals, is also considered in this report. To deal with dependent data sequence, common methods which aim …


Realistic Spiking Neuron Statistics In A Population Are Described By A Single Parametric Distribution, Lauren Crow 9370373 Jan 2015

Realistic Spiking Neuron Statistics In A Population Are Described By A Single Parametric Distribution, Lauren Crow 9370373

Undergraduate Research Posters

The spiking of activity of neurons throughout the cortex is random and complicated. This complicated activity requires theoretical formulations in order to understand the underlying principles of neural processing. A key aspect of theoretical investigations is characterizing the probability distribution of spiking activity. This study aims to better understand the statistics of the time between spikes, or interspike interval, in both real data and a spiking model with many time scales. Exploration of the interspike intervals of neural network activity can provide a better understanding of neural responses to different stimuli. We consider different parametric distribution fitting techniques to characterize …


Statistical Inference For The Mean Outcome Under A Possibly Non-Unique Optimal Treatment Strategy, Alexander R. Luedtke, Mark J. Van Der Laan Dec 2014

Statistical Inference For The Mean Outcome Under A Possibly Non-Unique Optimal Treatment Strategy, Alexander R. Luedtke, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider challenges that arise in the estimation of the value of an optimal individualized treatment strategy defined as the treatment rule that maximizes the population mean outcome, where the candidate treatment rules are restricted to depend on baseline covariates. We prove a necessary and sufficient condition for the pathwise differentiability of the optimal value, a key condition needed to develop a regular asymptotically linear (RAL) estimator of this parameter. The stated condition is slightly more general than the previous condition implied in the literature. We then describe an approach to obtain root-n rate confidence intervals for the optimal value …


Higher-Order Targeted Minimum Loss-Based Estimation, Marco Carone, Iván Díaz, Mark J. Van Der Laan Dec 2014

Higher-Order Targeted Minimum Loss-Based Estimation, Marco Carone, Iván Díaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Common approaches to parametric statistical inference often encounter difficulties in the context of infinite-dimensional models. The framework of targeted maximum likelihood estimation (TMLE), introduced in van der Laan & Rubin (2006), is a principled approach for constructing asymptotically linear and efficient substitution estimators in rich infinite-dimensional models. The mechanics of TMLE hinge upon first-order approximations of the parameter of interest as a mapping on the space of probability distributions. For such approximations to hold, a second-order remainder term must tend to zero sufficiently fast. In practice, this means an initial estimator of the underlying data-generating distribution with a sufficiently large …


Frameworks For Nurturing And Assessing Students’ Statistical Thinking In Regression Modelling, Wing Kin, Ken Li Dec 2014

Frameworks For Nurturing And Assessing Students’ Statistical Thinking In Regression Modelling, Wing Kin, Ken Li

Practical Social and Industrial Research Symposium

No abstract provided.


Cross-Design Synthesis For Extending The Applicability Of Trial Evidence When Treatment Effect Is Heterogeneous-I. Methodology, Ravi Varadhan, Carlos Weiss Nov 2014

Cross-Design Synthesis For Extending The Applicability Of Trial Evidence When Treatment Effect Is Heterogeneous-I. Methodology, Ravi Varadhan, Carlos Weiss

Johns Hopkins University, Dept. of Biostatistics Working Papers

Randomized controlled trials (RCTs) provide reliable evidence for approval of new treatments, informing clinical practice, and coverage decisions. The participants in RCTs are often not a representative sample of the larger at-risk population. Hence it is argued that the average treatment effect from the trial is not generalizable to the larger at-risk population. An essential premise of this argument is that there is significant heterogeneity in the treatment effect (HTE). We present a new method to extrapolate the treatment effect from a trial to a target group that is inadequately represented in the trial, when HTE is present. Our method …


Comparison Of Hazard, Odds And Risk Ratio In The Two-Sample Survival Problem, Benedict P. Dormitorio Aug 2014

Comparison Of Hazard, Odds And Risk Ratio In The Two-Sample Survival Problem, Benedict P. Dormitorio

Dissertations

Cox proportional hazards is the standard method for analyzing treatment efficacy when time-to-event data is available. In the absence of time-to-event, investigators may use logistic regression which only requires relative frequencies of events, or Poisson regression which requires only interval-summarized frequency tables of time-to-event. When event frequencies are used instead of time-to-events, does it always result in a loss in power?

We investigate the relative performance of the three methods. In particular, we compare the power of tests based on the respective effect-size estimates (1)hazard ratio (HR), (2)odds ratio (OR), and (3)risk ratio (RR). We use a variety of survival …


Interadapt -- An Interactive Tool For Designing And Evaluating Randomized Trials With Adaptive Enrollment Criteria, Aaron Joel Fisher, Harris Jaffee, Michael Rosenblum Jun 2014

Interadapt -- An Interactive Tool For Designing And Evaluating Randomized Trials With Adaptive Enrollment Criteria, Aaron Joel Fisher, Harris Jaffee, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

The interAdapt R package is designed to be used by statisticians and clinical investigators to plan randomized trials. It can be used to determine if certain adaptive designs offer tangible benefits compared to standard designs, in the context of investigators’ specific trial goals and constraints. Specifically, interAdapt compares the performance of trial designs with adaptive enrollment criteria versus standard (non-adaptive) group sequential trial designs. Performance is compared in terms of power, expected trial duration, and expected sample size. Users can either work directly in the R console, or with a user-friendly shiny application that requires no programming experience. Several added …


The Impact Of Student Performance On Large-Scale Assessments: A View Of Long-Term Health, Career, And Societal Outcomes, Roman Usatin Jun 2014

The Impact Of Student Performance On Large-Scale Assessments: A View Of Long-Term Health, Career, And Societal Outcomes, Roman Usatin

Seton Hall University Dissertations and Theses (ETDs)

This study examined the predictive power of student growth for large-scale assessments on meaningful life outcomes, focusing on the three categories of health, career, and societal involvement. Analysis was conducted using the NELS:88/00 dataset–a longitudinal study that followed a nationally-representative sample of over 12,000 eighth grade students from 1988 to 2000, until the students were 26 years old and entered into the work force. The large-scale assessment variables included math and reading performance in the 1988 cognitive batteries administered by NELS. To gauge growth levels, I generated Student Growth Percentiles (SGP) from tests administered by NELS from 1988 to 1992. …


How Sexism Makes The Man: Examining The Relationship Between Masculinity, Ambivalent Sexism, And Gender Stereotyping, Mariah L. Wilkerson Jun 2014

How Sexism Makes The Man: Examining The Relationship Between Masculinity, Ambivalent Sexism, And Gender Stereotyping, Mariah L. Wilkerson

Lawrence University Honors Projects

Masculinity is a precarious social status, meaning it can be lost through social and gender transgressions (Bosson & Vandello, 2011). Men often act in stereotypically masculine ways to reassert their masculinity and restore their social status after it has been threatened. The current study also examines masculinity in a new way, as a collective gender identity (e.g., Tajfel, 1982). I hypothesized that threatened men and men who identify as more masculine will display masculinity through more polarized attitudes towards traditional and nontraditional groups of men and women, endorsing traditional gender stereotypes, and intensified ambivalently sexist attitudes. Two empirical studies tested …


Targeted Maximum Likelihood Estimation Using Exponential Families, Iván Díaz, Michael Rosenblum Jun 2014

Targeted Maximum Likelihood Estimation Using Exponential Families, Iván Díaz, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

Targeted maximum likelihood estimation (TMLE) is a general method for estimating parameters in semiparametric and nonparametric models. Each iteration of TMLE involves fitting a parametric submodel that targets the parameter of interest. We investigate the use of exponential families to define the parametric submodel. This implementation of TMLE gives a general approach for estimating any smooth parameter in the nonparametric model. A computational advantage of this approach is that each iteration of TMLE involves estimation of a parameter in an exponential family, which is a convex optimization problem for which software implementing reliable and computationally efficient methods exists. We illustrate …


Musical Missteps: The Severity Of The Sophomore Slump In The Music Industry, Shane M. Zackery May 2014

Musical Missteps: The Severity Of The Sophomore Slump In The Music Industry, Shane M. Zackery

Scripps Senior Theses

This study looks at alternative models of follow-up album success in order to determine if there is a relationship between the decrease in Metascore ratings (assigned by Metacritic.com) between the first and second album for a musician or band and the 1) music genre or 2) the number of years between the first and second album release. The results support the dominant thought, which suggests that neither belonging to a certain genre of music nor waiting more or less time to drop the second album makes an artist more susceptible to the Sophomore Slump. This finding is important because it …


Variable Selection For Zero-Inflated And Overdispersed Data With Application To Health Care Demand In Germany, Zhu Wang, Shuangge Ma, Ching-Yun Wang May 2014

Variable Selection For Zero-Inflated And Overdispersed Data With Application To Health Care Demand In Germany, Zhu Wang, Shuangge Ma, Ching-Yun Wang

COBRA Preprint Series

In health services and outcome research, count outcomes are frequently encountered and often have a large proportion of zeros. The zero-inflated negative binomial (ZINB) regression model has important applications for this type of data. With many possible candidate risk factors, this paper proposes new variable selection methods for the ZINB model. We consider maximum likelihood function plus a penalty including the least absolute shrinkage and selection operator (LASSO), smoothly clipped absolute deviation (SCAD) and minimax concave penalty (MCP). An EM (expectation-maximization) algorithm is proposed for estimating the model parameters and conducting variable selection simultaneously. This algorithm consists of estimating penalized …


Nonparametric Identifiability Of Finite Mixture Models With Covariates For Estimating Error Rate Without A Gold Standard, Zheyu Wang, Xiao-Hua Zhou Apr 2014

Nonparametric Identifiability Of Finite Mixture Models With Covariates For Estimating Error Rate Without A Gold Standard, Zheyu Wang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Finite mixture models provide a flexible framework to study unobserved entities and have arisen in many statistical applications. The flexibility of these models in adapting various complicated structures makes it crucial to establish model identifiability when applying them in practice to ensure study validity and interpretation. However, researches to establish the identifiability of finite mixture model are limited and are usually restricted to a few specific model configurations. Conditions for model identifiability in the general case have not been established. In this paper, we provide conditions for both local identifiability and global identifiability of a finite mixture model. The former …


Improving The Design Of Cluster-Randomized Trials In Education: Informing The Selection Of Variance Design Parameter Values For Science Achievement Studies, Carl D. Westine Apr 2014

Improving The Design Of Cluster-Randomized Trials In Education: Informing The Selection Of Variance Design Parameter Values For Science Achievement Studies, Carl D. Westine

Dissertations

The purpose of this three-essay dissertation is to provide practical guidance to evaluators planning cluster-randomized trials (CRTs) of science achievement. In an educational setting, interventions are often administered at the cluster level, while outcomes are typically measured at the student level through standardized achievement testing. When evaluating an intervention, a CRT is appropriate because it allows for treatment to be modeled at a different level than the unit of analysis, and properly accounts for the violation of independence that occurs due to nesting. Accurately designing a CRT involves estimating variance parameters (i.e., intraclass correlations [ICCs] and percent of variance explained …


Harnessing Complexity: Analysis Methodology And Ethical Framework To Facilitate Utilization Of Video Data In Evaluations, Kurt A. Wilson Apr 2014

Harnessing Complexity: Analysis Methodology And Ethical Framework To Facilitate Utilization Of Video Data In Evaluations, Kurt A. Wilson

Dissertations

Most evaluations in the nonprofit and international development sectors are conducted in contexts of complexity; the specific intervention being evaluated is but one of many interrelated factors influencing the desired outcome. Video data, especially when directly generated by program participants, can provide both exceptionally rich qualitative data as well as contextually-relevant feedback within complex systems. Despite these unique strengths and opportunities, video data is underutilized in the field of evaluation. This dissertation addresses specific barriers associated with video data through three inter-related papers: Papers one and two (Chapters II and III) present the findings from two interrelated studies of an …


Comparing Partial Least Square Approaches In Gene-Or Region-Based Association Study For Multiple Quantitative Phenotypes, Zhongshang Yuan, Xiaoshuai Zhang, Fangyu Li, Jinghua Zhao, Fuzhong Xue Mar 2014

Comparing Partial Least Square Approaches In Gene-Or Region-Based Association Study For Multiple Quantitative Phenotypes, Zhongshang Yuan, Xiaoshuai Zhang, Fangyu Li, Jinghua Zhao, Fuzhong Xue

Human Biology Open Access Pre-Prints

On thinking quantitatively of complex diseases, there are at least three statistical strategies for association study: single SNP on single trait, gene-or region (with multiple SNPs) on single trait and on multiple traits. The third of which is the most general in dissecting the genetic mechanism underlying complex diseases underpinning multiple quantitative traits. Gene-or region association methods based on partial least square (PLS) approaches have been shown to have apparent power advantage. However, few attempts are developed for multiple quantitative phenotypes or traits underlying a condition or disease, and the performance of various PLS approaches used in association study for …


Global Resource Management Of Response Surface Methodology, Michael Chad Miller Mar 2014

Global Resource Management Of Response Surface Methodology, Michael Chad Miller

Dissertations and Theses

Statistical research can be more difficult to plan than other kinds of projects, since the research must adapt as knowledge is gained. This dissertation establishes a formal language and methodology for designing experimental research strategies with limited resources. It is a mathematically rigorous extension of a sequential and adaptive form of statistical research called response surface methodology. It uses sponsor-given information, conditions, and resource constraints to decompose an overall project into individual stages. At each stage, a "parent" decision-maker determines what design of experimentation to do for its stage of research, and adapts to the feedback from that research's potential …


Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan Jan 2014

Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In randomized trials, pair-matching is an intuitive design strategy to protect study validity and to potentially increase study power. In a common design, candidate units are identified, and their baseline characteristics used to create the best n/2 matched pairs. Within the resulting pairs, the intervention is randomized, and the outcomes measured at the end of follow-up. We consider this design to be adaptive, because the construction of the matched pairs depends on the baseline covariates of all candidate units. As consequence, the observed data cannot be considered as n/2 independent, identically distributed (i.i.d.) pairs of units, as current practice assumes. …


Adaptive Randomized Trial Designs That Cannot Be Dominated By Any Standard Design At The Same Total Sample Size, Michael Rosenblum Jan 2014

Adaptive Randomized Trial Designs That Cannot Be Dominated By Any Standard Design At The Same Total Sample Size, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

Prior work has shown that certain types of adaptive designs can always be dominated by a suitably chosen, standard, group sequential design. This applies to adaptive designs with rules for modifying the total sample size. A natural question is whether analogous results hold for other types of adaptive designs. We focus on adaptive enrichment designs, which involve preplanned rules for modifying enrollment criteria based on accrued data in a randomized trial. Such designs often involve multiple hypotheses, e.g., one for the total population and one for a predefined subpopulation, such as those with high disease severity at baseline. We fix …


Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly Jan 2014

Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly

CHIP Documents

This publication provides a contemporary treatment of the subject of meta-analysis in relation to social-personality psychology. Meta-analysis literally refers to the statistical pooling of the results of independent studies on a given subject, although in practice it refers as well to other steps of research synthesis, including defining the question under investigation, gathering all available research reports, coding of information about the studies and their effects, and interpretation/dissemination of results. Discussed as well are the hallmarks of high-quality meta-analyses.