Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2006

Discipline
Institution
Keyword
Publication
Publication Type

Articles 151 - 180 of 244

Full-Text Articles in Statistics and Probability

Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan Apr 2006

Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many statistical methods exist that can be used to learn a predictor based on observed data. Examples include decision trees, neural networks, support vector regression, least angle regression, Logic Regression, and the Deletion/Substitution/Addition algorithm. The optimal algorithm for prediction will vary depending on the underlying data-generating distribution. In this article, we introduce a "super learner," a prediction algorithm that applies any set of candidate learners and uses cross-validation to select among them. Theory shows that asymptotically the super learner performs essentially as well or better than any of the candidate learners. We briefly present the theory behind the super learner, …


Recurrent Event Models In The Presence Of A Terminal Event: Comparison, Inference And Data Analysis, Xianghua Luo, Mei-Cheng Wang Apr 2006

Recurrent Event Models In The Presence Of A Terminal Event: Comparison, Inference And Data Analysis, Xianghua Luo, Mei-Cheng Wang

Johns Hopkins University, Dept. of Biostatistics Working Papers

This article focuses on statistical implications of proportional rate models for recurrent event data in the presence of a terminal event. In such circumstances, various definitions of the recurrent rate function have been adopted in the proportional rate models. Although these rate functions have quite different interpretations, recognition of the differences has been lacking theoretically and practically. We compare three types of rate functions from both conceptual and quantitative perspectives; conclude that the inappropriate choice of a rate function may lead to misleading scientific conclusions. Simulations are conducted for comparisons of the focused models. Analysis of data from an AIDS …


Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard Apr 2006

Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

Statistical challenges arise in identifying meaningful patterns and structures from high dimensional genomic data sets. Relating HIV genotype (sequence of amino acids) to phenotypic resistance presents a typical problem. When the HIV virus is under antiretroviral drug pressure, unfavorable mutations of the target genes often lead to greatly increased resistance of the virus to drugs, including drugs the virus has not been exposed to. Identification of mutation combinations and their correlation to drug resistance is critical in guiding efficient prescription of HIV drugs. The identification of a subset of codons associated with drug resistance from a set of several hundreds …


Survival Analysis With Change Point Hazard Functions, Melody S. Goodman, Yi Li, Ram C. Tiwari Apr 2006

Survival Analysis With Change Point Hazard Functions, Melody S. Goodman, Yi Li, Ram C. Tiwari

Harvard University Biostatistics Working Paper Series

No abstract provided.


Modeling And Simulation Of Value -At -Risk In The Financial Market Area, Xiangyin Zheng Apr 2006

Modeling And Simulation Of Value -At -Risk In The Financial Market Area, Xiangyin Zheng

Doctoral Dissertations

Value-at-Risk (VaR) is a statistical approach to measure market risk. It is widely used by banks, securities firms, commodity and energy merchants, and other trading organizations. The main focus of this research is measuring and analyzing market risk by modeling and simulation of Value-at-Risk for portfolios in the financial market area. The objectives are (1) predicting possible future loss for a financial portfolio from VaR measurement, and (2) identifying how the distributions of the risk factors affect the distribution of the portfolio. Results from (1) and (2) provide valuable information for portfolio optimization and risk management.

The model systems chosen …


Measuring Inequality: Statistical Inference Theory With Applications, Mihaela Paun Apr 2006

Measuring Inequality: Statistical Inference Theory With Applications, Mihaela Paun

Doctoral Dissertations

In this dissertation we develop statistical inference for the Atkinson index, one of the measures of inequality used in studying economic inequality.

Specifically, we construct empirical estimators for the Atkinson index, both in the parametric and nonparametric case, and derive formulas for the asymptotic variances for the estimators. These statistics are used for testing hypothesis and constructing confidence intervals for the Atkinson index. We test the validity and the robustness of the asymptotic theory, by simulations (using R, a language and environment for statistical computing and graphics), in the case of one and two populations. In addition to proving asymptotic …


New Estimators Of A Circular Median, Sauwanit Ratanaruamkarn Apr 2006

New Estimators Of A Circular Median, Sauwanit Ratanaruamkarn

Dissertations

The specific properties of probability distributions on a circle require different definitions of several statistical concepts. For example, a median that can always be found for linear data not always exists on the circle. Several estimators of a circular median were proposed, lately by Otieno (2002), Otieno and Anderson - Cook (2003). Their work is, however, focused on the "preferred direction" which coincides with the median, mean and mode in the case of symmetric, unimodal distributions. This dissertation is focused on the estimators of a population median in a wider range of population distributions on a circle including distributions with …


Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan Mar 2006

Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

An important class of models in causal inference are the so-called marginal structural models which model the comparison between counterfactual outcome distributions corresponding with a static treatment intervention, conditional on user supplied baseline covariates, based on observing a longitudinal data structure on a sample of n independent and identically distributed experimental units. Identification of a static treatment regimen specific outcome distribution based on observational data requires beyond the so-called sequential randomization assumption that each experimental unit has positive probability of following the static treatment regimen. The latter assumption is called the experimental treatment assignment assumption (ETA) (which is parameter specific). …


Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui Mar 2006

Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui

UW Biostatistics Working Paper Series

Background: New health status instruments are described by psychometric properties, such as Reliability, Effect Size, and Responsiveness. For cluster-randomized trials, another important statistic is the Intraclass Correlation for the instrument within clusters. Studies using better instruments can be performed with smaller sample sizes, but better instruments may be more expensive in terms of dollars, lost opportunities, or poorer data quality due to the response burden of longer instruments. Investigators often need to estimate the psychometric properties of a new instrument, or of an established instrument in a new setting. Optimal sample sizes for estimating these properties have not been studied …


Detecting Pulsatile Hormone Secretion Events: A Bayesian Approach, Tim Johnson Mar 2006

Detecting Pulsatile Hormone Secretion Events: A Bayesian Approach, Tim Johnson

The University of Michigan Department of Biostatistics Working Paper Series

Many challenges arise in the analysis of pulsatile, or episodic, hormone concentration time series data. Among these challenges is the determination of the number and location of pulsatile events and the discrimination of events from noise. Analyses of these data are typically performed in two stages. In the first stage, the number and approximate location of the pulses are determined. In the second stage, a model (typically a deconvolution model) is fit to the data conditional on the number of pulses. Any error made in the first stage is carried over to the second stage. Furthermore, current methods, except two, …


Semiparametric Analysis For Correlated Recurrent And Terminal Events, Yining Ye, Jack Kalbfleisch, Doug E. Schaubel Mar 2006

Semiparametric Analysis For Correlated Recurrent And Terminal Events, Yining Ye, Jack Kalbfleisch, Doug E. Schaubel

The University of Michigan Department of Biostatistics Working Paper Series

In clinical and observational studies, recurrent event data (e.g. hospitalization) with a terminal event (e.g. death) are often encountered. In many instances, the terminal event is strongly correlated with the recurrent event process. In this article, we propose a semiparametric method to jointly model the recurrent and terminal event processes. The dependence is modeled by a shared gamma frailty that is included in both the recurrent event rate and terminal event hazard function. Marginal models are used to estimate the regression effects on the terminal and recurrent event processes and a Poisson model is used to estimate the dispersion of …


Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe Mar 2006

Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe

UW Biostatistics Working Paper Series

Recent scientific and technological innovations have produced an abundance of potential markers which are being investigated for their use in disease screen- ing and diagnosis. In evaluating these markers, it is often necessary to account for covariates which are associated with the marker of interest. These covariates may include subject characteristics, expertise of the test operator, test proce- dures, or aspects of specimen handling. In this paper, we propose the AROC, a covariate-adjusted measure of the classification accuracy. The AROC is the common covariate-specific ROC curve, when the covariate does not affect dis- crimination, and a weighted average of covariate-specific …


Censored Data Regression In High-Dimension And Low-Sample Size Settings For Genomic Applications, Hongzhe Li Mar 2006

Censored Data Regression In High-Dimension And Low-Sample Size Settings For Genomic Applications, Hongzhe Li

UPenn Biostatistics Working Papers

New high-throughput technologies are generating various types of high-dimensional genomic and proteomic data and meta-data (e.g., networks and pathways) in order to obtain a systems-level understanding of various complex diseases such as human cancers and cardiovascular diseases. As the amount and complexity of the data increase and as the questions being addressed become more sophisticated, we face the great challenge of how to model such data in order to draw valid statistical and biological conclusions. One important problem in genomic research is to relate these high-throughput genomic data to various clinical outcomes, including possibly censored survival outcomes such as age …


A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska Mar 2006

A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper proposes a statistical methodology for comparing the performance of evolutionary computation algorithms. A two-fold sampling scheme for collecting performance data is introduced, and these data are analyzed using bootstrap-based multiple hypothesis testing procedures. The proposed method is sufficiently flexible to allow the researcher to choose how performance is measured, does not rely upon distributional assumptions, and can be extended to analyze many other randomized numeric optimization routines. As a result, this approach offers a convenient, flexible, and reliable technique for comparing algorithms in a wide variety of applications.


The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun Mar 2006

The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun

UW Biostatistics Working Paper Series

The efficacy of an HIV vaccine to prevent infection is likely to depend on the genetic variation of the exposing virus. This paper addresses the problem of using data on the HIV sequences that infect vaccine efficacy trial participants to 1) test for vaccine efficacy more powerfully than procedures that ignore the sequence data; and 2) evaluate the dependence of vaccine efficacy on the divergence of infecting HIV strains from the HIV strain that is contained in the vaccine. Because hundreds of amino acid sites in each HIV genome are sequenced, it is natural to treat the divergence (defined in …


Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes Mar 2006

Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes

UW Biostatistics Working Paper Series

Consider a placebo-controlled preventive HIV vaccine efficacy trial. An HIV amino acid sequence is measured from each volunteer who acquires HIV, and these sequences are aligned together with the reference HIV sequence represented in the vaccine. We develop genome scanning methods to identify HIV positions at which the amino acids in sequences from infected vaccine recipients tend to be more divergent from the corresponding reference amino acid than the amino acids in sequences from infected placebo recipients. We consider five two-sample test statistics, based on Euclidean, Mahalanobis, and Kullback-Leibler divergence measures. Weights are incorporated to reflect biological information contained in …


On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger Mar 2006

On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Time series and case-crossover methods are often viewed as competing alternatives in environmental epidemiologic studies. Several recent studies have compared the time series and case-crossover methods. In this paper, we show that case-crossover using conditional logistic regression is a special case of time series analysis when there is a common exposure such as in air pollution studies. This equivalence provides computational convenience for case-crossover analyses and a better understanding of time series models. Time series log-linear regression accounts for over-dispersion of the Poisson variance, while case-crossover analyses typically do not. This equivalence also permits model checking for case-crossover data using …


Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei Mar 2006

Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei

Harvard University Biostatistics Working Paper Series

Suppose that we are interested in establishing simple, but reliable rules for predicting future t-year survivors via censored regression models. In this article, we present inference procedures for evaluating such binary classification rules based on various prediction precision measures quantified by the overall misclassification rate, sensitivity and specificity, and positive and negative predictive values. Specifically, under various working models we derive consistent estimators for the above measures via substitution and cross validation estimation procedures. Furthermore, we provide large sample approximations to the distributions of these nonsmooth estimators without assuming that the working model is correctly specified. Confidence intervals, for example, …


The Longitudinal Effect Of Self-Monitoring And Locus Of Control On Social Network Position In Friendship Networks, Gary J. Moore Mar 2006

The Longitudinal Effect Of Self-Monitoring And Locus Of Control On Social Network Position In Friendship Networks, Gary J. Moore

Theses and Dissertations

The purpose of this research was to identify how enduring personality characteristics predict a person's location in a network, locations which in turn affect outcomes such as performance. Specifically, this thesis examines how self-monitoring and locus of control influence an individual's location in a friendship social network over time. Hierarchical Linear Modeling (HLM) was used to analyze 28 groups of students and instructors at a military training course over six and one half weeks. Self-monitoring predicted betweenness centrality in five of six time periods while locus of control predicted betweenness centrality in three of six time periods. The moderation of …


Modeling The Performance Of A Baseball Player's Offensive Production, Michael Ross Smith Mar 2006

Modeling The Performance Of A Baseball Player's Offensive Production, Michael Ross Smith

Theses and Dissertations

This project addresses the problem of comparing the offensive abilities of players from different eras in Major League Baseball (MLB). We will study players from the perspective of an overall offensive summary statistic that is highly linked with scoring runs, or the Berry Value. We will build an additive model to estimate the innate ability of the player, the effect of the relative level of competition of each season, and the effect of age on performance using piecewise age curves. Using Hierarchical Bayes methodology with Gibbs sampling, we model each of these effects for each individual. The results of the …


Matrices And The Game Of Chutes And Ladders, Katherine Jedlicka Mar 2006

Matrices And The Game Of Chutes And Ladders, Katherine Jedlicka

Honors Capstones

Capstone submitted as a graduation requirement for the BSU Honors Program.


2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr Mar 2006

2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr

UW Biostatistics Working Paper Series

When a two-level design must be run in blocks of size two, there is a unique blocking scheme that enables estimation of all the main effects. Unfortunately this design does not enable estimation of any two-factor interactions. When the experimental goal is to estimate all main effects and two-factor interactions, it is necessary to combine replicates of the experiment that use different blocking schemes. In this paper we identify such designs for up to eight factors that enable estimation of all main effects and two-factor interactions with the fewest number of replications. In addition, we give a construction for general …


Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan Mar 2006

Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …


A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull Mar 2006

A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull

Harvard University Biostatistics Working Paper Series

We introduce a diagnostic test for the mixing distribution in a generalised linear mixed model. The test is based on the difference between the marginal maximum likelihood and conditional maximum likelihood estimates of a subset of the fixed effects in the model. We derive the asymptotic variance of this difference, and propose a test statistic that has a limiting chi-square distribution under the null hypothesis that the mixing distribution is correctly specified. For the important special case of the logistic regression model with random intercepts, we evaluate via simulation the power of the test in finite samples under several alternative …


Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng Mar 2006

Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng

UW Biostatistics Working Paper Series

Consider a continuous marker for predicting a binary outcome. For example, serum concentration of prostate specific antigen (PSA) may be used to calculate the risk of finding prostate cancer in a biopsy. In this paper we argue that the predictive capacity of a marker has to do with the population distribution of risk given the marker and suggest a graphical tool, the predictiveness curve, that displays this distribution. The display provides a common meaningful scale for comparing markers that may not be comparable on their original scales. Some existing measures of predictiveness are shown to be summary indices derived from …


Optimization Of A Multi-Echelon Repair System Via Generalized Pattern Search With Ranking And Selection: A Computational Study, Derek D. Tharaldson Mar 2006

Optimization Of A Multi-Echelon Repair System Via Generalized Pattern Search With Ranking And Selection: A Computational Study, Derek D. Tharaldson

Theses and Dissertations

With increasing developments in computer technology and available software, simulation is becoming a widely used tool to model, analyze, and improve a real world system or process. However, simulation in itself is not an optimization approach. Common optimization procedures require either an explicit mathematical formulation or numerous function evaluations at improving iterative points. Mathematical formulation is generally impossible for problems where simulation is relevant, which are characteristically the types of problems that arise in practical applications. Further complicating matters is the variability in the simulation response which can cause problems in iterative techniques using the simulation model as a function …


Significant Association Between Punitive And Compensatory Damages In Blockbuster Cases: A Methodological Primer, Theodore Eisenberg, Martin T. Wells Mar 2006

Significant Association Between Punitive And Compensatory Damages In Blockbuster Cases: A Methodological Primer, Theodore Eisenberg, Martin T. Wells

Cornell Law Faculty Publications

This article assesses the relation between punitive and compensatory damages in a data set, gathered by Hersch and Viscusi (H-V), consisting of all known punitive damages awards in excess of $100 million from 1985 through 2003. It shows that a strong, statistically significant relation exists between punitive and compensatory awards, a relation that replicates similar findings in nearly all other analyses of punitive and compensatory damages. H-V's claim that no significant relation exists between punitive and compensatory awards in these data appears to be an artifact of questionable regression methodology.


Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth Feb 2006

Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth

UW Biostatistics Working Paper Series

Objective: To compare performance of one-time health interventions to those that change the probability of transitioning from one health state to another. Study Design and Setting: We used multi-state life table methods to estimate the impact of eight types of interventions on several outcomes. Results: In a cohort beginning at age 65, curing all the sick persons at baseline would increase life expectancy by 0.23 years and increase years of healthy life by .54 years. An equal amount of improvement could be obtained with a 12% decrease in the probability of getting sick, a 16% increase in the probability of …


Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li Feb 2006

Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li

UPenn Biostatistics Working Papers

Mapping genes for complex human diseases is a challenging problem due to the fact that many such diseases are due to both genetic and enviromental risk factors and many also exhibit phenotypic heterogeneity, such as variable age of onset. Information on variable age of disease onset is often a good indicator for disease heterogeneity and incorporation of such information together with enviromental risk factors into genetic analysis should lead to more powerful tests for genetic analysis. Due to the problem of censoring, survival analysis methods have proved to be very useful for genetic analysis. In this paper, I review some …


On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam Feb 2006

On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam

UPenn Biostatistics Working Papers

It is well-known that the correlation among binary outcomes is constrained by the marginal means, yet approaches such as generalized estimating equations (GEE) do not check that the constraints for the correlations are satisfied. We explore this issue for Markovian dependence in the context of a GEE analysis of a clinical trial that compares Venlafaxine with Lithium in the prevention of major depressive episode. We obtain simplified expressions for the constraints for the logistic model and the equicorrelated and first-order autoregressive correlation structures. We then obtain the limiting values of the GEE and quasi-least squares (QLS) estimates of the correlation …