Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2011

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 271 - 300 of 406

Full-Text Articles in Statistics and Probability

Variable Selection And Parameter Estimation Using A Continuous And Differentiable Approximation To The L0 Penalty Function, Douglas Nielsen Vanderwerken Mar 2011

Variable Selection And Parameter Estimation Using A Continuous And Differentiable Approximation To The L0 Penalty Function, Douglas Nielsen Vanderwerken

Theses and Dissertations

L0 penalized likelihood procedures like Mallows' Cp, AIC, and BIC directly penalize for the number of variables included in a regression model. This is a straightforward approach to the problem of overfitting, and these methods are now part of every statistician's repertoire. However, these procedures have been shown to sometimes result in unstable parameter estimates as a result on the L0 penalty's discontinuity at zero. One proposed alternative, seamless-L0 (SELO), utilizes a continuous penalty function that mimics L0 and allows for stable estimates. Like other similar methods (e.g. LASSO and SCAD), SELO produces sparse solutions because the penalty function is …


Quantitative Interpretation Of A Genetic Model Of Carcinogenesis Using Computer Simulations, Donghai Dai, Brandon Beck, Xiaofang Wang, Cory Howk, Yi Li Mar 2011

Quantitative Interpretation Of A Genetic Model Of Carcinogenesis Using Computer Simulations, Donghai Dai, Brandon Beck, Xiaofang Wang, Cory Howk, Yi Li

Mathematics and Statistics Faculty Publications

The genetic model of tumorigenesis by Vogelstein et al. (V theory) and the molecular definition of cancer hallmarks by Hanahan and Weinberg (W theory) represent two of the most comprehensive and systemic understandings of cancer. Here, we develop a mathematical model that quantitatively interprets these seminal cancer theories, starting from a set of equations describing the short life cycle of an individual cell in uterine epithelium during tissue regeneration. The process of malignant transformation of an individual cell is followed and the tissue (or tumor) is described as a composite of individual cells in order to quantitatively account for intra-tumor …


Estimating Subject-Specific Treatment Differences For Risk-Benefit Assessment With Competing Risk Event-Time Data, Brian Claggett, Lihui Zhao, Lu Tian, Davide Castagno, L. J. Wei Mar 2011

Estimating Subject-Specific Treatment Differences For Risk-Benefit Assessment With Competing Risk Event-Time Data, Brian Claggett, Lihui Zhao, Lu Tian, Davide Castagno, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Hierarchical Bayesian Methods For Evaluation Of Traffic Project Efficacy, Andrew Nolan Olsen Mar 2011

Hierarchical Bayesian Methods For Evaluation Of Traffic Project Efficacy, Andrew Nolan Olsen

Theses and Dissertations

A main objective of Departments of Transportation is to improve the safety of the roadways over which they have jurisdiction. Safety projects, such as cable barriers and raised medians, are utilized to reduce both crash frequency and crash severity. The efficacy of these projects must be evaluated in order to use resources in the best way possible. Five models are proposed for the evaluation of traffic projects: (1) a Bayesian Poisson regression model; (2) a hierarchical Poisson regression model building on model (1) by adding hyperpriors; (3) a similar model correcting for overdispersion; (4) a dynamic linear model; and (5) …


Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan Mar 2011

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan

Johns Hopkins University, Dept. of Biostatistics Working Papers

We present a brief overview of targeted maximum likelihood for estimating the causal effect of a single time point treatment and of a two time point treatment. We focus on simple examples demonstrating how to apply the methodology developed in (van der Laan and Rubin, 2006; Moore and van der Laan, 2007; van der Laan, 2010a,b). We include R code for the single time point case.


Parameter Estimation For The Two-Parameter Weibull Distribution, Mark A. Nielsen Mar 2011

Parameter Estimation For The Two-Parameter Weibull Distribution, Mark A. Nielsen

Theses and Dissertations

The Weibull distribution, an extreme value distribution, is frequently used to model survival, reliability, wind speed, and other data. One reason for this is its flexibility; it can mimic various distributions like the exponential or normal. The two-parameter Weibull has a shape (γ) and scale (β) parameter. Parameter estimation has been an ongoing search to find efficient, unbiased, and minimal variance estimators. Through data analysis and simulation studies, the following three methods of estimation will be discussed and compared: maximum likelihood estimation (MLE), method of moments estimation (MME), and median rank regression (MRR). The analysis of wind speed data from …


Utilizing Universal Probability Of Expression Code (Upc) To Identify Disrupted Pathways In Cancer Samples, Michelle Rachel Withers Mar 2011

Utilizing Universal Probability Of Expression Code (Upc) To Identify Disrupted Pathways In Cancer Samples, Michelle Rachel Withers

Theses and Dissertations

Understanding the role of deregulated biological pathways in cancer samples has the potential to improve cancer treatment, making it more effective by selecting treatments that reverse the biological cause of the cancer. One of the challenges with pathway analysis is identifying a deregulated pathway in a given sample. This project develops the Universal Probability of Expression Code (UPC), a profile of a single deregulated biological path- way, and projects it into a cancer cell to determine if it is present. One of the benefits of this method is that rather than use information from a single over-expressed gene, it pro- …


A Note On The Interpretation Of Scale Values In Multidimensional Scaling Growth Analysis, Cody Ding Mar 2011

A Note On The Interpretation Of Scale Values In Multidimensional Scaling Growth Analysis, Cody Ding

Education Sciences and Professional Programs Faculty Works

No abstract provided.


Maximal Sensitive Dependence And The Optimal Path To Epidemic Extinction, Eric Forgoston, Simone Bianco, Leah B. Shaw, Ira B. Schwartz Mar 2011

Maximal Sensitive Dependence And The Optimal Path To Epidemic Extinction, Eric Forgoston, Simone Bianco, Leah B. Shaw, Ira B. Schwartz

Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works

Extinction of an epidemic or a species is a rare event that occurs due to a large, rare stochastic fluctuation. Although the extinction process is dynamically unstable, it follows an optimal path that maximizes the probability of extinction. We show that the optimal path is also directly related to the finite-time Lyapunov exponents of the underlying dynamical system in that the optimal path displays maximum sensitivity to initial conditions. We consider several stochastic epidemic models, and examine the extinction process in a dynamical systems framework. Using the dynamics of the finite-time Lyapunov exponents as a constructive tool, we demonstrate that …


Global Well-Posedness And Asymptotic Behavior Of A Class Of Initial-Boundary-Value Problems Of The Kdv Equation On A Finite Domain, Ivonne Rivas, Muhammad Usman, Bingyu Zhang Mar 2011

Global Well-Posedness And Asymptotic Behavior Of A Class Of Initial-Boundary-Value Problems Of The Kdv Equation On A Finite Domain, Ivonne Rivas, Muhammad Usman, Bingyu Zhang

Mathematics Faculty Publications

In this paper, we study a class of initial boundary value problem (IBVP) of the Korteweg- de Vries equation posed on a ?nite interval with nonhomogeneous boundary conditions. The IBVP is known to be locally well-posed, but its global L2 a priori estimate is not available and therefore it is not clear whether its solutions exist globally or blow up in finite time. It is shown in this paper that the solutions exist globally as long as their initial value and the associated boundary data are small, and moreover, those solutions decay exponentially if their boundary data decay exponentially.


An Approach To Nearest Neighboring Search For Multi-Dimensional Data, Yong Shi, Li Zhang, Lei Zhu Mar 2011

An Approach To Nearest Neighboring Search For Multi-Dimensional Data, Yong Shi, Li Zhang, Lei Zhu

Faculty Articles

Finding nearest neighbors in large multi-dimensional data has always been one of the research interests in data mining field. In this paper, we present our continuous research on similarity search problems. Previously we have worked on exploring the meaning of K nearest neighbors from a new perspective in PanKNN [20]. It redefines the distances between data points and a given query point Q, efficiently and effectively selecting data points which are closest to Q. It can be applied in various data mining fields. A large amount of real data sets have irrelevant or obstacle information which greatly affects the effectiveness …


Projective-Planar Graphs With No K3,4-Minor, John Maharry, Dan Slilaty Mar 2011

Projective-Planar Graphs With No K3,4-Minor, John Maharry, Dan Slilaty

Mathematics and Statistics Faculty Publications

An exact structure is described to classify the projective‐planar graphs that do not contain a K3, 4‐minor.


Characterizing The Arc, Paul Bankston Mar 2011

Characterizing The Arc, Paul Bankston

Mathematics, Statistics and Computer Science Faculty Research and Publications

We offer a mathematical logic framework for talking about when a given topological space is characterizable relative to a “class of its peers.” The framework involves a relational alphabet for which there is a natural way of assigning relational structures to spaces in the peer class; and a putative characterization of the given space consists of a set of first-order sentences over that alphabet, all true for the space. The attempted characterization is a success if any peer space satisfying all the sentences is inevitably homeomorphic to the given space.

For example, it has long been known that the arc …


Common Hypercyclic Vectors For The Conjugate Class Of A Hypercyclic Operator, Kit C. Chan, Rebecca Sanders Mar 2011

Common Hypercyclic Vectors For The Conjugate Class Of A Hypercyclic Operator, Kit C. Chan, Rebecca Sanders

Mathematics, Statistics and Computer Science Faculty Research and Publications

Given a separable, infinite dimensional Hilbert space, it was recently shown by the authors that there is a path of chaotic operators, which is dense in the operator algebra with the strong operator topology, and along which every operator has the exact same dense Gδ set of hypercyclic vectors. In the present work, we show that the conjugate set of any hypercyclic operator on a separable, infinite dimensional Banach space always contains a path of operators which is dense with the strong operator topology, and yet the set of common hypercyclic vectors for the entire path is a dense …


Signal And Noise In Complex-Valued Sense Mr Image Reconstruction, Daniel B. Rowe, Iain P. Bruce Mar 2011

Signal And Noise In Complex-Valued Sense Mr Image Reconstruction, Daniel B. Rowe, Iain P. Bruce

Mathematics, Statistics and Computer Science Faculty Research and Publications

In fMRI, brain images are not measured instantaneously and a volume of images can take two seconds to acquire at a low 64x64 resolution. Significant effort has been put forth on many fronts to decrease image acquisition time including parallel imaging. In parallel imaging, sub-sampled spatial frequency points are measured in parallel and combined to form a single image. Measurement time is decreased at the expense of increased image reconstruction difficulty and time. One significant parallel imaging technique known as SENSE utilizes a complex-valued regression coefficient estimation process with transposes replaced by conjugate transposes. However, in SENSE the noise structure …


Semidistributive Inverse Semigroups, Ii, Kyeong Hee Cheong, Peter R. Jones Mar 2011

Semidistributive Inverse Semigroups, Ii, Kyeong Hee Cheong, Peter R. Jones

Mathematics, Statistics and Computer Science Faculty Research and Publications

The description by Johnston-Thom and the second author of the inverse semigroups S for which the lattice LJ(S) of full inverse subsemigroups of S is join semidistributive is used to describe those for which (a) the lattice L(S) of all inverse subsemigroups or (b) the lattice lo(S) of convex inverse subsemigroups have that property. In contrast with the methods used by the authors to investigate lower semimodularity, the methods are based on decompositions via GS, the union of the subgroups of the semigroup (which is necessarily cryptic).


Some Problems And Solutions In The Experimental Science Of Technology: The Proper Use And Reporting Of Statistics In Computational Intelligence, With An Experimental Design From Computational Ethnomusicology, Mehmet Vurkaç Feb 2011

Some Problems And Solutions In The Experimental Science Of Technology: The Proper Use And Reporting Of Statistics In Computational Intelligence, With An Experimental Design From Computational Ethnomusicology, Mehmet Vurkaç

Systems Science Friday Noon Seminar Series

Statistics is the meta-science that lends validity and credibility to The Scientific Method. However, as a complex and advanced Science in itself, Statistics is often misunderstood and misused by scientists, engineers, medical and legal professionals and others. In the area of Computational Intelligence (CI), there have been numerous misuses of statistical techniques leading to the publishing of insupportable results, which, in addition to being a problem in itself, has also contributed to a degree of rift between the Statistics/Statistical Learning community and the Machine Learning/Computational Intelligence community. This talk surveys a number of misuses of statistical inference in CI settings, …


Efektivitas Dan Efisiensi Sistem Informasi Keluarga Berencana Di Puskesmas, Aragar Putri, Besral Besral Feb 2011

Efektivitas Dan Efisiensi Sistem Informasi Keluarga Berencana Di Puskesmas, Aragar Putri, Besral Besral

Kesmas

Efisiensi dan efektivitas sistem informasi keluarga berencana (KB)-kesehatan yang telah disosialisasikan sejak tahun 2007 dibandingkan dengan sistem yang lama belum diketahui. Suatu penelitian survei dilakukan di empat provinsi, yaitu DKI Jakarta, Lampung, Kalimatan Tengah, dan Bali. Di tiap provinsi dipilih dua kabupaten/kota dan pada tiap kabupaten/kota dipilih dua puskesmas (kecamatan) yang sudah menerapkan sistem informasi KB-kesehatan tersebut. Pengumpulan data dilakukan pada bulan Juni-September 2008. Penelitian ini menemukan bahwa efektivitas dan efisiensi sistem informasi KB yang baru cukup baik, 77,8% responden menyatakan lebih efektif atau sangat lebih efektif dan 66,7% responden menyatakan lebih efisien atau sangat lebih efisien dibandingkan dengan sistem …


Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers, Xiao-Hua Zhou, Yunbei Ma Feb 2011

Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers, Xiao-Hua Zhou, Yunbei Ma

UW Biostatistics Working Paper Series

In this paper, for time-to-event data, we propose a new statistical framework for casual inference in evaluating clinical utility of predictive biomarkers and in selecting an optimal treatment for a particular patient. This new casual framework is based on a new concept, called Biomarker Adjusted Treatment Effect (BATE) curve, which can be used to represent the clinical utility of a predictive biomarker and select an optimal treatment for one particular patient. We then propose semi-parametric methods for estimating the BATE curves of biomarkers and establish asymptotic results of the proposed estimators for the BATE curves. We also conduct extensive simulation …


Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan Feb 2011

Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (TMLE) presents an approach for construction of an efficient double-robust semi-parametric substitution estimator of a target feature of the data generating distribution, such as a statistical association measure or a causal effect parameter. tmle is a recently developed R package that implements TMLE for estimation of the effect of a binary treatment at a single point in time on an outcome of interest, controlling for user supplied covariates: the additive treatment effect, the relative risk, the odds ratio. The package allows outcome data with missingness, and experimental units that contribute repeated records of the point-treatment data …


Causal Inference Under Multiple Versions Of Treatment, Tyler J. Vanderweele, Miguel A. Hernan Feb 2011

Causal Inference Under Multiple Versions Of Treatment, Tyler J. Vanderweele, Miguel A. Hernan

COBRA Preprint Series

In this article we discuss the no-multiple-versions-of-treatment assumption and extend the potential outcomes framework to accommodate causal inference under violations of this assumption. A variety of examples are discussed in which the assumption may be violated. Identification results are provided for the overall treatment effect and the effect of treatment on the treated when multiple versions of treatment are present and also for the causal effect comparing a version of one treatment to some other version of the same or a different treatment. Further identification and interpretative results are given for cases in which a treatment variable is dichotomized to …


Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett Feb 2011

Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett

Johns Hopkins University, Dept. of Biostatistics Working Papers

In this manuscript, we use a two-stage decomposition for the analysis of func- tional magnetic resonance imaging (fMRI). In the first stage, spatial independent component analysis is applied to the group fMRI data to obtain common brain networks (spatial maps) and subject-specific mixing matrices (time courses). In the second stage, functional principal component analysis is utilized to decompose the mixing matrices into population- level eigenvectors and subject-specific loadings. Inference is performed using permutation-based exact conditional logistic regression for matched pairs data. Simulation studies suggest the ability of the decomposition methods to recover population brain networks and the major direction of …


Differential Gene Expression In Liver And Small Intestine From Lactating Rats Compared To Age-Matched Virgin Controls Detects Increased Mrna Of Cholesterol Biosynthetic Genes, Antony Athippozhy, Liping Huang, Clavia Ruth Wooton-Kee, Tianyong Zhao, Paiboon Jungsuwadee, Arnold J. Stromberg, Mary Vore Feb 2011

Differential Gene Expression In Liver And Small Intestine From Lactating Rats Compared To Age-Matched Virgin Controls Detects Increased Mrna Of Cholesterol Biosynthetic Genes, Antony Athippozhy, Liping Huang, Clavia Ruth Wooton-Kee, Tianyong Zhao, Paiboon Jungsuwadee, Arnold J. Stromberg, Mary Vore

Statistics Faculty Publications

BACKGROUND: Lactation increases energy demands four- to five-fold, leading to a two- to three-fold increase in food consumption, requiring a proportional adjustment in the ability of the lactating dam to absorb nutrients and to synthesize critical biomolecules, such as cholesterol, to meet the dietary needs of both the offspring and the dam. The size and hydrophobicity of the bile acid pool increases during lactation, implying an increased absorption and disposition of lipids, sterols, nutrients, and xenobiotics. In order to investigate changes at the transcriptomics level, we utilized an exon array and calculated expression levels to investigate changes in gene expression …


The (1,2)-Step Competition Graph Of A Tournament, Kim A. S. Factor, Sarah Merz Jan 2011

The (1,2)-Step Competition Graph Of A Tournament, Kim A. S. Factor, Sarah Merz

Mathematics, Statistics and Computer Science Faculty Research and Publications

The competition graph of a digraph, introduced by Cohen in 1968, has been extensively studied. More recently, in 2000, Cho, Kim, and Nam defined the m-step competition graph. In this paper, we offer another generalization of the competition graph. We define the (1,2)-step competition graph of a digraph D, denoted C1,2(D), as the graph on V(D) where {x,y}∈E(C1,2(D)) if and only if there exists a vertex z≠x,y, such that either dD−y( …


Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study, Xiao-Hua Zhou, Baojiang Chen Jan 2011

Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study, Xiao-Hua Zhou, Baojiang Chen

UW Biostatistics Working Paper Series

Identifying risk factors for transition rates among normal cognition, mildly cognitive impairment, dementia and death in an Alzheimer's disease study is very important. It is known that transition rates among these states are strongly time dependent. While Markov process models are often used to describe these disease progressions, the literature mainly focuses on time homogeneous processes, and limited tools are available for dealing with non-homogeneity. Further, patients may choose when they want to visit the clinics, which creates informative observations. In this paper, we develop methods to deal with non-homogeneous Markov processes through time scale transformation when observation times are …


Doubly Robust Estimates For Binary Longitudinal Data Analysis With Missing Response And Missing Covariates, Baojiang Chen, Xiao-Hua Zhou Jan 2011

Doubly Robust Estimates For Binary Longitudinal Data Analysis With Missing Response And Missing Covariates, Baojiang Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Longitudinal studies often feature incomplete response and covariate data. Likelihood-based methods such as the EM algorithm give consistent estimators for model parameters when data are missing at random provided that the response model and the missing covariate model are correctly specified; but we do not need to specify the missing data mechanism. An alternative method is the weighted estimating equation which gives consistent estimators if the missing data and response models are correctly specified; but we do not need to specify the distribution of the covariates that have missing values. In this paper we develop a doubly robust estimation method …


Semiparametric Estimation Of The Covariate-Specific Roc Curve In Presence Of Ignorable Verification Bias, Danping Liu, Xiao-Hua Zhou Jan 2011

Semiparametric Estimation Of The Covariate-Specific Roc Curve In Presence Of Ignorable Verification Bias, Danping Liu, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Covariate-specific ROC curves are often used to evaluate the classification accuracy of a medical diagnostic test or a biomarker, when the accuracy of the test is associated with certain covariates. In many large-scale screening tests, the gold standard is subject to missingness due to high cost or harmfulness to the patient. In this paper, we propose a semiparametric estimation method for the covariate-specific ROC curves with a partial missing gold standard. A location-scale model is constructed for the test result to model the covariates' effect, but the residual distributions are left unspecified. Thus the baseline and link functions of the …


Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou Jan 2011

Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

For many medical conditions there are several treatment options available to patients. We consider evaluating markers based on a simple treatment selection policy that incorporates information on the patient's marker value exceeding a threshold. Although traditional regression methods may assess the effect of the marker and treatment on outcomes, it is appealing to quantify more directly the potential impact on the population of using the marker to select treatment. A useful tool is the selection impact (SI) curve proposed by Song and Pepe (2004, \textit{Biometrics} \textbf{60}, 874--883) for binary outcomes. However, this approach does not deal with continuous outcomes, nor …


A Flexible Spatio-Temporal Model For Air Pollution: Allowing For Spatio-Temporal Covariates, Johan Lindstrom, Adam A. Szpiro, Paul D. Sampson, Lianne Sheppard, Assaf Oron, Mark Richards, Tim Larson Jan 2011

A Flexible Spatio-Temporal Model For Air Pollution: Allowing For Spatio-Temporal Covariates, Johan Lindstrom, Adam A. Szpiro, Paul D. Sampson, Lianne Sheppard, Assaf Oron, Mark Richards, Tim Larson

UW Biostatistics Working Paper Series

Given the increasing interest in the association between exposure to air pollution and adverse health outcomes, the development of models that provide accurate spatio-temporal predictions of air pollution concentrations at small spatial scales is of great importance when assessing potential health effects of air pollution. The methodology presented here has been developed as part of the Multi-Ethnic Study of Atherosclerosis and Air Pollution (MESA Air), a prospective cohort study funded by the US EPA to investigate the relationship between chronic exposure to air pollution and cardiovascular disease. We present a spatio-temporal framework that models and predicts ambient air pollution by …


Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu Jan 2011

Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We establish a fundamental equivalence between singular value decomposition (SVD) and functional principal components analysis (FPCA) models. The constructive relationship allows to deploy the numerical efficiency of SVD to fully estimate the components of FPCA, even for extremely high-dimensional functional objects, such as brain images. As an example, a functional mixed effect model is fitted to high-resolution morphometric (RAVENS) images. The main directions of morphometric variation in brain volumes are identified and discussed.