Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2006

Discipline
Institution
Keyword
Publication
Publication Type

Articles 91 - 120 of 244

Full-Text Articles in Statistics and Probability

Efficient Unbiased Estimating Equations For Analyzing Structured Correlation Matrices, Yihao Deng Jul 2006

Efficient Unbiased Estimating Equations For Analyzing Structured Correlation Matrices, Yihao Deng

Mathematics & Statistics Theses & Dissertations

Analysis of dependent continuous and discrete data has become an active area of research. For normal data, correlations fully quantify the dependence. And historically, maximum likelihood method has been very successful to estimate the correlations and unbiased estimating equation approach has become a popular alternative when there may be a departure from normality. In this thesis we show that the optimal unbiased estimating equation coincides with the likelihood equations for normal data. We then introduce a general class of weighted unbiased estimating equations to estimate parameters in a structured correlation matrix. We derive expressions for asymptotic covariance of the estimates, …


Computation Of Weights For Probabilistic Record Linkage Using The Em Algorithm, G. John Bauman Jun 2006

Computation Of Weights For Probabilistic Record Linkage Using The Em Algorithm, G. John Bauman

Theses and Dissertations

Record linkage is the process of combining information about a single individual from two or more records. Probabilistic record linkage gives weights to each field that is compared. The decision of whether the records should be linked is then determined by the sum of the weights, or “Score”, over all fields compared. Using methods similar to the simple versus simple most powerful test, an optimal record linkage decision rule can be established to minimize the number of unlinked records when the probability of false positive and false negative errors are specified. The weights needed for probabilistic record linkage necessitate linking …


A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky Jun 2006

A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky

Harvard University Biostatistics Working Paper Series

No abstract provided.


Longitudinal Nested Compliance Class Model In The Presence Of Time-Varying Noncompliance, Julia Y. Lin, Thomas R. Tenhave, Michael R. Elliott Jun 2006

Longitudinal Nested Compliance Class Model In The Presence Of Time-Varying Noncompliance, Julia Y. Lin, Thomas R. Tenhave, Michael R. Elliott

UPenn Biostatistics Working Papers

This article discusses a nested latent class model for analyzing longitudinal randomized trials when subjects do not always adhere to the treatment to which they are randomized. In the "Prevention of Suicide in Primary Care Elderly: Collaborative Trial" (PROSPECT) study, subjects were randomized to either the control treatment, where they received standard care, or to the intervention, where they received standard care in addition to meeting with depression health specialists. The health specialists educate patients, their families, and physicians about depression and monitor their treatment. Those randomized to the control treatment have no access to the health specialists; however, those …


Hierarchical Lévy Frailty Models And A Frailty Analysis Of Data On Infant Mortality In Norwegian Siblings, Tron Anders Moger, Odd O. Aalen Jun 2006

Hierarchical Lévy Frailty Models And A Frailty Analysis Of Data On Infant Mortality In Norwegian Siblings, Tron Anders Moger, Odd O. Aalen

UW Biostatistics Working Paper Series

Distributions determined by non-negative Lévy processes, which include the power variance function (PVF) distributions among others, are commonly used as frailty distributions to model dependent survival times in family data. We present a hierarchical frailty model constructed by randomizing scale parameters, corresponding to time parameters of Lévy processes, in the Lévy frailty distributions. In its simplest form, this yields a two-model with heterogeneity the individual and family level. The family level frailty is shared within families, creating dependence. In the more complex models, it is extended to allow for several levels of dependence. This yields models with nested dependence structures …


Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr Jun 2006

Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr

UW Biostatistics Working Paper Series

Introduction: Measures of health status and physical function do not usually include a specific code for death. This can cause problems in longitudinal studies because analyses limited to survivors may bias the results. One approach is to recode the status variables to include a reasonable value for death. One method that has been used is to replace each scale value with the estimated probability that a person with this value will be “healthy”. “Healthy” has been defined as being above a particular threshold on the variable of interest one year later, or alternatively as being in excellent, very good, or …


Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan Jun 2006

Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider random design nonparametric regression when the response variable is subject to right censoring. Following the work of Fan and Gijbels (1994), a common approach to this problem is to apply what has been termed a censoring unbiased transformation to the data to obtain surrogate responses, and then enter these surrogate responses with covariate data into standard smoothing algorithms. Existing censoring unbiased transformations generally depend on either the conditional survival function of the response of interest, or that of the censoring variable. We show that a mapping introduced in another statistical context is in fact a censoring unbiased transformation …


New Spiked-In Probe Sets For The Affymetrix Hgu-133a Latin Square Experiment, Monnie Mcgee, Zhongxue Chen Jun 2006

New Spiked-In Probe Sets For The Affymetrix Hgu-133a Latin Square Experiment, Monnie Mcgee, Zhongxue Chen

COBRA Preprint Series

The Affymetrix HGU-133A spike in data set has been used for determining the sensitivity and specificity of various methods for the analysis of microarray data. We show that there are 22 additional probe sets that detect spike in RNAs that should be considered as spike in probe sets. We assign each proposed spiked-in probe set to a concentration group within the Latin Square design, and examine the effects of the additional spiked-in probe sets on assessing the accuracy of analysis methods currently in use. We show that several popular preprocessing methods are more sensitive and specific when the new spike-ins …


A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan Jun 2006

A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider the standard multiple testing problem where many hypotheses are to be tested, each hypothesis is associated with a test statistic, and large test statistics provide evidence against the null hypotheses. One proposal to provide probabilistic control of Type-I errors is the use of procedures ensuring that the expected number of false positives does not exceed a user-supplied threshold. Among such multiple testing procedures, we derive the ``most powerful'' method, meaning the test statistic cutoffs that maximize the expected number of true positives. Unfortunately, these optimal cutoffs depend on the true unknown data generating distribution, so could never be used …


Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng Jun 2006

Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng

UW Biostatistics Working Paper Series

There are two popular statistical approaches to biomarker evaluation. One models the risk of disease (or disease outcome) using, for example, logistic regression. A marker is useful if it has a strong effect on risk. The second evaluates classification performance using measures such as sensitivity, specificity, predictive values and ROC curves. There is controversy about which approach is most appropriate. Moreover, the two approaches often give contradictory results on the same data. We present a new graphic, the predictiveness curve, that complements the risk modeling approach. It assesses the usefulness of a risk model when applied to the population. In …


Small Sample Confidence Intervals In Log Space Back-Transformed From Normal Space, Jason E. Tisdel Jun 2006

Small Sample Confidence Intervals In Log Space Back-Transformed From Normal Space, Jason E. Tisdel

Theses and Dissertations

The logarithmic transformation is commonly applied to a lognormal data set to improve symmetry, homoscedasticity, and linearity. Simple to implement and easy to understand, the logarithm function transforms the original data to closely resemble a normal distribution. Analysis in the normal space provides point estimates and confidence intervals, but transformation back to the original space using the naive approach yields confidence intervals of impractical width. The naive approach offers results that are often inadequate for practical purpose. We present an alternative approach that provides improved results in the form of decreased interval width, increased confidence level, or both. Our alternative …


Low Birth Weight, Very Low Birth Weight And Infant Mortality In San Bernardino County : A Secondary Analysis Of Maternal Factors, Rebecca D. Nanyonjo Jun 2006

Low Birth Weight, Very Low Birth Weight And Infant Mortality In San Bernardino County : A Secondary Analysis Of Maternal Factors, Rebecca D. Nanyonjo

Loma Linda University Electronic Theses, Dissertations & Projects

Purpose: National and state infant mortality rates have slowly declined over the last several years. Despite this reduction, San Bernardino County still has one of the highest infant mortality rates in California and racial disparities between Black and White infants not only persist but continue to widen. These disparities remain at the forefront of concern. Healthy People 2010 target objectives have yet to be reached, while national and state proposed plans have supported the statement that a community's largest health problem is initiated by its infant mortality. The purpose of this study was to investigate maternal factors through use of …


Comparing The Use Of A Manual And A Power Toothbrush By 3-To-4 Year-Old Children, Lorena Delgado Salcedo Jun 2006

Comparing The Use Of A Manual And A Power Toothbrush By 3-To-4 Year-Old Children, Lorena Delgado Salcedo

Loma Linda University Electronic Theses, Dissertations & Projects

PURPOSE: To compare the ability of three-to four-year-old children in performing toothbrushing with manual versus battery-powered toothbrushes with and without oral hygiene instructions.

METHODS: This study was a randomized, controlled, examiner-blind, 4-period crossover design performed on 50 healthy children, ages 3 to 4 years old. Children were assigned either a manual or power toothbrush at visit one and alternated between the two types of toothbrushes each week, for four weeks. Plaque was scored before and after brushing using the Greene and Vermillion -Simplified Oral Hygiene Index. 18,19 At their last two visits of the study, the children received brushing instructions …


Effects Of Diabetes Mellitus On The Healing Of The Dental Pulp, Stuart Evan Garber Jun 2006

Effects Of Diabetes Mellitus On The Healing Of The Dental Pulp, Stuart Evan Garber

Loma Linda University Electronic Theses, Dissertations & Projects

Diabetes Mellitus (DM) has been implicated as a factor affecting healing. The purpose of this study was to use the healing of experimentally exposed pulps subsequent to pulp capping as a model to determine the effect of DM on healing.

Twenty-two Sprague-Dawley rats were divided into two groups of eleven animals each. In one group, DM was induced by injection of 70 mg/KG of streptozotocin. In the other group, the animals were injected with sterile saline. Under anesthesia with Ketamine and Xylazine, the pulps of the maxillary first molars of all rats were exposed using a 1/16 round bur under …


The Effect Of Calcium Hydroxide Pastes On Root Dentin Fracture Resistance, Kurt W. Sturz Jun 2006

The Effect Of Calcium Hydroxide Pastes On Root Dentin Fracture Resistance, Kurt W. Sturz

Loma Linda University Electronic Theses, Dissertations & Projects

Calcium hydroxide is a common intracanal medicament used in the treatment of immature teeth that have been subjected to trauma or decay prior to root canal therapy. The effect of calcium hydroxide on immature root dentin is important. One area of concern is the effect that calcium hydroxide has on the fracture resistance of an immature tooth. It is the aim of this study to compare the effect of four different commercially available calcium hydroxide pastes on the fracture resistance of bovine teeth. Seventy-five freshly extracted, intact bovine incisors were prepared according to a modified Haapasalo and Orstavik technique. Each …


Using Profile Likelihood For Semiparametric Model Selection With Application To Proportional Hazards Mixed Models, Ronghui Xu, Anthony Gamst, Michael Donohue, Florin Vaida, David P. Harrington May 2006

Using Profile Likelihood For Semiparametric Model Selection With Application To Proportional Hazards Mixed Models, Ronghui Xu, Anthony Gamst, Michael Donohue, Florin Vaida, David P. Harrington

Harvard University Biostatistics Working Paper Series

No abstract provided.


Network Activity Arising From Optimal Diameters Of Neuronal Processes, Juliane Gansert May 2006

Network Activity Arising From Optimal Diameters Of Neuronal Processes, Juliane Gansert

Theses

Electrical coupling provides an important pathway for signal transmission between neurons. In several regions of the mammalian brain electrical synapses have been detected, and their role in the synchronization of neural networks and the generation of oscillations has been studied theoretically. Recently, it has been found that the amplitude of the postsynaptic potential is maximized for a specific diameter of the postsynaptic fiber.

In this thesis, the impact of the fiber's diameter on the success or failure of the action potential initiation and propagation is studied theoretically. Systems of two coupled neurons, as well as small networks, are investigated. The …


Comparative Analysis Of Parametric, Nonparametric And Permutation Methods For Differential Expression, Rahul Patil May 2006

Comparative Analysis Of Parametric, Nonparametric And Permutation Methods For Differential Expression, Rahul Patil

Theses

DNA microarrays permit us to study the expression of thousands of genes simultaneously. They are now used in many different contexts to compare mRNA levels between two or more samples of cells. Microarray experiments typically give us expression measurements on a large number of genes. Increasing popularity of microarray technology has resulted in a number of tests being proposed to detect differentials expression.

The purpose of study is to compare the parametric, non parametric and permutation tests when applied to microarray data for differential expression analysis. t test (parametric), Mann Whitney test (nonparametric) and Significance of analysis (permutation ) test …


A Data Gathering Toolkit For Biological Information Integration, Munira Lokhandwala May 2006

A Data Gathering Toolkit For Biological Information Integration, Munira Lokhandwala

Theses

SYSTERS is a biological information integration system containing protein sequences from many protein databases such as Swiss-Prot and TrEMBL and also protein sequences from complete genomes available at Ensembl, The Arabidopsis Information Resource, SGD and GeneDB. For some protein sequences their encoding nucleotide sequences can be found in their corresponding websites. However, for some protein sequences their encoding nucleotide sequences are missing.

The goal of this thesis is to. collect all nucleotide sequences for the protein sequences in SYSTERS and store them in a common database. There are two cases. The first case is that if the nucleotide sequences can …


Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan May 2006

Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Statistical methods have rarely been applied to learn individualized treatment rules, or rules for altering treatments over time in response to changes in individual covariates. Termed dynamic treatment regimes in the statistical literature, such individualized treatment rules are of primary importance in the practice of clinical medicine. History-Adjusted Marginal Structural Models (HA-MSM) estimate individualized treatment rules that assign, at each time point, the first action of the future static treatment plan that optimizes expected outcome given a patient's covariates. However, as we discuss here, the optimality of these rules can depend on the way in which treatment was assigned in …


A Marginalized Diffusion Model For Estimating Age At First Endoscopy Examination From Current Status Data, Diana Miglioretti, Elizabeth Brown May 2006

A Marginalized Diffusion Model For Estimating Age At First Endoscopy Examination From Current Status Data, Diana Miglioretti, Elizabeth Brown

UW Biostatistics Working Paper Series

We propose an approach for estimating the age at first endoscopy examination from current status data collected via two series of cross-sectional surveys. To model the national probability of ever having an endoscopy, we incorporate birth cohort effects into a mixed-influence diffusion model. We link a state-specific model to the national-level diffusion model using a marginalized modeling approach. In future research, results from our model will be used as microsimulation model inputs to estimate the contribution of endoscopy examinations to observed changes in colorectal cancer incidence and mortality.


Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha May 2006

Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha

Harvard University Biostatistics Working Paper Series

Generalized linear mixed models with semiparametric random effects are useful in a wide variety of Bayesian applications. When the random effects arise from a mixture of Dirichlet process (MDP) model, normal base measures and Gibbs sampling procedures based on the Pólya urn scheme are often used to simulate posterior draws. These algorithms are applicable in the conjugate case when (for a normal base measure) the likelihood is normal. In the non-conjugate case, the algorithms proposed by MacEachern and Müller (1998) and Neal (2000) are often applied to generate posterior samples. Some common problems associated with simulation algorithms for non-conjugate MDP …


Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie May 2006

Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider studies of cohorts of individuals after a critical event, such as an injury, with the following characteristics. First, the studies are designed to measure “input” variables, which describe the period before the critical event, and to characterize the distribution of the input variables in the cohort. Second, the studies are designed to measure “output” variables, primarily mortality after the critical event, and to characterize the predictive (conditional) distribution of mortality given the input variables in the cohort. Such studies often possess the complication that the input data are missing for those who die shortly after the critical event …


Bounded Search For De Novo Identification Of Degenerate Cis-Regulatory Elements, Jonathan M. Carlson, Arijit Chakravarty, Radhika S. Khetani, Robert H. Gross May 2006

Bounded Search For De Novo Identification Of Degenerate Cis-Regulatory Elements, Jonathan M. Carlson, Arijit Chakravarty, Radhika S. Khetani, Robert H. Gross

Dartmouth Scholarship

The identification of statistically overrepresented sequences in the upstream regions of coregulated genes should theoretically permit the identification of potential cis-regulatory elements. However, in practice many cis-regulatory elements are highly degenerate, precluding the use of an exhaustive word-counting strategy for their identification. While numerous methods exist for inferring base distributions using a position weight matrix, recent studies suggest that the independence assumptions inherent in the model, as well as the inability to reach a global optimum, limit this approach.


Food Shelf Life: Estimation And Experimental Design, Ross Allen Andrew Larsen May 2006

Food Shelf Life: Estimation And Experimental Design, Ross Allen Andrew Larsen

Theses and Dissertations

Shelf life is a parameter of the lifetime distribution of a food product, usually the time until a specified proportion (1-50%) of the product has spoiled according to taste. The data used to estimate shelf life typically come from a planned experiment with sampled food items observed at specified times. The observation times are usually selected adaptively using ‘staggered sampling.’ Ad-hoc methods based on linear regression have been recommended to estimate shelf life. However, other methods based on maximizing a likelihood (MLE) have been proposed, studied, and used. Both methods assume the Weibull distribution. The observed lifetimes in shelf life …


Multiple Imputation In The Presence Of Outliers, Michael Elliott May 2006

Multiple Imputation In The Presence Of Outliers, Michael Elliott

The University of Michigan Department of Biostatistics Working Paper Series

We consider the problem of obtaining population-based inference in the presence of missing data and outliers in the context of estimating obesity prevalence and body-mass index (BMI) measures from the Healthy For Life Study. Identifying multiple outliers in a multivariate setting is problematic because of problems such as masking, in which groups of outliers inflate the covariance matrix in a fashion that prevents their identification when included, and swamping, in which outliers skew covariances in a fashion that make non-outling observations appear to be outliers. We develop a latent class model that assumes each observation belongs to one of $K$ …


Bayesian Reference Inference On The Ratio Of Poisson Rates., Changbin Guo May 2006

Bayesian Reference Inference On The Ratio Of Poisson Rates., Changbin Guo

Electronic Theses and Dissertations

Bayesian reference analysis is a method of determining the prior under the Bayesian paradigm. It incorporates as little information as possible from the experiment. Estimation of the ratio of two independent Poisson rates is a common practical problem. In this thesis, the method of reference analysis is applied to derive the posterior distribution of the ratio of two independent Poisson rates, and then to construct point and interval estimates based on the reference posterior. In addition, the Frequentist coverage property of HPD intervals is verified through simulation.


An Analysis Of Financial Planning For Employees Of East Tennessee State University., Steven Roy Campbell May 2006

An Analysis Of Financial Planning For Employees Of East Tennessee State University., Steven Roy Campbell

Electronic Theses and Dissertations

The purpose of this study was to determine if East Tennessee State University provides its employees appropriate financial planning services. In particular, it is unknown to what degree employees of East Tennessee State University have actively engaged in financial planning.

The research was conducted during June and July, 2005. Data were gathered by surveying faculty, staff, and retirees of the university. Ten percent of the population responded to the study. The survey instrument covered the areas of retirement, other financial planning services, and attitudes toward financial planning.

The results of the data analysis gave insight into what degree employees of …


Identification Of Gene Expression Patterns Using Planned Linear Contrasts, Hao Li, Constance L. Wood, Yushu Liu, Thomas V. Getchell, Marilyn L. Getchell, Arnold J. Stromberg May 2006

Identification Of Gene Expression Patterns Using Planned Linear Contrasts, Hao Li, Constance L. Wood, Yushu Liu, Thomas V. Getchell, Marilyn L. Getchell, Arnold J. Stromberg

Statistics Faculty Publications

BACKGROUND: In gene networks, the timing of significant changes in the expression level of each gene may be the most critical information in time course expression profiles. With the same timing of the initial change, genes which share similar patterns of expression for any number of sampling intervals from the beginning should be considered co-expressed at certain level(s) in the gene networks. In addition, multiple testing problems are complicated in experiments with multi-level treatments when thousands of genes are involved.

RESULTS: To address these issues, we first performed an ANOVA F test to identify significantly regulated genes. The Benjamini and …


Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen May 2006

Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Bioequivalence trials are usually conducted to compare two or more formulations of a drug. Simultaneous assessment of bioequivalence on multiple endpoints is called multivariate bioequivalence. Despite the fact that some tests for multivariate bioequivalence are suggested, current practice usually involves univariate bioequivalence assessments ignoring the correlations between the endpoints such as AUC and Cmax. In this paper we develop a semiparametric Bayesian test for bioequivalence under multiple endpoints. Specifically, we show how the correlation between the endpoints can be incorporated in the analysis and how this correlation affects the inference. Resulting estimates and posterior probabilities ``borrow strength'' from one another …