Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Biostatistics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1921 - 1950 of 2512

Full-Text Articles in Statistics and Probability

Plasma S-Adenosylmethionine, Dnmt Polymorphisms, And Peripheral Blood Line-1 Methylation Among Healthy Chinese Adults In Singapore, Maki Inoue-Choi, Heather H. Nelson, Kim Robien, Erland Arning, Teodoro Bottiglieri, Woon-Puay Koh, Jian-Min Yuan Aug 2013

Plasma S-Adenosylmethionine, Dnmt Polymorphisms, And Peripheral Blood Line-1 Methylation Among Healthy Chinese Adults In Singapore, Maki Inoue-Choi, Heather H. Nelson, Kim Robien, Erland Arning, Teodoro Bottiglieri, Woon-Puay Koh, Jian-Min Yuan

Epidemiology Faculty Publications

Background

Global hypomethylation of repetitive DNA sequences is believed to occur early in tumorigenesis. There is a great interest in identifying factors that contribute to global DNA hypomethylation and associated cancer risk. We tested the hypothesis that plasma S-adenosylmethionine (SAM) level alone or in combination with genetic variation in DNA methyltransferases (DNMT1, DNMT3A andDNMT3B) was associated with global DNA methylation extent at long interspersed nucleotide element-1 (LINE-1) sequences.

Methods

Plasma SAM level and LINE-1 DNA methylation index were measured using stored blood samples collected from 440 healthy Singaporean Chinese adults during 1994-1999. Genetic polymorphisms of …


Net Reclassification Indices For Evaluating Risk Prediction Instruments: A Critical Review, Kathleen F. Kerr, Zheyu Wang, Holly Janes, Robyn Mcclelland, Bruce M. Psaty, Margaret S. Pepe Aug 2013

Net Reclassification Indices For Evaluating Risk Prediction Instruments: A Critical Review, Kathleen F. Kerr, Zheyu Wang, Holly Janes, Robyn Mcclelland, Bruce M. Psaty, Margaret S. Pepe

UW Biostatistics Working Paper Series

Background Net Reclassification Indices (NRI) have recently become popular statistics for measuring the prediction increment of new biomarkers.

Methods In this review, we examine the various types of NRI statistics and their correct interpretations. We evaluate the advantages and disadvantages of the NRI approach. For pre-defined risk categories, we relate NRI to existing measures of the prediction increment. We also consider statistical methodology for constructing confidence intervals for NRI statistics and evaluate the merits of NRI-based hypothesis testing.

Conclusions Investigators using NRI statistics should report them separately for events (cases) and nonevents (controls). When there are two risk categories, the …


Bayesian Statistical Methods In Gene-Environment And Gene-Gene Interaction Studies, Changlu Liu Aug 2013

Bayesian Statistical Methods In Gene-Environment And Gene-Gene Interaction Studies, Changlu Liu

Dissertations and Theses (Open Access)

Complex diseases such as cancer result from multiple genetic changes and environmental exposures. Due to the rapid development of genotyping and sequencing technologies, we are now able to more accurately assess causal effects of many genetic and environmental factors. Genome-wide association studies have been able to localize many causal genetic variants predisposing to certain diseases. However, these studies only explain a small portion of variations in the heritability of diseases. More advanced statistical models are urgently needed to identify and characterize some additional genetic and environmental factors and their interactions, which will enable us to better understand the causes of …


Correlates Of Hiv Acquisition In A Cohort Of Black Men Who Have Sex With Men In The United States: Hiv Prevention Trials Network (Hptn) 061, Beryl A. Koblin, Kenneth H. Mayer, Susan H. Eshleman, Lei Wang, Sharon B. Mannheimer, Carlos Del Rio, Steve Shoptaw, Manya Magnus, Susan Buchbinder, Leo Wilton, Ting-Yuan Liu, Vanessa Cummings, Estelle Piwowar-Manning, Sheldon D. Fields, Sam Griffith, Vanessa Elharrar, Darrell Wheeler Jul 2013

Correlates Of Hiv Acquisition In A Cohort Of Black Men Who Have Sex With Men In The United States: Hiv Prevention Trials Network (Hptn) 061, Beryl A. Koblin, Kenneth H. Mayer, Susan H. Eshleman, Lei Wang, Sharon B. Mannheimer, Carlos Del Rio, Steve Shoptaw, Manya Magnus, Susan Buchbinder, Leo Wilton, Ting-Yuan Liu, Vanessa Cummings, Estelle Piwowar-Manning, Sheldon D. Fields, Sam Griffith, Vanessa Elharrar, Darrell Wheeler

Epidemiology Faculty Publications

Background

Black men who have sex with men (MSM) in the United States (US) are affected by HIV at disproportionate rates compared to MSM of other race/ethnicities. Current HIV incidence estimates in this group are needed to appropriately target prevention efforts.

Methods

From July 2009 to October 2010, Black MSM reporting unprotected anal intercourse with a man in the past six months were enrolled and followed for one year in six US cities for a feasibility study of a multi-component intervention to reduce HIV infection. HIV incidence based on HIV seroconversion was calculated as number of events/100 person-years. Multivariate proportional …


Testing The Relative Performance Of Data Adaptive Prediction Algorithms: A Generalized Test Of Conditional Risk Differences, Benjamin A. Goldstein, Eric Polley, Farren Briggs, Mark J. Van Der Laan Jul 2013

Testing The Relative Performance Of Data Adaptive Prediction Algorithms: A Generalized Test Of Conditional Risk Differences, Benjamin A. Goldstein, Eric Polley, Farren Briggs, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In statistical medicine comparing the predictability or fit of two models can help to determine whether a set of prognostic variables contains additional information about medical outcomes, or whether one of two different model fits (perhaps based on different algorithms, or different set of variables) should be preferred for clinical use. Clinical medicine has tended to rely on comparisons of clinical metrics like C-statistics and more recently reclassification. Such metrics rely on the outcome being categorical and utilize a specific and often obscure loss function. In classical statistics one can use likelihood ratio tests and information based criterion if the …


Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen Jul 2013

Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

A framework is presented which allows an investigator to estimate the portion of the effect of one exposure that is attributable to an interaction with a second exposure. We show that when the two exposures are independent, the total effect of one exposure can be decomposed into a conditional effect of that exposure and a component due to interaction. The decomposition applies on difference or ratio scales. We discuss how the components can be estimated using standard regression models, and how these components can be used to evaluate the proportion of the total effect of the primary exposure attributable to …


The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk Jul 2013

The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk

Theses and Dissertations

Many continuous medical tests often rely on a threshold for diagnosis. There are two sequential testing strategies of interest: Believe the Positive (BP) and Believe the Negative (BN). BP classifies a patient positive if either the first test is greater than a threshold θ1 or negative on the first test and greater than θ2 on the second test. BN classifies a patient positive if the first test is greater than a threshold θ3 and greater than θ4 on the second test. Threshold pairs θ = (θ1, θ2) or (θ3, θ4), depending on strategy, are defined as optimal if they maximized …


Sample Size Considerations In The Design Of Cluster Randomized Trials Of Combination Hiv Prevention, Rui Wang, Ravi Goyal, Quanhong Lei, M. Essex, Victor Degruttola Jul 2013

Sample Size Considerations In The Design Of Cluster Randomized Trials Of Combination Hiv Prevention, Rui Wang, Ravi Goyal, Quanhong Lei, M. Essex, Victor Degruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Cardiovascular Outcome Trials In Type 2 Diabetes And The Sulphonylurea Controversy: Rationale For The Active-Comparator Carolina Trial, Julio Rosenstock, Nikolaus Marx, Steven E. Kahn, Bernard Zinman, John J. Kastelein, John M. Lachin, Erich Bluhmki, Sanjay Patel, Odd-Erik Johansen, Hans-Jurgen Woerle Jul 2013

Cardiovascular Outcome Trials In Type 2 Diabetes And The Sulphonylurea Controversy: Rationale For The Active-Comparator Carolina Trial, Julio Rosenstock, Nikolaus Marx, Steven E. Kahn, Bernard Zinman, John J. Kastelein, John M. Lachin, Erich Bluhmki, Sanjay Patel, Odd-Erik Johansen, Hans-Jurgen Woerle

Epidemiology Faculty Publications

Sulphonylureas (SUs) are widely used glucose-lowering agents in type 2 diabetes mellitus (T2DM) with apparent declining efficacy over time. Concerns have been raised from observational retrospective studies on the cardiovascular (CV) safety of SUs but there are few long-term data on CV outcomes from randomized controlled trials (RCTs) involving the use of this class of agents. Most of the observational studies and registry data are conflicting and vary with study population and methodology used for analyses. To address the SU controversy, we reviewed the recently published literature (until end of the year 2011) to evaluate the impact of SUs on …


Fast Covariance Estimation For High-Dimensional Functional Data, Luo Xiao, David Ruppert, Vadim Zipunnikov, Ciprian Crainiceanu Jun 2013

Fast Covariance Estimation For High-Dimensional Functional Data, Luo Xiao, David Ruppert, Vadim Zipunnikov, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

For smoothing covariance functions, we propose two fast algorithms that scale linearly with the number of observations per function. Most available methods and software cannot smooth covariance matrices of dimension J x J with J>500; the recently introduced sandwich smoother is an exception, but it is not adapted to smooth covariance matrices of large dimensions such as J \ge 10,000. Covariance matrices of order J=10,000, and even J=100,000$ are becoming increasingly common, e.g., in 2- and 3-dimensional medical imaging and high-density wearable sensor data. We introduce two new algorithms that can handle very large covariance matrices: 1) FACE: a …


Soft Null Hypotheses: A Case Study Of Image Enhancement Detection In Brain Lesions, Haochang Shou, Russell T. Shinohara, Han Liu, Daniel Reich, Ciprian Crainiceanu Jun 2013

Soft Null Hypotheses: A Case Study Of Image Enhancement Detection In Brain Lesions, Haochang Shou, Russell T. Shinohara, Han Liu, Daniel Reich, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

This work is motivated by a study of a population of multiple sclerosis (MS) patients using dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) to identify active brain lesions. At each visit, a contrast agent is administered intravenously to a subject and a series of images is acquired to reveal the location and activity of MS lesions within the brain. Our goal is to identify and quantify lesion enhancement location at the subject level and lesion enhancement patterns at the population level. With this example, we aim to address the difficult problem of transforming a qualitative scientific null hypothesis, such as "this …


Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola Jun 2013

Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh Jun 2013

Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider one observes n i.i.d. copies of a random variable with a probability distribution that is known to be an element of a particular statistical model. In order to define our statistical target we partition the sample in V equal size sub-samples, and use this partitioning to define V splits in estimation-sample (one of the V subsamples) and corresponding complementary parameter-generating sample that is used to generate a target parameter. For each of the V parameter-generating samples, we apply an algorithm that maps the sample in a target parameter mapping which represent the statistical target parameter generated by that parameter-generating …


When To Start Antiretroviral Therapy: The Need For An Evidence Base During Early Hiv Infection, James D. Lundgren, Abdel G. Babiker, Fred M. Gordin, Alvaro H. Borges, James D. Neaton Jun 2013

When To Start Antiretroviral Therapy: The Need For An Evidence Base During Early Hiv Infection, James D. Lundgren, Abdel G. Babiker, Fred M. Gordin, Alvaro H. Borges, James D. Neaton

Epidemiology Faculty Publications

Background

Strategies for use of antiretroviral therapy (ART) have traditionally focused on providing treatment to persons who stand to benefit immediately from initiating the therapy. There is global consensus that any HIV+ person with CD4 counts less than 350 cells/μl should initiate ART. However, it remains controversial whether ART is indicated in asymptomatic HIV-infected persons with CD4 counts above 350 cells/μl, or whether it is more advisable to defer initiation until the CD4 count has dropped to 350 cells/μl. The question of when the best time is to initiate ART during early HIV infection has always been vigorously debated. The …


Association Between Adverse Childhood Experiences And Diagnosis Of Cancer, Monique J. Brown, Leroy R. Thacker, Steven A. Cohen Jun 2013

Association Between Adverse Childhood Experiences And Diagnosis Of Cancer, Monique J. Brown, Leroy R. Thacker, Steven A. Cohen

Faculty Publications

Objective: Adverse childhood experiences (ACEs) are linked to multiple adverse health outcomes. This study examined the association between ACEs and cancer diagnosis.

Methods: Data from the 2010 Behavioral Risk Factor Surveillance System (BRFSS) survey were used. The BRFSS is the largest ongoing telephone health survey, conducted in all US states, the District of Columbia, Puerto Rico, Guam and the U.S. Virgin Islands, and provides data on a variety of health issues among the non-institutionalized adult population. Principal component analysis (PCA) was used to derive components for ACEs. Multivariable logistic regression models were used to provide adjusted odds ratios (OR) and …


Restricted Likelihood Ratio Tests For Functional Effects In The Functional Linear Model, Bruce J. Swihart, Jeff Goldsmith, Ciprian M. Crainiceanu Jun 2013

Restricted Likelihood Ratio Tests For Functional Effects In The Functional Linear Model, Bruce J. Swihart, Jeff Goldsmith, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The goal of our article is to provide a transparent, robust, and computationally feasible statistical approach for testing in the context of scalar-on-function linear regression models. In particular, we are interested in testing for the necessity of functional effects against standard linear models. Our methods are motivated by and applied to a large longitudinal study involving diffusion tensor imaging of intracranial white matter tracts in a susceptible cohort. In the context of this study, we conduct hypothesis tests that are motivated by anatomical knowledge and which support recent findings regarding the relationship between cognitive impairment and white matter demyelination. R-code …


Augmentation Of Propensity Scores For Medical Records-Based Research, Mikel Aickin Jun 2013

Augmentation Of Propensity Scores For Medical Records-Based Research, Mikel Aickin

COBRA Preprint Series

Therapeutic research based on electronic medical records suffers from the possibility of various kinds of confounding. Over the past 30 years, propensity scores have increasingly been used to try to reduce this possibility. In this article a gap is identified in the propensity score methodology, and it is proposed to augment traditional treatment-propensity scores with outcome-propensity scores, thereby removing all other aspects of common causes from the analysis of treatment effects.


A Versatile Test For Equality Of Two Survival Functions Based On Weighted Differences Of Kaplan-Meier Curves, Hajime Uno, Lu Tian, Brian Claggett, L. J. Wei May 2013

A Versatile Test For Equality Of Two Survival Functions Based On Weighted Differences Of Kaplan-Meier Curves, Hajime Uno, Lu Tian, Brian Claggett, L. J. Wei

Harvard University Biostatistics Working Paper Series

With censored event time observations, the logrank test is the most popular tool for testing the equality of two underlying survival distributions. Although this test is asymptotically distribution-free, it may not be powerful when the proportional hazards assumption is violated. Various other novel testing procedures have been proposed, which generally are derived by assuming a class of specific alternative hypotheses with respect to the hazard functions. The test considered by Pepe and Fleming (1989) is based on a linear combination of weighted differences of two Kaplan-Meier curves over time and is a natural tool to assess the difference of two …


Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny May 2013

Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny

U.C. Berkeley Division of Biostatistics Working Paper Series

Ensemble methods using the same underlying algorithm trained on different subsets of observations have recently received increased attention as practical prediction tools for massive datasets. We propose Subsemble: a general subset ensemble prediction method, which can be used for small, moderate, or large datasets. Subsemble partitions the full dataset into subsets of observations, fits a specified underlying algorithm on each subset, and uses a clever form of V-fold cross-validation to output a prediction function that combines the subset-specific fits. We give an oracle result that provides a theoretical performance guarantee for Subsemble. Through simulations, we demonstrate that Subsemble can be …


Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan May 2013

Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper describes a targeted maximum likelihood estimator (TMLE) for the parameters of longitudinal static and dynamic marginal structural models. We consider a longitudinal data structure consisting of baseline covariates, time-dependent intervention nodes, intermediate time-dependent covariates, and a possibly time dependent outcome. The intervention nodes at each time point can include a binary treatment as well as a right-censoring indicator. Given a class of dynamic or static interventions, a marginal structural model is used to model the mean of the intervention specific counterfactual outcome as a function of the intervention, time point, and possibly a subset of baseline covariates. Because …


Varying Index Coefficient Models, Shujie Ma, Peter Xuekun Song May 2013

Varying Index Coefficient Models, Shujie Ma, Peter Xuekun Song

The University of Michigan Department of Biostatistics Working Paper Series

It has been a long history of utilizing interactions in regression analysis to investigate interactive effects of covariates on response variables. In this paper we aim to address two kinds of new challenges resulted from the inclusion of such high-order effects in the regression model for complex data. The first kind arises from a situation where interaction effects of individual covariates are weak but those of combined covariates are strong, and the other kind pertains to the presence of nonlinear interactive effects. Generalizing the single index coefficient regression model (Xia and Li, 1999), we propose a new class of semiparametric …


Balancing Score Adjusted Targeted Minimum Loss-Based Estimation, Samuel D. Lendle, Bruce Fireman, Mark J. Van Der Laan May 2013

Balancing Score Adjusted Targeted Minimum Loss-Based Estimation, Samuel D. Lendle, Bruce Fireman, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Adjusting for a balancing score is sufficient for bias reduction when estimating causal effects including the average treatment effect and effect among the treated. Estimators that adjust for the propensity score in a nonparametric way, such as matching on an estimate of the propensity score, can be consistent when the estimated propensity score is not consistent for the true propensity score but converges to some other balancing score. We call this property the balancing score property, and discuss a class of estimators that have this property. We introduce a targeted minimum loss-based estimator (TMLE) for a treatment specific mean with …


Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming, Michael Rosenblum, Han Liu, En-Hsu Yen May 2013

Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming, Michael Rosenblum, Han Liu, En-Hsu Yen

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose new, optimal methods for analyzing randomized trials, when it is suspected that treatment effects may differ in two predefined subpopulations. Such sub-populations could be defined by a biomarker or risk factor measured at baseline. The goal is to simultaneously learn which subpopulations benefit from an experimental treatment, while providing strong control of the familywise Type I error rate. We formalize this as a multiple testing problem and show it is computationally infeasible to solve using existing techniques. Our solution involves a novel approach, in which we first transform the original multiple testing problem into a large, sparse linear …


Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan May 2013

Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many of the secondary outcomes in observational studies and randomized trials are rare. Methods for estimating causal effects and associations with rare outcomes, however, are limited, and this represents a missed opportunity for investigation. In this article, we construct a new targeted minimum loss-based estimator (TMLE) for the effect of an exposure or treatment on a rare outcome. We focus on the causal risk difference and statistical models incorporating bounds on the conditional risk of the outcome, given the exposure and covariates. By construction, the proposed estimator constrains the predicted outcomes to respect this model knowledge. Theoretically, this bounding provides …


Linking And Retaining Hiv Patients In Care: The Importance Of Provider Attitudes And Behaviors, Manya Magnus, Jane Herwehe, Michelli Murtaza-Rossini, Petera Reine, Damien Cuffie, Deann Gruber, Michael Kaiser May 2013

Linking And Retaining Hiv Patients In Care: The Importance Of Provider Attitudes And Behaviors, Manya Magnus, Jane Herwehe, Michelli Murtaza-Rossini, Petera Reine, Damien Cuffie, Deann Gruber, Michael Kaiser

Epidemiology Faculty Publications

Retention in HIV treatment may reduce morbidity and mortality, as well as slow the epidemic. Myriad barriers to retention include stigma, homophobia, structural barriers, transportation, and insurance. The purpose of this study was to evaluate patient perceptions of provider attitudes among HIV-infected persons within a state-wide public hospital system in Louisiana. A convenience sample of patients attending HIV clinics throughout the state participated in an anonymous interview. Factors associated with negative perceptions of care were evaluated in conjunction with a validated stigma measure. Factors associated with having a delayed entry into or break in care were evaluated in conjunction with …


Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis, Sihai Dave Zhao, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron May 2013

Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis, Sihai Dave Zhao, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron

Harvard University Biostatistics Working Paper Series

No abstract provided.


An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes, Ekaterina Eliseeva, Alan E. Hubbard, Ira B. Tager May 2013

An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes, Ekaterina Eliseeva, Alan E. Hubbard, Ira B. Tager

U.C. Berkeley Division of Biostatistics Working Paper Series

Analyses of epidemiological studies of the association between short-term changes in air pollution and health outcomes have not sufficiently discussed the degree to which the statistical models chosen for these analyses reflect what is actually known about the true data-generating distribution. We present a method to estimate population-level ambient air pollution (NO2) exposure-health (wheeze in children with asthma) response functions that is not dependent on assumptions about the data-generating function that underlies the observed data and which focuses on a specific scientific parameter of interest (the marginal adjusted association of exposure on probability of wheeze, over a grid of possible …


Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates, Wei Dai, Yi Li May 2013

Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates, Wei Dai, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Microarray technology has the potential to lead to a better understanding of biological processes and diseases such as cancer. When failure time outcomes are also available, one might be interested in relating gene expression profiles to the survival outcome such as time to cancer recurrence or time to death. This is statistically challenging because the number of covariates greatly exceeds the number of observations. While the majority of work has focused on regularized Cox regression model and accelerated failure time model, they may be restrictive in practice. We relax the model assumption and and consider a nonparametric transformation model that …


Compound Identification Using Penalized Linear Regression., Ruiqi Liu May 2013

Compound Identification Using Penalized Linear Regression., Ruiqi Liu

Electronic Theses and Dissertations

In this study, we propose a new method for compound identification using penalized linear regression. Compound identification is often achieved by matching the experimental mass spectra to the mass spectra stored in a reference library based on mass spectral similarity. In the context of the linear regression, the response variable is an experimental mass spectrum (i.e., query) and all the compounds in the reference library are the independent variables. However, the number of compounds in the reference library is much larger than the range of m/z values so that the data become high dimensional data with suffering from singularity. For …


Development Of Novel Methods To Minimize The Impact Of Sequencing Errors In The Next-Generation Sequencing Data Analysis, Xiaofeng Zheng May 2013

Development Of Novel Methods To Minimize The Impact Of Sequencing Errors In The Next-Generation Sequencing Data Analysis, Xiaofeng Zheng

Dissertations and Theses (Open Access)

Next-generation sequencing (NGS) technology has become a prominent tool in biological and biomedical research. However, NGS data analysis, such as de novo assembly, mapping and variants detection is far from maturity, and the high sequencing error-rate is one of the major problems. .

To minimize the impact of sequencing errors, we developed a highly robust and efficient method, MTM, to correct the errors in NGS reads. We demonstrated the effectiveness of MTM on both single-cell data with highly non-uniform coverage and normal data with uniformly high coverage, reflecting that MTM’s performance does not rely on the coverage of the sequencing …