Varying Index Coefficient Models,
2013
University of California - Riverside
Varying Index Coefficient Models, Shujie Ma, Peter Xuekun Song
The University of Michigan Department of Biostatistics Working Paper Series
It has been a long history of utilizing interactions in regression analysis to investigate interactive effects of covariates on response variables. In this paper we aim to address two kinds of new challenges resulted from the inclusion of such high-order effects in the regression model for complex data. The first kind arises from a situation where interaction effects of individual covariates are weak but those of combined covariates are strong, and the other kind pertains to the presence of nonlinear interactive effects. Generalizing the single index coefficient regression model (Xia and Li, 1999), we propose a new class of semiparametric …
Balancing Score Adjusted Targeted Minimum Loss-Based Estimation,
2013
University of California, Berkeley, School of Public Health, Division of Biostatistics
Balancing Score Adjusted Targeted Minimum Loss-Based Estimation, Samuel D. Lendle, Bruce Fireman, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Adjusting for a balancing score is sufficient for bias reduction when estimating causal effects including the average treatment effect and effect among the treated. Estimators that adjust for the propensity score in a nonparametric way, such as matching on an estimate of the propensity score, can be consistent when the estimated propensity score is not consistent for the true propensity score but converges to some other balancing score. We call this property the balancing score property, and discuss a class of estimators that have this property. We introduce a targeted minimum loss-based estimator (TMLE) for a treatment specific mean with …
Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming,
2013
Johns Hopkins Bloomberg School of Public Health, Department of Biostatistics
Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming, Michael Rosenblum, Han Liu, En-Hsu Yen
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose new, optimal methods for analyzing randomized trials, when it is suspected that treatment effects may differ in two predefined subpopulations. Such sub-populations could be defined by a biomarker or risk factor measured at baseline. The goal is to simultaneously learn which subpopulations benefit from an experimental treatment, while providing strong control of the familywise Type I error rate. We formalize this as a multiple testing problem and show it is computationally infeasible to solve using existing techniques. Our solution involves a novel approach, in which we first transform the original multiple testing problem into a large, sparse linear …
Estimating Effects On Rare Outcomes: Knowledge Is Power,
2013
UC Berkeley, School of Public Health-Division of Biostatistics
Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many of the secondary outcomes in observational studies and randomized trials are rare. Methods for estimating causal effects and associations with rare outcomes, however, are limited, and this represents a missed opportunity for investigation. In this article, we construct a new targeted minimum loss-based estimator (TMLE) for the effect of an exposure or treatment on a rare outcome. We focus on the causal risk difference and statistical models incorporating bounds on the conditional risk of the outcome, given the exposure and covariates. By construction, the proposed estimator constrains the predicted outcomes to respect this model knowledge. Theoretically, this bounding provides …
Linking And Retaining Hiv Patients In Care: The Importance Of Provider Attitudes And Behaviors,
2013
George Washington University
Linking And Retaining Hiv Patients In Care: The Importance Of Provider Attitudes And Behaviors, Manya Magnus, Jane Herwehe, Michelli Murtaza-Rossini, Petera Reine, Damien Cuffie, Deann Gruber, Michael Kaiser
Epidemiology Faculty Publications
Retention in HIV treatment may reduce morbidity and mortality, as well as slow the epidemic. Myriad barriers to retention include stigma, homophobia, structural barriers, transportation, and insurance. The purpose of this study was to evaluate patient perceptions of provider attitudes among HIV-infected persons within a state-wide public hospital system in Louisiana. A convenience sample of patients attending HIV clinics throughout the state participated in an anonymous interview. Factors associated with negative perceptions of care were evaluated in conjunction with a validated stigma measure. Factors associated with having a delayed entry into or break in care were evaluated in conjunction with …
Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis,
2013
University of Pennsylvania
Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis, Sihai Dave Zhao, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron
Harvard University Biostatistics Working Paper Series
No abstract provided.
An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes,
2013
University of California - Berkeley, Division of Biostatistics
An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes, Ekaterina Eliseeva, Alan E. Hubbard, Ira B. Tager
U.C. Berkeley Division of Biostatistics Working Paper Series
Analyses of epidemiological studies of the association between short-term changes in air pollution and health outcomes have not sufficiently discussed the degree to which the statistical models chosen for these analyses reflect what is actually known about the true data-generating distribution. We present a method to estimate population-level ambient air pollution (NO2) exposure-health (wheeze in children with asthma) response functions that is not dependent on assumptions about the data-generating function that underlies the observed data and which focuses on a specific scientific parameter of interest (the marginal adjusted association of exposure on probability of wheeze, over a grid of possible …
Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates,
2013
Harvard University
Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates, Wei Dai, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Microarray technology has the potential to lead to a better understanding of biological processes and diseases such as cancer. When failure time outcomes are also available, one might be interested in relating gene expression profiles to the survival outcome such as time to cancer recurrence or time to death. This is statistically challenging because the number of covariates greatly exceeds the number of observations. While the majority of work has focused on regularized Cox regression model and accelerated failure time model, they may be restrictive in practice. We relax the model assumption and and consider a nonparametric transformation model that …
Compound Identification Using Penalized Linear Regression.,
2013
University of Louisville
Compound Identification Using Penalized Linear Regression., Ruiqi Liu
Electronic Theses and Dissertations
In this study, we propose a new method for compound identification using penalized linear regression. Compound identification is often achieved by matching the experimental mass spectra to the mass spectra stored in a reference library based on mass spectral similarity. In the context of the linear regression, the response variable is an experimental mass spectrum (i.e., query) and all the compounds in the reference library are the independent variables. However, the number of compounds in the reference library is much larger than the range of m/z values so that the data become high dimensional data with suffering from singularity. For …
Development Of Novel Methods To Minimize The Impact Of Sequencing Errors In The Next-Generation Sequencing Data Analysis,
2013
The University of Texas Graduate School of Biomedical Sciences at Houston
Development Of Novel Methods To Minimize The Impact Of Sequencing Errors In The Next-Generation Sequencing Data Analysis, Xiaofeng Zheng
Dissertations and Theses (Open Access)
Next-generation sequencing (NGS) technology has become a prominent tool in biological and biomedical research. However, NGS data analysis, such as de novo assembly, mapping and variants detection is far from maturity, and the high sequencing error-rate is one of the major problems. .
To minimize the impact of sequencing errors, we developed a highly robust and efficient method, MTM, to correct the errors in NGS reads. We demonstrated the effectiveness of MTM on both single-cell data with highly non-uniform coverage and normal data with uniformly high coverage, reflecting that MTM’s performance does not rely on the coverage of the sequencing …
Integrative Biomarker Identification And Classification Using High Throughput Assays,
2013
The University of Texas Graduate School of Biomedical Sciences at Houston
Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong
Dissertations and Theses (Open Access)
It is well accepted that tumorigenesis is a multi-step procedure involving aberrant functioning of genes regulating cell proliferation, differentiation, apoptosis, genome stability, angiogenesis and motility. To obtain a full understanding of tumorigenesis, it is necessary to collect information on all aspects of cell activity. Recent advances in high throughput technologies allow biologists to generate massive amounts of data, more than might have been imagined decades ago. These advances have made it possible to launch comprehensive projects such as (TCGA) and (ICGC) which systematically characterize the molecular fingerprints of cancer cells using gene expression, methylation, copy number, microRNA and SNP microarrays …
Structured Functional Principal Component Analysis,
2013
Johns Hopkins Bloomberg School of Public Health
Structured Functional Principal Component Analysis, Haochang Shou, Vadim Zipunnikov, Ciprian Crainiceanu, Sonja Greven
Johns Hopkins University, Dept. of Biostatistics Working Papers
Motivated by modern observational studies, we introduce a class of functional models that expands nested and crossed designs. These models account for the natural inheritance of correlation structure from sampling design in studies where the fundamental sampling unit is a function or image. Inference is based on functional quadratics and their relationship with the underlying covariance structure of the latent processes. A computationally fast and scalable estimation procedure is developed for ultra-high dimensional data. Methods are illustrated in three examples: high-frequency accelerometer data for daily activity, pitch linguistic data for phonetic analysis, and EEG data for studying electrical brain activity …
Determinan Komplikasi Kronik Diabetes Melitus Pada Lanjut Usia,
2013
Departemen Biostatistika dan Ilmu Kependudukan Fakultas Kesehatan Masyarakat Universitas Indonesia
Determinan Komplikasi Kronik Diabetes Melitus Pada Lanjut Usia, Amrina Rosyada, Indang Trihandini
Kesmas
Indonesia menghadapi jumlah penduduk lanjut usia (lansia) yang semakin meningkat dan diikuti oleh peningkatan frekuensi penyakit tidak menular kronis atau multimorbiditas. Penelitian ini bertujuan untuk mengetahui prevalensi dan faktor yang berhubungan komplikasi kronis pada lansia penderita diabetes melitus. Penelitian ini menggunakan data Riset Kesehatan Dasar (Riskesdas) Tahun 2007 dengan desain cross sectional representatif Indonesia dan metode cluster 2 tahap untuk pengambilan sampel. Sampel adalah 1.565 lansia penderita diabetes melitus. Metode analisis yang digunakan meliputi analisis deskriptif dan multivariat. Hasil analisis menunjukkan bahwa prevalensi komplikasi kronis pada lansia adalah sekitar 73,1%, dengan hipertensi sebagai komplikasi terbanyak. Berdasarkan analisis multivariat diketahui pula …
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method,
2013
Virginia Commonwealth University
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Theses and Dissertations
Survival Analysis generally uses the median survival time as a common summary statistic. While the median possesses the desirable characteristic of being unbiased, there are times when it is not the best statistic to describe the data at hand. Royston and Parmar (2011) provide an argument that the restricted mean survival time should be the summary statistic used when the proportional hazards assumption is in doubt. Work in Restricted Means dates back to 1949 when J.O. Irwin developed a calculation for the standard error of the restricted mean using Greenwood’s formula. Since then the development of the restricted mean has …
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments,
2013
Virginia Commonwealth University
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Theses and Dissertations
Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …
Penalized Function-On-Function Regression,
2013
Department of Biostatistics, East Carolina University, Greenville, NC
Penalized Function-On-Function Regression, Andrada E. Ivanescu, Ana-Maria Staicu, Fabian Scheipl, Sonja Greven
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose a general framework for smooth regression of a functional response on one or multiple functional predictors. Using the mixed model representation of penalized regression expands the scope of function on function regression to many realistic scenarios. In particular, the approach can accommodate a densely or sparsely sampled functional response as well as multiple functional predictors that are observed: 1) on the same or different domains than the functional response; 2) on a dense or sparse grid; and 3) with or without noise. It also allows for seamless integration of continuous or categorical covariates and provides approximate confidence intervals …
Daily Walking And Life Expectancy Of Elderly People In The Iowa 65+ Rural Health Study,
2013
Georgia Southern University
Daily Walking And Life Expectancy Of Elderly People In The Iowa 65+ Rural Health Study, Hani M. Samawi
Biostatistics: Faculty Publications
The purpose of this paper is to investigate the hypothesis that outdoor daily walking, as an exercise, has an effect on the rate of mortality among those elderly people in the Iowa 65+ Rural Health Study (RHS). RHS is a prospective longitudinal cohort study of 8 years follow-up from 1981 to 1989. It consists of a random sample of 3,673 individuals (1,420 men and 2,253 women) aged 65 or older living in Washington and Iowa counties of the State of Iowa. Our analysis was conducted only on those non-institutional individuals who could without any help walk across a small room; …
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios,
2013
Virginia Commonwealth University
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Theses and Dissertations
In risk evaluation, the effect of mixtures of environmental chemicals on a common adverse outcome is of interest. However, due to the high dimensionality and inherent correlations among chemicals that occur together, the traditional methods (e.g. ordinary or logistic regression) are unsuitable. We extend and characterize a weighted quantile score (WQS) approach to estimating an index for a set of highly correlated components. In the case with environmental chemicals, we use the WQS to identify “bad actors” and estimate body burden. The accuracy of the WQS was evaluated through extensive simulation studies in terms of validity (ability of the WQS …
Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape,
2013
University of California - Berkeley
Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape, Brianna C. Heggeseth, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
Longitudinal studies play a prominent role in health, social and behavioral sciences as well as in the biological sciences, economics, and marketing. By following subjects over time, temporal changes in an outcome of interest can be directly observed and studied. An important question concerns the existence of distinct trajectory patterns. One way to determine these distinct patterns is through cluster analysis, which seeks to separate objects (subjects, patients, observational units) into homogeneous groups. Many methods have been adapted for longitudinal data, but almost all of them fail to explicitly group trajectories according to distinct pattern shapes. To fulfill the need …
Efficient Estimation Of Risk Ratios From Clustered Binary Data,
2013
Harvard School of Public Health
Efficient Estimation Of Risk Ratios From Clustered Binary Data, Matthew Cefalu, Eric Tchetgen Tchetgen
Harvard University Biostatistics Working Paper Series
No abstract provided.
