Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Life Sciences (62)
- Genetics and Genomics (56)
- Bioinformatics (44)
- Statistical Methodology (44)
- Computational Biology (39)
-
- Statistical Models (37)
- Statistical Theory (34)
- Biostatistics (31)
- Multivariate Analysis (31)
- Genetics (28)
- Medicine and Health Sciences (15)
- Applied Statistics (11)
- Genomics (9)
- Survival Analysis (9)
- Applied Mathematics (7)
- Mathematics (7)
- Numerical Analysis and Computation (7)
- Clinical Trials (6)
- Public Health (6)
- Design of Experiments and Sample Surveys (5)
- Diseases (5)
- Longitudinal Data Analysis and Time Series (5)
- Biometry (4)
- Computer Sciences (4)
- Epidemiology (4)
- Biochemistry, Biophysics, and Structural Biology (3)
- Cancer Biology (3)
- Institution
- Keyword
-
- Microarray (16)
- Gene expression (14)
- Genetics (10)
- Classification (5)
- Differential expression (5)
-
- Genomics (4)
- Microarrays (4)
- Bioinformatics (3)
- Cross-validation (3)
- False discovery rate (3)
- Mixture models (3)
- Model selection (3)
- Multiple comparisons (3)
- Multiple testing (3)
- Prediction (3)
- Survival analysis (3)
- Block design (2)
- Bootstrap (2)
- Clustering (2)
- Comparative genomic hybridization (2)
- Density estimation (2)
- Empirical Bayes (2)
- Feature selection (2)
- High-dimensional data (2)
- Normalization (2)
- Variable selection (2)
- (Quasi)separation (1)
- AdaBoost (1)
- Adaptive Dantzig variable selector; Censored linear regression; Buckley-James imputation; Model selection consistency; Asymptotic normality (1)
- Adjust p value (1)
- Publication Year
- Publication
-
- COBRA Preprint Series (15)
- UW Biostatistics Working Paper Series (15)
- Harvard University Biostatistics Working Paper Series (13)
- U.C. Berkeley Division of Biostatistics Working Paper Series (12)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (10)
-
- The University of Michigan Department of Biostatistics Working Paper Series (6)
- Theses and Dissertations (6)
- Bioconductor Project Working Papers (4)
- Electronic Theses and Dissertations (4)
- Pomona Faculty Publications and Research (4)
- Theses and Dissertations--Statistics (4)
- Dissertations and Theses (Open Access) (3)
- Graduate Theses and Dissertations (2)
- Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series (2)
- Yale Day of Data (2)
- Annual Symposium on Biomathematics and Ecology Education and Research (1)
- Appalachian Student Research Forum (1)
- Biostatistics Faculty Publications (1)
- Chemical Engineering Faculty Publications and Presentations (1)
- Graduate Research Posters (1)
- Mathematics & Statistics ETDs (1)
- Mathematics and Statistics Faculty Publications (1)
- SMU Data Science Review (1)
- Statistical Science Theses and Dissertations (1)
- UPenn Biostatistics Working Papers (1)
- Williams Honors College, Honors Research Projects (1)
- Publication Type
Articles 31 - 60 of 113
Full-Text Articles in Microarrays
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Theses and Dissertations
In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …
Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan
Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan
Theses and Dissertations--Statistics
This paper intend to find the proper hypothesis and test statistic for testing existence of bilaterally contamination when there exists nuisance parameter. The test statistic is based on method of moments estimators. Union-Intersection test is used for testing if the distribution of population can be implemented by a bilaterally contaminated normal model with unknown variance. This paper also developed a hierarchical normal mixture model (HNM) and applied it to birth weight data. EM algorithm is employed for parameter estimation and a singular Bayesian information criterion (sBIC) is applied to choose the number components. We also proposed a singular flexible information …
Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong
Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong
Dissertations and Theses (Open Access)
It is well accepted that tumorigenesis is a multi-step procedure involving aberrant functioning of genes regulating cell proliferation, differentiation, apoptosis, genome stability, angiogenesis and motility. To obtain a full understanding of tumorigenesis, it is necessary to collect information on all aspects of cell activity. Recent advances in high throughput technologies allow biologists to generate massive amounts of data, more than might have been imagined decades ago. These advances have made it possible to launch comprehensive projects such as (TCGA) and (ICGC) which systematically characterize the molecular fingerprints of cancer cells using gene expression, methylation, copy number, microRNA and SNP microarrays …
Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller
Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller
COBRA Preprint Series
We propose a methodological framework to assess heterogeneous patterns of association amongst components of a random vector expressed as a Gaussian directed acyclic graph. The proposed framework is likely to be useful when primary interest focuses on potential contrasts characterizing the association structure between known subgroups of a given sample. We provide inferential frameworks as well as an efficient computational algorithm to fit such a model and illustrate its validity through a simulation. We apply the model to Reverse Phase Protein Array data on Acute Myeloid Leukemia patients to show the contrast of association structure between refractory patients and relapsed …
A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg
A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg
COBRA Preprint Series
Identifying differentially expressed (DE) genes associated with a sample characteristic is the primary objective of many microarray studies. As more and more studies are carried out with observational rather than well controlled experimental samples, it becomes important to evaluate and properly control the impact of sample heterogeneity on DE gene finding. Typical methods for identifying DE genes require ranking all the genes according to a pre-selected statistic based on a single model for two or more group comparisons, with or without adjustment for other covariates. Such single model approaches unavoidably result in model misspecification, which can lead to increased error …
Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel
Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel
COBRA Preprint Series
In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …
Principled Sure Independence Screening For Cox Models With Ultra-High-Dimensional Covariates, Sihai Dave Zhao, Yi Li
Principled Sure Independence Screening For Cox Models With Ultra-High-Dimensional Covariates, Sihai Dave Zhao, Yi Li
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation, David R. Bickel
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation, David R. Bickel
COBRA Preprint Series
A general function to quantify the weight of evidence in a sample of data for one hypothesis over another is derived from the law of likelihood and from a statistical formalization of inference to the best explanation. For a fixed parameter of interest, the resulting weight of evidence that favors one composite hypothesis over another is the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function over the parameter of interest. Since the weight of evidence is generally only known up to a nuisance parameter, it is approximated by replacing the likelihood function with …
Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan
Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Super learning is a general loss based learning method that has been proposed and analyzed theoretically in van der Laan et al. (2007). In this article we consider super learning for prediction. The super learner is a prediction method designed to find the optimal combination of a collection of prediction algorithms. The super learner algorithm finds the combination of algorithms minimizing the cross-validated risk. The super learner framework is built on the theory of cross-validation and allows for a general class of prediction algorithms to be considered for the ensemble. Due to the previously established oracle results for the cross-validation …
Survival Prediction For Brain Tumor Patients Using Gene Expression Data, Vinicius Bonato
Survival Prediction For Brain Tumor Patients Using Gene Expression Data, Vinicius Bonato
Dissertations and Theses (Open Access)
Brain tumor is one of the most aggressive types of cancer in humans, with an estimated median survival time of 12 months and only 4% of the patients surviving more than 5 years after disease diagnosis. Until recently, brain tumor prognosis has been based only on clinical information such as tumor grade and patient age, but there are reports indicating that molecular profiling of gliomas can reveal subgroups of patients with distinct survival rates. We hypothesize that coupling molecular profiling of brain tumors with clinical information might improve predictions of patient survival time and, consequently, better guide future treatment decisions. …
A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao
A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao
Harvard University Biostatistics Working Paper Series
No abstract provided.
Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel
Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel
COBRA Preprint Series
Research on analyzing microarray data has focused on the problem of identifying differentially expressed genes to the neglect of the problem of how to integrate evidence that a gene is differentially expressed with information on the extent of its differential expression. Consequently, researchers currently prioritize genes for further study either on the basis of volcano plots or, more commonly, according to simple estimates of the fold change after filtering the genes with an arbitrary statistical significance threshold. While the subjective and informal nature of the former practice precludes quantification of its reliability, the latter practice is equivalent to using a …
The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin
The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …
A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …
Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit
Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
The multtest package is a standard Bioconductor package containing a suite of functions useful for executing, summarizing, and displaying the results from a wide variety of multiple testing procedures (MTPs). In addition to many popular MTPs, the central methodological focus of the multtest package is the implementation of powerful joint multiple testing procedures. Joint MTPs are able to account for the dependencies between test statistics by effectively making use of (estimates of) the test statistics joint null distribution. To this end, two additional bootstrap-based estimates of the test statistics joint null distribution have been developed for use in the …
Validation Of Differential Gene Expression Algorithms: Application Comparing Fold Change Estimation To Hypothesis Testing, David R. Bickel, Corey M. Yanofsky
Validation Of Differential Gene Expression Algorithms: Application Comparing Fold Change Estimation To Hypothesis Testing, David R. Bickel, Corey M. Yanofsky
COBRA Preprint Series
Sustained research on the problem of determining which genes are differentially expressed on the basis of microarray data has yielded a plethora of statistical algorithms, each justified by theory, simulation, or ad hoc validation and yet differing in practical results from equally justified algorithms. The widespread confusion on which method to use in practice has been exacerbated by the finding that simply ranking genes by their fold changes sometimes outperforms popular statistical tests.
Algorithms may be compared by quantifying each method's error in predicting expression ratios, whether such ratios are defined across microarray channels or between two independent groups. For …
Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry
Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Genome-wide association studies (GWAS) are used to discover genes underlying complex, heritable disorders for which less powerful study designs have failed in the past. The number of GWAS has skyrocketed recently with findings reported in top journals and the mainstream media. Mircorarrays are the genotype calling technology of choice in GWAS as they permit exploration of more than a million single nucleotide polymorphisms (SNPs)simultaneously. The starting point for the statistical analyses used by GWAS, to determine association between loci and disease, are genotype calls (AA, AB, or BB). However, the raw data, microarray probe intensities, are heavily processed before arriving …
Sparse Linear Discriminant Analysis For Simultaneous Testing For The Significance Of A Gene Set/Pathway And Gene Selection, Michael C. Wu, Lingson Zhang, Zhaoxi Wang, David C. Christiani, Xihong Lin
Sparse Linear Discriminant Analysis For Simultaneous Testing For The Significance Of A Gene Set/Pathway And Gene Selection, Michael C. Wu, Lingson Zhang, Zhaoxi Wang, David C. Christiani, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel
The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel
COBRA Preprint Series
The strength of the statistical evidence in a sample of data that favors one composite hypothesis over another may be quantified by the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function. Unlike the p-value and the Bayes factor, this measure of evidence is coherent in the sense that it cannot support a hypothesis over any hypothesis that it entails. Further, when comparing the hypothesis that the parameter lies outside a non-trivial interval to the hypotheses that it lies within the interval, the proposed measure of evidence almost always asymptotically favors the correct hypothesis …
Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin
Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein
A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein
Harvard University Biostatistics Working Paper Series
No abstract provided.
Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey
Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey
Harvard University Biostatistics Working Paper Series
No abstract provided.
Targeted Methods For Biomarker Discovery, The Search For A Standard, Catherine Tuglus, Mark J. Van Der Laan
Targeted Methods For Biomarker Discovery, The Search For A Standard, Catherine Tuglus, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
More often than not biomarker studies analyze large quantities of variables with complicated and generally unknown correlation structure. There are numerous statistical methods which attempt to unravel these variables and determine the underlying mechanism through identification of causally related biomarkers. Results from these methods are generally difficult to interpret and nearly impossible to compare across studies. The FDA has currently called for a standardization of methods and protocol for biomarker detection. In response, we propose targeted variable importance (tVIM) as a standardized method for biomarker discovery. Through the use of targeted Maximum Likelihood, tVIM provides double robust estimates of variable …
Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky
Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li
Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li
Harvard University Biostatistics Working Paper Series
Use of microarray technology often leads to high-dimensional and low- sample size data settings. Over the past several years, a variety of novel approaches have been proposed for variable selection in this context. However, only a small number of these have been adapted for time-to-event data where censoring is present. Among standard variable selection methods shown both to have good predictive accuracy and to be computationally efficient is the elastic net penalization approach. In this paper, adaptation of the elastic net approach is presented for variable selection both under the Cox proportional hazards model and under an accelerated failure time …
What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner
What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner
UW Biostatistics Working Paper Series
The reference design is a practical and popular choice for microarray studies using two-color platforms. In the reference design, the reference RNA uses half of all array resources, leading investigators to ask: What is the best reference RNA? We propose a novel method for evaluating reference RNAs and present the results of an experiment that was specially designed to evaluate three common choices of reference RNA. We found no compelling evidence in favor of any particular reference. In particular, a commercial reference showed no advantage in our data. Our experimental design also enabled a new way to test the effectiveness …
A Bayesian Hierarchical Model For Spot Fluorescence In Microarrays, Federico Mattia Stefanini
A Bayesian Hierarchical Model For Spot Fluorescence In Microarrays, Federico Mattia Stefanini
COBRA Preprint Series
Microarray experiments are characterized by the presence of many sources of experimental bias and a remarkably large technical variability. The assessment of differential expression for genes transcribed into a small number of mRNA copies heavily depends on the proper quantification of background fluorescence within spot. The rough model `observed = hybridization plus background' fluorescence is at first reformulated at spot level, then it is embedded into a Bayesian hierarchical model suited for fitting control spots. The novelties of the approach include the background correction performed on the latent mean of replicated spots, and an explicit model for outlying observations at …
On Comparing The Clustering Of Regression Models Method With K-Means Clustering, Li-Xuan Qin, Steven G. Self
On Comparing The Clustering Of Regression Models Method With K-Means Clustering, Li-Xuan Qin, Steven G. Self
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
Gene clustering is a common question addressed with microarray data. Previous methods, such as K-means clustering and hierarchical clustering, base gene clustering directly on the observed measurements. A new model-based clustering method, the clustering of regression models (CORM) method, bases the clustering of genes on their relationship to covariates. It explicitly models different sources of variations and bases gene clustering solely on the systematic variation. Both being partitional clustering, CORM is closely related to K-means clustering. In this paper, we discuss the relationship between the two clustering methods in terms of both model formulation and implications on other important aspects …
Conservative Estimation Of Optimal Multiple Testing Procedures, James E. Signorovitch
Conservative Estimation Of Optimal Multiple Testing Procedures, James E. Signorovitch
Harvard University Biostatistics Working Paper Series
No abstract provided.