Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Department of Statistics: Faculty Publications

Discipline
Keyword
Publication Year

Articles 61 - 90 of 162

Full-Text Articles in Statistics and Probability

Existing And Potential Statistical And Computational Approaches For The Analysis Of 3d Ct Images Of Plant Roots, Zheng Xu, Camilo Valdes, Jennifer Clarke Jan 2018

Existing And Potential Statistical And Computational Approaches For The Analysis Of 3d Ct Images Of Plant Roots, Zheng Xu, Camilo Valdes, Jennifer Clarke

Department of Statistics: Faculty Publications

Scanning technologies based on X-ray Computed Tomography (CT) have been widely used in many scientific fields including medicine, nanosciences and materials research. Considerable progress in recent years has been made in agronomic and plant science research thanks to X-ray CT technology. X-ray CT image-based phenotyping methods enable high-throughput and non-destructive measuring and inference of root systems, which makes downstream studies of complex mechanisms of plants during growth feasible. An impressive amount of plant CT scanning data has been collected, but how to analyze these data efficiently and accurately remains a challenge. We review statistical and computational approaches that have been …


Characterization Of Soybean Protein Adhesives Modified By Xanthan Gum, Chen Feng, Fang Wang, Zheng Xu, Huilin Sui, Yong Fang, Xiaozhi Tang, Xinchun Shen Jan 2018

Characterization Of Soybean Protein Adhesives Modified By Xanthan Gum, Chen Feng, Fang Wang, Zheng Xu, Huilin Sui, Yong Fang, Xiaozhi Tang, Xinchun Shen

Department of Statistics: Faculty Publications

The aim of this study was to provide a basis for the preparation of medical adhesives from soybean protein sources. Soybean protein (SP) adhesives mixed with different concentrations of xanthan gum (XG) were prepared. Their adhesive features were evaluated by physicochemical parameters and an in vitro bone adhesion assay. The results showed that the maximal adhesion strength was achieved in 5% SP adhesive with 0.5% XG addition, which was 2.6-fold higher than the SP alone. The addition of XG significantly increased the hydrogen bond and viscosity, as well as increased the β-sheet content but decreased the α-helix content in the …


Development Of 11-Plex Mol-Pcr Assay For The Rapid Screening Of Samples For Shiga Toxin-Producing Escherichia Coli, Travis A. Woods, Heather M. Mendez, Sandy Ortega, Xiaorong Shi, David Marx, Jianfa Bai, Rodney A. Moxley, T. G. Nagaraja, Steven W. Graves, Alina Deshpande Jan 2018

Development Of 11-Plex Mol-Pcr Assay For The Rapid Screening Of Samples For Shiga Toxin-Producing Escherichia Coli, Travis A. Woods, Heather M. Mendez, Sandy Ortega, Xiaorong Shi, David Marx, Jianfa Bai, Rodney A. Moxley, T. G. Nagaraja, Steven W. Graves, Alina Deshpande

Department of Statistics: Faculty Publications

Strains of Shiga toxin-producing Escherichia coli (STEC) are a serious threat to the health, with approximately half of the STEC related food-borne illnesses attributable to contaminated beef. We developed an assay that was able to screen samples for several important STEC associated serogroups (O26, O45, O103, O104, O111, O121, O145, O157) and three major virulence factors (eae, stx1, stx2) in a rapid and multiplexed format using the Multiplex oligonucleotide ligation-PCR (MOL-PCR) assay chemistry. This assay detected unique STEC DNA signatures and is meant to be used on samples from various sources related to beef production, providing a multiplex and high-throughput …


Application Of Transfer Learning For Cancer Drug Sensitivity Prediction, Saugato Rahman Dhruba, Raziur Rahman, Kevin Matlock, Souparno Ghosh, Ranadip Pal Jan 2018

Application Of Transfer Learning For Cancer Drug Sensitivity Prediction, Saugato Rahman Dhruba, Raziur Rahman, Kevin Matlock, Souparno Ghosh, Ranadip Pal

Department of Statistics: Faculty Publications

Background: In precision medicine, scarcity of suitable biological data often hinders the design of an appropriate predictive model. In this regard, large scale pharmacogenomics studies, like CCLE and GDSC hold the promise to mitigate the issue. However, one cannot directly employ data from multiple sources together due to the existing distribution shift in data. One way to solve this problem is to utilize the transfer learning methodologies tailored to fit in this specific context.

Results: In this paper, we present two novel approaches for incorporating information from a secondary database for improving the prediction in a target database. The first …


Investigation Of Model Stacking For Drug Sensitivity Prediction, Kevin Matlock, Carlos De Niz, Raziur Rahman, Souparno Ghosh, Ranadip Pal Jan 2018

Investigation Of Model Stacking For Drug Sensitivity Prediction, Kevin Matlock, Carlos De Niz, Raziur Rahman, Souparno Ghosh, Ranadip Pal

Department of Statistics: Faculty Publications

Background: A significant problem in precision medicine is the prediction of drug sensitivity for individual cancer cell lines. Predictive models such as Random Forests have shown promising performance while predicting from individual genomic features such as gene expressions. However, accessibility of various other forms of data types including information on multiple tested drugs necessitates the examination of designing predictive models incorporating the various data types.

Results: We explore the predictive performance of model stacking and the effect of stacking on the predictive bias and squarred error. In addition we discuss the analytical underpinnings supporting the advantages of stacking in reducing …


Heterogeneity Aware Random Forest For Drug Sensitivity Prediction, Raziur Rahman, Kevin Matlock, Souparno Ghosh, Ranadip Pal Sep 2017

Heterogeneity Aware Random Forest For Drug Sensitivity Prediction, Raziur Rahman, Kevin Matlock, Souparno Ghosh, Ranadip Pal

Department of Statistics: Faculty Publications

Samples collected in pharmacogenomics databases typically belong to various cancer types. For designing a drug sensitivity predictive model from such a database, a natural question arises whether a model trained on diverse inter-tumor heterogeneous samples will perform similar to a predictive model that takes into consideration the heterogeneity of the samples in model training and prediction. We explore this hypothesis and observe that ensemble model predictions obtained when cancer type is known out-perform predictions when that information is withheld even when the samples sizes for the former is considerably lower than the combined sample size. To incorporate the heterogeneity idea …


Assessing The Impact Of Retreat Mechanisms In A Simple Antarctic Ice Sheet Model Using Bayesian Calibration, Kelsey L. Ruckert, Gary Shaffer, David Pollard, Yawen Guan, Tony E. Wong, Chris E. Forest, Klaus Keller Jan 2017

Assessing The Impact Of Retreat Mechanisms In A Simple Antarctic Ice Sheet Model Using Bayesian Calibration, Kelsey L. Ruckert, Gary Shaffer, David Pollard, Yawen Guan, Tony E. Wong, Chris E. Forest, Klaus Keller

Department of Statistics: Faculty Publications

The response of the Antarctic ice sheet (AIS) to changing climate forcings is an important driver of sea-level changes. Anthropogenic climate change may drive a sizeable AIS tipping point response with subsequent increases in coastal flooding risks. Many studies analyzing flood risks use simple models to project the future responses of AIS and its sea-level contributions. These analyses have provided important new insights, but they are often silent on the effects of potentially important processes such as Marine Ice Sheet Instability (MISI) or Marine Ice Cliff Instability (MICI). These approximations can be well justified and result in more parsimonious and …


Perennial Warm-Season Grasses For Producing Biofuel And Enhancing Soil Properties: An Alternative To Corn Residue Removal, Humberto Blanco-Canqui, Robert B. Mitchell, Virginia L. Jin, Marty R. Schmer, Kent M. Eskridge Jan 2017

Perennial Warm-Season Grasses For Producing Biofuel And Enhancing Soil Properties: An Alternative To Corn Residue Removal, Humberto Blanco-Canqui, Robert B. Mitchell, Virginia L. Jin, Marty R. Schmer, Kent M. Eskridge

Department of Statistics: Faculty Publications

Removal of corn (Zea mays L.) residues at high rates for biofuel and other off-farm uses may negatively impact soil and the environment in the long term. Biomass removal from perennial warm-season grasses (WSGs) grown in marginally-productive lands could be an alternative to corn residue removal as biofuel feedstocks while controlling water and wind erosion, sequestering carbon (C), cycling water and nutrients, and enhancing other soil ecosystem services. We compared wind and water erosion potential, soil compaction, soil hydraulic properties, soil organic C (SOC), and soil fertility between biomass removal from WSGs and corn residue removal from rainfed no-till …


Impact Of Menthol Smoking On Nicotine Dependence For Diverse Racial/Ethnic Groups Of Daily Smokers, Julia N. Soulakova, Ryan R. Danczak Jan 2017

Impact Of Menthol Smoking On Nicotine Dependence For Diverse Racial/Ethnic Groups Of Daily Smokers, Julia N. Soulakova, Ryan R. Danczak

Department of Statistics: Faculty Publications

Introduction: The aims of this study were to evaluate whether menthol smoking and race/ethnicity are associated with nicotine dependence in daily smokers. Methods: The study used two subsamples of U.S. daily smokers who responded to the 2010–2011 Tobacco Use Supplement to the Current Population Survey. The larger subsample consisted of 18,849 non-Hispanic White (NHW), non-Hispanic Black (NHB), and Hispanic (HISP) smokers. The smaller subsample consisted of 1112 non-Hispanic American Indian/Alaska Native (AIAN), non-Hispanic Asian (ASIAN), non-Hispanic Hawaiian/Pacific Islander (HPI), and non-Hispanic Multiracial (MULT) smokers. Results: For larger (smaller) groups the rates were 45% (33%) for heavy smoking (16+ cig/day), 59% …


Evaluating Current Practices In Shelf Life Estimation, Robert Capen, David Christopher, Patrick Forenzo, Kim Huynh-Ba, David Leblond, Oscar Liu, John O'Neill, Nate Patterson, Michelle Quinlan, Radhika Rajagopalan, James Schwenke, Walter W. Stroup Jan 2017

Evaluating Current Practices In Shelf Life Estimation, Robert Capen, David Christopher, Patrick Forenzo, Kim Huynh-Ba, David Leblond, Oscar Liu, John O'Neill, Nate Patterson, Michelle Quinlan, Radhika Rajagopalan, James Schwenke, Walter W. Stroup

Department of Statistics: Faculty Publications

The current International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) methods for determining the supported shelf life of a drug product, described in ICH guidance documents Q1A and Q1E, are evaluated in this paper. To support this evaluation, an industry data set is used which is comprised of 26 individual stability batches of a common drug product where most batches are measured over a 24 month storage period. Using randomly sampled sets of 3 or 6 batches from the industry data set, the current ICH methods are assessed from three perspectives. First, the distributional properties …


Generalized Confidence Intervals Compatible With The Min Test For Simultaneous Comparisons Of One Subpopulation To Several Other Subpopulations, Julia N. Soulakova Jan 2017

Generalized Confidence Intervals Compatible With The Min Test For Simultaneous Comparisons Of One Subpopulation To Several Other Subpopulations, Julia N. Soulakova

Department of Statistics: Faculty Publications

A problem where one subpopulation is compared to several other subpopulations in terms of means with the goal of estimating the smallest difference between the means commonly arises in biology, medicine, and many other scientific fields. A generalization of Strassburger, Bretz and Hochberg (2004) approach for two comparisons is presented for cases with three and more comparisons. The method allows constructing an interval-estimator for the smallest mean difference, which is compatible with the Min test. An application to a fluency-disorder study is illustrated. Simulations confirmed adequate probability coverage for normally distributed outcomes for a number of designs.


Increasing Genomic-Enabled Prediction Accuracy By Modeling Genotype X Environment Interactions In Kansas Wheat, Diego Jarquin, Cristiano Lemas Da Silva, R. Chris Gaynor, Jesse Poland, Allan Fritz, Reka Howard, Sarah Battenfield, José Crossa Jan 2017

Increasing Genomic-Enabled Prediction Accuracy By Modeling Genotype X Environment Interactions In Kansas Wheat, Diego Jarquin, Cristiano Lemas Da Silva, R. Chris Gaynor, Jesse Poland, Allan Fritz, Reka Howard, Sarah Battenfield, José Crossa

Department of Statistics: Faculty Publications

Wheat (Triticum aestivum L.) breeding programs test experimental lines in multiple locations over multiple years to get an accurate assessment of grain yield and yield stability. Selections in early generations of the breeding pipeline are based on information from only one or few locations and thus materials are advanced with little knowledge of the genotype × environment interaction (G × E) effects. Later, large trials are conducted in several locations to assess the performance of more advanced lines across environments. Genomic selection (GS) models that include G × E covariates allow us to borrow information not only from related …


Application Of Response Surface Methods To Determine Conditions For Optimal Genomic Prediction, Reka Howard, Alicia L. Carriquiry, William D. Beavis Jan 2017

Application Of Response Surface Methods To Determine Conditions For Optimal Genomic Prediction, Reka Howard, Alicia L. Carriquiry, William D. Beavis

Department of Statistics: Faculty Publications

An epistatic genetic architecture can have a significant impact on prediction accuracies of genomic prediction (GP) methods. Machine learning methods predict traits comprised of epistatic genetic architectures more accurately than statistical methods based on additive mixed linear models. The differences between these types of GP methods suggest a diagnostic for revealing genetic architectures underlying traits of interest. In addition to genetic architecture, the performance of GP methods may be influenced by the sample size of the training population, the number of QTL, and the proportion of phenotypic variability due to genotypic variability (heritability). Possible values for these factors and the …


Trans-Ancestry Fine Mapping And Molecular Assays Identify Regulatory Variants At The Angptl8 Hdl-C Gwas Locus, Maren E. Cannon, Qing Duan, Ying Wu, Monica Zeynalzadeh, Zheng Xu, Antti J. Kangas, Pasi Soininen, Mika Ala-Korpela, Mete Civelek, Aldons J. Lusis, Johanna Kuusisto, Francis S. Collins, Michael Boehnke, Hua Tang, Markku Laakso, Yun Li, Karen L. Mohlke Jan 2017

Trans-Ancestry Fine Mapping And Molecular Assays Identify Regulatory Variants At The Angptl8 Hdl-C Gwas Locus, Maren E. Cannon, Qing Duan, Ying Wu, Monica Zeynalzadeh, Zheng Xu, Antti J. Kangas, Pasi Soininen, Mika Ala-Korpela, Mete Civelek, Aldons J. Lusis, Johanna Kuusisto, Francis S. Collins, Michael Boehnke, Hua Tang, Markku Laakso, Yun Li, Karen L. Mohlke

Department of Statistics: Faculty Publications

Recent genome-wide association studies (GWAS) have identified variants associated with highdensity lipoprotein cholesterol (HDL-C) located in or near the ANGPTL8 gene. Given the extensive sharing of GWAS loci across populations, we hypothesized that at least one shared variant at this locus affects HDL-C. The HDL-C–associated variants are coincident with expression quantitative trait loci for ANGPTL8 and DOCK6 in subcutaneous adipose tissue; however, only ANGPTL8 expression levels are associated with HDL-C levels. We identified a 400-bp promoter region of ANGPTL8 and enhancer regions within 5 kb that contribute to regulating expression in liver and adipose. To identify variants functionally responsible for …


A Bayes Interpretation Of Stacking For M-Complete And M-Open Settings, Tri Le, Bertrand S. Clarke Jan 2017

A Bayes Interpretation Of Stacking For M-Complete And M-Open Settings, Tri Le, Bertrand S. Clarke

Department of Statistics: Faculty Publications

In M-open problems where no true model can be conceptualized, it is common to back off from modeling and merely seek good prediction. Even in M-complete problems, taking a predictive approach can be very useful. Stacking is a model averaging procedure that gives a composite predictor by combining individual predictors from a list of models using weights that optimize a cross validation criterion. We show that the stacking weights also asymptotically minimize a posterior expected loss. Hence we formally provide a Bayesian justification for cross-validation. Often the weights are constrained to be positive and sum to one. For greater generality, …


Optimal Design Of Low-Density Snp Arrays For Genomic Prediction: Algorithm And Applications, Xiao-Lin Wu, Jiaqi Xu, Guofei Feng, George R. Wiggans, Jeremy F. Taylor, Jun He, Changsong Qian, Jiansheng Qiu, Barry Simpson, Jeremy Walker, Stewart Bauck Sep 2016

Optimal Design Of Low-Density Snp Arrays For Genomic Prediction: Algorithm And Applications, Xiao-Lin Wu, Jiaqi Xu, Guofei Feng, George R. Wiggans, Jeremy F. Taylor, Jun He, Changsong Qian, Jiansheng Qiu, Barry Simpson, Jeremy Walker, Stewart Bauck

Department of Statistics: Faculty Publications

Low-density (LD) single nucleotide polymorphism (SNP) arrays provide a cost-effective solution for genomic prediction and selection, but algorithms and computational tools are needed for the optimal design of LD SNP chips. A multiple-objective, local optimization (MOLO) algorithm was developed for design of optimal LD SNP chips that can be imputed accurately to medium-density (MD) or high-density (HD) SNP genotypes for genomic prediction. The objective function facilitates maximization of non-gap map length and system information for the SNP chip, and the latter is computed either as locus-averaged (LASE) or haplotype-averaged Shannon entropy (HASE) and adjusted for uniformity of the SNP distribution. …


How Often Are Antibiotic-Resistant Bacteria Said To “Evolve” In The News?, Nina Singh, Matthew T. Sit, Deanna M. Chung, Ana A. Lopez, Ranil Weerackoon, Pamela J. Yeh Jan 2016

How Often Are Antibiotic-Resistant Bacteria Said To “Evolve” In The News?, Nina Singh, Matthew T. Sit, Deanna M. Chung, Ana A. Lopez, Ranil Weerackoon, Pamela J. Yeh

Department of Statistics: Faculty Publications

Media plays an important role in informing the general public about scientific ideas.We examine whether the word “evolve,” sometimes considered controversial by the general public, is frequently used in the popular press. Specifically, we ask how often articles discussing antibiotic resistance use the word “evolve” (or its lexemes) as opposed to alternative terms such as “emerge” or “develop.” We chose the topic of antibiotic resistance because it is a medically important issue; bacterial evolution is a central player in human morbidity and mortality. We focused on the most widely-distributed newspapers written in English in the United States, United Kingdom, Canada, …


Systematic Evaluation Of The Impact Of Chip-Seq Read Designs On Genome Coverage, Peak Identification, And Allele-Specific Binding Detection, Qi Zhang, Xin Zeng, Sam Younkin, Trupti Kawli, Michael P. Snyder, Sündüz Kele Jan 2016

Systematic Evaluation Of The Impact Of Chip-Seq Read Designs On Genome Coverage, Peak Identification, And Allele-Specific Binding Detection, Qi Zhang, Xin Zeng, Sam Younkin, Trupti Kawli, Michael P. Snyder, Sündüz Kele

Department of Statistics: Faculty Publications

Background: Chromatin immunoprecipitation followed by sequencing (ChIP-seq) experiments revolutionized genome-wide profiling of transcription factors and histone modifications. Although maturing sequencing technologies allow these experiments to be carried out with short (36–50 bps), long (75–100 bps), single-end, or paired-end reads, the impact of these read parameters on the downstream data analysis are not well understood. In this paper, we evaluate the effects of different read parameters on genome sequence alignment, coverage of different classes of genomic features, peak identification, and allele-specific binding detection.

Results: We generated 101 bps paired-end ChIP-seq data for many transcription factors from human GM12878 and MCF7 cell …


The Impact Of Hair Coat Color On Longevity Of Holstein Cows In The Tropics, C. N. Lee, K. S. Baek, A. Parkhurst Jan 2016

The Impact Of Hair Coat Color On Longevity Of Holstein Cows In The Tropics, C. N. Lee, K. S. Baek, A. Parkhurst

Department of Statistics: Faculty Publications

Background: Over two decades of observations in the field in South East Asia and Hawai‘i suggest that majority of the commercial dairy herds are of black hair coat. Hence a simple study to determine the accuracy of the observation was conducted with two large dairy herds in Hawaii in the mid-1990s.

Methods: A retrospective study on longevity of Holstein cattle in the tropics was conducted using DairyComp-305 lactation information coupled with phenotypic evaluation of hair coat color in two large dairy farms. Cows were classified into 3 groups: a) black (B, >90%); b) black/white (BW, 50:50) and c) white (W, …


Sex-Specific Hippocampal 5-Hydroxymethylcytosine Is Disrupted In Response To Acute Stress, Ligia A. Papale, Sisi Li, Andy Madrid, Qi Zhang, Li Chen, Pankaj Chopra, Peng Jin, Sunduz Keles, Reid S. Alisch Jan 2016

Sex-Specific Hippocampal 5-Hydroxymethylcytosine Is Disrupted In Response To Acute Stress, Ligia A. Papale, Sisi Li, Andy Madrid, Qi Zhang, Li Chen, Pankaj Chopra, Peng Jin, Sunduz Keles, Reid S. Alisch

Department of Statistics: Faculty Publications

Environmental stress is among the most important contributors to increased susceptibility to develop psychiatric disorders. While it is well known that acute environmental stress alters gene expression, the molecular mechanisms underlying these changes remain largely unknown. 5-hydroxymethylcytosine (5hmC) is a novel environmentally sensitive epigenetic modification that is highly enriched in neurons and is associated with active neuronal transcription. Recently,we reported a genome-wide disruption of hippocampal 5hmCin male mice following acute stress that was correlated to altered transcript levels of genes in known stress related pathways. Since sex-specific endocrine mechanisms respond to environmental stimulus by altering the neuronal epigenome, we examined …


A Compendium Of Chromatin Contact Maps Reveals Spatially Active Regions In The Human Genome, Anthony D. Schmitt, Ming Hu, Inkyung Jung, Zheng Xu, Yunjiang Qiu, Catherine L. Tan, Yun Li, Shin Lin, Yiing Lin, Cathy L. Barr, Bing Ren Jan 2016

A Compendium Of Chromatin Contact Maps Reveals Spatially Active Regions In The Human Genome, Anthony D. Schmitt, Ming Hu, Inkyung Jung, Zheng Xu, Yunjiang Qiu, Catherine L. Tan, Yun Li, Shin Lin, Yiing Lin, Cathy L. Barr, Bing Ren

Department of Statistics: Faculty Publications

The three-dimensional configuration of DNA is integral to all nuclear processes in eukaryotes, yet our knowledge of the chromosome architecture is still limited. Genome-wide chromosome conformation capture studies have uncovered features of chromatin organization in cultured cells, but genome architecture in human tissues has yet to be explored. Here, we report the most comprehensive survey to date of chromatin organization in human tissues. Through integrative analysis of chromatin contact maps in 21 primary human tissues and cell types, we find topologically associating domains highly conserved in different tissues. We also discover genomic regions that exhibit unusually high levels of local …


Hiview: An Integrative Genome Browser To Leverage Hi‑C Results For The Interpretation Of Gwas Variants, Zheng Xu, Guosheng Zhang, Qing Duan, Shengjie Chai, Baqun Zhang, Cong Wu, Fulai Jin, Feng Yue, Yun Li, Ming Hu Jan 2016

Hiview: An Integrative Genome Browser To Leverage Hi‑C Results For The Interpretation Of Gwas Variants, Zheng Xu, Guosheng Zhang, Qing Duan, Shengjie Chai, Baqun Zhang, Cong Wu, Fulai Jin, Feng Yue, Yun Li, Ming Hu

Department of Statistics: Faculty Publications

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with complex traits and diseases. However, most of them are located in the non-protein coding regions, and therefore it is challenging to hypothesize the functions of these non-coding GWAS variants. Recent large efforts such as the ENCODE and Roadmap Epigenomics projects have predicted a large number of regulatory elements. However, the target genes of these regulatory elements remain largely unknown. Chromatin conformation capture based technologies such as Hi-C can directly measure the chromatin interactions and have generated an increasingly comprehensive catalog of the interactome between the distal regulatory elements …


A Bayesian Gwas Method Utilizing Haplotype Clusters For A Composite Breed Population, Danielle F. Wilson-Wells, Stephen D. Kachman Jan 2016

A Bayesian Gwas Method Utilizing Haplotype Clusters For A Composite Breed Population, Danielle F. Wilson-Wells, Stephen D. Kachman

Department of Statistics: Faculty Publications

Commercial beef cattle are often composites of multiple breeds. Current methods used to produce genomic predictors are based on the underlying assumption of animals being sampled from a homogeneous population. As a result, the predictors can perform poorly when used to predict the relative genetic merit of animals whose breed composition are different. In part, this is due to the changes in linkage disequilibrium between the markers and the quantitative trait loci as we move from one breed to the next. An alternative model based on breed specific haplotype clusters was developed to allow for differences in linkage disequilibrium across …


Design Of Probabilistic Random Forests With Applications To Anticancer Drug Sensitivity Prediction- 2016, Raziur Rahman, Saad Haider, Souparno Ghosh, Ranadip Pal Jan 2016

Design Of Probabilistic Random Forests With Applications To Anticancer Drug Sensitivity Prediction- 2016, Raziur Rahman, Saad Haider, Souparno Ghosh, Ranadip Pal

Department of Statistics: Faculty Publications

Random forests consisting of an ensemble of regression trees with equal weights are frequently used for design of predictive models. In this article, we consider an extension of the methodology by representing the regression trees in the form of probabilistic trees and analyzing the nature of heteroscedasticity. The probabilistic tree representation allows for analytical computation of confidence intervals (CIs), and the tree weight optimization is expected to provide stricter CIs with comparable performance in mean error. We approached the ensemble of probabilistic trees’ prediction from the perspectives of a mixture distribution and as a weighted sum of correlated random variables. …


Enscat: Clustering Of Categorical Data Via Ensembling, Bertrand S. Clarke, Saeid Amiri, Jennifer L. Clarke Jan 2016

Enscat: Clustering Of Categorical Data Via Ensembling, Bertrand S. Clarke, Saeid Amiri, Jennifer L. Clarke

Department of Statistics: Faculty Publications

Background: Clustering is a widely used collection of unsupervised learning techniques for identifying natural classes within a data set. It is often used in bioinformatics to infer population substructure. Genomic data are often categorical and high dimensional, e.g., long sequences of nucleotides. This makes inference challenging: The distance metric is often not well-defined on categorical data; running time for computations using high dimensional data can be considerable; and the Curse of Dimensionality often impedes the interpretation of the results. Up to the present, however, the literature and software addressing clustering for categorical data has not yet led to a standard …


Genomic Bayesian Prediction Model For Count Data With Genotype X Environment Interaction, Abelardo Montesinos-López, Osval A. Montesinos-López, José Crossa, Juan Burgueño, Kent M. Eskridge, Esteban Falconi-Castillo, Xinyao He, Pawan Singh, Karen Cichy Jan 2016

Genomic Bayesian Prediction Model For Count Data With Genotype X Environment Interaction, Abelardo Montesinos-López, Osval A. Montesinos-López, José Crossa, Juan Burgueño, Kent M. Eskridge, Esteban Falconi-Castillo, Xinyao He, Pawan Singh, Karen Cichy

Department of Statistics: Faculty Publications

Genomic tools allow the study of the whole genome, and facilitate the study of genotype-environment combinations and their relationship with phenotype. However, most genomic prediction models developed so far are appropriate for Gaussian phenotypes. For this reason, appropriate genomic prediction models are needed for count data, since the conventional regression models used on count data with a large sample size (nT ) and a small number of parameters (p) cannot be used for genomic-enabled prediction where the number of parameters (p) is larger than the sample size (nT ). Here, we propose a Bayesian mixed-negative binomial (BMNB) genomic …


A Genomic Bayesian Multi-Trait And Multi-Environment Model, Osval A. Montesinos-López, Abelardo Montesinos-López, José Crossa, Fernando Toledo, Oscar Pérez-Hernández, Kent M. Eskridge, Jessica Rutkoski Jan 2016

A Genomic Bayesian Multi-Trait And Multi-Environment Model, Osval A. Montesinos-López, Abelardo Montesinos-López, José Crossa, Fernando Toledo, Oscar Pérez-Hernández, Kent M. Eskridge, Jessica Rutkoski

Department of Statistics: Faculty Publications

When information on multiple genotypes evaluated in multiple environments is recorded, a multi-environment single trait model for assessing genotype × environment interaction (G×E) is usually employed. Comprehensive models that simultaneously take into account the correlated traits and trait × genotype × environment interaction (T×G×E) are lacking. In this research, we propose a Bayesian model for analyzing multiple traits and multiple environments for whole-genome prediction (WGP) model. For this model, we used Half-𝑡 priors on each standard deviation term and uniform priors on each correlation of the covariance matrix. These priors were not informative and led to posterior inferences that were …


Species Discovery And Diversity In Lobocriconema (Criconematidae: Nematoda) And Related Plant-Parasitic Nematodes From North American Ecoregions, Tom Powers, Ernest C. Bernard, T. Harris, Robert Higgins, M. Olson, S. Olson, M. Lodema, Julianne N. Matczyszyn, P. Mullin, L. Sutton, K.S Powers Jan 2016

Species Discovery And Diversity In Lobocriconema (Criconematidae: Nematoda) And Related Plant-Parasitic Nematodes From North American Ecoregions, Tom Powers, Ernest C. Bernard, T. Harris, Robert Higgins, M. Olson, S. Olson, M. Lodema, Julianne N. Matczyszyn, P. Mullin, L. Sutton, K.S Powers

Department of Statistics: Faculty Publications

There are many nematode species that, following formal description, are seldom mentioned again in the scientific literature. Lobocriconema thornei and L. incrassatum are two such species, described from North American forests, respectively 37 and 49 years ago. In the course of a 3-year nematode biodiversity survey of North American ecoregions, specimens resembling Lobocriconema species appeared in soil samples from both grassland and forested sites. Using a combination of molecular and morphological analyses, together with a set of species delimitation approaches, we have expanded the known range of these species, added to the species descriptions, and discovered a related group of …


Genomic-Enabled Prediction Of Ordinal Data With Bayesian Logistic Ordinal Regression, Osval A. Montesinos-López, Abelardo Montesinos-López, José Crossa, Juan Burgueño, Kent M. Eskridge Jan 2015

Genomic-Enabled Prediction Of Ordinal Data With Bayesian Logistic Ordinal Regression, Osval A. Montesinos-López, Abelardo Montesinos-López, José Crossa, Juan Burgueño, Kent M. Eskridge

Department of Statistics: Faculty Publications

Most genomic-enabled prediction models developed so far assume that the response variable is continuous and normally distributed. The exception is the probit model, developed for ordered categorical phenotypes. In statistical applications, because of the easy implementation of the Bayesian probit ordinal regression (BPOR) model, Bayesian logistic ordinal regression (BLOR) is implemented rarely in the context of genomic-enabled prediction [sample size (n) is much smaller than the number of parameters (p)]. For this reason, in this paper we propose a BLOR model using the Pólya-Gamma data augmentation approach that produces a Gibbs sampler with similar full conditional distributions of the BPORmodel …


Establishment And Persistence Of Yellow-Flowered Alfalfa No-Till Interseeded Into Crested Wheatgrass Stands, Christopher G. Misar, Lan Xu, Roger N. Gates, Arvid Boe, Patricia S. Johnson, Christopher S. Schauer, John R. Rickertsen, Walter Stroup Jan 2015

Establishment And Persistence Of Yellow-Flowered Alfalfa No-Till Interseeded Into Crested Wheatgrass Stands, Christopher G. Misar, Lan Xu, Roger N. Gates, Arvid Boe, Patricia S. Johnson, Christopher S. Schauer, John R. Rickertsen, Walter Stroup

Department of Statistics: Faculty Publications

Crested wheatgrass [Agropyron cristatum (L.) Gaertn., A. desertorum

(Fisch. ex Link) Schult., and related taxa] often exists

in near monoculture stands in the northern Great Plains.

Introducing locally adapted yellow-flowered alfalfa [Medicago

sativa L. subsp. falcata (L.) Arcang.] would complement crested

wheatgrass. Our objective was to evaluate effects of seeding

date, clethodim {(E) -2-[1-[[(3-chloro-2-propenyl)oxy]imino]

propyl]-5-[2-(ethylthio)propyl]-3-hydroxy-2-cyclohexen-1-one}

sod suppression, and seeding rate on initial establishment and

stand persistence of Falcata, a predominantly yellow-flowered

alfalfa, no-till interseeded into crested wheatgrass. Research was

initiated in August 2008 at Newcastle, WY; Hettinger, ND;

Fruitdale, SD; and Buffalo, SD. Effects of treatment …