Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Bioinformatics (398)
- Genomics (222)
- Physical Sciences and Mathematics (208)
- Genetics (172)
- Medicine and Health Sciences (155)
-
- Biochemistry, Biophysics, and Structural Biology (127)
- Molecular Genetics (105)
- Biology (104)
- Statistics and Probability (93)
- Ecology and Evolutionary Biology (87)
- Computer Sciences (86)
- Cell and Developmental Biology (81)
- Molecular Biology (68)
- Evolution (57)
- Microbiology (55)
- Biotechnology (51)
- Systems Biology (51)
- Biochemistry (43)
- Cell Biology (41)
- Microarrays (39)
- Structural Biology (39)
- Biostatistics (38)
- Medical Sciences (38)
- Engineering (37)
- Animal Sciences (33)
- Diseases (33)
- Statistical Methodology (32)
- Institution
-
- COBRA (113)
- The Texas Medical Center Library (49)
- Dartmouth College (47)
- Old Dominion University (40)
- University of Kentucky (32)
-
- City University of New York (CUNY) (31)
- Himmelfarb Health Sciences Library, The George Washington University (27)
- University of Nebraska - Lincoln (26)
- Thomas Jefferson University (22)
- Virginia Commonwealth University (19)
- Louisiana State University (13)
- The University of Southern Mississippi (13)
- California Polytechnic State University, San Luis Obispo (12)
- Clemson University (12)
- Illinois State University (9)
- Loyola University Chicago (9)
- University of Arkansas, Fayetteville (9)
- University of Connecticut (9)
- Augustana College (8)
- University of Louisville (8)
- Mississippi State University (7)
- West Virginia University (7)
- Western University (7)
- Chapman University (6)
- Munster Technological University (6)
- Swarthmore College (6)
- University of Montana (6)
- University of Nevada, Las Vegas (6)
- University of New Mexico (6)
- Harrisburg University of Science and Technology (5)
- Keyword
-
- Bioinformatics (70)
- Genetics (52)
- Computational biology (40)
- Genomics (37)
- Humans (36)
-
- Gene expression (32)
- Algorithms (24)
- Genome (19)
- Animals (18)
- Evolution (18)
- Machine learning (18)
- Models (16)
- Deep learning (15)
- Transcriptomics (15)
- Genetic (13)
- Computational Biology (12)
- Protein (12)
- Transcriptome (12)
- Phylogeny (11)
- Annotation (10)
- Computer simulation (10)
- Epigenetics (10)
- Machine Learning (10)
- Metabolism (10)
- Phylogenetics (10)
- Population genetics (10)
- Systems biology (10)
- Transcription factors (10)
- Cancer (9)
- RNA (9)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (48)
- Dartmouth Scholarship (36)
- Computer Science Faculty Publications (27)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
- Computational Biology Institute (24)
-
- Harvard University Biostatistics Working Paper Series (24)
- Theses and Dissertations (22)
- Dissertations, Theses, and Capstone Projects (18)
- COBRA Preprint Series (17)
- Electronic Theses and Dissertations (13)
- UW Biostatistics Working Paper Series (12)
- LSU Doctoral Dissertations (11)
- Publications and Research (11)
- Theses and Dissertations--Biology (11)
- Bioconductor Project Working Papers (10)
- Computational Medicine Center Faculty Papers (10)
- Dartmouth College Ph.D Dissertations (10)
- Dissertations (10)
- UPenn Biostatistics Working Papers (10)
- All Dissertations (8)
- Annual Symposium on Biomathematics and Ecology Education and Research (8)
- Honors Theses (8)
- Meiothermus ruber Genome Analysis Project (8)
- STAR Program Research Presentations (8)
- Bioinformatics Faculty Publications (7)
- Biology Faculty Publications (7)
- Department of Pathology, Anatomy, and Cell Biology Faculty Papers (7)
- Master's Theses (7)
- U.C. Berkeley Division of Biostatistics Working Paper Series (7)
- Biochemistry Publications (6)
- Publication Type
- File Type
Articles 331 - 360 of 756
Full-Text Articles in Computational Biology
Towards A Mathematical Model Of Motility Using Dictyostelium Discoideum: Proteins And Geometric Features That Regulate Bleb-Based Motility, Zully Santiago
Towards A Mathematical Model Of Motility Using Dictyostelium Discoideum: Proteins And Geometric Features That Regulate Bleb-Based Motility, Zully Santiago
Dissertations, Theses, and Capstone Projects
A variety of biological functions depend on actin organization. The organization of actin is tightly regulated by a plethora of extracellular and intracellular signaling, scaffolding, and actin-binding proteins. Dysfunctions in this regulation lead to immune diseases, increased susceptibility to pathogens, neurodegenerative diseases, developmental disorders, and cancer metastasis. A variety of actin-dependent processes, including cell motility, are regulated by several proteins of interest: Paxillin, a scaffolding protein; WASP, an actin nucleating protein; SCAR/WAVE, another WASP family actin nucleating protein; Talin, a cortex-to-membrane binding protein; Myosin II, an F-actin contracting motor protein; and Protein Kinase C, a protein kinase. D. discoideum cells …
Using Machine Learning For Antimicrobial Resistant Dna Identification, Jason I. Lingle, John Santerre
Using Machine Learning For Antimicrobial Resistant Dna Identification, Jason I. Lingle, John Santerre
SMU Data Science Review
In this paper, we present a novel machine learning-based methodology for identifying bacteria DNA sub-sequences that are associated with antimicrobial resistance. The dramatic rise in cases of antibiotic resistant bacteria has been an increasing threat across the globe as the existing treatments are rendered ineffective in treating most of these cases due to mutations of their DNA. Among the most recent bacteria to display antimicrobial resistance (AMR) is Neisseria Gonorrhea with the first global treatment failure taking place in 2016. In 2018, new cases of resistance to multiple, high levels of antibiotics were reported in the United Kingdom and Australia. …
Phylogenetic Trees And Networks Can Serve As Powerful And Complementary Approaches For Analysis Of Genomic Data, Christopher Blair, Cécile Ané
Phylogenetic Trees And Networks Can Serve As Powerful And Complementary Approaches For Analysis Of Genomic Data, Christopher Blair, Cécile Ané
Publications and Research
Genomic data have had a profound impact on nearly every biological discipline. In systematics and phylogenetics, the thousands of loci that are now being sequenced can be analyzed under the multispecies coalescent model (MSC) to explicitly account for gene tree discordance due to incomplete lineage sorting (ILS). However, the MSC assumes no gene flow post divergence, calling for additional methods that can accommodate this limitation. Explicit phylogenetic network methods have emerged, which can simultaneously account for ILS and gene flow by representing evolutionary history as a directed acyclic graph. In this point-of-view we highlight some of the strengths and limitations …
Effective Statistical Energy Function Based Protein Un/Structure Prediction, Avdesh Mishra
Effective Statistical Energy Function Based Protein Un/Structure Prediction, Avdesh Mishra
LSU New Orleans Theses and Dissertations
Proteins are an important component of living organisms, composed of one or more polypeptide chains, each containing hundreds or even thousands of amino acids of 20 standard types. The structure of a protein from the sequence determines crucial functions of proteins such as initiating metabolic reactions, DNA replication, cell signaling, and transporting molecules. In the past, proteins were considered to always have a well-defined stable shape (structured proteins), however, it has recently been shown that there exist intrinsically disordered proteins (IDPs), which lack a fixed or ordered 3D structure, have dynamic characteristics and therefore, exist in multiple states. Based on …
The Functional And Structural Analysis Of Drosophila Robo2 In Axon Guidance, Lafreda Janae Howard
The Functional And Structural Analysis Of Drosophila Robo2 In Axon Guidance, Lafreda Janae Howard
Graduate Theses and Dissertations
In animals with complex nervous systems such as mammals and insects, signaling pathways are responsible for guiding axons to their appropriate synaptic targets. Importantly, when this process is not successful during the development of an organism, outcomes include catastrophes such as human neurological diseases and disorders. It is vital to determine the underlying causes of such diseases by understanding the development of the nervous system. There are many pathways that have been identified to play a role in this, however, we lack an understanding of how these pathways can promote such diverse outcomes in different populations of neurons. These pathways …
Association Of Copy Number Variations With Chronic Hepatitis B In Chinese Population, Fang Niu
Association Of Copy Number Variations With Chronic Hepatitis B In Chinese Population, Fang Niu
Capstone Experience: Master of Public Health
With one third of the Hepatitis B virus (HBV) infection population of the world, chronic Hepatitis B (CHB) has become a top burden in China. CHB is a lifelong infection with HBV which can cause serious health problems, like cirrhosis, liver cancer or even death. HBV infection is known to result in various clinical conditions, including asymptomatic HBV carriers to chronic hepatitis and primary hepatocellular carcinoma. Several studies have shown that host genetic susceptibility could be an important factor that determines these various outcomes of HBV infection. Many Single Nucleotide Polymorphisms (SNPs) and Copy Number Variations (CNVs) have been associated …
Molecular Consequences Of High Taz Expression In Gliomas, Visweswaran Ravikumar
Molecular Consequences Of High Taz Expression In Gliomas, Visweswaran Ravikumar
Dissertations and Theses (Open Access)
Diffuse high grade gliomas are complex and lethal neoplasms of the adult central nervous system that are driven by a range of genetic and epigenetic alterations. Molecular classification of these tumors has identified different transcriptional subtypes, the most notable being Proneural (PN) and Mesenchymal (MES) classes. The most aggressive forms of the disease have a Mesenchymal expression signature, with reported PN-to-MES transition occurring with tumor progression. Master regulatory analysis has identified the transcriptional co-activator TAZ (WWTR1) as a major driver of the MES transition. Overexpression of this single protein in glioma stem cells has been shown to drive a transition …
The Development And Use Of Computational Tools In Forensic Science, Dennis E. Slice
The Development And Use Of Computational Tools In Forensic Science, Dennis E. Slice
Human Biology Open Access Pre-Prints
Modern computational resources make available a rich toolkit of statistical methods that can be applied to forensic questions. This toolkit is built on the foundation of statistical developments dating back to the 19th century. To fully and effectively exploit these developments, both the makers and users of software must be keenly aware of the quality, i.e., the accuracy and precision, of the data being modeled or analyzed, and end-users must be sufficiently familiar with the underlying theory to understand the process and results of any analysis or software they use. This is especially important for medico-legal personnel who might be …
High-Performance Computing Frameworks For Large-Scale Genome Assembly, Sayan Goswami
High-Performance Computing Frameworks For Large-Scale Genome Assembly, Sayan Goswami
LSU Doctoral Dissertations
Genome sequencing technology has witnessed tremendous progress in terms of throughput and cost per base pair, resulting in an explosion in the size of data. Typical de Bruijn graph-based assembly tools demand a lot of processing power and memory and cannot assemble big datasets unless running on a scaled-up server with terabytes of RAMs or scaled-out cluster with several dozens of nodes. In the first part of this work, we present a distributed next-generation sequence (NGS) assembler called Lazer, that achieves both scalability and memory efficiency by using partitioned de Bruijn graphs. By enhancing the memory-to-disk swapping and reducing the …
Microrna Profiling And Engineering Of Cho Cell Lines Stably Expressing Difficult-To-Express Lysosomal Protein, Ifeanyi Amadi
Microrna Profiling And Engineering Of Cho Cell Lines Stably Expressing Difficult-To-Express Lysosomal Protein, Ifeanyi Amadi
KGI Theses and Dissertations
Difficult-to-express (DTE) recombinant proteins like multi-specific proteins, DTE monoclonal antibodies and lysosomal enzymes, have seen difficulties in manufacturability using Chinese hamster ovary (CHO) cells and other mammalian cells as production platforms. CHO cells are preferably used for protein production because of their innate ability to secrete human-like recombinant proteins with post-translational modification, resistance to viral infection and familiarity with drug regulators. However, despite huge progress made in engineering CHO cells for high volumetric productivity, expression of DTE proteins like recombinant lysosomal sulfatase represent one of the poorly understood proteins. Furthermore, there are growing interest in the use of microRNAs (miRNAs) …
Computations Of Top-Down Attention By Modulating V1 Dynamics, David Berga, Xavier Otazu
Computations Of Top-Down Attention By Modulating V1 Dynamics, David Berga, Xavier Otazu
MODVIS Workshop
The human visual system processes information defining what is visually conspicuous (saliency) to our perception, guiding eye movements towards certain objects depending on scene context and its feature characteristics. However, attention has been known to be biased by top-down influences (relevance), which define voluntary eye movements driven by goal-directed behavior and memory. We propose a unified model of the visual cortex able to predict, among other effects, top-down visual attention and saccadic eye movements. First, we simulate activations of early mechanisms of the visual system (RGC/LGN), by processing distinct image chromatic opponencies with Gabor-like filters. Second, we use a cortical …
Topology And Dynamics Of Gene Regulatory Networks: A Meta-Analysis, Claus Kadelka
Topology And Dynamics Of Gene Regulatory Networks: A Meta-Analysis, Claus Kadelka
Biology and Medicine Through Mathematics Conference
No abstract provided.
Simplicity Diffexpress: A Bespoke Cloud-Based Interface For Rna-Seq Differential Expression Modeling And Analysis, Cintia C. Palu, Marcelo Ribeiro-Alves, Yanxin Wu, Brendan Lawlor, Pavel V. Baranov, Brian Kelly, Paul Walsh
Simplicity Diffexpress: A Bespoke Cloud-Based Interface For Rna-Seq Differential Expression Modeling And Analysis, Cintia C. Palu, Marcelo Ribeiro-Alves, Yanxin Wu, Brendan Lawlor, Pavel V. Baranov, Brian Kelly, Paul Walsh
Department of Computer Science Publications
One of the key challenges for transcriptomics-based research is not only the processing of large data but also modeling the complexity of features that are sources of variation across samples, which is required for an accurate statistical analysis. Therefore, our goal is to foster access for wet lab researchers to bioinformatics tools, in order to enhance their ability to explore biological aspects and validate hypotheses with robust analysis. In this context, user-friendly interfaces can enable researchers to apply computational biology methods without requiring bioinformatics expertise. Such bespoke platforms can improve the quality of the findings by allowing the researcher to …
Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang
Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang
Senior Honors Projects
Inspired by the landmark paper “Emergence of Scaling in Random Networks” by Barabási and Albert, the field of network science has focused heavily on the power law distribution in recent years. This distribution has been used to model everything from the popularity of sites on the World Wide Web to the number of citations received on a scientific paper. The feature of this distribution is highlighted by the fact that many nodes (websites or papers) have few connections (internet links or citations) while few “hubs” are connected to many nodes. These properties lead to two very important observed effects: the …
Fusarium Euwallacea: A Serious Threat To The Native And Ornamental Trees And Shrubs In Southern California, Greg Tyler, Yixing Zheng, Michael Kulinich, Hagop Atamian
Fusarium Euwallacea: A Serious Threat To The Native And Ornamental Trees And Shrubs In Southern California, Greg Tyler, Yixing Zheng, Michael Kulinich, Hagop Atamian
Student Scholar Symposium Abstracts and Posters
Fusarium Euwallacea is a fungus that has established symbiotic relationship with the beetle Euwallacea aff. fornicata. The beetle bores through the tree bark and into the sapwood making long tunnels inside the trees. The beetle carries the F. Euwallacea in a specialized structure on its body called mandibular mycangia and cultivates the fungus in the tunnels on which the beetle feeds to grow and reproduce. The growth of the fungus obstructs water and mineral transport in the plant xylem tissue, resulting in dieback, wilt and mortality of the host tree. Fungi are known to secrete proteins called effectors in …
Computational Genomic Models For Spatio-Temporal Investigation Of Early Lung Cancer Pathology, Smruthy Sivakumar
Computational Genomic Models For Spatio-Temporal Investigation Of Early Lung Cancer Pathology, Smruthy Sivakumar
Dissertations and Theses (Open Access)
Lung cancer, of which non-small cell lung cancer (NSCLC) is the most common form, is the second most prevalent cancer and the leading cause of cancer-related deaths. NSCLCs primarily comprise adenocarcinomas (LUAD) and squamous cell carcinomas (LUSC). Advances in early detection and prevention have been limited by the lack of early-stage biomarkers and targets. A comprehensive molecular characterization of premalignant lesions and tumor-adjacent normal tissue can aid in better understanding NSCLC pathogenesis. However, these investigations are further challenged by limited tissue availability and low cellular fractions of detectable somatic mutations.
Therefore, there is a dearth of knowledge about the pathogenesis …
Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang
Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang
Biostatistics Faculty Publications
To analyze gene expression data with sophisticated grouping structures and to extract hidden patterns from such data, feature selection is of critical importance. It is well known that genes do not function in isolation but rather work together within various metabolic, regulatory, and signaling pathways. If the biological knowledge contained within these pathways is taken into account, the resulting method is a pathway-based algorithm. Studies have demonstrated that a pathway-based method usually outperforms its gene-based counterpart in which no biological knowledge is considered. In this article, a pathway-based feature selection is firstly divided into three major categories, namely, pathway-level selection, …
Fys: Ethics And Technology (Phil 07/Cpsc 15) Syllabus, Ameet Soni, Krista Karbowski Thomason
Fys: Ethics And Technology (Phil 07/Cpsc 15) Syllabus, Ameet Soni, Krista Karbowski Thomason
Digital Humanities Curricular Development
There has been an accelerated shift in the influence of computing technology and the use of algorithms in our daily lives. With this technology comes serious ethical questions. Philosophers are often well-equipped to wrestle with ethical questions, but less well-equipped to wrestle with questions of technology itself. Computer scientists are well-equipped to deal with the problems and challenges of technology, but less well-equipped to deal with the ethical problems and challenges that technology can pose. In this co-taught course, we bring together the two fields to address ethical questions involving social media, data mining, self-driving cars, artificial intelligence, and other …
Lab Practicum For Bias In Algorithms, Ameet Soni, Krista Karbowski Thomason
Lab Practicum For Bias In Algorithms, Ameet Soni, Krista Karbowski Thomason
Digital Humanities Curricular Development
This is a course assignment to demonstrate potential biases encoded in algorithms (this can be linked more specifically to natural language processing, machine learning, or artificial intelligence) using the Word Embedding Association Test. In lab, students will work with programs that demonstrate the usefulness of word embedding algorithms in finding relationships between words. Then, students will use an implementation of the algorithm in "Semantics derived automatically from language corpora contain human-like biases" by Caliskan et al. to detect gender and racial bias encoded in word embeddings. The assignment has students design and run an experiment using the WEAT algorithm to …
Genome-Wide Association Studies In Maize And Sorghum, Preston Hurst
Genome-Wide Association Studies In Maize And Sorghum, Preston Hurst
Department of Agronomy and Horticulture: Dissertations, Theses, and Student Research
Genome-wide association studies are used to identify genetic variants associated with a particular phenotype. GWAS has been used in a variety of taxa, from humans, to fish to plants . The present analysis is focused on two species important to the human species: maize and sorghum. A GWAS in maize was carried out on the modification of the Ga1-s allele. The Ga1 locus has long been studied as being involved in a unilateral crossing barrier . However, it has long been suspected that the locus is modified by background genetic factors . GWAS was used to observe candidates for this …
Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang
Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang
Biostatistics Faculty Publications
With the rapid evolution of high-throughput technologies, time series/longitudinal high-throughput experiments have become possible and affordable. However, the development of statistical methods dealing with gene expression profiles across time points has not kept up with the explosion of such data. The feature selection process is of critical importance for longitudinal microarray data. In this study, we proposed aggregating a gene’s expression values across time into a single value using the sign average method, thereby degrading a longitudinal feature selection process into a classic one. Regularized logistic regression models with pseudogenes (i.e., the sign average of genes across time as predictors) …
Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan
Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan
COBRA Preprint Series
One of the major goals in large-scale genomic studies is to identify genes with a prognostic impact on time-to-event outcomes which provide insight into the disease's process. With rapid developments in high-throughput genomic technologies in the past two decades, the scientific community is able to monitor the expression levels of tens of thousands of genes and proteins resulting in enormous data sets where the number of genomic features is far greater than the number of subjects. Methods based on univariate Cox regression are often used to select genomic features related to survival outcome; however, the Cox model assumes proportional hazards …
Supervised Dimension Reduction For Large-Scale "Omics" Data With Censored Survival Outcomes Under Possible Non-Proportional Hazards, Lauren Spirko-Burns, Karthik Devarajan
Supervised Dimension Reduction For Large-Scale "Omics" Data With Censored Survival Outcomes Under Possible Non-Proportional Hazards, Lauren Spirko-Burns, Karthik Devarajan
COBRA Preprint Series
The past two decades have witnessed significant advances in high-throughput ``omics" technologies such as genomics, proteomics, metabolomics, transcriptomics and radiomics. These technologies have enabled simultaneous measurement of the expression levels of tens of thousands of features from individual patient samples and have generated enormous amounts of data that require analysis and interpretation. One specific area of interest has been in studying the relationship between these features and patient outcomes, such as overall and recurrence-free survival, with the goal of developing a predictive ``omics" profile. Large-scale studies often suffer from the presence of a large fraction of censored observations and potential …
Cophosk: A Method For Comprehensive Kinase Substrate Annotation Using Co-Phosphorylation Analysis, Marzieh Ayati, Danica D. Wiredja, Daniela M. Schlatzer, Sean Maxwell, Ming Li, Mehmet Koyutürk, Mark R. Chance
Cophosk: A Method For Comprehensive Kinase Substrate Annotation Using Co-Phosphorylation Analysis, Marzieh Ayati, Danica D. Wiredja, Daniela M. Schlatzer, Sean Maxwell, Ming Li, Mehmet Koyutürk, Mark R. Chance
Faculty Scholarship
We present CoPhosK to predict kinase-substrate associations for phosphopeptide substrates detected by mass spectrometry (MS). The tool utilizes a Naïve Bayes framework with priors of known kinase-substrate associations (KSAs) to generate its predictions. Through the mining of MS data for the collective dynamic signatures of the kinases’ substrates revealed by correlation analysis of phosphopeptide intensity data, the tool infers KSAs in the data for the considerable body of substrates lacking such annotations. We benchmarked the tool against existing approaches for predicting KSAs that rely on static information (e.g. sequences, structures and interactions) using publically available MS data, including breast, colon, …
Mrub_3019 Casa Gene Is An Ortholog To E. Coli B2760, Kelsey Heiland, Dr. Lori Scott
Mrub_3019 Casa Gene Is An Ortholog To E. Coli B2760, Kelsey Heiland, Dr. Lori Scott
Meiothermus ruber Genome Analysis Project
This research is part of the Meiothermus ruber genome annotation project which aims to predict gene function with various bioinformatics tools. We investigated the function of Mrub_3019, which encodes the CasA protein involved in the multi-subunit effector complex for the CRISPR-Cas immunity system and predicted it to be an ortholog of E. coli K12 MG1655 b2760 (casA). We predicted that Mrub_3019 encodes the protein CasA, which is involved in PAM recognition of CRISPR interference pathway. Foreign DNA will bind to CasA, which signals Cas3 for helicase-mediated DNA degradation. Our hypothesis is supported by low E-values for pairwise alignment in NCBI …
Mrub_3015 Is Orthologous To The B2757 Gene Found In Escherichia Coli Coding For Casd, Ramona Collins, Dr. Lori Scott
Mrub_3015 Is Orthologous To The B2757 Gene Found In Escherichia Coli Coding For Casd, Ramona Collins, Dr. Lori Scott
Meiothermus ruber Genome Analysis Project
This project is part of the Meiothermus ruber genome analysis project, which uses a collection of online bioinformatics tools to predict gene function. We investigated the biological function of the gene Mrub_3015, which we hypothesize is a component of the CRISPR-Cas prokaryotic defense system. We predict that Mrub_3015 (DNA coordinates 3055550...3056245) encodes the the CRISPR-associated protein cas5, which is integral in maintaining the crRNA-DNA structure, keeping the complex from base pairing with the target phage DNA. Our hypothesis is supported by identical hits for Mrub_3015 and b2527 to the KEGG, Pfam, TIGRfam, CDD and PDB databases as well as a …
Mrub_3018 Is Orthologous To E. Coli B2759 (Casb), Kyle Parker, Dr. Lori Scott
Mrub_3018 Is Orthologous To E. Coli B2759 (Casb), Kyle Parker, Dr. Lori Scott
Meiothermus ruber Genome Analysis Project
This project is part of the Meiothermus ruber genome analysis project, which uses a collection of online bioinformatics tools to predict gene function. We studied the biological activity of the Mrub_3018 gene, which we hypothesize is orthologous to E. coli gene B2759. We predicted that Mrub_3018(DNA coordinates 3057916… 3058524) encodes the protein CasB. CasB is a protein in the CRISPR CASCADE that will function as a structural protein. When the rest of the proteins form an “S” formation CasB will connect the front and back of the “S” creating a back bone for the structure. It will help bind DNA …
Iris-Eda: An Integrated Rna-Seq Interpretation System For Gene Expression Data Analysis, Brandon Monier, Adam Mcdermaid, Cankun Wang, Jing Zhao, Allison Miller, Anne Fennell, Qin Ma
Iris-Eda: An Integrated Rna-Seq Interpretation System For Gene Expression Data Analysis, Brandon Monier, Adam Mcdermaid, Cankun Wang, Jing Zhao, Allison Miller, Anne Fennell, Qin Ma
Agronomy, Horticulture and Plant Science Faculty Publications
Next-Generation Sequencing has made available substantial amounts of large-scale Omics data, providing unprecedented opportunities to understand complex biological systems. Specifically, the value of RNA-Sequencing (RNA-Seq) data has been confirmed in inferring how gene regulatory systems will respond under various conditions (bulk data) or cell types (single-cell data). RNA-Seq can generate genome-scale gene expression profiles that can be further analyzed using correlation analysis, co-expression analysis, clustering, differential gene expression (DGE), among many other studies. While these analyses can provide invaluable information related to gene expression, integration and interpretation of the results can prove challenging. Here we present a tool called IRIS-EDA, …
Phylogenetic History Of The Amy Gene Cluster In Catarrhines, Christian M. Gagnon
Phylogenetic History Of The Amy Gene Cluster In Catarrhines, Christian M. Gagnon
Theses and Dissertations
This study phylogenetically analyzed 30 AMY-related genes from 11 primates. The results show the gradual expansion of the AMY gene family which could have allowed primates to adapt to various ecological landscapes and maximize energy intake from starch-rich foods in periods of food scarcity.
A Novel Pathway-Based Distance Score Enhances Assessment Of Disease Heterogeneity In Gene Expression, Yunqing Liu, Xiting Yan
A Novel Pathway-Based Distance Score Enhances Assessment Of Disease Heterogeneity In Gene Expression, Yunqing Liu, Xiting Yan
Yale Day of Data
Distance-based unsupervised clustering of gene expression data is commonly used to identify heterogeneity in biologic samples. However, high noise levels in gene expression data and the relatively high correlation between genes are often encountered, so traditional distances such as Euclidean distance may not be effective at discriminating the biological differences between samples. In this study, we developed a novel computational method to assess the biological differences based on pathways by assuming that ontologically defined biological pathways in biologically similar samples have similar behavior. Application of this distance score results in more accurate, robust, and biologically meaningful clustering results in both …