Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Bioinformatics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 241 - 270 of 985

Full-Text Articles in Computer Sciences

Computational Methods For Prediction And Classification Of G Protein-Coupled Receptors, Khodeza Begum Jan 2017

Computational Methods For Prediction And Classification Of G Protein-Coupled Receptors, Khodeza Begum

Open Access Theses & Dissertations

G protein-coupled receptors (GPCRs) constitute the largest group of membrane receptor proteins in eukaryotes. Due to their significant roles in many physiological processes such as vision, smell, and inflammation, GPCRs are the targets of many prescribed drugs. However, the functional and structural diversity of GPCRs has kept their prediction and classification based on amino acid sequence data as a challenging bioinformatics problem. As existing computational methods to predict and classify GPCRs are focused on mammalian (mostly human) data, the ultimate goal of our project is to establish an ensemble approach and implement a web-based software that can be used reliably …


A Semantics-Based Measure Of Emoji Similarity, Sanjaya Wijeratne, Lakshika Balasuriya, Amit Sheth, Derek Doran Jan 2017

A Semantics-Based Measure Of Emoji Similarity, Sanjaya Wijeratne, Lakshika Balasuriya, Amit Sheth, Derek Doran

Kno.e.sis Publications

Emoji have grown to become one of the most important forms of communication on the web. With its widespread use, measuring the similarity of emoji has become an important problem for contemporary text processing since it lies at the heart of sentiment analysis, search, and interface design tasks. This paper presents a comprehensive analysis of the semantic similarity of emoji through embedding models that are learned over machine-readable emoji meanings in the EmojiNet knowledge base. Using emoji descriptions, emoji sense labels and emoji sense definitions, and with different training corpora obtained from Twitter and Google News, we develop and test …


Identifying Depressive Disorder In The Twitter Population, Goonmeet Bajaj, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth Jan 2017

Identifying Depressive Disorder In The Twitter Population, Goonmeet Bajaj, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth

Kno.e.sis Publications

Depression is a highly prevalent public health challenge and a major cause of disability across the globe.

  • Annually 6.7% of Americans (that is, more than 16 million).
  • Traditional approaches to curb depression involve survey·based methods via phone or online questionnaires.
  • Large temporal gaps and cognitive bias.

Social media provides a method for learning users' feelings, emotions, behaviors, and decisions in real-time.


Novel Neuroevolution Techniques For The Life Science Domain, Timothy Manning Jan 2017

Novel Neuroevolution Techniques For The Life Science Domain, Timothy Manning

Theses

The life science domain is a high value research area, both in terms of the benefits in increased knowledge and in societal impact. Much of the research funding has focused on wet lab based approaches to increase visibility into biological processes and producing maximal relevant information on which to make decisions. Given the complexity of biological functions, in many cases this has led to an information overload. Researchers are now able to routinely generate and access petabytes of data as a result of high throughput experiments, and this capability is growing. This data can be difficult to interpret and intractable …


Biosimp: Using Software Testing Techniques For Sampling And Inference In Biological Organisms, Mikaela Cashman, Jennie L. Catlett, Myra B. Cohen, Nicole R. Buan, Zahmeeth Sakkaff, Massimiliano Pierobon, Christine A. Kelley Jan 2017

Biosimp: Using Software Testing Techniques For Sampling And Inference In Biological Organisms, Mikaela Cashman, Jennie L. Catlett, Myra B. Cohen, Nicole R. Buan, Zahmeeth Sakkaff, Massimiliano Pierobon, Christine A. Kelley

School of Computing: Conference and Workshop Papers

Years of research in software engineering have given us novel ways to reason about, test, and predict the behavior of complex software systems that contain hundreds of thousands of lines of code. Many of these techniques have been inspired by nature such as genetic algorithms, swarm intelligence, and ant colony optimization. In this paper we reverse the direction and present BioSIMP, a process that models and predicts the behavior of biological organisms to aid in the emerging field of systems biology. It utilizes techniques from testing and modeling of highly-configurable software systems. Using both experimental and simulation data we show …


Horizontal And Vertical Integration Of Bio-Molecular Data, Tin Chi Nguyen Jan 2017

Horizontal And Vertical Integration Of Bio-Molecular Data, Tin Chi Nguyen

Wayne State University Dissertations

Modern biomedical research lies at the crossroads of data gathering, interpretation, and hypothesis testing. Due to noise, study bias, or too small changes in biological signals between disease and healthy, individual studies often fail to identify the true phenomenon. Data integration is the key to obtaining the power needed to pinpoint the biological mechanisms of disease states. Given this, we tried to make important contributions in both horizontal and vertical integration of high-throughput data; the former is meta-analysis of independent studies, while the latter is the integration of multi-omics data.

For horizontal meta-analysis, we developed two frameworks: DANUBE and the …


Preliminary Investigation Of Walking Motion Using A Combination Of Image And Signal Processing, Bradley Schneider, Tanvi Banerjee Dec 2016

Preliminary Investigation Of Walking Motion Using A Combination Of Image And Signal Processing, Bradley Schneider, Tanvi Banerjee

Kno.e.sis Publications

We present the results of analyzing gait motion in first-person video taken from a commercially available wearable camera embedded in a pair of glasses. The video is analyzed with three different computer vision methods to extract motion vectors from different gait sequences from four individuals for comparison against a manually annotated ground truth dataset. Using a combination of signal processing and computer vision techniques, gait features are extracted to identify the walking pace of the individual wearing the camera as well as validated using the ground truth dataset. Our preliminary results indicate that the extraction of activity from the video …


A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis Dec 2016

A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis

Open Access Dissertations

Mass spectrometry (MS) imaging is a powerful investigation technique for a wide range of biological applications such as molecular histology of tissue, whole body sections, and bacterial films , and biomedical applications such as cancer diagnosis. MS imaging visualizes the spatial distribution of molecular ions in a sample by repeatedly collecting mass spectra across its surface, resulting in complex, high-dimensional imaging datasets. Two of the primary goals of statistical analysis of MS imaging experiments are classification (for supervised experiments), i.e. assigning pixels to pre-defined classes based on their spectral profiles, and segmentation (for unsupervised experiments), i.e. assigning pixels to newly …


Network Inference Driven Drug Discovery, Gergely Zahoránszky-Kőhalmi, Tudor I. Oprea, Cristian G. Bologa, Subramani Mani, Oleg Ursu Nov 2016

Network Inference Driven Drug Discovery, Gergely Zahoránszky-Kőhalmi, Tudor I. Oprea, Cristian G. Bologa, Subramani Mani, Oleg Ursu

Biomedical Sciences ETDs

The application of rational drug design principles in the era of network-pharmacology requires the investigation of drug-target and target-target interactions in order to design new drugs. The presented research was aimed at developing novel computational methods that enable the efficient analysis of complex biomedical data and to promote the hypothesis generation in the context of translational research. The three chapters of the Dissertation relate to various segments of drug discovery and development process.

The first chapter introduces the integrated predictive drug discovery platform „SmartGraph”. The novel collaborative-filtering based algorithm „Target Based Recommender (TBR)” was developed in the framework of this …


Fractal Analysis Of Dna Sequences, Christian G. Arias, Pedro Antonio Moreno Phd, Carlos Tellez Oct 2016

Fractal Analysis Of Dna Sequences, Christian G. Arias, Pedro Antonio Moreno Phd, Carlos Tellez

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Fedrr: Fast, Exhaustive Detection Of Redundant Hierarchical Relations For Quality Improvement Of Large Biomedical Ontologies, Guangming Xing, Guo-Qiang Zhang, Licong Cui Oct 2016

Fedrr: Fast, Exhaustive Detection Of Redundant Hierarchical Relations For Quality Improvement Of Large Biomedical Ontologies, Guangming Xing, Guo-Qiang Zhang, Licong Cui

Institute for Biomedical Informatics Faculty Publications

Background: Redundant hierarchical relations refer to such patterns as two paths from one concept to another, one with length one (direct) and the other with length greater than one (indirect). Each redundant relation represents a possibly unintended defect that needs to be corrected in the ontology quality assurance process. Detecting and eliminating redundant relations would help improve the results of all methods relying on the relevant ontological systems as knowledge source, such as the computation of semantic distance between concepts and for ontology matching and alignment.

Results: This paper introduces a novel and scalable approach, called FEDRR – Fast, Exhaustive …


Bayesian Networks To Assess The Newborn Stool Microbiome, William E. Bennett Jr. Aug 2016

Bayesian Networks To Assess The Newborn Stool Microbiome, William E. Bennett Jr.

McKelvey School of Engineering Graduate Student Theses & Dissertations

In human stool, a large population of bacterial genes and transcripts from hundreds of genera coexist with host genes and transcripts. Assessments of the metagenome and transcriptome are particularly challenging, since there is a great deal of sequence overlap among related species and related genes. We sequenced the total RNA content from stool samples in a neonate using previously-described methods. We then performed stepwise alignment of different populations of RNA sequence reads to different indices, including ribosomal databases, the human genome, and all sequenced bacterial genomes. Each pool of RNA at each alignment step was subjected to compression to assess …


Analyzing Clinical Depressive Symptoms In Twitter, Amir Hossein Yazdavar, Hussein S. Al-Olimat, Tanvi Banerjee, Krishnaprasad Thirunarayan, Amit P. Sheth Aug 2016

Analyzing Clinical Depressive Symptoms In Twitter, Amir Hossein Yazdavar, Hussein S. Al-Olimat, Tanvi Banerjee, Krishnaprasad Thirunarayan, Amit P. Sheth

Kno.e.sis Publications

350 million people are suffering from clinical depression worldwide.


Protein Residue-Residue Contact Prediction Using Stacked Denoising Autoencoders, Joseph Bailey Luttrell Iv Aug 2016

Protein Residue-Residue Contact Prediction Using Stacked Denoising Autoencoders, Joseph Bailey Luttrell Iv

Honors Theses

Protein residue-residue contact prediction is one of many areas of bioinformatics research that aims to assist researchers in the discovery of structural features of proteins. Predicting the existence of such structural features can provide a starting point for studying the tertiary structures of proteins. This has the potential to be useful in applications such as drug design where tertiary structure predictions may play an important role in approximating the interactions between drugs and their targets without expending the monetary resources necessary for preliminary experimentation. Here, four different methods involving deep learning, support vector machines (SVMs), and direct coupling analysis were …


Incremental Phylogenetics By Repeated Insertions: An Evolutionary Tree Algorithm, Peter Revesz, Zhiqiang Li Aug 2016

Incremental Phylogenetics By Repeated Insertions: An Evolutionary Tree Algorithm, Peter Revesz, Zhiqiang Li

School of Computing: Faculty Publications

We introduce the idea of constructing hypothetical evolutionary trees using an incremental algorithm that inserts species one-by-one into the current evolutionary tree. The method of incremental phylogenetics by repeated insertions lead to an algorithm that can be used on DNA, RNA and amino acid sequences. According to experimental results on both synthetic and biological data, the new algorithm generates more accurate evolutionary trees than the UPGMA and the Neighbor Joining algorithms.


Use Of Clustering Techniques For Protein Domain Analysis, Eric Rodene Jul 2016

Use Of Clustering Techniques For Protein Domain Analysis, Eric Rodene

School of Computing: Dissertations, Theses, and Student Research

Next-generation sequencing has allowed many new protein sequences to be identified. However, this expansion of sequence data limits the ability to determine the structure and function of most of these newly-identified proteins. Inferring the function and relationships between proteins is possible with traditional alignment-based phylogeny. However, this requires at least one shared subsequence. Without such a subsequence, no meaningful alignments between the protein sequences are possible. The entire protein set (or proteome) of an organism contains many unrelated proteins. At this level, the necessary similarity does not occur. Therefore, an alternative method of understanding relationships within diverse sets of proteins …


Ciliate Codon Translator Program Manual, Quentin D. Altemose Jul 2016

Ciliate Codon Translator Program Manual, Quentin D. Altemose

Mathematics Summer Fellows

Understanding the evolutionary history of organisms allows us to better comprehend selective pressures and their effects on larger populations. In our study, we focused on analyzing the DNA of ciliate groups, which are single celled protozoans characterized by the presence of cilia on their outer membrane. We utilized the DNA of the organisms to analyze the changes in population genotype over time. We tested existing evolutionary models (designed to represent natural genetic variation over time in populations) against our data to identify the model with the best fit and likelihood. From the DNA and the evolutionary model with the highest …


What Motivates High School Students To Take Precautions Against The Spread Of Influenza? A Data Science Approach To Latent Modeling Of Compliance With Preventative Practice, William L. Romine, Tanvi Banerjee, William R. Folk, Lloyd H. Barrow Jul 2016

What Motivates High School Students To Take Precautions Against The Spread Of Influenza? A Data Science Approach To Latent Modeling Of Compliance With Preventative Practice, William L. Romine, Tanvi Banerjee, William R. Folk, Lloyd H. Barrow

Kno.e.sis Publications

– This study focuses on a central question: What key behavioral factors influence high school students’ compliance with preventative measures against the transmission of influenza? We use multilevel logistic regression to equate logit measures for eight precautions to students’ latent compliance levels on a common scale. Using linear regression, we explore the efficacy of knowledge of influenza, affective perceptions about influenza and its prevention, prior illness, and gender in predicting compliance. Hand washing and respiratory etiquette are the easiest precautions for students, and hand sanitizer use and keeping the hands away from the face are the most difficult. Perceptions of …


Machine Learning Methods For Brain Image Analysis, Ahmed Fakhry Jul 2016

Machine Learning Methods For Brain Image Analysis, Ahmed Fakhry

Computer Science Theses & Dissertations

Understanding how the brain functions and quantifying compound interactions between complex synaptic networks inside the brain remain some of the most challenging problems in neuroscience. Lack or abundance of data, shortage of manpower along with heterogeneity of data following from various species all served as an added complexity to the already perplexing problem. The ability to process vast amount of brain data need to be performed automatically, yet with an accuracy close to manual human-level performance. These automated methods essentially need to generalize well to be able to accommodate data from different species. Also, novel approaches and techniques are becoming …


A Dynamic Run-Profile Energy-Aware Approach For Scheduling Computationally Intensive Bioinformatics Applications, Sachin Pawaskar, Hesham Ali Jul 2016

A Dynamic Run-Profile Energy-Aware Approach For Scheduling Computationally Intensive Bioinformatics Applications, Sachin Pawaskar, Hesham Ali

Computer Science Faculty Proceedings & Presentations

High Performance Computing (HPC) resources are housed in large datacenters, which consume exorbitant amounts of energy and are quickly demanding attention from businesses as they result in high operating costs. On the other hand HPC environments have been very useful to researchers in many emerging areas in life sciences such as Bioinformatics and Medical Informatics. In an earlier work, we introduced a dynamic model for energy aware scheduling (EAS) in a HPC environment; the model is domain agnostic and incorporates both the deadline parameter as well as energy parameters for computationally intensive applications. Our proposed EAS model incorporates 2-phases. In …


Gene Set Enrichment And Projection: A Computational Tool For Knowledge Discovery In Transcriptomes, Karl Douglas Stamm Jul 2016

Gene Set Enrichment And Projection: A Computational Tool For Knowledge Discovery In Transcriptomes, Karl Douglas Stamm

Dissertations (1934 -)

Explaining the mechanism behind a genetic disease involves two phases, collecting and analyzing data associated to the disease, then interpreting those data in the context of biological systems. The objective of this dissertation was to develop a method of integrating complementary datasets surrounding any single biological process, with the goal of presenting the response to a signal in terms of a set of downstream biological effects. This dissertation specifically tests the hypothesis that computational projection methods overlaid with domain expertise can direct research towards relevant systems-level signals underlying complex genetic disease. To this end, I developed a software algorithm named …


A Computational Framework For Learning From Complex Data: Formulations, Algorithms, And Applications, Wenlu Zhang Jul 2016

A Computational Framework For Learning From Complex Data: Formulations, Algorithms, And Applications, Wenlu Zhang

Computer Science Theses & Dissertations

Many real-world processes are dynamically changing over time. As a consequence, the observed complex data generated by these processes also evolve smoothly. For example, in computational biology, the expression data matrices are evolving, since gene expression controls are deployed sequentially during development in many biological processes. Investigations into the spatial and temporal gene expression dynamics are essential for understanding the regulatory biology governing development. In this dissertation, I mainly focus on two types of complex data: genome-wide spatial gene expression patterns in the model organism fruit fly and Allen Brain Atlas mouse brain data. I provide a framework to explore …


Gene Network Understanding And Analysis, Maria E. Somoza May 2016

Gene Network Understanding And Analysis, Maria E. Somoza

Theses

Gene regulatory network (GRN) is a collection of regulators that interact with each other in the cell to govern the gene expression levels of mRNA and proteins. These regulators can either be DNA, RNA, protein and their complex. Transcriptional gene regulation is an important mechanisms in which an in-depth study can lead to various practical applications, and a greater understanding of how organisms control their cellular behavior. One of the most widely studied organisms in gene regulatory networks are the Mycobacterium tuberculosis and Corynebacterium glutamicum ATCC 13032.

Gene co-expression networks are of biological interests due to co-expressed genes which are …


Uusing The Kdj As A Trading Strategy On Biotech Companies, Shijie Zha May 2016

Uusing The Kdj As A Trading Strategy On Biotech Companies, Shijie Zha

Theses

Mean Reversion is the most commonly used model in quantitative trading. This model is associated with several factors, like ma5 and ma10 line. These factors are the most significant in stock markets. However, the disadvantages of this model are lag and inaccuracy.

In this research, we get the historical and current stock data by web crawler, analyze the quantitative data and build a new model involved with the KDJ. Taking biotech companies marketed in the United States and B-share marketed in China as the research subjects, the result shows increased profits compared with the Mean Reversion model. It also shows …


2016-01-A3dsrinp-Csc-Sta-Cmb-522-Bps-542, Raymond Pulver, Neal Buxton, Xiaodong Wang, John Lucci, Jean Yves Hervé, Lenore Martin May 2016

2016-01-A3dsrinp-Csc-Sta-Cmb-522-Bps-542, Raymond Pulver, Neal Buxton, Xiaodong Wang, John Lucci, Jean Yves Hervé, Lenore Martin

Bioinformatics Software Design Projects

Cholesterol is carried and transported through bloodstream by lipoproteins. There are two types of lipoproteins: low density lipoprotein, or LDL, and high density lipoprotein, or HDL. LDL cholesterol is considered “bad” cholesterol because it can form plaque and hard deposit leading to arteries clog and make them less flexible. Heart attack or stroke will happen if the hard deposit blocks a narrowed artery. HDL cholesterol helps to remove LDL from the artery back to the liver.

Traditionally, particle counts of LDL and HDL plays an important role to understanding and prediction of heart disease risk. But recently research suggested that …


Chipathlon: A Competitive Assessment For Gene Regulation Tools, Avi Knecht, Adam Caprez, Istvan Ladunga Apr 2016

Chipathlon: A Competitive Assessment For Gene Regulation Tools, Avi Knecht, Adam Caprez, Istvan Ladunga

UCARE: Research Products

When gene regulation of the cell cycle malfunctions, it frequently causes cancer.

Adult, differentiated cells can be reprogrammed to induced pluripotent stem cell; which can then be reprogrammed to heart muscle, skin, etc, to repair damaged tissue (to limited extent in clinical practice).

ChIPathlon: Evaluate the performance of all transcription factor mapping (peak calling) methods. To this end, we will develop a scalable and easy to use super computing pipeline to stage data, compare many different peak calling and differential binding site tools, and store all results into a single database.


A Mitochondrial Dna-Based Computational Model Of The Spread Of Human Populations, Peter Revesz Mar 2016

A Mitochondrial Dna-Based Computational Model Of The Spread Of Human Populations, Peter Revesz

School of Computing: Faculty Publications

This paper presents a mitochondrial DNA-based computational model of the spread of human populations. The computation model is based on a new measure of the relatedness of two populations that may be both heterogeneous in terms of their set of mtDNA haplogroups. The measure gives an exponentially increasing weight for the similarity of two haplogroups with the number of levels shared in the mtDNA classification tree. In an experiment, the computational model is applied to the study of the relatedness of seven human populations ranging from the Neolithic through the Bronze Age to the present. The human populations included in …


Semantic, Cognitive, And Perceptual Computing: Paradigms That Shape Human Experience, Amit P. Sheth, Pramod Anantharam, Cory Henson Mar 2016

Semantic, Cognitive, And Perceptual Computing: Paradigms That Shape Human Experience, Amit P. Sheth, Pramod Anantharam, Cory Henson

Kno.e.sis Publications

Unlike machine-centric computing, in which efficient data processing takes precedence over contextual tailoring, human-centric computation provides a personalized data interpretation that most users find highly relevant to their needs. The authors show how semantic, cognitive, and perceptual computing paradigms work together to produce actionable information.


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulous, Konstantin Läufer, George K. Thiruvathukal, Catherine Putonti Jan 2016

A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulous, Konstantin Läufer, George K. Thiruvathukal, Catherine Putonti

Computer Science: Faculty Publications and Other Works

As sequencing technologies continue to drop in price and increase in throughput, new challenges emerge for the management and accessibility of genomic sequence data. We have developed a pipeline for facilitating the storage, retrieval, and subsequent analysis of molecular data, integrating both sequence and metadata. Taking a polyglot approach involving multiple languages, libraries, and persistence mechanisms, sequence data can be aggregated from publicly available and local repositories. Data are exposed in the form of a RESTful web service, formatted for easy querying, and retrieved for downstream analyses. As a proof of concept, we have developed a resource for annotated HIV-1 …