Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Wright State University (632)
- The Texas Medical Center Library (43)
- New Jersey Institute of Technology (31)
- Old Dominion University (22)
- University of Kentucky (20)
-
- University of Nebraska - Lincoln (19)
- University of Nebraska at Omaha (18)
- City University of New York (CUNY) (15)
- Wayne State University (14)
- University of Missouri, St. Louis (11)
- Singapore Management University (9)
- San Jose State University (8)
- University of Texas at El Paso (6)
- Brigham Young University (5)
- LSU New Orleans (5)
- Loyola University Chicago (5)
- University of South Florida (5)
- Louisiana Tech University (4)
- Missouri University of Science and Technology (4)
- Nova Southeastern University (4)
- University of Louisville (4)
- University of Massachusetts Boston (4)
- Virginia Commonwealth University (4)
- California Polytechnic State University, San Luis Obispo (3)
- Central Washington University (3)
- Chinese Academy of Sciences (3)
- Dartmouth College (3)
- Karbala International Journal of Modern Science (3)
- Kennesaw State University (3)
- Munster Technological University (3)
- Keyword
-
- Semantic Web (46)
- Bioinformatics (42)
- Machine learning (28)
- Ontology (27)
- Humans (26)
-
- Machine Learning (23)
- Semantic Sensor Web (22)
- Deep learning (20)
- Twitter (15)
- Artificial intelligence (14)
- RDF (14)
- Artificial Intelligence (12)
- Clustering (12)
- Social Media (12)
- SSW (11)
- Deep Learning (10)
- Algorithms (9)
- Data mining (9)
- Ontologies (9)
- Semantic Analytics (9)
- Semantic Web Services (8)
- Linked Data (7)
- SAWSDL (7)
- Semantic web (7)
- Semantics (7)
- Social Networks (7)
- Cloud Computing (6)
- Computational Biology (6)
- Computer science (6)
- Correlation networks (6)
- Publication Year
- Publication
-
- Kno.e.sis Publications (540)
- Computer Science and Engineering Faculty Publications (91)
- Faculty, Staff and Student Publications (42)
- Theses (30)
- Computer Science Theses & Dissertations (15)
-
- Wayne State University Dissertations (11)
- Computer Science Faculty Publications (9)
- Research Collection School Of Computing and Information Systems (9)
- Interdisciplinary Informatics Faculty Proceedings & Presentations (8)
- Dissertations (7)
- Theses and Dissertations--Computer Science (7)
- Chemistry & Biochemistry Faculty Works (6)
- Open Educational Resources (6)
- School of Computing: Dissertations, Theses, and Student Research (6)
- Doctoral Dissertations (5)
- Electronic Theses and Dissertations (5)
- Open Access Theses & Dissertations (5)
- Publications and Research (5)
- USF Tampa Graduate Theses and Dissertations (5)
- CCAC Theses and Dissertations (4)
- Dissertations, Theses, and Capstone Projects (4)
- Faculty Publications (4)
- Faculty Publications, Computer Science (4)
- LSU New Orleans Theses and Dissertations (4)
- Master's Projects (4)
- Master's Theses (4)
- School of Computing: Faculty Publications (4)
- Theses and Dissertations (4)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (3)
- Computer Science Faculty Proceedings & Presentations (3)
- Publication Type
- File Type
Articles 241 - 270 of 985
Full-Text Articles in Computer Sciences
Computational Methods For Prediction And Classification Of G Protein-Coupled Receptors, Khodeza Begum
Computational Methods For Prediction And Classification Of G Protein-Coupled Receptors, Khodeza Begum
Open Access Theses & Dissertations
G protein-coupled receptors (GPCRs) constitute the largest group of membrane receptor proteins in eukaryotes. Due to their significant roles in many physiological processes such as vision, smell, and inflammation, GPCRs are the targets of many prescribed drugs. However, the functional and structural diversity of GPCRs has kept their prediction and classification based on amino acid sequence data as a challenging bioinformatics problem. As existing computational methods to predict and classify GPCRs are focused on mammalian (mostly human) data, the ultimate goal of our project is to establish an ensemble approach and implement a web-based software that can be used reliably …
A Semantics-Based Measure Of Emoji Similarity, Sanjaya Wijeratne, Lakshika Balasuriya, Amit Sheth, Derek Doran
A Semantics-Based Measure Of Emoji Similarity, Sanjaya Wijeratne, Lakshika Balasuriya, Amit Sheth, Derek Doran
Kno.e.sis Publications
Emoji have grown to become one of the most important forms of communication on the web. With its widespread use, measuring the similarity of emoji has become an important problem for contemporary text processing since it lies at the heart of sentiment analysis, search, and interface design tasks. This paper presents a comprehensive analysis of the semantic similarity of emoji through embedding models that are learned over machine-readable emoji meanings in the EmojiNet knowledge base. Using emoji descriptions, emoji sense labels and emoji sense definitions, and with different training corpora obtained from Twitter and Google News, we develop and test …
Identifying Depressive Disorder In The Twitter Population, Goonmeet Bajaj, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth
Identifying Depressive Disorder In The Twitter Population, Goonmeet Bajaj, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth
Kno.e.sis Publications
Depression is a highly prevalent public health challenge and a major cause of disability across the globe.
- Annually 6.7% of Americans (that is, more than 16 million).
- Traditional approaches to curb depression involve survey·based methods via phone or online questionnaires.
- Large temporal gaps and cognitive bias.
Social media provides a method for learning users' feelings, emotions, behaviors, and decisions in real-time.
Novel Neuroevolution Techniques For The Life Science Domain, Timothy Manning
Novel Neuroevolution Techniques For The Life Science Domain, Timothy Manning
Theses
The life science domain is a high value research area, both in terms of the benefits in increased knowledge and in societal impact. Much of the research funding has focused on wet lab based approaches to increase visibility into biological processes and producing maximal relevant information on which to make decisions. Given the complexity of biological functions, in many cases this has led to an information overload. Researchers are now able to routinely generate and access petabytes of data as a result of high throughput experiments, and this capability is growing. This data can be difficult to interpret and intractable …
Biosimp: Using Software Testing Techniques For Sampling And Inference In Biological Organisms, Mikaela Cashman, Jennie L. Catlett, Myra B. Cohen, Nicole R. Buan, Zahmeeth Sakkaff, Massimiliano Pierobon, Christine A. Kelley
Biosimp: Using Software Testing Techniques For Sampling And Inference In Biological Organisms, Mikaela Cashman, Jennie L. Catlett, Myra B. Cohen, Nicole R. Buan, Zahmeeth Sakkaff, Massimiliano Pierobon, Christine A. Kelley
School of Computing: Conference and Workshop Papers
Years of research in software engineering have given us novel ways to reason about, test, and predict the behavior of complex software systems that contain hundreds of thousands of lines of code. Many of these techniques have been inspired by nature such as genetic algorithms, swarm intelligence, and ant colony optimization. In this paper we reverse the direction and present BioSIMP, a process that models and predicts the behavior of biological organisms to aid in the emerging field of systems biology. It utilizes techniques from testing and modeling of highly-configurable software systems. Using both experimental and simulation data we show …
Horizontal And Vertical Integration Of Bio-Molecular Data, Tin Chi Nguyen
Horizontal And Vertical Integration Of Bio-Molecular Data, Tin Chi Nguyen
Wayne State University Dissertations
Modern biomedical research lies at the crossroads of data gathering, interpretation, and hypothesis testing. Due to noise, study bias, or too small changes in biological signals between disease and healthy, individual studies often fail to identify the true phenomenon. Data integration is the key to obtaining the power needed to pinpoint the biological mechanisms of disease states. Given this, we tried to make important contributions in both horizontal and vertical integration of high-throughput data; the former is meta-analysis of independent studies, while the latter is the integration of multi-omics data.
For horizontal meta-analysis, we developed two frameworks: DANUBE and the …
Preliminary Investigation Of Walking Motion Using A Combination Of Image And Signal Processing, Bradley Schneider, Tanvi Banerjee
Preliminary Investigation Of Walking Motion Using A Combination Of Image And Signal Processing, Bradley Schneider, Tanvi Banerjee
Kno.e.sis Publications
We present the results of analyzing gait motion in first-person video taken from a commercially available wearable camera embedded in a pair of glasses. The video is analyzed with three different computer vision methods to extract motion vectors from different gait sequences from four individuals for comparison against a manually annotated ground truth dataset. Using a combination of signal processing and computer vision techniques, gait features are extracted to identify the walking pace of the individual wearing the camera as well as validated using the ground truth dataset. Our preliminary results indicate that the extraction of activity from the video …
A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis
A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis
Open Access Dissertations
Mass spectrometry (MS) imaging is a powerful investigation technique for a wide range of biological applications such as molecular histology of tissue, whole body sections, and bacterial films , and biomedical applications such as cancer diagnosis. MS imaging visualizes the spatial distribution of molecular ions in a sample by repeatedly collecting mass spectra across its surface, resulting in complex, high-dimensional imaging datasets. Two of the primary goals of statistical analysis of MS imaging experiments are classification (for supervised experiments), i.e. assigning pixels to pre-defined classes based on their spectral profiles, and segmentation (for unsupervised experiments), i.e. assigning pixels to newly …
Network Inference Driven Drug Discovery, Gergely Zahoránszky-Kőhalmi, Tudor I. Oprea, Cristian G. Bologa, Subramani Mani, Oleg Ursu
Network Inference Driven Drug Discovery, Gergely Zahoránszky-Kőhalmi, Tudor I. Oprea, Cristian G. Bologa, Subramani Mani, Oleg Ursu
Biomedical Sciences ETDs
The application of rational drug design principles in the era of network-pharmacology requires the investigation of drug-target and target-target interactions in order to design new drugs. The presented research was aimed at developing novel computational methods that enable the efficient analysis of complex biomedical data and to promote the hypothesis generation in the context of translational research. The three chapters of the Dissertation relate to various segments of drug discovery and development process.
The first chapter introduces the integrated predictive drug discovery platform „SmartGraph”. The novel collaborative-filtering based algorithm „Target Based Recommender (TBR)” was developed in the framework of this …
Fractal Analysis Of Dna Sequences, Christian G. Arias, Pedro Antonio Moreno Phd, Carlos Tellez
Fractal Analysis Of Dna Sequences, Christian G. Arias, Pedro Antonio Moreno Phd, Carlos Tellez
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Fedrr: Fast, Exhaustive Detection Of Redundant Hierarchical Relations For Quality Improvement Of Large Biomedical Ontologies, Guangming Xing, Guo-Qiang Zhang, Licong Cui
Fedrr: Fast, Exhaustive Detection Of Redundant Hierarchical Relations For Quality Improvement Of Large Biomedical Ontologies, Guangming Xing, Guo-Qiang Zhang, Licong Cui
Institute for Biomedical Informatics Faculty Publications
Background: Redundant hierarchical relations refer to such patterns as two paths from one concept to another, one with length one (direct) and the other with length greater than one (indirect). Each redundant relation represents a possibly unintended defect that needs to be corrected in the ontology quality assurance process. Detecting and eliminating redundant relations would help improve the results of all methods relying on the relevant ontological systems as knowledge source, such as the computation of semantic distance between concepts and for ontology matching and alignment.
Results: This paper introduces a novel and scalable approach, called FEDRR – Fast, Exhaustive …
Bayesian Networks To Assess The Newborn Stool Microbiome, William E. Bennett Jr.
Bayesian Networks To Assess The Newborn Stool Microbiome, William E. Bennett Jr.
McKelvey School of Engineering Graduate Student Theses & Dissertations
In human stool, a large population of bacterial genes and transcripts from hundreds of genera coexist with host genes and transcripts. Assessments of the metagenome and transcriptome are particularly challenging, since there is a great deal of sequence overlap among related species and related genes. We sequenced the total RNA content from stool samples in a neonate using previously-described methods. We then performed stepwise alignment of different populations of RNA sequence reads to different indices, including ribosomal databases, the human genome, and all sequenced bacterial genomes. Each pool of RNA at each alignment step was subjected to compression to assess …
Analyzing Clinical Depressive Symptoms In Twitter, Amir Hossein Yazdavar, Hussein S. Al-Olimat, Tanvi Banerjee, Krishnaprasad Thirunarayan, Amit P. Sheth
Analyzing Clinical Depressive Symptoms In Twitter, Amir Hossein Yazdavar, Hussein S. Al-Olimat, Tanvi Banerjee, Krishnaprasad Thirunarayan, Amit P. Sheth
Kno.e.sis Publications
350 million people are suffering from clinical depression worldwide.
Protein Residue-Residue Contact Prediction Using Stacked Denoising Autoencoders, Joseph Bailey Luttrell Iv
Protein Residue-Residue Contact Prediction Using Stacked Denoising Autoencoders, Joseph Bailey Luttrell Iv
Honors Theses
Protein residue-residue contact prediction is one of many areas of bioinformatics research that aims to assist researchers in the discovery of structural features of proteins. Predicting the existence of such structural features can provide a starting point for studying the tertiary structures of proteins. This has the potential to be useful in applications such as drug design where tertiary structure predictions may play an important role in approximating the interactions between drugs and their targets without expending the monetary resources necessary for preliminary experimentation. Here, four different methods involving deep learning, support vector machines (SVMs), and direct coupling analysis were …
Incremental Phylogenetics By Repeated Insertions: An Evolutionary Tree Algorithm, Peter Revesz, Zhiqiang Li
Incremental Phylogenetics By Repeated Insertions: An Evolutionary Tree Algorithm, Peter Revesz, Zhiqiang Li
School of Computing: Faculty Publications
We introduce the idea of constructing hypothetical evolutionary trees using an incremental algorithm that inserts species one-by-one into the current evolutionary tree. The method of incremental phylogenetics by repeated insertions lead to an algorithm that can be used on DNA, RNA and amino acid sequences. According to experimental results on both synthetic and biological data, the new algorithm generates more accurate evolutionary trees than the UPGMA and the Neighbor Joining algorithms.
Use Of Clustering Techniques For Protein Domain Analysis, Eric Rodene
Use Of Clustering Techniques For Protein Domain Analysis, Eric Rodene
School of Computing: Dissertations, Theses, and Student Research
Next-generation sequencing has allowed many new protein sequences to be identified. However, this expansion of sequence data limits the ability to determine the structure and function of most of these newly-identified proteins. Inferring the function and relationships between proteins is possible with traditional alignment-based phylogeny. However, this requires at least one shared subsequence. Without such a subsequence, no meaningful alignments between the protein sequences are possible. The entire protein set (or proteome) of an organism contains many unrelated proteins. At this level, the necessary similarity does not occur. Therefore, an alternative method of understanding relationships within diverse sets of proteins …
Ciliate Codon Translator Program Manual, Quentin D. Altemose
Ciliate Codon Translator Program Manual, Quentin D. Altemose
Mathematics Summer Fellows
Understanding the evolutionary history of organisms allows us to better comprehend selective pressures and their effects on larger populations. In our study, we focused on analyzing the DNA of ciliate groups, which are single celled protozoans characterized by the presence of cilia on their outer membrane. We utilized the DNA of the organisms to analyze the changes in population genotype over time. We tested existing evolutionary models (designed to represent natural genetic variation over time in populations) against our data to identify the model with the best fit and likelihood. From the DNA and the evolutionary model with the highest …
What Motivates High School Students To Take Precautions Against The Spread Of Influenza? A Data Science Approach To Latent Modeling Of Compliance With Preventative Practice, William L. Romine, Tanvi Banerjee, William R. Folk, Lloyd H. Barrow
What Motivates High School Students To Take Precautions Against The Spread Of Influenza? A Data Science Approach To Latent Modeling Of Compliance With Preventative Practice, William L. Romine, Tanvi Banerjee, William R. Folk, Lloyd H. Barrow
Kno.e.sis Publications
– This study focuses on a central question: What key behavioral factors influence high school students’ compliance with preventative measures against the transmission of influenza? We use multilevel logistic regression to equate logit measures for eight precautions to students’ latent compliance levels on a common scale. Using linear regression, we explore the efficacy of knowledge of influenza, affective perceptions about influenza and its prevention, prior illness, and gender in predicting compliance. Hand washing and respiratory etiquette are the easiest precautions for students, and hand sanitizer use and keeping the hands away from the face are the most difficult. Perceptions of …
Machine Learning Methods For Brain Image Analysis, Ahmed Fakhry
Machine Learning Methods For Brain Image Analysis, Ahmed Fakhry
Computer Science Theses & Dissertations
Understanding how the brain functions and quantifying compound interactions between complex synaptic networks inside the brain remain some of the most challenging problems in neuroscience. Lack or abundance of data, shortage of manpower along with heterogeneity of data following from various species all served as an added complexity to the already perplexing problem. The ability to process vast amount of brain data need to be performed automatically, yet with an accuracy close to manual human-level performance. These automated methods essentially need to generalize well to be able to accommodate data from different species. Also, novel approaches and techniques are becoming …
A Dynamic Run-Profile Energy-Aware Approach For Scheduling Computationally Intensive Bioinformatics Applications, Sachin Pawaskar, Hesham Ali
A Dynamic Run-Profile Energy-Aware Approach For Scheduling Computationally Intensive Bioinformatics Applications, Sachin Pawaskar, Hesham Ali
Computer Science Faculty Proceedings & Presentations
High Performance Computing (HPC) resources are housed in large datacenters, which consume exorbitant amounts of energy and are quickly demanding attention from businesses as they result in high operating costs. On the other hand HPC environments have been very useful to researchers in many emerging areas in life sciences such as Bioinformatics and Medical Informatics. In an earlier work, we introduced a dynamic model for energy aware scheduling (EAS) in a HPC environment; the model is domain agnostic and incorporates both the deadline parameter as well as energy parameters for computationally intensive applications. Our proposed EAS model incorporates 2-phases. In …
Gene Set Enrichment And Projection: A Computational Tool For Knowledge Discovery In Transcriptomes, Karl Douglas Stamm
Gene Set Enrichment And Projection: A Computational Tool For Knowledge Discovery In Transcriptomes, Karl Douglas Stamm
Dissertations (1934 -)
Explaining the mechanism behind a genetic disease involves two phases, collecting and analyzing data associated to the disease, then interpreting those data in the context of biological systems. The objective of this dissertation was to develop a method of integrating complementary datasets surrounding any single biological process, with the goal of presenting the response to a signal in terms of a set of downstream biological effects. This dissertation specifically tests the hypothesis that computational projection methods overlaid with domain expertise can direct research towards relevant systems-level signals underlying complex genetic disease. To this end, I developed a software algorithm named …
A Computational Framework For Learning From Complex Data: Formulations, Algorithms, And Applications, Wenlu Zhang
A Computational Framework For Learning From Complex Data: Formulations, Algorithms, And Applications, Wenlu Zhang
Computer Science Theses & Dissertations
Many real-world processes are dynamically changing over time. As a consequence, the observed complex data generated by these processes also evolve smoothly. For example, in computational biology, the expression data matrices are evolving, since gene expression controls are deployed sequentially during development in many biological processes. Investigations into the spatial and temporal gene expression dynamics are essential for understanding the regulatory biology governing development. In this dissertation, I mainly focus on two types of complex data: genome-wide spatial gene expression patterns in the model organism fruit fly and Allen Brain Atlas mouse brain data. I provide a framework to explore …
Gene Network Understanding And Analysis, Maria E. Somoza
Gene Network Understanding And Analysis, Maria E. Somoza
Theses
Gene regulatory network (GRN) is a collection of regulators that interact with each other in the cell to govern the gene expression levels of mRNA and proteins. These regulators can either be DNA, RNA, protein and their complex. Transcriptional gene regulation is an important mechanisms in which an in-depth study can lead to various practical applications, and a greater understanding of how organisms control their cellular behavior. One of the most widely studied organisms in gene regulatory networks are the Mycobacterium tuberculosis and Corynebacterium glutamicum ATCC 13032.
Gene co-expression networks are of biological interests due to co-expressed genes which are …
Uusing The Kdj As A Trading Strategy On Biotech Companies, Shijie Zha
Uusing The Kdj As A Trading Strategy On Biotech Companies, Shijie Zha
Theses
Mean Reversion is the most commonly used model in quantitative trading. This model is associated with several factors, like ma5 and ma10 line. These factors are the most significant in stock markets. However, the disadvantages of this model are lag and inaccuracy.
In this research, we get the historical and current stock data by web crawler, analyze the quantitative data and build a new model involved with the KDJ. Taking biotech companies marketed in the United States and B-share marketed in China as the research subjects, the result shows increased profits compared with the Mean Reversion model. It also shows …
2016-01-A3dsrinp-Csc-Sta-Cmb-522-Bps-542, Raymond Pulver, Neal Buxton, Xiaodong Wang, John Lucci, Jean Yves Hervé, Lenore Martin
2016-01-A3dsrinp-Csc-Sta-Cmb-522-Bps-542, Raymond Pulver, Neal Buxton, Xiaodong Wang, John Lucci, Jean Yves Hervé, Lenore Martin
Bioinformatics Software Design Projects
Cholesterol is carried and transported through bloodstream by lipoproteins. There are two types of lipoproteins: low density lipoprotein, or LDL, and high density lipoprotein, or HDL. LDL cholesterol is considered “bad” cholesterol because it can form plaque and hard deposit leading to arteries clog and make them less flexible. Heart attack or stroke will happen if the hard deposit blocks a narrowed artery. HDL cholesterol helps to remove LDL from the artery back to the liver.
Traditionally, particle counts of LDL and HDL plays an important role to understanding and prediction of heart disease risk. But recently research suggested that …
Chipathlon: A Competitive Assessment For Gene Regulation Tools, Avi Knecht, Adam Caprez, Istvan Ladunga
Chipathlon: A Competitive Assessment For Gene Regulation Tools, Avi Knecht, Adam Caprez, Istvan Ladunga
UCARE: Research Products
When gene regulation of the cell cycle malfunctions, it frequently causes cancer.
Adult, differentiated cells can be reprogrammed to induced pluripotent stem cell; which can then be reprogrammed to heart muscle, skin, etc, to repair damaged tissue (to limited extent in clinical practice).
ChIPathlon: Evaluate the performance of all transcription factor mapping (peak calling) methods. To this end, we will develop a scalable and easy to use super computing pipeline to stage data, compare many different peak calling and differential binding site tools, and store all results into a single database.
A Mitochondrial Dna-Based Computational Model Of The Spread Of Human Populations, Peter Revesz
A Mitochondrial Dna-Based Computational Model Of The Spread Of Human Populations, Peter Revesz
School of Computing: Faculty Publications
This paper presents a mitochondrial DNA-based computational model of the spread of human populations. The computation model is based on a new measure of the relatedness of two populations that may be both heterogeneous in terms of their set of mtDNA haplogroups. The measure gives an exponentially increasing weight for the similarity of two haplogroups with the number of levels shared in the mtDNA classification tree. In an experiment, the computational model is applied to the study of the relatedness of seven human populations ranging from the Neolithic through the Bronze Age to the present. The human populations included in …
Semantic, Cognitive, And Perceptual Computing: Paradigms That Shape Human Experience, Amit P. Sheth, Pramod Anantharam, Cory Henson
Semantic, Cognitive, And Perceptual Computing: Paradigms That Shape Human Experience, Amit P. Sheth, Pramod Anantharam, Cory Henson
Kno.e.sis Publications
Unlike machine-centric computing, in which efficient data processing takes precedence over contextual tailoring, human-centric computation provides a personalized data interpretation that most users find highly relevant to their needs. The authors show how semantic, cognitive, and perceptual computing paradigms work together to produce actionable information.
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
COBRA Preprint Series
Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …
A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulous, Konstantin Läufer, George K. Thiruvathukal, Catherine Putonti
A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulous, Konstantin Läufer, George K. Thiruvathukal, Catherine Putonti
Computer Science: Faculty Publications and Other Works
As sequencing technologies continue to drop in price and increase in throughput, new challenges emerge for the management and accessibility of genomic sequence data. We have developed a pipeline for facilitating the storage, retrieval, and subsequent analysis of molecular data, integrating both sequence and metadata. Taking a polyglot approach involving multiple languages, libraries, and persistence mechanisms, sequence data can be aggregated from publicly available and local repositories. Data are exposed in the form of a RESTful web service, formatted for easy querying, and retrieved for downstream analyses. As a proof of concept, we have developed a resource for annotated HIV-1 …