Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Bioinformatics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 211 - 240 of 985

Full-Text Articles in Computer Sciences

Khealth Digital Personalized Healthcare Technology For Pediatric Asthma, Utkarshani Jaimini, Hong Y. Yip, Revathy Venkataramanan, Dipesh Kadariya, Vaikunth Sridharan, Tanvi Banerjee, Krishnaprasad Thirunarayan, Maninder Kalra, Amit Sheth Jan 2018

Khealth Digital Personalized Healthcare Technology For Pediatric Asthma, Utkarshani Jaimini, Hong Y. Yip, Revathy Venkataramanan, Dipesh Kadariya, Vaikunth Sridharan, Tanvi Banerjee, Krishnaprasad Thirunarayan, Maninder Kalra, Amit Sheth

Kno.e.sis Publications

Episodic:Traditional Clinician Centric Healthcare

Questions to be answered:

1.Can we reduce the number of asthma attacks through continuous monitoring of the patient's health condition?

2.Can we predict the asthma attack based on the data collected from the patient?

3.Can we predict the asthma vulnerability score for a patient?

4.Can we predict the asthma severity level of a patient?

5.Can we understand the casual relationship between the asthma symptom and the possible factors responsible for it?


Personalized Health Knowledge Graph, Amelia Gyrard, Manas Gaur, Saeedeh Shekarpour, Krishnaprasad Thirunarayan, Amit Sheth Jan 2018

Personalized Health Knowledge Graph, Amelia Gyrard, Manas Gaur, Saeedeh Shekarpour, Krishnaprasad Thirunarayan, Amit Sheth

Kno.e.sis Publications

Our current health applications do not adequately take into account contextual and personalized knowledge about patients. In order to design “Personalized Coach for Healthcare” applications to manage chronic diseases, there is a need to create a Personalized Healthcare Knowledge Graph (PHKG) that takes into consideration a patient’s health condition (personalized knowledge) and enriches that with contextualized knowledge from environmental sensors and Web of Data (e.g., symptoms and treatments for diseases). To develop PHKG, aggregating knowledge from various heterogeneous sources such as the Internet of Things (IoT) devices, clinical notes, and Electronic Medical Records (EMRs) is necessary. In this paper, we …


Scalable Feature Selection And Extraction With Applications In Kinase Polypharmacology, Derek Jones Jan 2018

Scalable Feature Selection And Extraction With Applications In Kinase Polypharmacology, Derek Jones

Theses and Dissertations--Computer Science

In order to reduce the time associated with and the costs of drug discovery, machine learning is being used to automate much of the work in this process. However the size and complex nature of molecular data makes the application of machine learning especially challenging. Much work must go into the process of engineering features that are then used to train machine learning models, costing considerable amounts of time and requiring the knowledge of domain experts to be most effective. The purpose of this work is to demonstrate data driven approaches to perform the feature selection and extraction steps in …


Ultra-Fast And Memory-Efficient Lookups For Cloud, Networked Systems, And Massive Data Management, Ye Yu Jan 2018

Ultra-Fast And Memory-Efficient Lookups For Cloud, Networked Systems, And Massive Data Management, Ye Yu

Theses and Dissertations--Computer Science

Systems that process big data (e.g., high-traffic networks and large-scale storage) prefer data structures and algorithms with small memory and fast processing speed. Efficient and fast algorithms play an essential role in system design, despite the improvement of hardware. This dissertation is organized around a novel algorithm called Othello Hashing. Othello Hashing supports ultra-fast and memory-efficient key-value lookup, and it fits the requirements of the core algorithms of many large-scale systems and big data applications. Using Othello hashing, combined with domain expertise in cloud, computer networks, big data, and bioinformatics, I developed the following applications that resolve several major …


Algorithms For Reconstruction Of Gene Regulatory Networks From High -Throughput Gene Expression Data, Wenping Deng Jan 2018

Algorithms For Reconstruction Of Gene Regulatory Networks From High -Throughput Gene Expression Data, Wenping Deng

Dissertations, Master's Theses and Master's Reports

Understanding gene interactions in complex living systems is one of the central tasks in system biology. With the availability of microarray and RNA-Seq technologies, a multitude of gene expression datasets has been generated towards novel biological knowledge discovery through statistical analysis and reconstruction of gene regulatory networks (GRN). Reconstruction of GRNs can reveal the interrelationships among genes and identify the hierarchies of genes and hubs in networks. The new algorithms I developed in this dissertation are specifically focused on the reconstruction of GRNs with increased accuracy from microarray and RNA-Seq high-throughput gene expression data sets.

The first algorithm (Chapter 2) …


Qualitative Change Detection Approach For Preventive Therapies, Cristina Mitrea Jan 2018

Qualitative Change Detection Approach For Preventive Therapies, Cristina Mitrea

Wayne State University Dissertations

Currently, most diseases are diagnosed only after disease-associated changes have occurred. In this PhD dissertation, we propose a paradigm shift from treating the disease to maintaining the healthy state. The proposed approach is able to identify when systemic qualitative changes in biological systems happen, thus opening the possibility of therapeutic interventions before the occurrence of symptoms. The change detection method exploits knowledge from biological networks and longitudinal data using a system impact analysis approach. This approach is validated on eight datasets, for seven different model organisms and eight biological phenomena. On these data, our proposed method performs well, consistently identifying …


Machine Learning Techniques Implementation In Power Optimization, Data Processing, And Bio-Medical Applications, Khalid Khairullah Mezied Al-Jabery Jan 2018

Machine Learning Techniques Implementation In Power Optimization, Data Processing, And Bio-Medical Applications, Khalid Khairullah Mezied Al-Jabery

Doctoral Dissertations

"The rapid progress and development in machine-learning algorithms becomes a key factor in determining the future of humanity. These algorithms and techniques were utilized to solve a wide spectrum of problems extended from data mining and knowledge discovery to unsupervised learning and optimization. This dissertation consists of two study areas. The first area investigates the use of reinforcement learning and adaptive critic design algorithms in the field of power grid control. The second area in this dissertation, consisting of three papers, focuses on developing and applying clustering algorithms on biomedical data. The first paper presents a novel modelling approach for …


Using Data To Improve Services For Infants With Hearing Loss: Linking Newborn Hearing Screening Records With Early Intervention Records, Maria Gonzalez, Lori Iarossi, Yan Wu, Ying Huang, Kirsten Siegenthaler Nov 2017

Using Data To Improve Services For Infants With Hearing Loss: Linking Newborn Hearing Screening Records With Early Intervention Records, Maria Gonzalez, Lori Iarossi, Yan Wu, Ying Huang, Kirsten Siegenthaler

Journal of Early Hearing Detection and Intervention

The purpose of this study was to match records of infants with permanent hearing loss from the New York Early Hearing Detection and Intervention Information System (NYEHDI-IS) to records of infants with permanent hearing loss receiving early intervention services from the New York State Early Intervention Program (NYSEIP) to identify areas in the state where hearing screening, diagnostic evaluations and referrals to the NYSEIP were not being made or documented in a timely manner. Data from 2014-2016 NYEHDI-IS and NYEIS information systems were matched using The Link King. There were 274 infants documented in NYEIS Information System as receiving early …


Ordinal Convolutional Neural Networks For Predicting Rdoc Positive Valence Psychiatric Symptom Severity Scores, Anthony Rios, Ramakanth Kavuluru Nov 2017

Ordinal Convolutional Neural Networks For Predicting Rdoc Positive Valence Psychiatric Symptom Severity Scores, Anthony Rios, Ramakanth Kavuluru

Computer Science Faculty Publications

Background—The CEGS N-GRID 2016 Shared Task in Clinical Natural Language Processing (NLP) provided a set of 1000 neuropsychiatric notes to participants as part of a competition to predict psychiatric symptom severity scores. This paper summarizes our methods, results, and experiences based on our participation in the second track of the shared task.

Objective—Classical methods of text classification usually fall into one of three problem types: binary, multi-class, and multi-label classification. In this effort, we study ordinal regression problems with text data where misclassifications are penalized differently based on how far apart the ground truth and model predictions are …


Predicting Mental Conditions Based On "History Of Present Illness" In Psychiatric Notes With Deep Neural Networks, Tung Tran, Ramakanth Kavuluru Nov 2017

Predicting Mental Conditions Based On "History Of Present Illness" In Psychiatric Notes With Deep Neural Networks, Tung Tran, Ramakanth Kavuluru

Computer Science Faculty Publications

Background—Applications of natural language processing to mental health notes are not common given the sensitive nature of the associated narratives. The CEGS N-GRID 2016 Shared Task in Clinical Natural Language Processing (NLP) changed this scenario by providing the first set of neuropsychiatric notes to participants. This study summarizes our efforts and results in proposing a novel data use case for this dataset as part of the third track in this shared task.

Objective—We explore the feasibility and effectiveness of predicting a set of common mental conditions a patient has based on the short textual description of patient’s history …


Big Data: Computational Biology Opens A New Window On The World's Challenges For Colby Scientists, Kate Carlisle Oct 2017

Big Data: Computational Biology Opens A New Window On The World's Challenges For Colby Scientists, Kate Carlisle

Colby Magazine

"What makes us 'us' and not a plant? Not a bacteria, or a virus," asks Andrea Tilden, the J. Warren Merrill Associate Professor of Biology and a genomics expert. "Any one genome has six thousand novels worth of information. Computational biology is the tool we use to read them."


Aberrant Coordination Geometries Discovered In Most Abundant Metalloproteins, Sen Yao, Robert M. Flight, Eric C. Rouchka, Hunter N. B. Moseley Oct 2017

Aberrant Coordination Geometries Discovered In Most Abundant Metalloproteins, Sen Yao, Robert M. Flight, Eric C. Rouchka, Hunter N. B. Moseley

Commonwealth Computational Summit

Metalloproteins play crucial biochemical roles in our body and are essential across all domains of life. The structural environment around a metal ion, especially the coordination geometry (CG), is both sequentially and functionally relevant. Studies of the metalloprotein’s CG will greatly help alleviate the imbalance between the ample sequence data available and the insufficient knowledge on protein functions. Current methodologies in characterizing metalloproteins’ CG consider only previously reported CG (canonical CG) models based primarily on nonbiological chemical context. Exceptions to these canonical CG models can greatly hamper the ability to characterize metalloproteins both structurally and functionally.


Adversarial Discriminative Domain Adaptation For Extracting Protein-Protein Interactions From Text, Anthony Rios, Ramakanth Kavuluru, Zhiyong Lu Oct 2017

Adversarial Discriminative Domain Adaptation For Extracting Protein-Protein Interactions From Text, Anthony Rios, Ramakanth Kavuluru, Zhiyong Lu

Commonwealth Computational Summit

Relation extraction is the process of extracting structured information from unstructured text. Recently, neural networks (NNs) have produced state-of-art results in extracting protein-protein interactions (PPIs) from text. While multiple corpora have been created to extract PPIs from text, most methods have shown poor cross-corpora generalization. In other words, models trained on one dataset perform poorly on other datasets for the same task. In the case of PPI, the F1 has been shown to vary by as much as 30% between different datasets. In this work, we utilize adversarial discriminative domain adaptation (ADDA) to improve the generalization between the source and …


A Combinatorial Framework For Multiple Rna Interaction Prediction, Syed Ali Ahmed Sep 2017

A Combinatorial Framework For Multiple Rna Interaction Prediction, Syed Ali Ahmed

Dissertations, Theses, and Capstone Projects

The interaction of two RNA molecules involves a complex interplay between folding and binding that warranted recent developments in RNA-RNA interaction algorithms. However, biological mechanisms in which more than two RNAs take part in an interaction also exist.

A typical algorithmic approach to such problems is to find the minimum energy structure. Often the computationally optimal solution does not represent the biologically correct structure of the interaction. In addition, different biological structures may be observed, depending on several factors. Furthermore, scoring techniques often miss critical details about dependencies within different parts of the structure, which typically leads to lower scores …


Morphogenesis And Growth Driven By Selection Of Dynamical Properties, Yuri Cantor Sep 2017

Morphogenesis And Growth Driven By Selection Of Dynamical Properties, Yuri Cantor

Dissertations, Theses, and Capstone Projects

Organisms are understood to be complex adaptive systems that evolved to thrive in hostile environments. Though widely studied, the phenomena of organism development and growth, and their relationship to organism dynamics is not well understood. Indeed, the large number of components, their interconnectivity, and complex system interactions all obscure our ability to see, describe, and understand the functioning of biological organisms.

Here we take a synthetic and computational approach to the problem, abstracting the organism as a cellular automaton. Such systems are discrete digital models of real-world environments, making them more accessible and easier to study then their physical world …


Machine Learning Based Protein Sequence To (Un)Structure Mapping And Interaction Prediction, Sumaiya Iqbal Aug 2017

Machine Learning Based Protein Sequence To (Un)Structure Mapping And Interaction Prediction, Sumaiya Iqbal

LSU New Orleans Theses and Dissertations

Proteins are the fundamental macromolecules within a cell that carry out most of the biological functions. The computational study of protein structure and its functions, using machine learning and data analytics, is elemental in advancing the life-science research due to the fast-growing biological data and the extensive complexities involved in their analyses towards discovering meaningful insights. Mapping of protein’s primary sequence is not only limited to its structure, we extend that to its disordered component known as Intrinsically Disordered Proteins or Regions in proteins (IDPs/IDRs), and hence the involved dynamics, which help us explain complex interaction within a cell that …


Testing The Independence Hypothesis Of Accepted Mutations For Pairs Of Adjacent Amino Acids In Protein Sequences, Jyotsna Ramanan, Peter Revesz Jul 2017

Testing The Independence Hypothesis Of Accepted Mutations For Pairs Of Adjacent Amino Acids In Protein Sequences, Jyotsna Ramanan, Peter Revesz

School of Computing: Faculty Publications

Evolutionary studies usually assume that the genetic mutations are independent of each other. However, that does not imply that the observed mutations are independent of each other because it is possible that when a nucleotide is mutated, then it may be biologically beneficial if an adjacent nucleotide mutates too. With a number of decoded genes currently available in various genome libraries and online databases, it is now possible to have a large-scale computer-based study to test whether the independence assumption holds for pairs of adjacent amino acids. Hence the independence question also arises for pairs of adjacent amino acids within …


Discovering Explanatory Models To Identify Relevant Tweets On Zika, Roopteja Muppalla, Michele Miller, Tanvi Banerjee, William L. Romine Jul 2017

Discovering Explanatory Models To Identify Relevant Tweets On Zika, Roopteja Muppalla, Michele Miller, Tanvi Banerjee, William L. Romine

Kno.e.sis Publications

Zika virus has caught the worlds attention, and has led people to share their opinions and concerns on social media like Twitter. Using text-based features, extracted with the help of Parts of Speech (POS) taggers and N-gram, a classifier was built to detect Zika related tweets from Twitter. With a simple logistic classifier, the system was successful in detecting Zika related tweets from Twitter with a 92% accuracy. Moreover, key features were identified that provide deeper insights on the content of tweets relevant to Zika. This system can be leveraged by domain experts to perform sentiment analysis, and understand the …


A Knowledge Graph Framework For Detecting Traffic Events Using Stationary Cameras, Roopteja Muppalla, Sarasi Lalithsena, Tanvi Banerjee, Amit Sheth Jun 2017

A Knowledge Graph Framework For Detecting Traffic Events Using Stationary Cameras, Roopteja Muppalla, Sarasi Lalithsena, Tanvi Banerjee, Amit Sheth

Kno.e.sis Publications

With the rapid increase in urban development, it is critical to utilize dynamic sensor streams for traffic understanding, especially in larger cities where route planning or infrastructure planning is more critical. This creates a strong need to understand traffic patterns using ubiquitous sensors to allow city officials to be better informed when planning urban construction and to provide an understanding of the traffic dynamics in the city. In this study, we propose our framework ITSKG (Imagery-based Traffic Sensing Knowledge Graph) which utilizes the stationary traffic camera information as sensors to understand the traffic patterns. The proposed system extracts image-based features …


Software Development For Genome Sequence Analysis, David Farr May 2017

Software Development For Genome Sequence Analysis, David Farr

Symposium Of University Research and Creative Expression (SOURCE)

The cost of genome sequencing has decreased rapidly, expanding availability for many biological applications (Muir 2016). For example, researchers can now obtain genome sequences from multiple populations under different types of selection. Comparison of these sequences allows for identification of chromosome regions and specific genes associated with adaptive evolution (Kelly 2013). As an increasing number of researchers engage in this type of inquiry, many have created in-house computer scripts to analyze the raw sequence data (e.g., Kelly 2013), creating a gap in both continuity and standardization.

Using a test dataset and preliminary results from an ongoing artificial selection experiment in …


Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane Apr 2017

Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane

Theses

Alzheimer Disease (AD) is difficult to diagnose by using genetic testing or other traditional methods. Unlike diseases with simple genetic risk components, there exists no single marker determining as to whether someone will develop AD. Furthermore, AD is highly heterogeneous and different subgroups of individuals develop the disease due to differing factors. Traditional diagnostic methods using perceivable cognitive deficiencies are often too little too late due to the brain having suffered damage from decades of disease progression. In order to observe AD at early stages prior to the observation of cognitive deficiencies, biomarkers with greater accuracy are required. By using …


A Proposed Frequency-Based Feature Selection Method For Cancer Classification, Yi Pan Apr 2017

A Proposed Frequency-Based Feature Selection Method For Cancer Classification, Yi Pan

Masters Theses & Specialist Projects

Feature selection method is becoming an essential procedure in data preprocessing step. The feature selection problem can affect the efficiency and accuracy of classification models. Therefore, it also relates to whether a classification model can have a reliable performance. In this study, we compared an original feature selection method and a proposed frequency-based feature selection method with four classification models and three filter-based ranking techniques using a cancer dataset. The proposed method was implemented in WEKA which is an open source software. The performance is evaluated by two evaluation methods: Recall and Receiver Operating Characteristic (ROC). Finally, we found the …


What Are People Tweeting About Zika? An Exploratory Study Concerning Its Symptoms, Treatment, Transmission, And Prevention, Michele Miller, Tanvi Banerjee, Roopteja Muppalla, William L. Romine, Amit Sheth Apr 2017

What Are People Tweeting About Zika? An Exploratory Study Concerning Its Symptoms, Treatment, Transmission, And Prevention, Michele Miller, Tanvi Banerjee, Roopteja Muppalla, William L. Romine, Amit Sheth

Kno.e.sis Publications

Background: In order to harness what people are tweeting about Zika, there needs to be a computational framework that leverages machine learning techniques to recognize relevant Zika tweets and, further, categorize these into disease-specific categories to address specific societal concerns related to the prevention, transmission, symptoms, and treatment of Zika virus.

Objective: The purpose of this study was to determine the relevancy of the tweets and what people were tweeting about the 4 disease characteristics of Zika: symptoms, transmission, prevention, and treatment.

Methods: A combination of natural language processing and machine learning techniques was used to determine what people were …


Eassistant: Cognitive Assistance For Identification And Auto-Triage Of Actionable Conversations, Hamid R. Motahari Nezhad, Kalpa Gunaratna, Juan Cappi Apr 2017

Eassistant: Cognitive Assistance For Identification And Auto-Triage Of Actionable Conversations, Hamid R. Motahari Nezhad, Kalpa Gunaratna, Juan Cappi

Kno.e.sis Publications

The browser and screen have been the main user interfaces of the Web and mobile apps. The notification mechanism is an evolution in the user interaction paradigm by keeping users updated without checking applications. Conversational agents are posed to be the next revolution in user interaction paradigms. However, without intelligence on the triage of content served by the interaction and content differentiation in applications, interaction paradigms may still place the burden of information overload on users. In this paper, we focus on the problem of intelligent identification of actionable information in the content served by applications, and in particular in …


Road Accidents Bigdata Mining And Visualization Using Support Vector Machines, Usha Lokala, Srinivas Nowduri, Prabhakar K. Sharma Jan 2017

Road Accidents Bigdata Mining And Visualization Using Support Vector Machines, Usha Lokala, Srinivas Nowduri, Prabhakar K. Sharma

Kno.e.sis Publications

Useful information has been extracted from the road accident data in United Kingdom (UK), using data analytics method, for avoiding possible accidents in rural and urban areas. This analysis make use of several methodologies such as data integration, support vector machines (SVM), correlation machines and multinomial goodness. The entire datasets have been imported from the traffic department of UK with due permission. The information extracted from these huge datasets forms a basis for several predictions, which in turn avoid unnecessary memory lapses. Since data is expected to grow continuously over a period of time, this work primarily proposes a new …


Relatedness-Based Multi-Entity Summarization, Kalpa Gunaratna, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth, Gong Cheng Jan 2017

Relatedness-Based Multi-Entity Summarization, Kalpa Gunaratna, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit Sheth, Gong Cheng

Kno.e.sis Publications

Representing world knowledge in a machine processable format is important as entities and their descriptions have fueled tremendous growth in knowledge-rich information processing platforms, services, and systems. Prominent applications of knowledge graphs include search engines (e.g., Google Search and Microsoft Bing), email clients (e.g., Gmail), and intelligent personal assistants (e.g., Google Now, Amazon Echo, and Apple’s Siri). In this paper, we present an approach that can summarize facts about a collection of entities by analyzing their relatedness in preference to summarizing each entity in isolation. Specifically, we generate informative entity summaries by selecting: (i) inter-entity facts that are similar and …


An Out-Of-Core Gpu Based Dimensionality Reduction Algorithm For Big Mass Spectrometry Data And Its Application In Bottom-Up Proteomics, Muaaz Awan, Fahad Saeed Jan 2017

An Out-Of-Core Gpu Based Dimensionality Reduction Algorithm For Big Mass Spectrometry Data And Its Application In Bottom-Up Proteomics, Muaaz Awan, Fahad Saeed

Parallel Computing and Data Science Lab Technical Reports

Modern high resolution Mass Spectrometry instruments can generate millions of spectra in a single systems biology experiment. Each spectrum consists of thousands of peaks but only a small number of peaks actively contribute to deduction of peptides. Therefore, pre-processing of MS data to detect noisy and non-useful peaks are an active area of research. Most of the sequential noise reducing algorithms are impractical to use as a pre-processing step due to high time-complexity. In this paper, we present a GPU based dimensionality-reduction algorithm, called G-MSR, for MS2 spectra. Our proposed algorithm uses novel data structures which optimize the memory and …


Gpu-Pcc: A Gpu Based Technique To Compute Pairwise Pearson’S Correlation Coefficients For Big Fmri Data, Taban Eslami, Muaaz Gul Awan, Fahad Saeed Jan 2017

Gpu-Pcc: A Gpu Based Technique To Compute Pairwise Pearson’S Correlation Coefficients For Big Fmri Data, Taban Eslami, Muaaz Gul Awan, Fahad Saeed

Parallel Computing and Data Science Lab Technical Reports

Functional Magnetic Resonance Imaging (fMRI) is a non-invasive brain imaging technique for studying the brain’s functional activities. Pearson’s Correlation Coefficient is an important measure for capturing dynamic behaviors and functional connectivity between brain components. One bottleneck in computing Correlation Coefficients is the time it takes to process big fMRI data. In this paper, we propose GPU-PCC, a GPU based algorithm based on vector dot product, which is able to compute pairwise Pearson’s Correlation Coefficients while performing computation once for each pair. Our method is able to compute Correlation Coefficients in an ordered fashion without the need to do post-processing reordering …


A Novel Approach For Classifying Gene Expression Data Using Topic Modeling, Soon Jye Kho, Himi Yalamanchili, Michael L. Raymer, Amit Sheth Jan 2017

A Novel Approach For Classifying Gene Expression Data Using Topic Modeling, Soon Jye Kho, Himi Yalamanchili, Michael L. Raymer, Amit Sheth

Kno.e.sis Publications

Understanding the role of differential gene expression in cancer etiology and cellular process is a complex problem that continues to pose a challenge due to sheer number of genes and inter-related biological processes involved. In this paper, we employ an unsupervised topic model, Latent Dirichlet Allocation (LDA) to mitigate overfitting of high-dimensionality gene expression data and to facilitate understanding of the associated pathways. LDA has been recently applied for clustering and exploring genomic data but not for classification and prediction. Here, we proposed to use LDA inclustering as well as in classification of cancer and healthy tissues using lung cancer …


K-Mer Analysis Pipeline For Classification Of Dna Sequences From Metagenomic Samples, Russell Kaehler Jan 2017

K-Mer Analysis Pipeline For Classification Of Dna Sequences From Metagenomic Samples, Russell Kaehler

Graduate Student Theses, Dissertations, & Professional Papers

Biological sequence datasets are increasing at a prodigious rate. The volume of data in these datasets surpasses what is observed in many other fields of science. New developments wherein metagenomic DNA from complex bacterial communities is recovered and sequenced are producing a new kind of data known as metagenomic data, which is comprised of DNA fragments from many genomes. Developing a utility to analyze such metagenomic data and predict the sample class from which it originated has many possible implications for ecological and medical applications. Within this document is a description of a series of analytical techniques used to process …