Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Bioinformatics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 121 - 150 of 985

Full-Text Articles in Computer Sciences

Tagtools: Software Supporting Biologging Research, Sam Fynewever, Racheal Tejevbo, Stacy L. De Ruiter Jan 2021

Tagtools: Software Supporting Biologging Research, Sam Fynewever, Racheal Tejevbo, Stacy L. De Ruiter

Summer Research

Observing animals in obscure habitats has long posed a challenge in biology. Biologging addresses this broad problem, obtaining data about animals like their acceleration, GPS position, etc. using electronic tags. Data from tags can lead to important knowledge of animal behavior, as DeRuiter et al. (2013) have shown1 . However, data is often shared raw, causing inconsistencies. This project maintains TagTools, software addressing this specific problem. TagToolsfunctions (in Matlab, Octave & R) calibrate data and help analyze behavior. TagTools software is traditionally taught at in-person workshops. Our old documents (practicals) were hosted online, but had bugs & were very long, …


The Role Of Software Engineering In Bioinformatics, Brendan Sean Lawlor Jan 2021

The Role Of Software Engineering In Bioinformatics, Brendan Sean Lawlor

Theses

This thesis proposes that by applying state-of-the-art software engineering tools, techniques and frameworks to currently recognised challenges in bioinformatics, improved outcomes can be attained in that field. It begins by decomposing software engineering into two categories, namely process and architecture, and choosing two key challenges in the practice of bioinformatics: reproducibility and scalability. The body of the thesis is an exploration of the intersection between these two software engineering categories and these two bioinformatics challenges. The question is asked: Can best practices in professional software engineering be applied to address key issues in the bioinformatics domain, creating positive outcomes? And …


Soda: An Open-Source Library For Visualizing Biological Sequence Annotation, Jack W. Roddy, Travis J. Wheeler Jan 2021

Soda: An Open-Source Library For Visualizing Biological Sequence Annotation, Jack W. Roddy, Travis J. Wheeler

Graduate Student Theses, Dissertations, & Professional Papers

Genome annotation is the process of identifying and labeling known genetic sequences or features within a genome. Across the various subfields within modern molecular biology, there is a common need for the visualization of such annotations. Genomic data is often visualized on web browser platforms, providing users with easy access to visualization tools without the need for installing any software or, in many cases, underlying datasets. While there exists a broad range of web-based visualization tools, there is, to my knowledge, no lightweight, modern library tailored towards the visualization of genomic data. Instead, developers charged with the task of producing …


Neural Network Supervised And Reinforcement Learning For Neurological, Diagnostic, And Modeling Problems, Donald Wunsch Iii Jan 2021

Neural Network Supervised And Reinforcement Learning For Neurological, Diagnostic, And Modeling Problems, Donald Wunsch Iii

Masters Theses

“As the medical world becomes increasingly intertwined with the tech sphere, machine learning on medical datasets and mathematical models becomes an attractive application. This research looks at the predictive capabilities of neural networks and other machine learning algorithms, and assesses the validity of several feature selection strategies to reduce the negative effects of high dataset dimensionality. Our results indicate that several feature selection methods can maintain high validation and test accuracy on classification tasks, with neural networks performing best, for both single class and multi-class classification applications. This research also evaluates a proof-of-concept application of a deep-Q-learning network (DQN) to …


A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh Jan 2021

A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh

Dartmouth College Master’s Theses

Computational methods for predicting binding interfaces between antigens and antibodies (epitopes and paratopes) are faster and cheaper than traditional experimental structure determination methods. A sufficiently reliable computational predictor that could scale to large sets of available antibody sequence data could thus inform and expedite many biomedical pursuits, such as better understanding immune responses to vaccination and natural infection and developing better drugs and vaccines. However, current state-of-the-art predictors produce discontiguous predictions, e.g., predicting the epitope in many different spots on an antigen, even though in reality they typically comprise a single localized region. We seek to produce contiguous predicted epitopes, …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke Jan 2021

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Bioinformatics Metadata Extraction For Machine Learning Analysis, Zachary Tom Dec 2020

Bioinformatics Metadata Extraction For Machine Learning Analysis, Zachary Tom

Master's Projects

Next generation sequencing (NGS) has revolutionized the biological sciences. Today, entire genomes can be rapidly sequenced, enabling advancements in personalized medicine, genetic diseases, and more. The National Center for Biotechnology Information (NCBI) hosts the Sequence Read Archive (SRA) containing vast amounts of valuable NGS data. Recently, research has shown that sequencing errors in conventional NGS workflows are key confounding factors for detecting mutations. Various steps such as sample handling and library preparation can introduce artifacts that affect the accuracy of calling rare mutations. Thus, there is a need for more insight into the exact relationship between various steps of the …


Extending Import Detection Algorithms For Concept Import From Two To Three Biomedical Terminologies, Vipina K. Keloth, James Geller, Yan Chen, Julia Xu Dec 2020

Extending Import Detection Algorithms For Concept Import From Two To Three Biomedical Terminologies, Vipina K. Keloth, James Geller, Yan Chen, Julia Xu

Publications and Research

Background: While enrichment of terminologies can be achieved in different ways, filling gaps in the IS-A hierarchy backbone of a terminology appears especially promising. To avoid difficult manual inspection, we started a research program in 2014, investigating terminology densities, where the comparison of terminologies leads to the algorithmic discovery of potentially missing concepts in a target terminology. While candidate concepts have to be approved for import by an expert, the human effort is greatly reduced by algorithmic generation of candidates. In previous studies, a single source terminology was used with one target terminology.

Methods: In this paper, we are extending …


Development Of Computational To Ols To Target Microrna, Luo Song Dec 2020

Development Of Computational To Ols To Target Microrna, Luo Song

Dissertations and Theses (Open Access)

MicroRNAs (a.k.a, miRNAs) play an important role in disease development. However, few of their structures have been determined and structure-based computational methods remain challenging in accurately predicting their interactions with small molecules. To address this issue, my thesis is to develop integrated approaches to screening for novel inhibitors by targeting specific structure motifs in miRNAs. The project starts with implementing a tool to find potential miRNA targets with desired motifs. I combined both sequence information of miRNAs and known RNA structure data from Protein Data Bank (PDB) to predict the miRNA structure and identify the motif to target, then I …


New Methods For Deep Learning Based Real-Valued Inter-Residue Distance Prediction, Jacob Barger Nov 2020

New Methods For Deep Learning Based Real-Valued Inter-Residue Distance Prediction, Jacob Barger

Theses

Background: Much of the recent success in protein structure prediction has been a result of accurate protein contact prediction--a binary classification problem. Dozens of methods, built from various types of machine learning and deep learning algorithms, have been published over the last two decades for predicting contacts. Recently, many groups, including Google DeepMind, have demonstrated that reformulating the problem as a multi-class classification problem is a more promising direction to pursue. As an alternative approach, we recently proposed real-valued distance predictions, formulating the problem as a regression problem. The nuances of protein 3D structures make this formulation appropriate, allowing predictions …


Discrete Models And Algorithms For Analyzing Dna Rearrangements, Jasper Braun Nov 2020

Discrete Models And Algorithms For Analyzing Dna Rearrangements, Jasper Braun

USF Tampa Graduate Theses and Dissertations

In this work, language and tools are introduced, which model many-to-many mappings that comprise DNA rearrangements in nature. Existing theoretical models and data processing methods depend on the premise that DNA segments in the rearrangement precursor are in a clear one-to-one correspondence with their destinations in the recombined product. However, ambiguities in the rearrangement maps obtained from the ciliate species Oxytricha trifallax violate this assumption demonstrating a necessity for the adaptation of theory and practice.

In order to take into account the ambiguities in the rearrangement maps, generalizations of existing recombination models are proposed. Edges in an ordered graph model …


Literature Retrieval For Precision Medicine With Neural Matching And Faceted Summarization, Jiho Noh, Ramakanth Kavuluru Nov 2020

Literature Retrieval For Precision Medicine With Neural Matching And Faceted Summarization, Jiho Noh, Ramakanth Kavuluru

Institute for Biomedical Informatics Faculty Publications

Information retrieval (IR) for precision medicine (PM) often involves looking for multiple pieces of evidence that characterize a patient case. This typically includes at least the name of a condition and a genetic variation that applies to the patient. Other factors such as demographic attributes, comorbidities, and social determinants may also be pertinent. As such, the retrieval problem is often formulated as ad hoc search but with multiple facets (e.g., disease, mutation) that may need to be incorporated. In this paper, we present a document reranking approach that combines neural query-document matching and text summarization toward such retrieval scenarios. Our …


Deepfrag-K: A Fragment-Based Deep Learning Approach For Protein Fold Recognition, Wessam Elhefnawy, Min Li, Jianxin Wang, Yaohang Li Nov 2020

Deepfrag-K: A Fragment-Based Deep Learning Approach For Protein Fold Recognition, Wessam Elhefnawy, Min Li, Jianxin Wang, Yaohang Li

Computer Science Faculty Publications

Background: One of the most essential problems in structural bioinformatics is protein fold recognition. In this paper, we design a novel deep learning architecture, so-called DeepFrag-k, which identifies fold discriminative features at fragment level to improve the accuracy of protein fold recognition. DeepFrag-k is composed of two stages: the first stage employs a multi-modal Deep Belief Network (DBN) to predict the potential structural fragments given a sequence, represented as a fragment vector, and then the second stage uses a deep convolutional neural network (CNN) to classify the fragment vector into the corresponding fold.

Results: Our results show that DeepFrag-k yields …


Integrated Multiparametric Radiomics And Informatics System For Characterizing Breast Tumor Characteristics With The Oncotypedx Gene Assay, Michael A. Jacobs, Christopher B. Umbricht, Vishwa S. Parekh, Riham H. El Khouli, Leslie Cope, Katarzyna J. Macura, Susan Harvey, Antonio C. Wolff Sep 2020

Integrated Multiparametric Radiomics And Informatics System For Characterizing Breast Tumor Characteristics With The Oncotypedx Gene Assay, Michael A. Jacobs, Christopher B. Umbricht, Vishwa S. Parekh, Riham H. El Khouli, Leslie Cope, Katarzyna J. Macura, Susan Harvey, Antonio C. Wolff

Radiology Faculty Publications

Optimal use of multiparametric magnetic resonance imaging (mpMRI) can identify key MRI parameters and provide unique tissue signatures defining phenotypes of breast cancer. We have developed and implemented a new machine-learning informatic system, termed Informatics Radiomics Integration System (IRIS) that integrates clinical variables, derived from imaging and electronic medical health records (EHR) with multiparametric radiomics (mpRad) for identifying potential risk of local or systemic recurrence in breast cancer patients. We tested the model in patients (n = 80) who had Estrogen Receptor positive disease and underwent OncotypeDX gene testing, radiomic analysis, and breast mpMRI. The IRIS method was trained …


Feature Selection Via Random Subsets Of Uncorrelated Features, Long Kim Dang Sep 2020

Feature Selection Via Random Subsets Of Uncorrelated Features, Long Kim Dang

USF Tampa Graduate Theses and Dissertations

The role of feature selection is crucial in many applications. A few of these include computational biology, image classification and risk management. In biology, gene expression micro array data sets have been used extensively in many areas of research. These data sets typically suffer from an important problem: the ratio between the number of features over the number of examples is very high. This problem mainly affects prediction accuracy because it is best to collect more labeled examples than features. A correlation based random subspace ensemble feature selector (CCC_RSM) was proposed to handle this problem [5]. In this approach, first …


Timing Of Maximal Weight Reduction Following Bariatric Surgery: A Study In Chinese Patients, Ting Xu, Chen Wang, Hongwei Zhang, Xiaodong Han, Weijie Liu, Junfeng Han, Haoyong Yu, Jin Chen, Pin Zhang, Jianzhong Di Sep 2020

Timing Of Maximal Weight Reduction Following Bariatric Surgery: A Study In Chinese Patients, Ting Xu, Chen Wang, Hongwei Zhang, Xiaodong Han, Weijie Liu, Junfeng Han, Haoyong Yu, Jin Chen, Pin Zhang, Jianzhong Di

Computer Science Faculty Publications

Introduction: Bariatric surgery is a well-received treatment for obesity with maximal weight loss at 12–36 months postoperatively. We investigated the effect of early bariatric surgery on weight reduction of Chinese patients in accordance with their preoperation characteristics.

Materials and Methods: Altogether, 409 patients with obesity from a prospective cohort in a single bariatric center were enrolled retrospectively and evaluated for up to 4 years. Measurements obtained included surgery type, duration of diabetic condition, besides the usual body mass index data tuple. Weight reduction was expressed as percent total weight loss (%TWL) and percent excess weight loss (%EWL).

Results: RYGB or …


Machine Learning Applications For Drug Repurposing, Hansaim Lim Sep 2020

Machine Learning Applications For Drug Repurposing, Hansaim Lim

Dissertations, Theses, and Capstone Projects

The cost of bringing a drug to market is astounding and the failure rate is intimidating. Drug discovery has been of limited success under the conventional reductionist model of one-drug-one-gene-one-disease paradigm, where a single disease-associated gene is identified and a molecular binder to the specific target is subsequently designed. Under the simplistic paradigm of drug discovery, a drug molecule is assumed to interact only with the intended on-target. However, small molecular drugs often interact with multiple targets, and those off-target interactions are not considered under the conventional paradigm. As a result, drug-induced side effects and adverse reactions are often neglected …


Enrichment Of Ontologies Using Machine Learning And Summarization, Hao Liu Aug 2020

Enrichment Of Ontologies Using Machine Learning And Summarization, Hao Liu

Dissertations

Biomedical ontologies are structured knowledge systems in biomedicine. They play a major role in enabling precise communications in support of healthcare applications, e.g., Electronic Healthcare Records (EHR) systems. Biomedical ontologies are used in many different contexts to facilitate information and knowledge management. The most widely used clinical ontology is the SNOMED CT. Placing a new concept into its proper position in an ontology is a fundamental task in its lifecycle of curation and enrichment.

A large biomedical ontology, which typically consists of many tens of thousands of concepts and relationships, can be viewed as a complex network with concepts as …


Β-Amyloid And Tau Drive Early Alzheimer's Disease Decline While Glucose Hypometabolism Drives Late Decline, Tyler C. Hammond, Xin Xing, Chris Wang, David Ma, Kwangsik Nho, Paul K. Crane, Fanny Elahi, David A. Ziegler, Gongbo Liang, Qiang Cheng, Lucille M. Yanckello, Nathan Jacobs, Ai-Ling Lin Jul 2020

Β-Amyloid And Tau Drive Early Alzheimer's Disease Decline While Glucose Hypometabolism Drives Late Decline, Tyler C. Hammond, Xin Xing, Chris Wang, David Ma, Kwangsik Nho, Paul K. Crane, Fanny Elahi, David A. Ziegler, Gongbo Liang, Qiang Cheng, Lucille M. Yanckello, Nathan Jacobs, Ai-Ling Lin

Sanders-Brown Center on Aging Faculty Publications

Clinical trials focusing on therapeutic candidates that modify β-amyloid (Aβ) have repeatedly failed to treat Alzheimer’s disease (AD), suggesting that Aβ may not be the optimal target for treating AD. The evaluation of Aβ, tau, and neurodegenerative (A/T/N) biomarkers has been proposed for classifying AD. However, it remains unclear whether disturbances in each arm of the A/T/N framework contribute equally throughout the progression of AD. Here, using the random forest machine learning method to analyze participants in the Alzheimer’s Disease Neuroimaging Initiative dataset, we show that A/T/N biomarkers show varying importance in predicting AD development, with elevated biomarkers of Aβ …


Introduction To The R-Package: Usdampr, Elliott James Dennis, Bowen Chen Jun 2020

Introduction To The R-Package: Usdampr, Elliott James Dennis, Bowen Chen

Extension Farm and Ranch Management News

Why the Need for the Package? In the 1990’s, concern over growing packer concentration and a hog industry market shock resulted in discontent among producers and packers. As a result, the United States Congress passed the Livestock Mandatory Reporting Act of 1999 (1999 Act) [Pub. L. 106-78, Title IX] which is required to be reauthorized every five years. See here for a full history of the Livestock Mandatory Reporting Background.

Market reports were publicly issued in the form of .txt files with varying frequency from April 2000 to April 2020. Current and historical data were also housed in a USDA-AMS …


Comparative Analysis Of Metabolic Pathways Of Bacteria Used In Fermented Food, Keanu Hoang, Kiran Bastola May 2020

Comparative Analysis Of Metabolic Pathways Of Bacteria Used In Fermented Food, Keanu Hoang, Kiran Bastola

Theses/Capstones/Creative Projects

This study presents a novel methodology for analyzing metabolic pathways. Utilizing KEGG REST API through a Biopython package and file parser, data about whether or not a bacteria has an enzyme or not was extracted. The results found that differences in metabolic pathway enrichment values follow along the lines of genera and pathway type. In particular, bacteria found in food spoilage and commercial nitrogen fixing products had high values of enrichment.


Cylindrical Similarity Measurement For Helices In Medium-Resolution Cryo-Electron Microscopy Density Maps, Salim Sazzed, Peter Scheible, Maytha Alshammari, Willy Wriggers, Jing He Apr 2020

Cylindrical Similarity Measurement For Helices In Medium-Resolution Cryo-Electron Microscopy Density Maps, Salim Sazzed, Peter Scheible, Maytha Alshammari, Willy Wriggers, Jing He

College of Sciences Posters

Cryo-electron microscopy (cryo-EM) density maps at medium resolution (5-10 Å) reveal secondary structural features such as α-helices and β-sheets, but they lack the side chains details that would enable a direct structure determination. Among the more than 800 entries in the Electron Microscopy Data Bank (EMDB) of medium-resolution density maps that are associated with atomic models, a wide variety of similarities can be observed between maps and models. To validate such atomic models and to classify structural features, a local similarity criterion, the F1 score, is proposed and evaluated in this study. The F1 score is theoretically normalized to a …


Inflammatory Bowel Disease Diagnosis Using Metagenomic Classification, Michael Riggle Apr 2020

Inflammatory Bowel Disease Diagnosis Using Metagenomic Classification, Michael Riggle

Masters Theses & Specialist Projects

Inflammatory bowel disease (IBD) is a set of disorders that involve chronic inflammation of digestive tracts, e.g., Crohn's disease (CD) and ulcerative colitis (UC). Millions of people around the world have inflammatory bowel disease. However, it is still difficult to treat IBD due to its unknown cause. In fact, accurately diagnosing inflammatory bowel disease (IBD) can be very challenging too since some of IBD symptoms can mimic those of other conditions. In this work, we apply classification methods to help improve the success rate of diagnosis. We study four formulations of IBD classification: i) IBD and non-IBD (binary classification), ii) …


Multi-Label Model For Toxicity Prediction, Xiu Huan Yap, Michael L. Raymer Apr 2020

Multi-Label Model For Toxicity Prediction, Xiu Huan Yap, Michael L. Raymer

Celebration of Undergraduate & Graduate Research, Scholarship, and Creative Activities Materials

Most computational predictive models are specifically trained for a single toxicity endpoint. Since more than 1300 toxicity assays have been reported in the TOXCAST dashboard, achieving high coverage over this growing number of toxicity endpoints remains challenging. Furthermore, single-endpoint models lack the ability to learn dependencies between endpoints, such as those targeting similar biological pathways, which may be used to boost model performance. In this study, we characterize the performance of 3 multi-label classification (MLC) models, namely Classifier Chains (CC), Label Powersets (LP) and Stacking (SBR), on Tox21 challenge data. These MLC models employ the Problem Transformation approach, which is …


Statistical Analysis Of Social Network Change, Teresa D. Schmidt Jan 2020

Statistical Analysis Of Social Network Change, Teresa D. Schmidt

Systems Science Friday Noon Seminar Series

We explore two statistical methods that infer social network structures and statistically test those structures for change over time: regression-based differential network analysis (R-DNA) and information theory-based differential network analysis (I-DNA). RDNA is adapted from bioinformatics and I-DNA employs reconstructability analysis. Both methods are used to analyze Medicaid claims data from one-year periods before and after the formation of the Health Share of Oregon Coordinated Care Organization (CCO). We hypothesized that Health Share’s CCO formation would be followed by several changes in the healthcare delivery network.

Application of R-DNA and I-DNA to claims data involves three steps: (a) the inference …


Using Cuda To Enhance Data Processing Of Variant Call Format Files For Statistical Genetic Analysis, Heather Mckinnon Jan 2020

Using Cuda To Enhance Data Processing Of Variant Call Format Files For Statistical Genetic Analysis, Heather Mckinnon

All Graduate Projects

Utilizing the power of GPU parallel processing with CUDA can speed up the processing of Variant Call Format (VCF) files and statistical analysis of genomic data. A software package designed toward this purpose would be beneficial to genetic researchers by saving them time which they could spend on other aspects of their research. A data set containing genetics from a study of trichome production in Mimulus guttatus, or yellow monkey flower, was used to develop a package to test the effectiveness of GPU parallel processing versus serial executions. After a serial version of the code was generated and benchmarked, OpenACC …


Machine Learning Methods For The Analysis Of Metagenomes, Vito Adrian Cantu Alessio Robles Jan 2020

Machine Learning Methods For The Analysis Of Metagenomes, Vito Adrian Cantu Alessio Robles

CGU Theses & Dissertations

As of October 2020, there are 18.6 × 1015 DNA base pairs publicly available in the Sequence Read Archive and this number is growing at an exponential rate. As DNA sequencing prices continue to drop, many research groups around the world have incorporated high throughput sequencing in their research, giving us access to sequences from many distinct ecosystems. This has revolutionized the field of metagenomics, which aims to fully characterize all organisms and their interactions in a particular system. Nevertheless, the plethora of available data has made its analysis difficult as traditional techniques such as genome assembly or sequence alignment …


Representation Learning With Autoencoders For Electronic Health Records, Najibesadat Sadatijafarkalaei Jan 2020

Representation Learning With Autoencoders For Electronic Health Records, Najibesadat Sadatijafarkalaei

Wayne State University Theses

Increasing volume of Electronic Health Records (EHR) in recent years provides great opportunities for data scientists to collaborate on different aspects of healthcare research by applying advanced analytics to these EHR clinical data. A key requirement however

is obtaining meaningful insights from high dimensional, sparse and complex clinical data. Data science approaches typically address this challenge by performing feature learning in order to build more reliable and informative feature representations from clinical data followed by supervised learning. In this research, we propose a predictive modeling approach based on deep feature representations and word embedding techniques. Our method uses different deep …


Classifying Relations Using Recurrent Neural Network With Ontological-Concept Embedding, Mario J. Lorenzo Jan 2020

Classifying Relations Using Recurrent Neural Network With Ontological-Concept Embedding, Mario J. Lorenzo

CCAC Theses and Dissertations

Relation extraction and classification represents a fundamental and challenging aspect of Natural Language Processing (NLP) research which depends on other tasks such as entity detection and word sense disambiguation. Traditional relation extraction methods based on pattern-matching using regular expressions grammars and lexico-syntactic pattern rules suffer from several drawbacks including the labor involved in handcrafting and maintaining large number of rules that are difficult to reuse. Current research has focused on using Neural Networks to help improve the accuracy of relation extraction tasks using a specific type of Recurrent Neural Network (RNN). A promising approach for relation classification uses an RNN …


Repositories For Taxonomic Data: Where We Are And What Is Missing, Aurélian Miralles, Teddy Bruy, Katherine Wolcott, Mark D. Scherz, Dominik Begerow, Bank Beszteri, Michael Bonkowski, Janine Felden, Birgit Gemeinholzer, Frank Glaw, Frank Oliver Glöckner, Oliver Hawlitschek, Ivaylo Kostadinov, Tim W. Nattkemper, Christian Printzen, Jasmin Renz, Nataliya Rybalka, Marc Stadler, Tanja Weibulat, Thomas Wilke, Susanne S. Renner, Miguel Vences Jan 2020

Repositories For Taxonomic Data: Where We Are And What Is Missing, Aurélian Miralles, Teddy Bruy, Katherine Wolcott, Mark D. Scherz, Dominik Begerow, Bank Beszteri, Michael Bonkowski, Janine Felden, Birgit Gemeinholzer, Frank Glaw, Frank Oliver Glöckner, Oliver Hawlitschek, Ivaylo Kostadinov, Tim W. Nattkemper, Christian Printzen, Jasmin Renz, Nataliya Rybalka, Marc Stadler, Tanja Weibulat, Thomas Wilke, Susanne S. Renner, Miguel Vences

Harold W. Manter Laboratory of Parasitology: Library Materials

Natural history collections are leading successful large-scale projects of specimen digitization (images, metadata, DNA barcodes), thereby transforming taxonomy into a big data science. Yet, little effort has been directed towards safeguarding and subsequently mobilizing the considerable amount of original data generated during the process of naming 15,000–20,000 species every year. From the perspective of alpha-taxonomists, we provide a review of the properties and diversity of taxonomic data, assess their volume and use, and establish criteria for optimizing data repositories. We surveyed 4,113 alpha-taxonomic studies in representative journals for 2002, 2010, and 2018, and found an increasing yet comparatively limited use …