Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

University of Kentucky

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 38

Full-Text Articles in Data Science

A Major Update And Improved Validation Functionality In The Mwtab Python Library And The Metabolomics Workbench File Status Website, P. Travis Thompson, Hunter N. B. Moseley Jan 2026

A Major Update And Improved Validation Functionality In The Mwtab Python Library And The Metabolomics Workbench File Status Website, P. Travis Thompson, Hunter N. B. Moseley

Markey Cancer Center Faculty Publications

Background: The Metabolomics Workbench (MW) is a public scientific data repository consisting of experimental data and metadata from metabolomics studies collected with mass spectroscopy (MS) and nuclear magnetic resonance (NMR) analyses. Although not as rapidly as in the past, MW has steadily evolved, updating its mwTab and JSON deposition text file formats and its web-based infrastructure. However, the growth of MW has been exponential since its inception in 2013 and continues to be exponential, with the number of datasets hosted on the repository increasing by 50% since April 2024. As part of regular maintenance to keep up with changes to …


Temporal Machine Learning For Predicting Accidents And Violations In The Mining Industry, Nathan T. Kelley Jan 2025

Temporal Machine Learning For Predicting Accidents And Violations In The Mining Industry, Nathan T. Kelley

Theses and Dissertations--Mining Engineering

This thesis examines the predictive capability of a temporal machine learning model for forecasting future accidents and violations at individual mines, based on historical data. Mine accidents were categorized by accident classification and violations were categorized by the Part Section. The primary datasets utilized were the mine safety and health administration’s (MSHA’s) Accident Injuries and Violations datasets. The available datasets were cleaned and organized by mine type and commodity, then divided into separate subsets for training, validating, and testing. Different models, cutoff metrics, learning rates, number of hidden layers, data processing methods, data processing divisions, number of points observed …


A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci Jan 2024

A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci

Theses and Dissertations--Civil Engineering

Researchers and practitioners studied the effects ride-hailing had in cities before the covid-19 pandemic. Previous research found ride-hailing to produce negative externalities, such as reducing transit ridership and increasing congestion in various cities. Since the pandemic, ride-hailing ridership has nearly recovered to pre-pandemic levels in Chicago. Ride-hailing ridership has grown steadily since the pandemic while a rider’s willingness to share their trip stagnated. Ride-hailing ridership nearly recovering to pre-covid levels in Chicago suggests that transportation planners, and policy makers, will need to continue assessing the impacts ride-hailing trips have in their cities.

Pickup and drop off locations in the Chicago …


Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta Jan 2024

Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta

Theses and Dissertations--Computer Science

End-to-end relation extraction (E2ERE) is a crucial task in natural language processing (NLP) that involves identifying and classifying semantic relationships between entities in text. This thesis compares three paradigms for end-to-end relation extraction (E2ERE) in biomedicine, focusing on rare diseases with discontinuous and nested entities. We evaluate Named Entity Recognition (NER) to Relation Extraction (RE) pipelines, sequence-to-sequence models, and generative pre-trained transformer (GPT) models using the RareDis information extraction dataset. Our findings indicate that pipeline models are the most effective, followed closely by sequence-to-sequence models. GPT models, despite having eight times as many parameters, perform worse than sequence-to-sequence models and …


Molecular Understanding And Design Of Deep Eutectic Solvents And Proteins Using Computer Simulations And Machine Learning, Usman Lame Abbas Jan 2024

Molecular Understanding And Design Of Deep Eutectic Solvents And Proteins Using Computer Simulations And Machine Learning, Usman Lame Abbas

Theses and Dissertations--Chemical and Materials Engineering

Hydrophobic deep eutectic solvents (DESs) have emerged as excellent extractants. A major challenge is the lack of an efficient tool to discover DES candidates. Currently, the search relies heavily on the researchers’ intuition or a trial-and-error process, which leads to a low success rate or bypassing of promising candidates. DES performance depends on the heterogeneous hydrogen bond environment formed by multiple hydrogen bond donors and acceptors. Understanding this heterogeneous hydrogen bond environment can help develop principles for designing high performance DESs for extraction and other separation applications. This work investigates the structure and dynamics of hydrogen bonds in hydrophobic DESs …


Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian Jan 2023

Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian

Theses and Dissertations--Computer Science

As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally - using only a measure of task performance as feedback--can violate societal norms for acceptable behavior or cause harm. Consequently, it becomes necessary to prioritize task performance and ensure that AI actions do not have detrimental effects. Value alignment is a property of intelligent agents, wherein they solely pursue goals and activities that are non-harmful and beneficial to humans. Current approaches to value alignment largely depend on imitation learning or learning from demonstration methods. However, the dynamic nature …


Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan Jan 2023

Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan

Theses and Dissertations--Statistics

Neural networks have experienced widespread adoption and have become integral in cutting-edge domains like computer vision, natural language processing, and various contemporary fields. However, addressing the statistical aspects of neural networks has been a persistent challenge, with limited satisfactory results. In my research, I focused on exploring statistical intervals applied to neural networks, specifically confidence intervals and tolerance intervals. I employed variance estimation methods, such as direct estimation and resampling, to assess neural networks and their performance under outlier scenarios. Remarkably, when outliers were present, the resampling method with infinitesimal jackknife estimation yielded confidence intervals that closely aligned with nominal …


High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang Jan 2023

High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang

Theses and Dissertations--Statistics

This dissertation focuses on the problem of high dimensional data analysis, which arises in many fields including genomics, finance, and social sciences. In such settings, the number of features or variables is much larger than the number of observations, posing significant challenges to traditional statistical methods.

To address these challenges, this dissertation proposes novel methods for variable screening and inference. The first part of the dissertation focuses on variable screening, which aims to identify a subset of important variables that are strongly associated with the response variable. Specifically, we propose a robust nonparametric screening method to effectively select the predictors …


Clustering Hospital Performance Using Group-Based Multi-Trajectory Modeling With Singular Bayesian Information Criterion, Gaixin Du Jan 2023

Clustering Hospital Performance Using Group-Based Multi-Trajectory Modeling With Singular Bayesian Information Criterion, Gaixin Du

Theses and Dissertations--Epidemiology and Biostatistics

Hospital performance is complex and patient-experience oriented. Currently, the Centers for Medicare and Medicaid Services (CMS) evaluate hospitals yearly with a single score of one to five ("Star Rating") using composite measures from five domains. However, a single composite score cannot fully describe it, and alternative measures should be considered. Healthcare quality improvement needs long-term data to validate effectiveness. Group-based multi-trajectory modeling (GBMTM) estimates probabilities of latent group membership based on longitudinal profiles from multiple outcomes. We use GBMTM to identify groups of hospitals with similar performance in SAS PROC TRAJ.

We downloaded Medicare-eligible hospitals (N=5,111) that provided patient care …


The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson Jan 2023

The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson

Theses and Dissertations--Computer Science

We introduce a novel approach for learning behaviors using human-provided feedback that is subject to systematic bias. Our method, known as BASIL, models the feedback signal as a combination of a heuristic evaluation of an action's utility and a probabilistically-drawn bias value, characterized by unknown parameters. We present both the general framework for our technique and specific algorithms for biases drawn from a normal distribution. We evaluate our approach across various environments and tasks, comparing it to interactive and non-interactive machine learning methods, including deep learning techniques, using human trainers and a synthetic oracle with feedback distorted to varying degrees. …


The Performance Optimization Of Asp Solving Based On Encoding Rewriting And Encoding Selection, Liu Liu Jan 2022

The Performance Optimization Of Asp Solving Based On Encoding Rewriting And Encoding Selection, Liu Liu

Theses and Dissertations--Computer Science

Answer set programming (ASP) has long been used for modeling and solving hard search problems. These problems are modeled in ASP as encodings, a collection of rules that declaratively describe the logic of the problem without explicitly listing how to solve it. It is common that the same problem has several different but equivalent encodings in ASP. Experience shows that the performance of these ASP encodings may vary greatly from instance to instance when processed by current state-of-the-art ASP grounder/solver systems. In particular, it is rarely the case that one encoding outperforms all others. Moreover, running an ASP system on …


Batch Normalization Preconditioning For Neural Network Training, Susanna Luisa Gertrude Lange Jan 2022

Batch Normalization Preconditioning For Neural Network Training, Susanna Luisa Gertrude Lange

Theses and Dissertations--Mathematics

Batch normalization (BN) is a popular and ubiquitous method in deep learning that has been shown to decrease training time and improve generalization performance of neural networks. Despite its success, BN is not theoretically well understood. It is not suitable for use with very small mini-batch sizes or online learning. In this work, we propose a new method called Batch Normalization Preconditioning (BNP). Instead of applying normalization explicitly through a batch normalization layer as is done in BN, BNP applies normalization by conditioning the parameter gradients directly during training. This is designed to improve the Hessian matrix of the loss …


Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su Jan 2022

Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su

Theses and Dissertations--Statistics

When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …


Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton Jan 2022

Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton

Theses and Dissertations--Epidemiology and Biostatistics

In the United States, the prevalence of long-term exposure to opioid drugs, for both medically and nonmedically indicated purposes, has increased considerably since the mid-1990’s. Concerns have emerged about the potential health effects of opioid use. There is also growing interest in other possible connections with opioid use including cardiovascular disease. Electronic health records (EHR) contain information about patient care in the form of structured codes and unstructured notes. Natural language processing (NLP) provides a tool for processing unstructured textual data in EHR clinical notes and extracts useful information for research with structured formats. The purpose of this dissertation was …


Exploration Of Lignin-Based Superabsorbent Polymers (Hydrogels) For Soil Water Management And As A Carrier For Delivering Rhizobium Spp., Toby Adjuik Jan 2022

Exploration Of Lignin-Based Superabsorbent Polymers (Hydrogels) For Soil Water Management And As A Carrier For Delivering Rhizobium Spp., Toby Adjuik

Theses and Dissertations--Biosystems and Agricultural Engineering

Superabsorbent polymers (hydrogels) as soil amendments may improve soil hydraulic properties and act as carrier materials beneficial to soil microorganisms. Researchers have mostly explored synthetic hydrogels which may not be environmentally sustainable. This dissertation focused on the development and application of lignin-based hydrogels as sustainable soil amendments. This dissertation also explores the development of pedotransfer transfer functions (PTFs) for predicting saturated hydraulic conductivity using statistical and machine learning methods with a publicly available large data set. A lignin-based hydrogel was synthesized, and its impact on soil water retention was determined in silt loam and loamy fine sand soils. Hydrogel treatment …


Characterizing Long Covid: Deep Phenotype Of A Complex Condition, Rachel R. Deer, Madeline A. Rock, Nicole Vasilevsky, Leigh Carmody, Halie Rando, Alfred J. Anzalone, Marc D. Basson, Tellen D. Bennett, Timothy Bergquist, Eilis A. Boudreau, Carolyn T. Bramante, James Brian Byrd, Tiffany J. Callahan, Lauren E. Chan, Haitao Chu, Christopher G. Chute, Ben D. Coleman, Hannah E. Davis, Joel Gagnier, Casey S. Greene, Ramakanth Kavuluru Nov 2021

Characterizing Long Covid: Deep Phenotype Of A Complex Condition, Rachel R. Deer, Madeline A. Rock, Nicole Vasilevsky, Leigh Carmody, Halie Rando, Alfred J. Anzalone, Marc D. Basson, Tellen D. Bennett, Timothy Bergquist, Eilis A. Boudreau, Carolyn T. Bramante, James Brian Byrd, Tiffany J. Callahan, Lauren E. Chan, Haitao Chu, Christopher G. Chute, Ben D. Coleman, Hannah E. Davis, Joel Gagnier, Casey S. Greene, Ramakanth Kavuluru

Institute for Biomedical Informatics Faculty Publications

BACKGROUND: Numerous publications describe the clinical manifestations of post-acute sequelae of SARS-CoV-2 (PASC or "long COVID"), but they are difficult to integrate because of heterogeneous methods and the lack of a standard for denoting the many phenotypic manifestations. Patient-led studies are of particular importance for understanding the natural history of COVID-19, but integration is hampered because they often use different terms to describe the same symptom or condition. This significant disparity in patient versus clinical characterization motivated the proposed ontological approach to specifying manifestations, which will improve capture and integration of future long COVID studies.

METHODS: The Human Phenotype Ontology …


Awegnn: Auto-Parametrized Weighted Element-Specific Graph Neural Networks For Molecules., Timothy Szocinski, Duc Duy Nguyen, Guo-Wei Wei Jul 2021

Awegnn: Auto-Parametrized Weighted Element-Specific Graph Neural Networks For Molecules., Timothy Szocinski, Duc Duy Nguyen, Guo-Wei Wei

Mathematics Faculty Publications

While automated feature extraction has had tremendous success in many deep learning algorithms for image analysis and natural language processing, it does not work well for data involving complex internal structures, such as molecules. Data representations via advanced mathematics, including algebraic topology, differential geometry, and graph theory, have demonstrated superiority in a variety of biomolecular applications, however, their performance is often dependent on manual parametrization. This work introduces the auto-parametrized weighted element-specific graph neural network, dubbed AweGNN, to overcome the obstacle of this tedious parametrization process while also being a suitable technique for automated feature extraction on these internally complex …


Big Data: Ethics, Resources, And Potential Collaboration, Matthew Zook Feb 2021

Big Data: Ethics, Resources, And Potential Collaboration, Matthew Zook

Geography Presentations

This presentation goes over 10 simple rules for responsible big data research.


Challenges When Identifying Migration From Geo-Located Twitter Data, Caitrin Armstrong, Ate Poorthuis, Matthew Zook, Derek Ruths, Thomas Soehl Jan 2021

Challenges When Identifying Migration From Geo-Located Twitter Data, Caitrin Armstrong, Ate Poorthuis, Matthew Zook, Derek Ruths, Thomas Soehl

Geography Faculty Publications

Given the challenges in collecting up-to-date, comparable data on migrant populations the potential of digital trace data to study migration and migrants has sparked considerable interest among researchers and policy makers. In this paper we assess the reliability of one such data source that is heavily used within the research community: geolocated tweets. We assess strategies used in previous work to identify migrants based on their geolocation histories. We apply these approaches to infer the travel history of a set of Twitter users who regularly posted geolocated tweets between July 2012 and June 2015. In a second step we hand-code …


Machine Learning And Bioinformatic Insights Into Key Enzymes For A Bio-Based Circular Economy, Japheth E. Gado Jan 2021

Machine Learning And Bioinformatic Insights Into Key Enzymes For A Bio-Based Circular Economy, Japheth E. Gado

Theses and Dissertations--Chemical and Materials Engineering

The world is presently faced with a sustainability crisis; it is becoming increasingly difficult to meet the energy and material needs of a growing global population without depleting and polluting our planet. Greenhouse gases released from the continuous combustion of fossil fuels engender accelerated climate change, and plastic waste accumulates in the environment. There is need for a circular economy, where energy and materials are renewably derived from waste items, rather than by consuming limited resources. Deconstruction of the recalcitrant linkages in natural and synthetic polymers is crucial for a circular economy, as deconstructed monomers can be used to manufacture …


Multi-Stream Longitudinal Data Analysis Using Deep Learning, Sajjad Fouladvand Jan 2021

Multi-Stream Longitudinal Data Analysis Using Deep Learning, Sajjad Fouladvand

Theses and Dissertations--Computer Science

Longitudinal healthcare data encompasses all tasks where patients information are collected at multiple follow-up times. Analyzing this data is critical in addressing many real world problems in healthcare such as disease prediction and prevention. In this thesis, technical challenges in analyzing longitudinal administrative claims data are addressed and novel deep learning based models are proposed for multi-stream data analysis and disease prediction tasks. These algorithms and frameworks are assessed mainly on substance use disorders prediction tasks and specifically designed to tackled these disorders. Substance use disorder is a public health crisis costing the US an estimated $740 billion annually in …


Dimension Reduction Techniques In Regression, Pei Wang Jan 2021

Dimension Reduction Techniques In Regression, Pei Wang

Theses and Dissertations--Statistics

Because of the advances of modern technology, the size of the collected data nowadays is larger and the structure is more complex. To deal with such kinds of data, sufficient dimension reduction (SDR) and reduced rank (RR) regression are two powerful tools. This dissertation focuses on these two tools and it is composed of three projects. In the first project, we introduce a new SDR method through a novel approach of feature filter to recover the central mean subspace exhaustively along with a method to determine the dimension, two variable selection methods, and extensions to multivariate response and large p …


Neural Representations Of Concepts And Texts For Biomedical Information Retrieval, Jiho Noh Jan 2021

Neural Representations Of Concepts And Texts For Biomedical Information Retrieval, Jiho Noh

Theses and Dissertations--Computer Science

Information retrieval (IR) methods are an indispensable tool in the current landscape of exponentially increasing textual data, especially on the Web. A typical IR task involves fetching and ranking a set of documents (from a large corpus) in terms of relevance to a user's query, which is often expressed as a short phrase. IR methods are the backbone of modern search engines where additional system-level aspects including fault tolerance, scale, user interfaces, and session maintenance are also addressed. In addition to fetching documents, modern search systems may also identify snippets within the documents that are potentially most relevant to the …


Revisiting Absolute Pose Regression, Hunter Blanton Jan 2021

Revisiting Absolute Pose Regression, Hunter Blanton

Theses and Dissertations--Computer Science

Images provide direct evidence for the position and orientation of the camera in space, known as camera pose. Traditionally, the problem of estimating the camera pose requires reference data for determining image correspondence and leveraging geometric relationships between features in the image. Recent advances in deep learning have led to a new class of methods that regress the pose directly from a single image.

This thesis proposes methods for absolute camera pose regression. Absolute pose regression estimates the pose of a camera from a single image as the output of a fixed computation pipeline. These methods have many practical benefits …


Viral Data, Agnieszka Leszczynski, Matthew Zook Nov 2020

Viral Data, Agnieszka Leszczynski, Matthew Zook

Geography Faculty Publications

We are experiencing a historical moment characterized by unprecedented conditions of virality: a viral pandemic, the viral diffusion of misinformation and conspiracy theories, the viral momentum of ongoing Hong Kong protests, and the viral spread of #BlackLivesMatter demonstrations and related efforts to defund policing. These co-articulations of crises, traumas, and virality both implicate and are implicated by big data practices occurring in a present that is pervasively mediated by data materialities, deeply rooted dataist ideologies that entrench processes of datafication as granting objective access to truth and attendant practices of tracking, data analytics, algorithmic prediction, and data-driven targeting of individuals …


A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. Mackenzie, Christopher I. Amos, Jiang Gui Oct 2020

A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. Mackenzie, Christopher I. Amos, Jiang Gui

Markey Cancer Center Faculty Publications

BACKGROUND: Genome-wide association studies (GWAS) have proven successful in predicting genetic risk of disease using single-locus models; however, identifying single nucleotide polymorphism (SNP) interactions at the genome-wide scale is limited due to computational and statistical challenges. We addressed the computational burden encountered when detecting SNP interactions for survival analysis, such as age of disease-onset. To confront this problem, we developed a novel algorithm, called the Efficient Survival Multifactor Dimensionality Reduction (ES-MDR) method, which used Martingale Residuals as the outcome parameter to estimate survival outcomes, and implemented the Quantitative Multifactor Dimensionality Reduction method to identify significant interactions associated with age of …


Imaging Data On Characterization Of Retinal Autofluorescent Lesions In A Mouse Model Of Juvenile Neuronal Ceroid Lipofuscinosis (Cln3 Disease), Qing Jun Wang, Kyung Sik Jung, Kabhilan Mohan, Mark E. Kleinman Oct 2020

Imaging Data On Characterization Of Retinal Autofluorescent Lesions In A Mouse Model Of Juvenile Neuronal Ceroid Lipofuscinosis (Cln3 Disease), Qing Jun Wang, Kyung Sik Jung, Kabhilan Mohan, Mark E. Kleinman

Ophthalmology and Visual Science Faculty Publications

Juvenile neuronal ceroid lipofuscinosis (JNCL, aka. juvenile Batten disease or CLN3 disease), a lethal pediatric neurodegenerative disease without cure, often presents with vision impairment and characteristic ophthalmoscopic features including focal areas of hyper-autofluorescence. In the associated research article “Loss of CLN3, the gene mutated in juvenile neuronal ceroid lipofuscinosis, leads to metabolic impairment and autophagy induction in retinal pigment epithelium” (Zhong et al., 2020) [1], we reported ophthalmoscopic observations of focal autofluorescent lesions or puncta in the Cln3Δex7/8 mouse retina at as young as 8 month old. In this data article, we performed differential interference contrast and …


A Tree Frog (Boana Pugnax) Dataset Of Skin Transcriptome For The Identification Of Biomolecules With Potential Antimicrobial Activities, Yamil Liscano Martinez, Claudia Marcela Arenas Gómez, Jeramiah J. Smith, Jean Paul Delgado Oct 2020

A Tree Frog (Boana Pugnax) Dataset Of Skin Transcriptome For The Identification Of Biomolecules With Potential Antimicrobial Activities, Yamil Liscano Martinez, Claudia Marcela Arenas Gómez, Jeramiah J. Smith, Jean Paul Delgado

Biology Faculty Publications

Increases in the prevalence of multiply resistant microbes have necessitated the search for new molecules with antimicrobial properties. One noteworthy avenue in this search is inspired by the presence of native antimicrobial peptides in the skin of amphibians. Having the second highest diversity of frogs worldwide, Colombian anurans represent an extensive natural reservoir that could be tapped in this search. Among this diversity, species such as Boana pugnax (the Chirique-Flusse Treefrog) are particularly notable, in that they thrive in a diversity of marginal habitats, utilize both aquatic and arboreal habitats, and are members of one of few genera that are …


Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro Aug 2020

Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro

Horticulture Faculty Publications

This dataset presents shotgun metagenomic sequencing of sunflower rhizosphere microbiome in Bloemhof, South Africa. Data were collected to decipher the structure and function in the sunflower microbial community. Illumina HiSeq platform using next generation sequencing of the DNA was carried out. The metagenome comprised 8,991,566 sequences totaling 1,607,022,279 bp size and 66% GC content. The metagenome was deposited into the NCBI database and can be accessed with the SRA accession number SRR10418054. An online metagenome server (MG RAST) using the subsystem database revealed bacteria had the highest taxonomical representation with 98.47%, eukaryote at 1.23%, and archaea at 0.20%. The most …


Covid-19 Is Spatial: Ensuring That Mobile Big Data Is Used For Social Good, Age Poom, Olle Järv, Matthew Zook, Tuuli Toivonen Jul 2020

Covid-19 Is Spatial: Ensuring That Mobile Big Data Is Used For Social Good, Age Poom, Olle Järv, Matthew Zook, Tuuli Toivonen

Geography Faculty Publications

The mobility restrictions related to COVID-19 pandemic have resulted in the biggest disruption to individual mobilities in modern times. The crisis is clearly spatial in nature, and examining the geographical aspect is important in understanding the broad implications of the pandemic. The avalanche of mobile Big Data makes it possible to study the spatial effects of the crisis with spatiotemporal detail at the national and global scales. However, the current crisis also highlights serious limitations in the readiness to take the advantage of mobile Big Data for social good, both within and beyond the interests of health sector. We propose …