Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Old Dominion University (7)
- Harrisburg University of Science and Technology (5)
- University of Nebraska - Lincoln (5)
- California Polytechnic State University, San Luis Obispo (2)
- Missouri University of Science and Technology (2)
-
- Purdue University (2)
- University of Kentucky (2)
- Boise State University (1)
- Clemson University (1)
- Cleveland State University (1)
- Duquesne University (1)
- Louisiana State University (1)
- Michigan Technological University (1)
- The University of Southern Mississippi (1)
- University of Arkansas, Fayetteville (1)
- University of Nevada, Las Vegas (1)
- Virginia Commonwealth University (1)
- Washington University in St. Louis (1)
- Western Michigan University (1)
- Keyword
-
- Genetics (5)
- Bioinformatics (3)
- Hybridization (3)
- Computational biology (2)
- Evolution (2)
-
- Machine learning (2)
- Molecules (2)
- ADHD (1)
- AI (1)
- Academic subjects (1)
- Acetylation in metabolism (1)
- Adipogenesis (1)
- Agent-based modeling (1)
- Aggregation (1)
- Alzheimer (1)
- Amino acid (1)
- Anabaena (1)
- Applied computing (1)
- Artificial intelligence (1)
- Autoencoder (1)
- Auxiliary metabolic genes (1)
- Axis (1)
- Basal subtype (1)
- Bio-medical science (1)
- Biogeography-based optimization (1)
- Biomarkers (1)
- Bioremediation (1)
- Blastn (1)
- Breast Cancer (1)
- Breast cancer (1)
- Publication Year
- Publication
-
- Computer Science Faculty Publications (6)
- Faculty Works (5)
- Engineering Management and Systems Engineering Faculty Research & Creative Works (2)
- School of Computing: Conference and Workshop Papers (2)
- All Dissertations (1)
-
- Biological Sciences Faculty Publications (1)
- Biomedical Engineering Undergraduate Honors Theses (1)
- Boise State University Theses and Dissertations (1)
- Civil and Environmental Engineering and Construction Faculty Research (1)
- Department of Computer Electronics and Engineering: Dissertations, Theses, and Student Research (1)
- Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research (1)
- Dissertations (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Electrical and Computer Engineering Faculty Publications (1)
- Electronic Theses and Dissertations (1)
- Graduate Industrial Research Symposium (1)
- LSU Doctoral Dissertations (1)
- MODVIS Workshop (1)
- Master's Theses (1)
- McKelvey School of Engineering Graduate Student Theses & Dissertations (1)
- Parallel Computing and Data Science Lab Technical Reports (1)
- Plant Pathology Faculty Publications (1)
- STAR Program Research Presentations (1)
- School of Computing: Dissertations, Theses, and Student Research (1)
- Theses and Dissertations (1)
- Theses and Dissertations--Plant and Soil Sciences (1)
- Publication Type
Articles 1 - 30 of 37
Full-Text Articles in Computational Biology
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Parkinson's disease (PD) is a complex neurodegenerative disorder with a significant genetic component. While genome-wide association studies (GWAS) have been instrumental in identifying genetic variants associated with PD, the reliance on large sample sizes and population-level analyses may overlook variants with lower minor allele frequencies or individual-specific relevance. Individualized Bayesian Inference (IBI) offers a promising method to complement GWAS by identifying and prioritizing candidate genetic markers at both the individual and patients-like-me subgroup levels. This study evaluates the application of IBI to PD genetics, using GWAS as a baseline for comparison. We analyzed genetic data from the Fox Insight online …
Large Scale Kmer-Based Proteomic Analysis: An Application Towards Evolutionary Constraint Discovery, Matthew Chak
Large Scale Kmer-Based Proteomic Analysis: An Application Towards Evolutionary Constraint Discovery, Matthew Chak
Master's Theses
Large protein databases now make it possible to study short peptides across natural protein sequence space at unprecedented scale, but exhaustively counting k-mers across billions of protein sequences remains computationally difficult. This thesis develops an exact amino-acid k-mer counting method based on direct addressing, in which fixed-length amino-acid strings are encoded as base-20 integers and updated with a sliding-window recurrence. By avoiding key storage and collision resolution, this approach removes overhead inherent to hash-map-based methods when the k-mer space is sufficiently dense. A memory analysis shows when direct addressing is preferable to open-addressing hash tables, and expected-saturation calculations motivate its …
Privacy-Preserving Federated Learning With Optimized Ensemble Weighting And Knowledge Distillation For Covid-19 Detection From Non-Iid Medical Imaging Data, Richard Annan, Hong Qin, Robert Newman, Madhuri Siddula, Letu Qingge
Privacy-Preserving Federated Learning With Optimized Ensemble Weighting And Knowledge Distillation For Covid-19 Detection From Non-Iid Medical Imaging Data, Richard Annan, Hong Qin, Robert Newman, Madhuri Siddula, Letu Qingge
Computer Science Faculty Publications
Medical imaging enables rapid and accurate diagnosis of COVID-19, with CT scans proving especially effective. However, data privacy concerns limit collaborative model development across hospitals. To address this issue, we introduce a novel federated learning framework. It is referred to as Independent Knowledge Distillation with post-Ensemble Federated Learning (IKDEFL). Differential Privacy (DP) is integrated into the framework to improve privacy guarantees. Three DP mechanisms are evaluated. These include Fixed Gaussian, Gaussian Adaptive, and Tree Adaptive. The evaluation has been conducted on heterogeneous and Non-Independent and Identically Distributed (Non-IID) datasets. These datasets reflect real-world hospital scenarios. Results show that IKDEFL significantly …
Computational And Ai Frameworks For Identifying Key Regulatory Genes And Their Target Genes In Plants And Humans, Md Khairul Islam
Computational And Ai Frameworks For Identifying Key Regulatory Genes And Their Target Genes In Plants And Humans, Md Khairul Islam
Dissertations, Master's Theses and Master's Reports
This dissertation presents computational and AI-driven frameworks for identifying key regulatory genes and their downstream targets across plant and human biological systems. Three studies address distinct challenges in genomic regulation using advanced machine learning and bioinformatics approaches.
The first study introduces DyGAF (Dynamic Gene Attention Focus), a dual-attention transformer framework that identifies and ranks disease-relevant biomarker genes by simultaneously modeling independent molecular responses and interdependent regulatory network behavior. Two attention models provide complementary perspectives on gene importance and are fused through a novel combination metric. Applied to COVID-19 nasopharyngeal swab profiles, the attention-weighted representations achieved 94.23% classification accuracy, high sensitivity, …
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Cancer is mainly caused by a relatively small portion of somatic genome alterations (SGAs), called cancer drivers. Despite success in identifying a good number of cancer drivers, many more remain to be discovered to explain various cancers. Moreover, limited tools are available to identify potential interactions among cancer drivers for a better understanding of oncogenesis. To tackle these challenges, we have developed a novel approach called individualized Bayesian inference using a decision tree (IBI-DT). IBI-DT recognizes the genetic heterogeneity among cancer patients, where different individuals or patient subgroups of distinct genomic makeup may have different drivers. IBI-DT works by constructing …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Hidden Markov Model For Identifying Local Variants In Human Genomes Using Simulated Data, Scott Mccallum
Hidden Markov Model For Identifying Local Variants In Human Genomes Using Simulated Data, Scott Mccallum
Electronic Theses and Dissertations
Identifying adaptive mutations in genetic data is challenging due to the low frequency of occurrence of such events, and because signatures of selection are intertwined with the footprints of various other evolutionary forces that shape our genomes. Even when a larger region appears to be under selection, genomic sites that are linked to adaptive mutations have similar statistical signals, and thus can obfuscate the identification of the actual adaptive mutation. The new method described here uses a Hidden Markov Model that allows for classification of neutral, linked, and sweep (adaptive mutation) genomic sites. This model is general and can be …
Transcriptomic Profiling Of Engineered Human Gene Integration In The Mouse Genome Following In Vitro Gene Editing, Ethan Potts, Made Harumi Padmaswari, Christopher Nelson
Transcriptomic Profiling Of Engineered Human Gene Integration In The Mouse Genome Following In Vitro Gene Editing, Ethan Potts, Made Harumi Padmaswari, Christopher Nelson
Biomedical Engineering Undergraduate Honors Theses
Gene replacement is a promising method of therapy for genetic diseases. However, safety and efficacy are areas that need more research. This experiment aims to use RNA sequencing and bioinformatic techniques to provide answers to these questions and provide direction to future studies to develop a gene replacement therapeutic. C2C12 mouse myoblast cells were transfected with a vector containing a CRISPR-Cas9 system and the Human Factor IX (hF9) gene in order to hijack target genes and integrate the hF9 gene. The two target genes, myoglobin (Mb) and creatine kinase (Ckm), were chosen for their high rate of expression and low …
A Machine Learning Model Of Perturb-Seq Data For Use In Space Flight Gene Expression Profile Analysis, Liam F. Johnson, James Casaletto, Lauren Sanders, Sylvain Costes
A Machine Learning Model Of Perturb-Seq Data For Use In Space Flight Gene Expression Profile Analysis, Liam F. Johnson, James Casaletto, Lauren Sanders, Sylvain Costes
Graduate Industrial Research Symposium
The genetic perturbations caused by spaceflight on biological systems tend to have a system-wide effect which is often difficult to deconvolute it into individual signals with specific points of origin. Single cell multi-omic data can provide a profile of the perturbational effects, but does not necessarily indicate the initial point of interference within the network. The objective of this project is to take advantage of large scale and genome-wide perturbational datasets by using them to train a tuned machine learning model that is capable of predicting the effects of unseen perturbations in new data. Perturb-Seq datasets are large libraries of …
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Biological Sciences Faculty Publications
Chemical risk assessment plays a pivotal role in safeguarding public health and environmental safety by evaluating the potential hazards and risks associated with chemical exposures. In recent years, the convergence of artificial intelligence (AI), machine learning (ML), and omics technologies has revolutionized the field of chemical risk assessment, offering new insights into toxicity mechanisms, predictive modeling, and risk management strategies. This perspective review explores the synergistic potential of AI/ML and omics in deciphering clastogen-induced genomic instability for carcinogenic risk prediction. We provide an overview of key findings, challenges, and opportunities in integrating AI/ML and omics technologies for chemical risk assessment, …
Missing Value Imputation For Single Omics And Multi-Omics Data, Meng Song
Missing Value Imputation For Single Omics And Multi-Omics Data, Meng Song
Dissertations
The integration analyses of multi-omics data have the advantages of extending our understanding of biological system across multiple omics layers, unraveling the functional mechanism of complex disease development, and refining the discovery of novel drug targets. However, multi-omics studies often face challenges such as data heterogeneity, missing values problem, interpretability, and imbalance classes. Among these challenges, the missing values problem is a critical issue for large cohort studies as not all samples will get a complete measurement for all the omics layers. To address the problem of missing values in multi-omics data, I focused on the imputation of completely missing …
A Single-Cell Atlas Of Bovine Skeletal Muscle Reveals Mechanisms Regulating Intramuscular Adipogenesis And Fibrogenesis, Leshan Wang, Peidong Gao, Chaoyang Li, Qianglin Liu, Zeyang Yao, Yuxia Li, Xujia Zhang, Jiangwen Sun, Constantine Simintiras, Matthew Welborn, Kenneth Mcmillin, Stephanie Oprescu, Shihuan Kuang, Xing Fu
A Single-Cell Atlas Of Bovine Skeletal Muscle Reveals Mechanisms Regulating Intramuscular Adipogenesis And Fibrogenesis, Leshan Wang, Peidong Gao, Chaoyang Li, Qianglin Liu, Zeyang Yao, Yuxia Li, Xujia Zhang, Jiangwen Sun, Constantine Simintiras, Matthew Welborn, Kenneth Mcmillin, Stephanie Oprescu, Shihuan Kuang, Xing Fu
Computer Science Faculty Publications
Background
Intramuscular fat (IMF) and intramuscular connective tissue (IMC) are often seen in human myopathies and are central to beef quality. The mechanisms regulating their accumulation remain poorly understood. Here, we explored the possibility of using beef cattle as a novel model for mechanistic studies of intramuscular adipogenesis and fibrogenesis.
Methods
Skeletal muscle single-cell RNAseq was performed on three cattle breeds, including Wagyu (high IMF), Brahman (abundant IMC but scarce IMF), and Wagyu/Brahman cross. Sophisticated bioinformatics analyses, including clustering analysis, gene set enrichment analyses, gene regulatory network construction, RNA velocity, pseudotime analysis, and cell-cell communication analysis, were performed to elucidate …
Dfhic: A Dilated Full Convolution Model To Enhance The Resolution Of Hi-C Data, Bin Wang, Kun Liu, Yaohang Li, Jianxin Wang
Dfhic: A Dilated Full Convolution Model To Enhance The Resolution Of Hi-C Data, Bin Wang, Kun Liu, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation: Hi-C technology has been the most widely used chromosome conformation capture(3C) experiment that measures the frequency of all paired interactions in the entire genome, which is a powerful tool for studying the 3D structure of the genome. The fineness of the constructed genome structure depends on the resolution of Hi-C data. However, due to the fact that high-resolution Hi-C data require deep sequencing and thus high experimental cost, most available Hi-C data are in low-resolution. Hence, it is essential to enhance the quality of Hi-C data by developing the effective computational methods.
Results: In this work, we propose …
Sequence-Based Bioinformatics Approaches To Predict Virus–Host Relationships In Archaea And Eukaryotes, Yingshan Li
Sequence-Based Bioinformatics Approaches To Predict Virus–Host Relationships In Archaea And Eukaryotes, Yingshan Li
School of Computing: Dissertations, Theses, and Student Research
Viral metagenomics is independent of lab culturing and capable of investigating viromes of virtually any given environmental niches. While numerous sequences of viral genomes have been assembled from metagenomic studies over the past years, the natural hosts for the majority of these viral contigs have not been determined. Different computational approaches have been developed to predict hosts of bacteria phages. Nevertheless, little progress has been made in the virus-host prediction, especially for viruses that infect eukaryotes and archaea. In this study, by analyzing all documented viruses with known eukaryotic and archaeal hosts, we assessed the predictive power of four computational …
Large Genomes Assembly Using Mapreduce Framework, Yuehua Zhang
Large Genomes Assembly Using Mapreduce Framework, Yuehua Zhang
All Dissertations
Knowing the genome sequence of an organism is the essential step toward understanding its genomic and genetic characteristics. Currently, whole genome shotgun (WGS) sequencing is the most widely used genome sequencing technique to determine the entire DNA sequence of an organism. Recent advances in next-generation sequencing (NGS) techniques have enabled biologists to generate large DNA sequences in a high-throughput and low-cost way. However, the assembly of NGS reads faces significant challenges due to short reads and an enormously high volume of data. Despite recent progress in genome assembly, current NGS assemblers cannot generate high-quality results or efficiently handle large genomes …
The Low Abundance Of Cpg In The Sars-Cov-2 Genome Is Not An Evolutionarily Signature Of Zap, Ali Afrasiabi, Hamid Alinejad-Rokny, Azad Khosh, Mostafa Rahnama, Nigel Lovell, Zhenming Xu, Diako Ebrahimi
The Low Abundance Of Cpg In The Sars-Cov-2 Genome Is Not An Evolutionarily Signature Of Zap, Ali Afrasiabi, Hamid Alinejad-Rokny, Azad Khosh, Mostafa Rahnama, Nigel Lovell, Zhenming Xu, Diako Ebrahimi
Plant Pathology Faculty Publications
The zinc finger antiviral protein (ZAP) is known to restrict viral replication by binding to the CpG rich regions of viral RNA, and subsequently inducing viral RNA degradation. This enzyme has recently been shown to be capable of restricting SARS-CoV-2. These data have led to the hypothesis that the low abundance of CpG in the SARS-CoV-2 genome is due to an evolutionary pressure exerted by the host ZAP. To investigate this hypothesis, we performed a detailed analysis of many coronavirus sequences and ZAP RNA binding preference data. Our analyses showed neither evidence for an evolutionary pressure acting specifically on CpG …
A Novel Jumbo Phage Phima05 Inhibits Harmful Microcystis Sp., Ampapan Naknaen, Oramas Suttinun, Komwit Surachat, Eakalak Khan, Rattanaruji Pomwised
A Novel Jumbo Phage Phima05 Inhibits Harmful Microcystis Sp., Ampapan Naknaen, Oramas Suttinun, Komwit Surachat, Eakalak Khan, Rattanaruji Pomwised
Civil and Environmental Engineering and Construction Faculty Research
Microcystis poses a concern because of its potential contribution to eutrophication and production of microcystins (MCs). Phage treatment has been proposed as a novel biocontrol method for Microcystis. Here, we isolated a lytic cyanophage named PhiMa05 with high efficiency against MCs-producing Microcystis strains. Its burst size was large, with approximately 127 phage particles/infected cell, a short latent period (1 day), and high stability to broad salinity, pH and temperature ranges. The PhiMa05 structure was composed of an icosahedral capsid (100 nm) and tail (120 nm), suggesting that the PhiMa05 belongs to the Myoviridae family. PhiMa05 inhibited both planktonic and aggregated …
Fmri Feature Extraction Model For Adhd Classification Using Convolutional Neural Network, Senuri De Silva, Sanuwani Udara Dayarathna, Gangani Ariyarathne, Dulani Meedeniya, Sampath Jayarathna
Fmri Feature Extraction Model For Adhd Classification Using Convolutional Neural Network, Senuri De Silva, Sanuwani Udara Dayarathna, Gangani Ariyarathne, Dulani Meedeniya, Sampath Jayarathna
Computer Science Faculty Publications
Biomedical intelligence provides a predictive mechanism for the automatic diagnosis of diseases and disorders. With the advancements of computational biology, neuroimaging techniques have been used extensively in clinical data analysis. Attention deficit hyperactivity disorder (ADHD) is a psychiatric disorder, with the symptomology of inattention, impulsivity, and hyperactivity, in which early diagnosis is crucial to prevent unwelcome outcomes. This study addresses ADHD identification using functional magnetic resonance imaging (fMRI) data for the resting state brain by evaluating multiple feature extraction methods. The features of seed-based correlation (SBC), fractional amplitude of low-frequency fluctuation (fALFF), and regional homogeneity (ReHo) are comparatively applied to …
Leveraging Chemical And Computational Biology To Probe The Cellulose Synthase Complex, B. Kirtley Amos
Leveraging Chemical And Computational Biology To Probe The Cellulose Synthase Complex, B. Kirtley Amos
Theses and Dissertations--Plant and Soil Sciences
Cellular expansion in plants is a complex process driven by the constraint of internal cellular turgor pressure by an expansible cell wall. The main structural element of the cell wall is cellulose. Cellulose is vital to plant fitness and the protein complex that creates it is an excellent target for small molecule inhibition to create herbicides. In the following thesis many small molecules (SMs) from a diverse library were screened in search of new cellulose biosynthesis inhibitors (CBI). Loss of cellular expansion was the primary phenotype used to search for putative CBIs. As such, this was approached in a forward …
Metabolic Network Analysis Of Filamentous Cyanobacteria, Daniel Alexis Norena-Caro
Metabolic Network Analysis Of Filamentous Cyanobacteria, Daniel Alexis Norena-Caro
LSU Doctoral Dissertations
Cyanobacteria were the first organisms to use oxygenic photosynthesis, converting CO2 into useful organic chemicals. However, the chemical industry has historically relied on fossil raw materials to produce organic precursors, which has contributed to global warming. Thus, cyanobacteria have emerged as sustainable stakeholders for biotechnological production. The filamentous cyanobacterium Anabaena sp. UTEX 2576 can metabolize multiple sources of Nitrogen and was studied as a platform for biotechnological production of high-value chemicals (i.e., pigments, antioxidants, vitamins and secondary metabolites). From a Chemical engineering perspective, the biomass generation in this organism was thoroughly studied by interpreting the cell as a microbial …
Mathematical Models Of Cellular Signaling And Supramolecular Self-Assembly, Pratip Rana
Mathematical Models Of Cellular Signaling And Supramolecular Self-Assembly, Pratip Rana
Theses and Dissertations
Synthetic biologists endeavor to predict how the increasing complexity of multi-step signaling cascades impacts the fidelity of molecular signaling, whereby cellular state information is often transmitted with proteins diffusing by a pseudo-one-dimensional stochastic process. We address this problem by using a one-dimensional drift-diffusion model to derive an approximate lower bound on the degree of facilitation needed to achieve single-bit informational efficiency in signaling cascades as a function of their length. We find that a universal curve of the Shannon-Hartley form describes the information transmitted by a signaling chain of arbitrary length and depends upon only a small number of physically …
Computations Of Top-Down Attention By Modulating V1 Dynamics, David Berga, Xavier Otazu
Computations Of Top-Down Attention By Modulating V1 Dynamics, David Berga, Xavier Otazu
MODVIS Workshop
The human visual system processes information defining what is visually conspicuous (saliency) to our perception, guiding eye movements towards certain objects depending on scene context and its feature characteristics. However, attention has been known to be biased by top-down influences (relevance), which define voluntary eye movements driven by goal-directed behavior and memory. We propose a unified model of the visual cortex able to predict, among other effects, top-down visual attention and saccadic eye movements. First, we simulate activations of early mechanisms of the visual system (RGC/LGN), by processing distinct image chromatic opponencies with Gabor-like filters. Second, we use a cortical …
Acetylation Profiles Of Histone And Non-Histone Proteins In Breast Cancer, Alla Karpova
Acetylation Profiles Of Histone And Non-Histone Proteins In Breast Cancer, Alla Karpova
McKelvey School of Engineering Graduate Student Theses & Dissertations
This study evaluates the impact of protein acetylation on breast cancer gene expression and the regulation of metabolism. Acetylation is the second abundant post-translational modification after phosphorylation, regulating protein activity and function. The alterations in acetylation of both histone and non-histone proteins is known to be related to many human diseases, including cancer. Acetylation and deacetylation of histones is closely associated with the regulation of gene expression, while acetylation of non-histone proteins may have a broad effect on major cellular processes, such as proliferation, metabolism, cell cycle and apoptosis, imbalanced regulation of which is essential for cancer development. Therefore, it’s …
An Out-Of-Core Gpu Based Dimensionality Reduction Algorithm For Big Mass Spectrometry Data And Its Application In Bottom-Up Proteomics, Muaaz Awan, Fahad Saeed
An Out-Of-Core Gpu Based Dimensionality Reduction Algorithm For Big Mass Spectrometry Data And Its Application In Bottom-Up Proteomics, Muaaz Awan, Fahad Saeed
Parallel Computing and Data Science Lab Technical Reports
Modern high resolution Mass Spectrometry instruments can generate millions of spectra in a single systems biology experiment. Each spectrum consists of thousands of peaks but only a small number of peaks actively contribute to deduction of peptides. Therefore, pre-processing of MS data to detect noisy and non-useful peaks are an active area of research. Most of the sequential noise reducing algorithms are impractical to use as a pre-processing step due to high time-complexity. In this paper, we present a GPU based dimensionality-reduction algorithm, called G-MSR, for MS2 spectra. Our proposed algorithm uses novel data structures which optimize the memory and …
Comparing An Atomic Model Or Structure To A Corresponding Cryo-Electron Microscopy Image At The Central Axis Of A Helix, Stephanie Zeil, Julio Kovacs, Willy Wriggers, Jing He
Comparing An Atomic Model Or Structure To A Corresponding Cryo-Electron Microscopy Image At The Central Axis Of A Helix, Stephanie Zeil, Julio Kovacs, Willy Wriggers, Jing He
Computer Science Faculty Publications
Three-dimensional density maps of biological specimens from cryo-electron microscopy (cryo-EM) can be interpreted in the form of atomic models that are modeled into the density, or they can be compared to known atomic structures. When the central axis of a helix is detectable in a cryo-EM density map, it is possible to quantify the agreement between this central axis and a central axis calculated from the atomic model or structure. We propose a novel arc-length association method to compare the two axes reliably. This method was applied to 79 helices in simulated density maps and six case studies using cryo-EM …
A Computational Model Of The Spread Of Ancient Human Populations Based On Mitochondrial Dna Samples, Peter Revesz
A Computational Model Of The Spread Of Ancient Human Populations Based On Mitochondrial Dna Samples, Peter Revesz
School of Computing: Conference and Workshop Papers
The extraction of mitochondrial DNA (mtDNA) from ancient human population samples provides important data for the reconstruction of population influences, spread and evolution from the Neolithic to the present. This paper presents a mtDNA-based similarity measure between pairs of human populations and a computational model for the evolution of human populations. In a computational experiment, the paper studies the mtDNA information from five Neolithic and Bronze Age populations, namely the Andronovo, the Bell Beaker, the Minoan, the Rössen and the Únětice populations. In the past these populations were identified as separate cultural groups based on geographic location, age and the …
Mutations Of Adjacent Amino Acid Pairs Are Not Always Independent, Jyotsna Ramanan, Peter Revesz
Mutations Of Adjacent Amino Acid Pairs Are Not Always Independent, Jyotsna Ramanan, Peter Revesz
School of Computing: Conference and Workshop Papers
Evolutionary studies usually assume that the genetic mutations are independent of each other. This paper tests the independence hypothesis for genetic mutations with regard to protein coding regions. According to the new experimental results the independence assumption generally holds, but there are certain exceptions. In particular, the coding regions that represent two adjacent amino acids seem to change in ways that sometimes deviate significantly from the expected theoretical probability under the independence assumption.
Evolutionary Search For Models Of Planarian Regeneration Using Experimental Data, Marianna Viktorovna Budnikova
Evolutionary Search For Models Of Planarian Regeneration Using Experimental Data, Marianna Viktorovna Budnikova
Boise State University Theses and Dissertations
The ability of science to produce experimental data greatly surpasses our current ability to effectively visualize, conceptualize, and integrate the vast volumes of available data into a unified understanding of how complex biological systems work. This inability is a hindrance to scientific progress, and is particularly daunting when one considers multidimensional and shape-based observations as in the field of regenerative biology. For example, for at least the last 200 years, scientists have been interested in the exceptional ability of Planaria to regenerate lost tissues from damage, and there is a large amount of experimental data available on this organism. However, …
Creating A Package In R, Brit Schneiders, Eric Archer
Creating A Package In R, Brit Schneiders, Eric Archer
STAR Program Research Presentations
In a time of increasingly efficient technology and data production, scientists are producing data faster than it can be analyzed. Therefore, user accessibility to data analysis is becoming more and more critical. In general, researchers have a set of raw data and want an efficient means to their final analysis. A package serves as that means by creating a set of functions and making them accessible to the user. Often, a user has a small piece of code to run (a single R script, for example), and that script requires the use of certain functions, which are contained in a …
Classification Of Genomic Sequences By Latent Semantic Analysis, Samuel F. Way
Classification Of Genomic Sequences By Latent Semantic Analysis, Samuel F. Way
Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research
Evolutionary distance measures provide a means of identifying and organizing related organisms by comparing their genomic sequences. As such, techniques that quantify the level of similarity between DNA sequences are essential in our efforts to decipher the genetic code in which they are written.
Traditional methods for estimating the evolutionary distance separating two genomic sequences often require that the sequences first be aligned before they are compared. Unfortunately, this preliminary step imposes great computational burden, making this class of techniques impractical for applications involving a large number of sequences. Instead, we desire new methods for differentiating genomic sequences that eliminate …