Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Life Sciences (10)
- Bioinformatics (8)
- Genetics and Genomics (6)
- Medicine and Health Sciences (6)
- Computer Sciences (5)
-
- Computational Biology (4)
- Data Science (4)
- Statistical Models (4)
- Applied Mathematics (3)
- Applied Statistics (3)
- Biochemistry, Biophysics, and Structural Biology (3)
- Engineering (3)
- Genetics (3)
- Survival Analysis (3)
- Biochemistry (2)
- Biomedical Engineering and Bioengineering (2)
- Cell and Developmental Biology (2)
- Mathematics (2)
- Molecular Biology (2)
- Numerical Analysis and Scientific Computing (2)
- Physics (2)
- Public Health (2)
- Statistical Methodology (2)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (1)
- Artificial Intelligence and Robotics (1)
- Astrophysics and Astronomy (1)
- Atomic, Molecular and Optical Physics (1)
- Institution
-
- The Texas Medical Center Library (2)
- University of South Dakota (2)
- COBRA (1)
- Duquesne University (1)
- Michigan Technological University (1)
-
- Munster Technological University (1)
- New Jersey Institute of Technology (1)
- The University of Akron (1)
- University of Kentucky (1)
- University of Louisville (1)
- University of Montana (1)
- University of New Mexico (1)
- University of South Carolina (1)
- University of Texas at El Paso (1)
- Virginia Commonwealth University (1)
- Wayne State University (1)
- Publication Year
- Publication
-
- Dissertations and Theses (2)
- Dissertations and Theses (Open Access) (2)
- Electronic Theses and Dissertations (2)
- Theses and Dissertations (2)
- COBRA Preprint Series (1)
-
- Department of Mathematics Publications (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Mathematics & Statistics ETDs (1)
- Open Access Theses & Dissertations (1)
- Pathology and Laboratory Medicine Faculty Publications (1)
- Theses (1)
- Wayne State University Theses (1)
- Williams Honors College, Honors Research Projects (1)
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Biostatistics
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
Electronic Theses and Dissertations
As single-cell RNA sequencing (scRNA-seq) data expands, robust methods for integrating diverse datasets are critical. This dissertation applies Persistent Homology (PH), a technique from Topological Data Analysis (TDA), to a collection of scRNA-seq datasets spanning eight tissue types to quantify how data integration affects topological features and biological interpretability. We assessed global topological structure using Betti curves, Euler characteristics, and persistence landscapes across raw, normalized, and integrated data representations. Our analysis revealed a performance inversion: while conventional methods excelled on unintegrated data, high-granularity topological methods, particularly those sensitive to global data structure, became superior after integration. This suggests a synergy …
Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng
Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng
Theses and Dissertations
Lesion-symptom mapping (LSM) studies offer insight into the brain areas involved in various aspects of cognition. This is commonly done via behavioral testing in patients with a naturally occurring brain injury or lesions (e.g., strokes or brain tumors). This results in high-dimensional observational data where lesion status (present/absent) is non-uniformly distributed, with some voxels having lesions in very few (or no) subjects. In this situation, mass univariate hypothesis tests have severe power heterogeneity where many tests are known a priori to have little to no power. Additionally, high-dimensional observational data can be grouped according to brain anatomical structure.
In this …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …
Interactions Of The Sars-Cov-2 Viral Genome 3’-Untranslated Region With Viral And Host Rnas, Caleb Frye, Mihaela Rita Mihailescu
Interactions Of The Sars-Cov-2 Viral Genome 3’-Untranslated Region With Viral And Host Rnas, Caleb Frye, Mihaela Rita Mihailescu
Electronic Theses and Dissertations
This dissertation focuses on the characterization of RNA-RNA interactions within the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) genome and with host microRNAs. As the causative agent of coronavirus disease 2019 (COVID-19), SARS-CoV-2 has evolved rapidly since its appearance. This has warranted prompt characterization of the virus particularly of its single stranded RNA (ssRNA) genome. By using a combination of bioinformatics, biophysics, and/or biological assays, we analyzed the SARS-CoV-2 viral genomic RNA and uncovered interactions of genomic RNA with host RNAs, highlighting an underutilized method of targeting RNA viruses. We showed here that the conserved elements in the viral genomic …
Assessing The Utility Of Breast Cancer Polygenic Risk Scores And Association With Clinical Factors In A Population Of Breast Cancer Patients, John L. Slunecka
Assessing The Utility Of Breast Cancer Polygenic Risk Scores And Association With Clinical Factors In A Population Of Breast Cancer Patients, John L. Slunecka
Dissertations and Theses
INTRODUCTION: Breast cancer (BC) is the most common cancer among women and is classified as a complex disease. Advances in population genomics have led to the development of polygenic risk scores (PRSs) with the potential to enhance current risk models, but replication is often limited. OBJECTIVE: We sought to assess the predictive capabilities of two high-powered BC PRSs in a sample population selected for breast cancer. In addition, the capacity of the PRSs to predict clinical variables that could improve BC screening and treatments was explored. METHODS: Two published PRS algorithms (313 vs 3820) were used to score female subjects …
Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang
Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang
Dissertations and Theses (Open Access)
In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.
My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the …
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Mathematics & Statistics ETDs
The piezoelectric response has been a measure of interest in density functional theory (DFT) for micro-electromechanical systems (MEMS) since the inception of MEMS technology. Piezoelectric-based MEMS devices find wide applications in automobiles, mobile phones, healthcare devices, and silicon chips for computers, to name a few. Piezoelectric properties of doped aluminum nitride (AlN) have been under investigation in materials science for piezoelectric thin films because of its wide range of device applicability. In this research using rigorous DFT calculations, high throughput ab-initio simulations for 23 AlN alloys are generated.
This research is the first to report strong enhancements of piezoelectric properties …
Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman
Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman
Dissertations and Theses
One approach to interrogating the complexities of human systems in their well-regulated and dysregulated states is through the use of digital twins. Digital twins are virtual representations of physical systems that are descriptive of an individual's state of health, an object fundamentally related to precision medicine. A key element for building a functional digital twin type for a disease or predicting the therapeutic efficacy of a potential treatment is harmonized, machine-parsable domain knowledge. Hypothesis-driven investigations are the gold standard for representing subsystems, but their results encompass a limited knowledge of the full biosystem. Multi-omics data is one rich source of …
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Open Access Theses & Dissertations
With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Tumor Heterogeneity As A Predictor Of Response To Neoadjuvant Chemotherapy In Locally Advanced Rectal Cancer, Alissa Greenbaum, David R. Martin, Therese J. Bocklage, Ji-Hyun Lee, Scott A. Ness, Ashwani Rajput
Tumor Heterogeneity As A Predictor Of Response To Neoadjuvant Chemotherapy In Locally Advanced Rectal Cancer, Alissa Greenbaum, David R. Martin, Therese J. Bocklage, Ji-Hyun Lee, Scott A. Ness, Ashwani Rajput
Pathology and Laboratory Medicine Faculty Publications
BACKGROUND: Neoadjuvant chemoradiotherapy (nCRT) is the standard of care for locally advanced adenocarcinoma of the rectum, but it is currently unknown which patients have disease that will respond. This study tested the correlation between response to nCRT and intratumoral heterogeneity using next-generation sequencing assays.
PATIENTS AND METHODS: DNA was extracted from formalin-fixed, paraffin-embedded biopsy samples from a cohort of patients with locally advanced rectal adenocarcinoma (T3/4 or N1/2 disease) who received nCRT. High read-depth sequencing of > 400 cancer-relevant genes was performed. Tumor mutations and variant allele frequencies were used to calculate mutant-allele tumor heterogeneity (MATH) scores as measures of intratumoral …
Integrative Pathway Analysis Pipeline For Mirna And Mrna Data, Diana Mabel Diaz Herrera
Integrative Pathway Analysis Pipeline For Mirna And Mrna Data, Diana Mabel Diaz Herrera
Wayne State University Theses
The identification of pathways that are involved in a particular phenotype helps us understand the underlying biological processes. Traditional pathway analysis techniques aim to infer the impact on individual pathways using only mRNA levels. However, recent studies showed that gene expression alone is unable to capture the whole picture of biological phenomena. At the same time, MicroRNAs (miRNAs) are newly discovered gene regulators that have shown to play an important role in diagnosis, and prognosis for different types of diseases. Current pathway analysis techniques do not take miRNAs into consideration. In this project, we investigate the effect of integrating miRNA …
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
COBRA Preprint Series
Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …
Preparedness Of Hospitals In The Republic Of Ireland For An Influenza Pandemic, An Infection Control Perspective, Mary Reidy, Fiona Ryan, Dervla Hogan, Seán Lacey, Claire Buckley
Preparedness Of Hospitals In The Republic Of Ireland For An Influenza Pandemic, An Infection Control Perspective, Mary Reidy, Fiona Ryan, Dervla Hogan, Seán Lacey, Claire Buckley
Department of Mathematics Publications
When an influenza pandemic occurs most of the population is susceptible and attack rates can range as high as 40–50 %. The most important failure in pandemic planning is the lack of standards or guidelines regarding what it means to be ‘prepared’. The aim of this study was to assess the preparedness of acute hospitals in the Republic of Ireland for an influenza pandemic from an infection control perspective.
Evaluation Of The Signature Molecular Descriptor With Blosum62 And An All-Atom Description For Use In Sequence Alignment Of Proteins, Lindsay M. Aichinger
Evaluation Of The Signature Molecular Descriptor With Blosum62 And An All-Atom Description For Use In Sequence Alignment Of Proteins, Lindsay M. Aichinger
Williams Honors College, Honors Research Projects
This Honors Project focused on a few aspects of this topic. The second is comparing the molecular signature kernels to three of the BLOSUM matrices (30, 62, and 90) to test the accuracy of the mathematical model. The kernel matrix was manipulated in order to improve the relationship by focusing on side groups and also by changing how the structure was represented in the matrix by increasing the initial height distance from the central atom (Height 1 and Height 2 included).
There were multiple design constraints for this project. The first was the comparison with the BLOSUM matrices (30, 62, …
The Association Between The Il-1 Pathway, Isaac C. Wun
The Association Between The Il-1 Pathway, Isaac C. Wun
Dissertations and Theses (Open Access)
Cutaneous malignant melanoma (CMM) is a potentially lethal malignancy that warrants attention and further research, as it is known to that there is an increasing rate of incidence in theUnited States, and it is also known that exposure to UV light is its most crucial risk factor, and family history of melanoma is also an important risk factor. Melanoma is an aggressive and lethal cancer in humans. There are an estimated new 132,000 melanoma cases annually worldwide, and the trend has doubled in the past 20 years. However, attempts to treat melanoma have encountered considerable resistance and remained ineffective. The …
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Theses and Dissertations
Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …
An Application In Bioinformatics : A Comparison Of Affymetrix And Compugen Human Genome Microarrays, Milind Misra
An Application In Bioinformatics : A Comparison Of Affymetrix And Compugen Human Genome Microarrays, Milind Misra
Theses
The human genome microarrays from Compugen® and Affymetrix® were compared in the context of the emerging field of computational biology. The two premier database servers for genomic sequence data, the National Center for Biotechnology Information and the European Bioinformatics Institute, were described in detail. The various databases and data mining tools available through these data servers were also discussed. Microarrays were examined from a historical perspective and their main current applications-expression analysis, mutation analysis, and comparative genomic hybridization-were discussed. The two main types of microarrays, cDNA spotted microarrays and high-density spotted microarrays were analyzed by exploring the human genome microarray …