Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Keyword
-
- Feature Selection (2)
- Lasso (2)
- Bioinformatics (1)
- C-Index (1)
- COVID-19 (1)
-
- Clinicogenomic (1)
- Cognition (1)
- Comparison-wise Power (1)
- Computational biology (1)
- Count Data Modeling (1)
- DNA repair (1)
- Drug discovery (1)
- Ensemble Learning (1)
- Excursion (1)
- Familywise Error Rate (1)
- Functional Independent Measure (FIM) (1)
- Fused Regularization (1)
- Generalized Linear Latent Mixed Models (GLLAMM) (1)
- Hadoop (1)
- Health Data Analysis (1)
- High-Dimensional Data (1)
- High-dimensional data (1)
- Inversion (1)
- Machine Learning (1)
- Machine learning (1)
- Microtubule proteins (1)
- Molecular dynamics (1)
- Motor (1)
- Multiple Comparisons (1)
- Multiple-Inflation Poisson Model (1)
Articles 1 - 10 of 10
Full-Text Articles in Biostatistics
Simultaneous Selection Of Inflations And Variables In Multiple Inflations Poisson Model (Mip), John Koomson
Simultaneous Selection Of Inflations And Variables In Multiple Inflations Poisson Model (Mip), John Koomson
Open Access Theses & Dissertations
Count data frequently arise in biomedical, economic, and social science research and are often characterized by structural excesses at specific count levels. To accommodate such patterns, Su et al. (2013), among others, introduced the Multiple-Inflation Poisson (MIP) model, which allows for multiple inflated counts within the distribution. However, two critical challenges remain in modeling such data: (i) identifying the true inflation points where excess counts occur, and (ii) selecting the relevant covariates that explain variation in the inflation and count process. This dissertation addresses these issues by advancing the MIP model through a novel methodology that enables the simultaneous selection …
Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta
Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta
Open Access Theses & Dissertations
Prostate cancer (PrCa) remains a critical challenge in precision oncology due to several reasons including its apparent heterogenous condition, recurrence following treatment and rapid progressive forms. Therefore, identifying patients at risk of progression is essential to fast-track therapeutic decisions and improve outcomes. Despite recent advances in genomic and molecular profiling, conventional PrCa risk assessment tools heavily rely on a few clinical parameters, neglecting the prognostic potential of genomic biomarkers in the presence of clinical biomarkers. This study presents a computational pipeline to harmonize and evaluate the prognostic value of clinicogenomic profiles of patients in modelling progression free survival (PFS). PFS, …
Random Forest For High-Dimensional Data, George Ekow Quaye
Random Forest For High-Dimensional Data, George Ekow Quaye
Open Access Theses & Dissertations
The exponential growth of data has led to a rapid increase in high-dimensional datasets across various domains, presenting significant challenges in data analysis, particularly in predictive modeling tasks. Traditional Random Forest (RF), while robust, often struggles with datasets filled with numerous noisy or non-informative features, compromising both performance and accuracy. This study introduces an advanced algorithm, High-Dimensional Random Forests (HDRF), designed to address these challenges by integrating robust multivariate feature selection techniques directly into the decision tree construction process. Unlike standard RF, HDRF incorporates ridge regression-based variable screening at each decision split, enhancing its ability to identify and utilize the …
Developing And Applying Computational Algorithms To Reveal Health-Related Biomolecular Interactions, Yixin Xie
Developing And Applying Computational Algorithms To Reveal Health-Related Biomolecular Interactions, Yixin Xie
Open Access Theses & Dissertations
Computational biology is an interdisciplinary area that applies computational approaches in biological big data, including protein amino acid sequences, genetic sequences, etc., which is widely used to analyze protein-protein interactions, make predictions in drug discovery, develop vaccines, etc. Popular methods include mathematical modeling, molecular dynamics simulations, data science mythology, etc. With the help of computational algorithms and applications, drug development is much faster than traditional processes, as it reduces risks early on in a drug discovery process and helps researchers select target candidates that have the highest potential for success. In my doctoral research, I applied multi-scale computational approaches to …
The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun
The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun
Open Access Theses & Dissertations
Fast and accurate calculations of the electrostatic features for highly charged biomolecules such as DNA, RNA, highly charged proteins, are crucial but challenging tasks. Traditional implicit solvent methods calculate the electrostatic features fast, but they are not able to balance the high net charges in the biomolecules effectively. Explicit solvent methods add unbalanced ions to neutralize the highly charged biomolecules in molecular dynamic simulations, which require more expensive computing resources. Here we developed a novel method, the Hybridizing Ions Treatment (HIT) method, which hybridizes the implicit solvent method with the explicit method to realistically calculate the electrostatic potential for highly …
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Open Access Theses & Dissertations
With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …
Combination Of Resampling Based Lasso Feature Selection And Ensembles Of Regularized Regression Models, Abhijeet R. Patil
Combination Of Resampling Based Lasso Feature Selection And Ensembles Of Regularized Regression Models, Abhijeet R. Patil
Open Access Theses & Dissertations
In high-dimensional data, the performance of various classiers is largely dependent on the selection of important features. Most of the individual classiers using existing feature selection (FS) methods do not perform well for highly correlated data. Obtaining important
features using the FS method and selecting the best performing classier is a challenging task in high throughput data. In this research, we propose a combination of resampling based least absolute shrinkage and selection operator (LASSO) feature selection (RLFS)
and ensembles of regularized regression models (ERRM) capable of handling data with the high correlation structures. The ERRM boosts the prediction accuracy with …
Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson
Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson
Open Access Theses & Dissertations
In typical genomics studies involving numerous association tests of gene mutations with a disease, error rate control via multiplicity adjustment is paramount because even if all genes were to be non-differentially associated, we would still make some false positives. Many methods exist that incorporate the control of multiplicity for normally distributed endpoints in sample size estimation, but none addresses the issue for non-normally correlated endpoints.
One common practice in the literature is to assume an equal correlation among all differentially associated or expressed genes, thereby using the generalized binomial or beta-binomial model to compute the comparison-wise power of detecting these …
Generalized Linear Latent Mixed Modeling Of Functional Independent Measures And Patient Outcomes, Maduranga Kasun Dassanayake
Generalized Linear Latent Mixed Modeling Of Functional Independent Measures And Patient Outcomes, Maduranga Kasun Dassanayake
Open Access Theses & Dissertations
The Functional Independent Measure (FIM) is one of the most widely accepted functional assessment measures used in the rehabilitation community. Past research studies have investigated the relationship between place of discharge, admission FIM scores or FIM difference scores, and patients' characteristics and found relationships between those variables. However, most of these studies fail to account for the multi-layered multidimensionality of the FIM and the measurement error associated with the FIM items. This study utilizes Generalized Linear Latent Mixed Models (GLLAMM) and Structural Equation Models (SEM) to assess which patient characteristics are associated with FIM difference scores and the structural relationship …
Secondary Structure Prediction Of Long Rna Sequences Based On Inversion Excursions And A Modularized Mapreduce Framework, Daniel Tesfai Yehdego
Secondary Structure Prediction Of Long Rna Sequences Based On Inversion Excursions And A Modularized Mapreduce Framework, Daniel Tesfai Yehdego
Open Access Theses & Dissertations
Ribonucleic acid (RNA) molecules and their secondary structures play important roles in many biological processes including gene expression and regulation. The genomes of many viruses are also RNA molecules. Since secondary structures are crucial for RNA functionality, computational predictions of the RNA secondary structures have been widely studied. However, the tremendous demands on computer memory and computing time for complex secondary structures limit the capability of existing thermodynamically based algorithms for structure predictions to handling only short RNA sequences with a few hundred bases. One approach to overcome this limitation is by first cutting long RNA sequences into shorter, non-overlapping …