Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Classification

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 76

Full-Text Articles in Statistics and Probability

A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari May 2026

A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari

Theses and Dissertations

Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …


Empirical Benchmarks For Interpreting Effect Sizes In Violent Crime Interventions, Kohta James Matsukawa Hansen Apr 2026

Empirical Benchmarks For Interpreting Effect Sizes In Violent Crime Interventions, Kohta James Matsukawa Hansen

Theses and Dissertations

Criminal justice researchers apply Cohen’s (1988) benchmarks to classify effect sizes as small, medium, or large, despite these standards never being meant for broad, decontextualized use (Cohen, 1988; Gies et al., 2024; Goulet-Pelletier & Cousineau, 2018; Lakens, 2013; Milner et al., 2023). Repeatedly doing so may weaken statistical validity, distort findings, and impede effective policymaking. This study introduces the first effect size benchmarks specifically designed for violent crime interventions.

Using a quasi-meta-analysis framework, 1,605 effect sizes from 104 violent crime intervention studies from the CrimeSolutions clearinghouse were converted to Cohen’s d. Three new discrete benchmarking methods were created using a …


An Ensemble Classifier For Ordinal Outcomes In High-Dimensional Genomics Data, Heranga K. Rathnasekara, Sinjini Sikdar Jan 2026

An Ensemble Classifier For Ordinal Outcomes In High-Dimensional Genomics Data, Heranga K. Rathnasekara, Sinjini Sikdar

Mathematics & Statistics Faculty Publications

Analysis of genomics data for predicting disease outcomes is a fast-growing field in medical research. There often exist categorical, specifically, ordinal outcomes that need to be predicted based on genomic profiles. This has led to recent development of some high-dimensional ordinal classification methods that can address the large dimensionality of the genomic covariate set. These high-dimensional ordinal models tend to vary widely in their performance depending on the data they are applied to and the evaluation criteria used. In this article, we outline an ensemble ordinal classifier that integrates different ordinal modeling approaches through bootstrap-based model evaluation, multi-metric performance assessment, …


Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May Aug 2025

Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May

All Graduate Theses and Dissertations, Fall 2023 to Present

Classification tasks are fundamental in statistical machine learning. In classification tasks, a general goal is to build or select a model that can correctly classify data with as few errors as possible. However, for a particular dataset, the minimal number of errors achievable is seldom zero since overlap in the data makes errors unavoidable. As a result, it is often difficult for machine learning practitioners and data scientists to know whether classification errors can be reduced through further refinement. A potential solution to this lies in the Bayes error rate (BER). The BER is the lowest error rate achievable for …


Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez Jun 2025

Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez

SAML-25 Workshop on Statistical and Machine Learning

The accurate classification of nanoparticles (NPs) based on their shapes is crucial for understanding their physical-chemical properties and predict their bioactivity. Nowadays, synthesis method are able to produce a broad range of shapes, such as spheres, cubes and branched NPs and commonly these NP shapes are only described qualitative. This study presents NP descriptors obtained from NPs contours extracted from electron microscopy images. Descriptors such as Fourier descriptors, aspect ratio, and compactness are then used as input for machine learning classifiers. In particular, XGBoost, Random Forest, and neural networks are explored and the their performances are compared and discussed.


From Data To Insight: A Machine Learning Approach In Classifying Dairy Cow Productivity Level And Identifying Important Influencing Variables, Fatkhurokhman Fauzi, Achmad Fauzan, Rhendy K P Widiyanto, Khairil Anwar Notodiputro, Bagus Sartono Jan 2025

From Data To Insight: A Machine Learning Approach In Classifying Dairy Cow Productivity Level And Identifying Important Influencing Variables, Fatkhurokhman Fauzi, Achmad Fauzan, Rhendy K P Widiyanto, Khairil Anwar Notodiputro, Bagus Sartono

Knowledge Engineering and Data Science

Identifying influential predictor variables is crucial for enhancing model interpretability in supervised classification. This study applies Permutation Variable Importance (PVI), a model-agnostic approach, to evaluate variable relevance after model fitting. Using data from the 2024 Indonesia Dairy Cow Productivity Survey, this research investigates five classification techniques: (1) Support Vector Machine (SVM), (2) Neural Network (NN), (3) k-Nearest Neighbors (kNN), (4) Naïve Bayes Classifier (NB), and (5) Logistic Regression (LR), to identify which method(s) yield the best performance based on evaluation metrics such as accuracy, sensitivity, and specificity. PVI is employed to identify the most influential predictor variables within the best-performing …


Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier Dec 2024

Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

The purpose of this research is to augment linear and kernelized ordinal distance metric learning (L/KODML) techniques with a proposed variable selection methodology that integrates the Sequential Multi-Response Feature Selection (SMuRFS) algorithm. Additionally, we aim to embed ordinal triplet constraints into a deep learning architecture, and to propose a general framework for deep learning-based ordinal classification. A variety of simulation studies and real data experiments were conducted to evaluate the various methodologies. For the distance metric learning and variable selection, results showed that the integration of SMuRFS performed effective variable selection and improved prediction accuracy. For the triplet constraints, incorporating …


Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi Aug 2024

Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi

Dissertations

In many machine learning applications, such as image tagging, document classi-fication, and medical diagnosis, a data instance can be associated with multiple classes in parallel so that each instance is associated with multiple response variables simultaneously defining multi-label classification. Standard multi-label classification methods that provide point predictions have been developed. They lack in quantifying the uncertainty of predictions. These methods also lack in accounting for label dependencies and are very computationally expensive. This dissertation develops two methods of multi-label classification using conformal prediction that quantify the uncertainty of predictions. Chapter 1 introduces notations and tools that have been used in …


A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang Aug 2024

A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang

Rose-Hulman Undergraduate Mathematics Journal

Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …


Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn May 2024

Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn

SMU Data Science Review

As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …


The Geometry Of Dynamic Time-Dependent Best-Worst Choice Pairs, Sasanka Adikari, Norou Diawara, Haim Bar Jan 2024

The Geometry Of Dynamic Time-Dependent Best-Worst Choice Pairs, Sasanka Adikari, Norou Diawara, Haim Bar

Mathematics & Statistics Faculty Publications

There has been increasing interest in best–worst discrete choice experiments (BWDCEs) in health economics, transportation research, and other fields over the last few years. BWDCEs have distinct advantages compared to other measurement approaches in discrete choice experiments (DCEs). A systematic study of best–worst (BW) choice pairs can be traced back to the 1990s. Recently, new ideas have been introduced to the subject. Calculating utility helps measure the attractiveness of BW choices. The goal of this paper is twofold. First, we extend the idea of the BW choice pair to include dynamic, time-dependent transition probability and capture utility at each time …


Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre Dec 2023

Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre

SMU Data Science Review

Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …


Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz Sep 2023

Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz

Beyond: Undergraduate Research Journal

When it comes to registering to vote, Hispanic voters can only register as “Hispanic” in the “Race/Ethnicity” category, causing difficulties when analyzing voting trends amongst the Hispanic community. Upon the recent idea that not all Hispanic Groups vote the same, the goal is to create a model that can possibly identify a voter’s Hispanic Group with the information provided on the public Florida voter file. This is accomplished using name and zip code data for all voters in Palm Beach, Florida. This paper will explore the model implemented, its findings and limitations. Palm Beach, Florida, is met with low confidence …


Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu Aug 2023

Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu

Open Access Theses & Dissertations

The goal of classification is to develop a model that can be used to accurately assign new observations to labeled classes based on the patterns learned from the training data. K-nearest Neighbors algorithm (KNN) is a popular and widely used algorithm for classification, however, its performance can be adversely affected by the presence of outliers in a dataset. In this study we have modified this existing KNN algorithm that can alleviate the effect of outliers in a dataset, thereby improving the performance of the KNN algorithm. We compared the performances of the Modified KNN method and the Existing KNN algorithm …


Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley Apr 2023

Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley

Senior Theses

In Major League Baseball (MLB), the outcome of a stolen base attempt has important implications. Success moves the runner closer to scoring, while failure records an out and removes the runner from the basepaths altogether. Therefore, it is important that the decision by a coach or player to steal a base is well-informed. In this thesis, I explore a statistical approach to making this decision. I train logistic regression and random forest models, using data about the game situation and about the runner, pitcher, and catcher involved in the stolen base attempt, to estimate the probability that a stolen base …


Integrating And Optimizing Genomic, Weather, And Secondary Trait Data For Multiclass Classification, Vamsi Manthena, Diego Jarquín, Reka Howard Mar 2023

Integrating And Optimizing Genomic, Weather, And Secondary Trait Data For Multiclass Classification, Vamsi Manthena, Diego Jarquín, Reka Howard

Department of Statistics: Faculty Publications

Modern plant breeding programs collect several data types such as weather, images, and secondary or associated traits besides the main trait (e.g., grain yield). Genomic data is high-dimensional and often over-crowds smaller data types when naively combined to explain the response variable. There is a need to develop methods able to effectively combine different data types of differing sizes to improve predictions. Additionally, in the face of changing climate conditions, there is a need to develop methods able to effectively combine weather information with genotype data to predict the performance of lines better. In this work, we develop a novel …


Analyzing Relationships With Machine Learning, Oscar Ko Feb 2023

Analyzing Relationships With Machine Learning, Oscar Ko

Dissertations, Theses, and Capstone Projects

Procedurally, this project aims to take a dataset, analyze it, and offer insights to the audience in an easy-to-digest format. Conceptually, this project will seek to explore questions like: “Do couples that meet through online dating or dating apps have higher or lower quality relationships?”, “Can any features in this dataset help predict how a subject would rate their relationship quality?”, and “What other insights can I derive from using machine learning for exploratory analysis?” The intended audience for this project is anyone interested in romantic relationships or machine learning.

The dataset is from a Stanford University survey, “How Couples …


Prediction Of Rapid Early Progression And Survival Risk With Pre-Radiation Mri In Who Grade 4 Glioma Patients, Walia Farzana, Mustafa M. Basree, Norou Diawara, Zeina Shboul, Sagel Dubey, Marie M. Lockheart, Mohamed Hamza, Joshua D. Palmer, Khan Iftekharuddin Jan 2023

Prediction Of Rapid Early Progression And Survival Risk With Pre-Radiation Mri In Who Grade 4 Glioma Patients, Walia Farzana, Mustafa M. Basree, Norou Diawara, Zeina Shboul, Sagel Dubey, Marie M. Lockheart, Mohamed Hamza, Joshua D. Palmer, Khan Iftekharuddin

Electrical & Computer Engineering Faculty Publications

Rapid early progression (REP) has been defined as increased nodular enhancement at the border of the resection cavity, the appearance of new lesions outside the resection cavity, or increased enhancement of the residual disease after surgery and before radiation. Patients with REP have worse survival compared to patients without REP (non-REP). Therefore, a reliable method for differentiating REP from non-REP is hypothesized to assist in personlized treatment planning. A potential approach is to use the radiomics and fractal texture features extracted from brain tumors to characterize morphological and physiological properties. We propose a random sampling-based ensemble classification model. The proposed …


Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia Sep 2022

Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia

SMU Data Science Review

In this paper, machine learning techniques are used to reconstruct particle collision pathways. CERN (Conseil européen pour la recherche nucléaire) uses a massive underground particle collider, called the Large Hadron Collider or LHC, to produce particle collisions at extremely high speeds. There are several layers of detectors in the collider that track the pathways of particles as they collide. The data produced from collisions contains an extraneous amount of background noise, i.e., decays from known particle collisions produce fake signal. Particularly, in the first layer of the detector, the pixel tracker, there is an overwhelming amount of background noise that …


Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel Sep 2022

Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel

SMU Data Science Review

Since the pandemic started, researchers have been trying to find a way to detect COVID-19 which is a cost-effective, fast, and reliable way to keep the economy viable and running. This research details how chest X-ray radiography can be utilized to detect the infection. This can be for implementation in Airports, Schools, and places of business. Currently, Chest imaging is not a first-line test for COVID-19 due to low diagnostic accuracy and confounding with other viral pneumonia. Different pre-trained algorithms were fine-tuned and applied to the images to train the model and the best model obtained was fine-tuned InceptionV3 model …


Combining Phenotypic And Genomic Data To Improve Prediction Of Binary Traits, Diego Jarquin, Arkaprava Roy, Bertrand S. Clarke, Subhashis Ghosal Sep 2022

Combining Phenotypic And Genomic Data To Improve Prediction Of Binary Traits, Diego Jarquin, Arkaprava Roy, Bertrand S. Clarke, Subhashis Ghosal

Department of Statistics: Faculty Publications

Plant breeders want to develop cultivars that outperform existing genotypes. Some characteristics (here ‘main traits’) of these cultivars are categorical and difficult to measure directly. It is important to predict the main trait of newly developed genotypes accurately. In addition to marker data, breeding programs often have information on secondary traits (or ‘phenotypes’) that are easy to measure. Our goal is to improve prediction of main traits with interpretable relations by combining the two data types using variable selection techniques. However, the genomic characteristics can overwhelm the set of secondary traits, so a standard technique may fail to select any …


A Strategy To Identify Event Specific Hospitalizations In Large Health Claims Databases, Joshua Lambert, Harpal Sandhu, Emily Kean, Teenu Xavier, Aviv Brokman, Zachary Steckler, Lee Park, Arnold Stromberg May 2022

A Strategy To Identify Event Specific Hospitalizations In Large Health Claims Databases, Joshua Lambert, Harpal Sandhu, Emily Kean, Teenu Xavier, Aviv Brokman, Zachary Steckler, Lee Park, Arnold Stromberg

Statistics Faculty Publications

Background: Health insurance claims data offer a unique opportunity to study disease distribution on a large scale. Challenges arise in the process of accurately analyzing these raw data. One important challenge to overcome is the accurate classification of study outcomes. For example, using claims data, there is no clear way of classifying hospitalizations due to a specific event. This is because of the inherent disjointedness and lack of context that typically come with raw claims data.

Methods: In this paper, we propose a framework for classifying hospitalizations due to a specific event. We then tested this framework in …


Split Classification Model For Complex Clustered Data, Katherine Gerot Mar 2022

Split Classification Model For Complex Clustered Data, Katherine Gerot

Honors Program: Senior Projects (Public)

Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.


Shape-Based Classification Of Partially Observed Curves, With Applications To Anthropology, Gregory J. Matthews, Karthik Bharath, Sebastian Kurtek, Juliet K. Brophy, George K. Thiruvathukal, Ofer Harel Oct 2021

Shape-Based Classification Of Partially Observed Curves, With Applications To Anthropology, Gregory J. Matthews, Karthik Bharath, Sebastian Kurtek, Juliet K. Brophy, George K. Thiruvathukal, Ofer Harel

Computer Science: Faculty Publications and Other Works

We consider the problem of classifying curves when they are observed only partially on their parameter domains. We propose computational methods for (i) completion of partially observed curves; (ii) assessment of completion variability through a nonparametric multiple imputation procedure; (iii) development of nearest neighbor classifiers compatible with the completion techniques. Our contributions are founded on exploiting the geometric notion of shape of a curve, defined as those aspects of a curve that remain unchanged under translations, rotations and reparameterizations. Explicit incorporation of shape information into the computational methods plays the dual role of limiting the set of all possible completions …


Bayesian Nonparametric Model For Functional Data Analysis, Tahmidul Islam Apr 2021

Bayesian Nonparametric Model For Functional Data Analysis, Tahmidul Islam

Theses and Dissertations

Functional data analysis (FDA) experienced a burst of growth after Ramsay and Silverman published their textbook in 1997. Functional data analysis interests researchers because of the challenges it adds to well-established multivariate analysis. Unlike finite dimensional random vectors, we visualize infinite dimensional random functions; for example, curves, images, brain scans, etc. A vast amount of literature have been dedicated to developing models for functional data. The ideas are mostly based on basis function representations and kernel-based nonparametric methods. In this dissertation, we propose a Bayesian treatment of nonparametric functional data analysis by introducing a Gaussian process (GP) over the space …


Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo Aug 2020

Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo

Dissertations

In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.

First, to improve the prediction accuracy of learning …


Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam Aug 2020

Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam

Student Works (2020-2029)

One of the approaches for structural health monitoring (SHM) consists of two major components, i.e. a network of sensors to collect the response data and an extraction method to obtain information on the structural health condition. Data mining (DM) is a novel data extraction technology which can employ for development of inverse analysis. Implementation of DM techniques in different areas of civil engineering has recently given very good results. However, application of DM in SHM is not used as much as expected, thus, many challenges are still ahead. Therefore, it is necessary to develop the applicability of DM in SHM. …


A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley Jul 2020

A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley

Theses and Dissertations

According to the Centers for Disease Control and Prevention, about 18.2 million adults age 20 and older have Coronary Artery Disease in the United States. Early diagnosis is therefore of crucial importance to help prevent debilitating consequences, and principally death for many patients. In this study we use data containing gene expression values from peripheral blood samples in 198 non-diabetic patients, with the goal of developing an age and sex gene expression model for diagnosis of Coronary Artery Disease. We employ machine learning methods to obtain a classification based on genetic information, age and sex. Our implementation uses feed forward …


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya May 2020

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


Gradient Boosting For Survival Analysis With Applications In Oncology, Nam Phuong Nguyen Jan 2020

Gradient Boosting For Survival Analysis With Applications In Oncology, Nam Phuong Nguyen

USF Tampa Graduate Theses and Dissertations

Cancer is one of the most deadly diseases that the world has been fighting against over decades. An enormous number of research has been conducted, via a wide scale of approaches, raging from genetic analysis to mathematical modeling. Survival analysis is a well-performed methodology frequently used to estimate the survival probability of a patient. Although there has been a large number of methods for survival analysis, efficient exploration of a high-dimensional feature space has been challenging due to its computational cost and complexity. This thesis adapts the component-wise gradient boosting algorithms for cancer survival analysis, and also proposes a new …