Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (16)
- Computer Sciences (15)
- Medicine and Health Sciences (14)
- Multivariate Analysis (13)
- Statistical Models (13)
-
- Data Science (12)
- Statistical Methodology (11)
- Mathematics (10)
- Public Health (9)
- Statistical Theory (9)
- Biostatistics (8)
- Categorical Data Analysis (8)
- Clinical Epidemiology (8)
- Life Sciences (7)
- Social and Behavioral Sciences (7)
- Other Statistics and Probability (6)
- Applied Mathematics (5)
- Microarrays (5)
- Epidemiology (4)
- Genetics and Genomics (4)
- Artificial Intelligence and Robotics (3)
- Arts and Humanities (3)
- Clinical Trials (3)
- Engineering (3)
- Medical Specialties (3)
- Other Computer Sciences (3)
- Probability (3)
- Theory and Algorithms (3)
- Institution
-
- COBRA (14)
- Southern Methodist University (6)
- Utah State University (5)
- University of Nebraska - Lincoln (4)
- University of South Florida (4)
-
- Old Dominion University (3)
- University of South Carolina (3)
- Wayne State University (3)
- Brigham Young University (2)
- Loyola University Chicago (2)
- Virginia Commonwealth University (2)
- City University of New York (CUNY) (1)
- Department of Primary Industries and Regional Development, Western Australia (1)
- Embry-Riddle Aeronautical University (1)
- Florida Institute of Technology (1)
- Georgia Southern University (1)
- Illinois State University (1)
- Kennesaw State University (1)
- Louisiana Tech University (1)
- Marquette University (1)
- Minnesota State University, Mankato (1)
- Missouri University of Science and Technology (1)
- New Jersey Institute of Technology (1)
- Nova Southeastern University (1)
- Purdue University (1)
- Rose-Hulman Institute of Technology (1)
- Technological University Dublin (1)
- The Texas Medical Center Library (1)
- The University of Southern Mississippi (1)
- Universitas Negeri Malang (1)
- Publication Year
- Publication
-
- UW Biostatistics Working Paper Series (11)
- Theses and Dissertations (9)
- SMU Data Science Review (6)
- USF Tampa Graduate Theses and Dissertations (4)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (3)
-
- Journal of Modern Applied Statistical Methods (3)
- Department of Statistics: Faculty Publications (2)
- Dissertations (2)
- Electronic Theses and Dissertations (2)
- Mathematics & Statistics Faculty Publications (2)
- U.C. Berkeley Division of Biostatistics Working Paper Series (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
- Arts & Sciences Graduate Student Theses and Dissertations (1)
- Beyond: Undergraduate Research Journal (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computer Science: Faculty Publications and Other Works (1)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (1)
- Dissertations and Theses (Open Access) (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctoral Dissertations (1)
- Electrical & Computer Engineering Faculty Publications (1)
- Honors Program: Senior Projects (Public) (1)
- Journal of the Department of Agriculture, Western Australia, Series 4 (1)
- Knowledge Engineering and Data Science (1)
- Mahurin Honors College Capstone Experience/Thesis Projects (1)
- Mathematics and Statistics Faculty Research & Creative Works (1)
- Mathematics and Statistics: Faculty Publications and Other Works (1)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (1)
- Publication Type
Articles 1 - 30 of 76
Full-Text Articles in Statistics and Probability
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
Theses and Dissertations
Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …
Empirical Benchmarks For Interpreting Effect Sizes In Violent Crime Interventions, Kohta James Matsukawa Hansen
Empirical Benchmarks For Interpreting Effect Sizes In Violent Crime Interventions, Kohta James Matsukawa Hansen
Theses and Dissertations
Criminal justice researchers apply Cohen’s (1988) benchmarks to classify effect sizes as small, medium, or large, despite these standards never being meant for broad, decontextualized use (Cohen, 1988; Gies et al., 2024; Goulet-Pelletier & Cousineau, 2018; Lakens, 2013; Milner et al., 2023). Repeatedly doing so may weaken statistical validity, distort findings, and impede effective policymaking. This study introduces the first effect size benchmarks specifically designed for violent crime interventions.
Using a quasi-meta-analysis framework, 1,605 effect sizes from 104 violent crime intervention studies from the CrimeSolutions clearinghouse were converted to Cohen’s d. Three new discrete benchmarking methods were created using a …
An Ensemble Classifier For Ordinal Outcomes In High-Dimensional Genomics Data, Heranga K. Rathnasekara, Sinjini Sikdar
An Ensemble Classifier For Ordinal Outcomes In High-Dimensional Genomics Data, Heranga K. Rathnasekara, Sinjini Sikdar
Mathematics & Statistics Faculty Publications
Analysis of genomics data for predicting disease outcomes is a fast-growing field in medical research. There often exist categorical, specifically, ordinal outcomes that need to be predicted based on genomic profiles. This has led to recent development of some high-dimensional ordinal classification methods that can address the large dimensionality of the genomic covariate set. These high-dimensional ordinal models tend to vary widely in their performance depending on the data they are applied to and the evaluation criteria used. In this article, we outline an ensemble ordinal classifier that integrates different ordinal modeling approaches through bootstrap-based model evaluation, multi-metric performance assessment, …
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
All Graduate Theses and Dissertations, Fall 2023 to Present
Classification tasks are fundamental in statistical machine learning. In classification tasks, a general goal is to build or select a model that can correctly classify data with as few errors as possible. However, for a particular dataset, the minimal number of errors achievable is seldom zero since overlap in the data makes errors unavoidable. As a result, it is often difficult for machine learning practitioners and data scientists to know whether classification errors can be reduced through further refinement. A potential solution to this lies in the Bayes error rate (BER). The BER is the lowest error rate achievable for …
Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez
Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez
SAML-25 Workshop on Statistical and Machine Learning
The accurate classification of nanoparticles (NPs) based on their shapes is crucial for understanding their physical-chemical properties and predict their bioactivity. Nowadays, synthesis method are able to produce a broad range of shapes, such as spheres, cubes and branched NPs and commonly these NP shapes are only described qualitative. This study presents NP descriptors obtained from NPs contours extracted from electron microscopy images. Descriptors such as Fourier descriptors, aspect ratio, and compactness are then used as input for machine learning classifiers. In particular, XGBoost, Random Forest, and neural networks are explored and the their performances are compared and discussed.
From Data To Insight: A Machine Learning Approach In Classifying Dairy Cow Productivity Level And Identifying Important Influencing Variables, Fatkhurokhman Fauzi, Achmad Fauzan, Rhendy K P Widiyanto, Khairil Anwar Notodiputro, Bagus Sartono
From Data To Insight: A Machine Learning Approach In Classifying Dairy Cow Productivity Level And Identifying Important Influencing Variables, Fatkhurokhman Fauzi, Achmad Fauzan, Rhendy K P Widiyanto, Khairil Anwar Notodiputro, Bagus Sartono
Knowledge Engineering and Data Science
Identifying influential predictor variables is crucial for enhancing model interpretability in supervised classification. This study applies Permutation Variable Importance (PVI), a model-agnostic approach, to evaluate variable relevance after model fitting. Using data from the 2024 Indonesia Dairy Cow Productivity Survey, this research investigates five classification techniques: (1) Support Vector Machine (SVM), (2) Neural Network (NN), (3) k-Nearest Neighbors (kNN), (4) Naïve Bayes Classifier (NB), and (5) Logistic Regression (LR), to identify which method(s) yield the best performance based on evaluation metrics such as accuracy, sensitivity, and specificity. PVI is employed to identify the most influential predictor variables within the best-performing …
Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier
Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
The purpose of this research is to augment linear and kernelized ordinal distance metric learning (L/KODML) techniques with a proposed variable selection methodology that integrates the Sequential Multi-Response Feature Selection (SMuRFS) algorithm. Additionally, we aim to embed ordinal triplet constraints into a deep learning architecture, and to propose a general framework for deep learning-based ordinal classification. A variety of simulation studies and real data experiments were conducted to evaluate the various methodologies. For the distance metric learning and variable selection, results showed that the integration of SMuRFS performed effective variable selection and improved prediction accuracy. For the triplet constraints, incorporating …
Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi
Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi
Dissertations
In many machine learning applications, such as image tagging, document classi-fication, and medical diagnosis, a data instance can be associated with multiple classes in parallel so that each instance is associated with multiple response variables simultaneously defining multi-label classification. Standard multi-label classification methods that provide point predictions have been developed. They lack in quantifying the uncertainty of predictions. These methods also lack in accounting for label dependencies and are very computationally expensive. This dissertation develops two methods of multi-label classification using conformal prediction that quantify the uncertainty of predictions. Chapter 1 introduces notations and tools that have been used in …
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
Rose-Hulman Undergraduate Mathematics Journal
Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …
Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn
Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn
SMU Data Science Review
As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …
The Geometry Of Dynamic Time-Dependent Best-Worst Choice Pairs, Sasanka Adikari, Norou Diawara, Haim Bar
The Geometry Of Dynamic Time-Dependent Best-Worst Choice Pairs, Sasanka Adikari, Norou Diawara, Haim Bar
Mathematics & Statistics Faculty Publications
There has been increasing interest in best–worst discrete choice experiments (BWDCEs) in health economics, transportation research, and other fields over the last few years. BWDCEs have distinct advantages compared to other measurement approaches in discrete choice experiments (DCEs). A systematic study of best–worst (BW) choice pairs can be traced back to the 1990s. Recently, new ideas have been introduced to the subject. Calculating utility helps measure the attractiveness of BW choices. The goal of this paper is twofold. First, we extend the idea of the BW choice pair to include dynamic, time-dependent transition probability and capture utility at each time …
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
SMU Data Science Review
Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Beyond: Undergraduate Research Journal
When it comes to registering to vote, Hispanic voters can only register as “Hispanic” in the “Race/Ethnicity” category, causing difficulties when analyzing voting trends amongst the Hispanic community. Upon the recent idea that not all Hispanic Groups vote the same, the goal is to create a model that can possibly identify a voter’s Hispanic Group with the information provided on the public Florida voter file. This is accomplished using name and zip code data for all voters in Palm Beach, Florida. This paper will explore the model implemented, its findings and limitations. Palm Beach, Florida, is met with low confidence …
Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu
Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu
Open Access Theses & Dissertations
The goal of classification is to develop a model that can be used to accurately assign new observations to labeled classes based on the patterns learned from the training data. K-nearest Neighbors algorithm (KNN) is a popular and widely used algorithm for classification, however, its performance can be adversely affected by the presence of outliers in a dataset. In this study we have modified this existing KNN algorithm that can alleviate the effect of outliers in a dataset, thereby improving the performance of the KNN algorithm. We compared the performances of the Modified KNN method and the Existing KNN algorithm …
Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley
Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley
Senior Theses
In Major League Baseball (MLB), the outcome of a stolen base attempt has important implications. Success moves the runner closer to scoring, while failure records an out and removes the runner from the basepaths altogether. Therefore, it is important that the decision by a coach or player to steal a base is well-informed. In this thesis, I explore a statistical approach to making this decision. I train logistic regression and random forest models, using data about the game situation and about the runner, pitcher, and catcher involved in the stolen base attempt, to estimate the probability that a stolen base …
Integrating And Optimizing Genomic, Weather, And Secondary Trait Data For Multiclass Classification, Vamsi Manthena, Diego Jarquín, Reka Howard
Integrating And Optimizing Genomic, Weather, And Secondary Trait Data For Multiclass Classification, Vamsi Manthena, Diego Jarquín, Reka Howard
Department of Statistics: Faculty Publications
Modern plant breeding programs collect several data types such as weather, images, and secondary or associated traits besides the main trait (e.g., grain yield). Genomic data is high-dimensional and often over-crowds smaller data types when naively combined to explain the response variable. There is a need to develop methods able to effectively combine different data types of differing sizes to improve predictions. Additionally, in the face of changing climate conditions, there is a need to develop methods able to effectively combine weather information with genotype data to predict the performance of lines better. In this work, we develop a novel …
Analyzing Relationships With Machine Learning, Oscar Ko
Analyzing Relationships With Machine Learning, Oscar Ko
Dissertations, Theses, and Capstone Projects
Procedurally, this project aims to take a dataset, analyze it, and offer insights to the audience in an easy-to-digest format. Conceptually, this project will seek to explore questions like: “Do couples that meet through online dating or dating apps have higher or lower quality relationships?”, “Can any features in this dataset help predict how a subject would rate their relationship quality?”, and “What other insights can I derive from using machine learning for exploratory analysis?” The intended audience for this project is anyone interested in romantic relationships or machine learning.
The dataset is from a Stanford University survey, “How Couples …
Prediction Of Rapid Early Progression And Survival Risk With Pre-Radiation Mri In Who Grade 4 Glioma Patients, Walia Farzana, Mustafa M. Basree, Norou Diawara, Zeina Shboul, Sagel Dubey, Marie M. Lockheart, Mohamed Hamza, Joshua D. Palmer, Khan Iftekharuddin
Prediction Of Rapid Early Progression And Survival Risk With Pre-Radiation Mri In Who Grade 4 Glioma Patients, Walia Farzana, Mustafa M. Basree, Norou Diawara, Zeina Shboul, Sagel Dubey, Marie M. Lockheart, Mohamed Hamza, Joshua D. Palmer, Khan Iftekharuddin
Electrical & Computer Engineering Faculty Publications
Rapid early progression (REP) has been defined as increased nodular enhancement at the border of the resection cavity, the appearance of new lesions outside the resection cavity, or increased enhancement of the residual disease after surgery and before radiation. Patients with REP have worse survival compared to patients without REP (non-REP). Therefore, a reliable method for differentiating REP from non-REP is hypothesized to assist in personlized treatment planning. A potential approach is to use the radiomics and fractal texture features extracted from brain tumors to characterize morphological and physiological properties. We propose a random sampling-based ensemble classification model. The proposed …
Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia
Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia
SMU Data Science Review
In this paper, machine learning techniques are used to reconstruct particle collision pathways. CERN (Conseil européen pour la recherche nucléaire) uses a massive underground particle collider, called the Large Hadron Collider or LHC, to produce particle collisions at extremely high speeds. There are several layers of detectors in the collider that track the pathways of particles as they collide. The data produced from collisions contains an extraneous amount of background noise, i.e., decays from known particle collisions produce fake signal. Particularly, in the first layer of the detector, the pixel tracker, there is an overwhelming amount of background noise that …
Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel
Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel
SMU Data Science Review
Since the pandemic started, researchers have been trying to find a way to detect COVID-19 which is a cost-effective, fast, and reliable way to keep the economy viable and running. This research details how chest X-ray radiography can be utilized to detect the infection. This can be for implementation in Airports, Schools, and places of business. Currently, Chest imaging is not a first-line test for COVID-19 due to low diagnostic accuracy and confounding with other viral pneumonia. Different pre-trained algorithms were fine-tuned and applied to the images to train the model and the best model obtained was fine-tuned InceptionV3 model …
Combining Phenotypic And Genomic Data To Improve Prediction Of Binary Traits, Diego Jarquin, Arkaprava Roy, Bertrand S. Clarke, Subhashis Ghosal
Combining Phenotypic And Genomic Data To Improve Prediction Of Binary Traits, Diego Jarquin, Arkaprava Roy, Bertrand S. Clarke, Subhashis Ghosal
Department of Statistics: Faculty Publications
Plant breeders want to develop cultivars that outperform existing genotypes. Some characteristics (here ‘main traits’) of these cultivars are categorical and difficult to measure directly. It is important to predict the main trait of newly developed genotypes accurately. In addition to marker data, breeding programs often have information on secondary traits (or ‘phenotypes’) that are easy to measure. Our goal is to improve prediction of main traits with interpretable relations by combining the two data types using variable selection techniques. However, the genomic characteristics can overwhelm the set of secondary traits, so a standard technique may fail to select any …
A Strategy To Identify Event Specific Hospitalizations In Large Health Claims Databases, Joshua Lambert, Harpal Sandhu, Emily Kean, Teenu Xavier, Aviv Brokman, Zachary Steckler, Lee Park, Arnold Stromberg
A Strategy To Identify Event Specific Hospitalizations In Large Health Claims Databases, Joshua Lambert, Harpal Sandhu, Emily Kean, Teenu Xavier, Aviv Brokman, Zachary Steckler, Lee Park, Arnold Stromberg
Statistics Faculty Publications
Background: Health insurance claims data offer a unique opportunity to study disease distribution on a large scale. Challenges arise in the process of accurately analyzing these raw data. One important challenge to overcome is the accurate classification of study outcomes. For example, using claims data, there is no clear way of classifying hospitalizations due to a specific event. This is because of the inherent disjointedness and lack of context that typically come with raw claims data.
Methods: In this paper, we propose a framework for classifying hospitalizations due to a specific event. We then tested this framework in …
Split Classification Model For Complex Clustered Data, Katherine Gerot
Split Classification Model For Complex Clustered Data, Katherine Gerot
Honors Program: Senior Projects (Public)
Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.
Shape-Based Classification Of Partially Observed Curves, With Applications To Anthropology, Gregory J. Matthews, Karthik Bharath, Sebastian Kurtek, Juliet K. Brophy, George K. Thiruvathukal, Ofer Harel
Shape-Based Classification Of Partially Observed Curves, With Applications To Anthropology, Gregory J. Matthews, Karthik Bharath, Sebastian Kurtek, Juliet K. Brophy, George K. Thiruvathukal, Ofer Harel
Computer Science: Faculty Publications and Other Works
We consider the problem of classifying curves when they are observed only partially on their parameter domains. We propose computational methods for (i) completion of partially observed curves; (ii) assessment of completion variability through a nonparametric multiple imputation procedure; (iii) development of nearest neighbor classifiers compatible with the completion techniques. Our contributions are founded on exploiting the geometric notion of shape of a curve, defined as those aspects of a curve that remain unchanged under translations, rotations and reparameterizations. Explicit incorporation of shape information into the computational methods plays the dual role of limiting the set of all possible completions …
Bayesian Nonparametric Model For Functional Data Analysis, Tahmidul Islam
Bayesian Nonparametric Model For Functional Data Analysis, Tahmidul Islam
Theses and Dissertations
Functional data analysis (FDA) experienced a burst of growth after Ramsay and Silverman published their textbook in 1997. Functional data analysis interests researchers because of the challenges it adds to well-established multivariate analysis. Unlike finite dimensional random vectors, we visualize infinite dimensional random functions; for example, curves, images, brain scans, etc. A vast amount of literature have been dedicated to developing models for functional data. The ideas are mostly based on basis function representations and kernel-based nonparametric methods. In this dissertation, we propose a Bayesian treatment of nonparametric functional data analysis by introducing a Gaussian process (GP) over the space …
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Dissertations
In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.
First, to improve the prediction accuracy of learning …
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Student Works (2020-2029)
One of the approaches for structural health monitoring (SHM) consists of two major components, i.e. a network of sensors to collect the response data and an extraction method to obtain information on the structural health condition. Data mining (DM) is a novel data extraction technology which can employ for development of inverse analysis. Implementation of DM techniques in different areas of civil engineering has recently given very good results. However, application of DM in SHM is not used as much as expected, thus, many challenges are still ahead. Therefore, it is necessary to develop the applicability of DM in SHM. …
A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley
A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley
Theses and Dissertations
According to the Centers for Disease Control and Prevention, about 18.2 million adults age 20 and older have Coronary Artery Disease in the United States. Early diagnosis is therefore of crucial importance to help prevent debilitating consequences, and principally death for many patients. In this study we use data containing gene expression values from peripheral blood samples in 198 non-diabetic patients, with the goal of developing an age and sex gene expression model for diagnosis of Coronary Artery Disease. We employ machine learning methods to obtain a classification based on genetic information, age and sex. Our implementation uses feed forward …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Gradient Boosting For Survival Analysis With Applications In Oncology, Nam Phuong Nguyen
Gradient Boosting For Survival Analysis With Applications In Oncology, Nam Phuong Nguyen
USF Tampa Graduate Theses and Dissertations
Cancer is one of the most deadly diseases that the world has been fighting against over decades. An enormous number of research has been conducted, via a wide scale of approaches, raging from genetic analysis to mathematical modeling. Survival analysis is a well-performed methodology frequently used to estimate the survival probability of a patient. Although there has been a large number of methods for survival analysis, efficient exploration of a high-dimensional feature space has been challenging due to its computational cost and complexity. This thesis adapts the component-wise gradient boosting algorithms for cancer survival analysis, and also proposes a new …