Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (3)
- Data Science (3)
- Medicine and Health Sciences (3)
- Clinical Trials (2)
- Multivariate Analysis (2)
-
- Other Computer Sciences (2)
- Public Health (2)
- Statistical Methodology (2)
- Statistical Models (2)
- Analysis (1)
- Applied Statistics (1)
- Art and Design (1)
- Arts and Humanities (1)
- Biostatistics (1)
- Clinical Epidemiology (1)
- Digital Humanities (1)
- Diseases (1)
- Interactive Arts (1)
- Mathematics (1)
- Medical Immunology (1)
- Medical Sciences (1)
- Medical Specialties (1)
- Microarrays (1)
- Oncology (1)
- Other Physical Sciences and Mathematics (1)
- Other Public Health (1)
- Other Statistics and Probability (1)
- Institution
- Publication
- Publication Type
Articles 1 - 8 of 8
Full-Text Articles in Categorical Data Analysis
Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier
Variable Selection In Distance Metric Learning And Triplet Constraints For Deep Learning Based Ordinal Classification, James D. Clothier
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
The purpose of this research is to augment linear and kernelized ordinal distance metric learning (L/KODML) techniques with a proposed variable selection methodology that integrates the Sequential Multi-Response Feature Selection (SMuRFS) algorithm. Additionally, we aim to embed ordinal triplet constraints into a deep learning architecture, and to propose a general framework for deep learning-based ordinal classification. A variety of simulation studies and real data experiments were conducted to evaluate the various methodologies. For the distance metric learning and variable selection, results showed that the integration of SMuRFS performed effective variable selection and improved prediction accuracy. For the triplet constraints, incorporating …
Analyzing Relationships With Machine Learning, Oscar Ko
Analyzing Relationships With Machine Learning, Oscar Ko
Dissertations, Theses, and Capstone Projects
Procedurally, this project aims to take a dataset, analyze it, and offer insights to the audience in an easy-to-digest format. Conceptually, this project will seek to explore questions like: “Do couples that meet through online dating or dating apps have higher or lower quality relationships?”, “Can any features in this dataset help predict how a subject would rate their relationship quality?”, and “What other insights can I derive from using machine learning for exploratory analysis?” The intended audience for this project is anyone interested in romantic relationships or machine learning.
The dataset is from a Stanford University survey, “How Couples …
Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel
Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel
SMU Data Science Review
Since the pandemic started, researchers have been trying to find a way to detect COVID-19 which is a cost-effective, fast, and reliable way to keep the economy viable and running. This research details how chest X-ray radiography can be utilized to detect the infection. This can be for implementation in Airports, Schools, and places of business. Currently, Chest imaging is not a first-line test for COVID-19 due to low diagnostic accuracy and confounding with other viral pneumonia. Different pre-trained algorithms were fine-tuned and applied to the images to train the model and the best model obtained was fine-tuned InceptionV3 model …
Split Classification Model For Complex Clustered Data, Katherine Gerot
Split Classification Model For Complex Clustered Data, Katherine Gerot
Honors Program: Senior Projects (Public)
Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Identification Of Slums In Mumbai, India: Unsupervised Classification Techniques, Frankie St. Amand
Identification Of Slums In Mumbai, India: Unsupervised Classification Techniques, Frankie St. Amand
Thinking Matters Symposium Archive
Slums are contiguous settlements. Inhabitants lack access to safe water, sanitation and sewage infrastructure, secure housing tenure, uncrowded living space, and permanent, durable housing. Addressing these problematic trends begins with identifying contiguous settlements within Mumbai’s urban fabric. Classifications can be performed using satellite images and remote sensing techniques to yield accurate results. Through literature reviews, socio-cultural analysis, and examination of high resolution satellite imagery, this project aims to develop a systematic, accessible, and reproducible method of classifying Mumbai’s slums.
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
UW Biostatistics Working Paper Series
The Net Reclassification Index (NRI) is a very popular measure for evaluating the improvement in prediction performance gained by adding a marker to a set of baseline predictors. However, the statistical properties of this novel measure have not been explored in depth. We demonstrate the alarming result that the NRI statistic calculated on a large test dataset using risk models derived from a training set is likely to be positive even when the new marker has no predictive information. A related theoretical example is provided in which a miscalibrated risk model that includes an uninformative marker is proven to erroneously …
Data Mining Methods For Malware Detection, Muazzam Siddiqui
Data Mining Methods For Malware Detection, Muazzam Siddiqui
Electronic Theses and Dissertations
This research investigates the use of data mining methods for malware (malicious programs) detection and proposed a framework as an alternative to the traditional signature detection methods. The traditional approaches using signatures to detect malicious programs fails for the new and unknown malwares case, where signatures are not available. We present a data mining framework to detect malicious programs. We collected, analyzed and processed several thousand malicious and clean programs to find out the best features and build models that can classify a given program into a malware or a clean class. Our research is closely related to information retrieval …