Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (45)
- Data Science (36)
- Statistical Models (32)
- Artificial Intelligence and Robotics (28)
- Applied Statistics (27)
-
- Mathematics (18)
- Engineering (17)
- Applied Mathematics (14)
- Numerical Analysis and Scientific Computing (13)
- Statistical Methodology (12)
- Longitudinal Data Analysis and Time Series (11)
- Categorical Data Analysis (10)
- Medicine and Health Sciences (10)
- Probability (10)
- Biostatistics (9)
- Multivariate Analysis (9)
- Life Sciences (8)
- Social and Behavioral Sciences (8)
- Business (7)
- Databases and Information Systems (7)
- Computer Engineering (6)
- Other Statistics and Probability (6)
- Theory and Algorithms (6)
- Electrical and Computer Engineering (5)
- Analysis (4)
- Bioinformatics (4)
- Business Analytics (4)
- Earth Sciences (4)
- Institution
-
- Southern Methodist University (11)
- University of South Florida (6)
- Utah State University (6)
- Claremont Colleges (5)
- Technological University Dublin (4)
-
- West Virginia University (4)
- California Polytechnic State University, San Luis Obispo (3)
- Georgia Southern University (3)
- University of Arkansas, Fayetteville (3)
- University of Kentucky (3)
- COBRA (2)
- California State University, San Bernardino (2)
- City University of New York (CUNY) (2)
- East Tennessee State University (2)
- Embry-Riddle Aeronautical University (2)
- Kennesaw State University (2)
- Michigan Technological University (2)
- Missouri University of Science and Technology (2)
- Northern Illinois University (2)
- Purdue University (2)
- The University of Akron (2)
- University at Albany, State University of New York (2)
- University of North Florida (2)
- University of Texas at El Paso (2)
- Belmont University (1)
- Bowling Green State University (1)
- Chapman University (1)
- Clemson University (1)
- Dartmouth College (1)
- Duquesne University (1)
- Publication Year
- Publication
-
- SMU Data Science Review (11)
- USF Tampa Graduate Theses and Dissertations (6)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (4)
- Electronic Theses and Dissertations (4)
- CMC Senior Theses (3)
-
- College of Graduate Studies: Theses & Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- Master's Theses (3)
- SAML-25 Workshop on Statistical and Machine Learning (3)
- Theses and Dissertations (3)
- Data Science Undergraduate Honors Theses (2)
- Dissertations, Master's Theses and Master's Reports (2)
- Dissertations, Theses, and Capstone Projects (2)
- Doctor of Data Science and Analytics Dissertations (2)
- Electronic Theses, Projects, and Dissertations (2)
- Graduate Research Theses & Dissertations (2)
- Open Access Theses & Dissertations (2)
- U.C. Berkeley Division of Biostatistics Working Paper Series (2)
- UNF Graduate Theses and Dissertations (2)
- Williams Honors College, Honors Research Projects (2)
- All Dissertations (1)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (1)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- Beyond: Undergraduate Research Journal (1)
- CBER Conference (1)
- CGU Theses & Dissertations (1)
- Capstone Projects (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Conference papers (1)
- Dartmouth College Ph.D Dissertations (1)
- Publication Type
Articles 1 - 30 of 101
Full-Text Articles in Statistics and Probability
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Dissertations, Theses, and Capstone Projects
Jamaica Bay, located along the southeastern coast of New York City, acts as a biodiverse estuary of wetlands, meadows, and salt marsh islands. The purpose of this study is to analyze the water quality conditions of the region over time, comparing locations around the bay to identify hyperlocal features that influence larger trends in the hydrological system. Ten variables were used as water quality indicators, including total Kjeldahl nitrogen, salinity, pH, Secchi disk depth, and total phosphorus, among others, across five stations in the bay, between 1994 and 2024. After data cleaning and standardization methods were applied, principal component analysis …
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Dissertations, Theses, and Projects
The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci
Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci
Theses and Dissertations--Electrical and Computer Engineering
Fine-grained Temporal Action Segmentation (TAS) has become a cornerstone of video understanding, offering dense frame-level predictions essential for clinical assessment, surgical skill evaluation, and human-computer interaction. While TAS methods have delivered strong results on coarse-grained benchmarks, two fundamental challenges persist: (1) global attention mechanisms dilute boundary information critical for subsecond precision, a phenomenon we term the temporal granularity bottleneck, and (2) dense frame-level annotation remains prohibitively expensive, with most datasets requiring exhaustive labeling of lengthy untrimmed videos. These challenges are particularly pronounced in medical domains, where sub-second primitives define clinical outcomes while expert annotation remains scarce. In this dissertation, we …
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
UNF Graduate Theses and Dissertations
Liver cirrhosis is associated with substantial morbidity and mortality, making one-year mortality prediction a clinically relevant problem. Using a liver cirrhosis dataset as the motivating application, this thesis evaluates five machine learning classifiers—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—under five class-imbalance handling strategies: Baseline learning, Random Oversampling, SMOTE-NC, ADASYN, and Cost-Sensitive Learning. Hyperparameter tuning was conducted using randomized search, and predictive performance was assessed over 200 iterations of Monte Carlo Cross-Validation using Accuracy, Precision, Recall, Fl-score, and ROC-AUC.
The results suggest that imbalance-handling strategies can materially affect predictive performance, particularly recall. Because the outcome of interest is death within …
Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams
Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams
UNF Graduate Theses and Dissertations
We compare five numerical approaches for approximating solutions to the Black–Scholes partial differential equation for pricing European call options: FTCS, BTCS, Crank– Nicolson, Monte Carlo simulation, and a physics–informed neural network (PINN). These methods span finite difference techniques, probabilistic simulation, and machine learning. Performance is evaluated based on computational efficiency and accuracy relative to the analytical Black–Scholes solution.
Among the methods, Crank–Nicolson and the PINN demonstrated the strongest overall performance. Crank–Nicolson achieved the highest accuracy but exhibited increased runtime as the number of underlying stock price grid points grew. In contrast, the PINN produced slightly less accurate results but with …
Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik
Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik
Honors Projects
Liver Transplantations are crucial treatment for end-stage liver disease. However, a persistent deficit of donor organs necessitates maximizing the utility of each available graft to minimize failure rates. We evaluated whether donor–recipient molecular immunogenicity metrics - Electrostatic and Hydrophobic Mismatch Scores (HMS/EMS) and eplet-based counts - improve post–liver-transplant survival prediction. The analytic cohort comprised adult, first time, single-organ deceased-donor transplants drawn from Scientific Registry of Transplant Recipients; follow-up was truncated at five years, and the endpoint was all-cause graft failure (earliest of graft failure or death; otherwise, censored). HLA variables were derived via high- resolution conversion and molecular mismatch computations …
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Discovery Undergraduate Interdisciplinary Research Internship
Understanding and accurately predicting crop yield is becoming increasingly important today in the face of global food security challenges, and thus, the availability of standardized data and scalable models is the need of the hour. To support this, researchers have developed CY-Bench (Crop Yield Benchmark), a comprehensive dataset that helps forecast maize and wheat yields on a global scale. This research project primarily involved working with the CY-Bench dataset aiming to improve crop yield prediction through machine learning. Initially, papers explaining the CY-Bench dataset and other papers for agriculture modeling were studied and analyzed in detail. The research then progressed …
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
All Graduate Theses and Dissertations, Fall 2023 to Present
Classification tasks are fundamental in statistical machine learning. In classification tasks, a general goal is to build or select a model that can correctly classify data with as few errors as possible. However, for a particular dataset, the minimal number of errors achievable is seldom zero since overlap in the data makes errors unavoidable. As a result, it is often difficult for machine learning practitioners and data scientists to know whether classification errors can be reduced through further refinement. A potential solution to this lies in the Bayes error rate (BER). The BER is the lowest error rate achievable for …
Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang
Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang
Lingnan Theses (MPhil & PhD)
Accurate forecasting of metro passenger flow is vital for efficient urban transportation management and optimal resource allocation in modern cities. Traditional ARIMA-based models effectively capture regular, cyclical patterns but struggle with sudden, nonlinear fluctuations caused by random events such as weather disruptions, special events, or service interruptions. Moreover, existing research predominantly focuses on individual stations, overlooking the complex cross-station interactions inherent in networked metro systems where passenger flows are interconnected across the entire network.
To address these critical limitations, we propose the Three-Stage Latent Dynamics Forecasting (T-LDF) Framework, a novel approach that systematically integrates temporal decomposition, latent dynamics extraction, and …
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
Beyond: Undergraduate Research Journal
Flight time prediction plays a crucial role in modern air travel, benefiting airlines and passengers alike. Accurate predictions enable airlines to optimize schedules, allocate resources effectively, and ensure passenger safety and satisfaction. In recent years, machine learning models, such as neural networks and XGBoost, have gained popularity for predicting flight times. This study aims to compare the performance of neural network and XGBoost models in predicting flight times, considering factors such as weather conditions, air traffic control, and aircraft performance. The results indicate that both models are effective, with XGBoost achieving slightly higher accuracy. However, neural networks offer advantages in …
Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh
Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh
University Honors Theses
This study evaluates ChatGPT's ability to forecast influenza rates, such as the number of flu cases, hospitalizations, and death during peak season periods using CDC data, and comparing forecasts against actual results to calculate statistical accuracy and consistency. Influenza forecasting is essential for public health planning, but traditional methods may not always provide timely or accurate predictions. In this research study, ChatGPT was utilized to predict the influenza rates for the following week based on the previous week's data obtained from the FluView surveillance system. The predicted rates were compared to the actual influenza rates to assess the model's overall …
Ai-Driven Personalized Radiotherapy Planning, Nithin Venkatesh, Marco Pota, Maged Shaban
Ai-Driven Personalized Radiotherapy Planning, Nithin Venkatesh, Marco Pota, Maged Shaban
SAML-25 Workshop on Statistical and Machine Learning
The planning of radiation oncology treatment is made more dynamic and individualized by Artificial Intelligence (AI). Routine radiotherapy practice applies normative procedures indifferent to patient-specific parameters such as tumor volume, patient anatomy, and heterogeneity in the delineation of treatment response. Inadequate and over-radiation treatment is the most prevalent outcome. Further, with the inclusion of AI, it can facilitate enhancing the healthcare industry through optimizing radiotherapy using an array of patient information such as molecular profiles and imaging data. The product offers an end-to-end AI-driven solution to all aspects of radiotherapy, from initial consultation (diagnosis) to adaptive treatment planning. All the …
Interpretable Ai In Education: A Comparison Of Glass-Box Models For Predicting Student Success, Jan Glazenborg
Interpretable Ai In Education: A Comparison Of Glass-Box Models For Predicting Student Success, Jan Glazenborg
SAML-25 Workshop on Statistical and Machine Learning
This Master’s thesis addresses early identification of first-year Computer Science students at risk of underperformance by comparing inherently interpretable (“glass-box”) predictive models with the existing Naïve Bayes–based PreSS tool. The PreSS dataset was originally compiled by Quille & Bergin from 692 first-year CS1 students across eleven institutions in Ireland and Denmark, who completed surveys on programming and mathematics backgrounds, gaming habits and a short programming test four to six hours into the course. Seventeen normalized features capturing demographic, academic and behavioural factors were extracted. In this thesis, four machine learning models are evaluated: Naïve Bayes, explainable boosting machines, automatic piecewise …
Benchmarking Energy And Performance Of Parallel Machine Learning Models Using Hardware And Software Power Meters, Urooj Asgher, Tania Malik
Benchmarking Energy And Performance Of Parallel Machine Learning Models Using Hardware And Software Power Meters, Urooj Asgher, Tania Malik
SAML-25 Workshop on Statistical and Machine Learning
The growing reliance on machine learning algorithms across domains such as healthcare, transportation, and finance has led to their increased deployment on high-performance computing platforms. While performance optimization remains a central concern, energy efficiency is emerging as a critical design consideration, particularly in light of global sustainability goals. This study presents a comparative analysis of the energy consumption and performance of serial and parallel implementations of four machine learning algorithms, K-means clustering, Ant Colony Optimization, Logistic Regression, and Random Search. Experiments were conducted on an HPC testbed using both hardware-based and software-based power meters to measure energy consumption. The results …
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Dartmouth College Ph.D Dissertations
In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Master's Theses
The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Capstone Projects
This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta
Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta
Open Access Theses & Dissertations
Prostate cancer (PrCa) remains a critical challenge in precision oncology due to several reasons including its apparent heterogenous condition, recurrence following treatment and rapid progressive forms. Therefore, identifying patients at risk of progression is essential to fast-track therapeutic decisions and improve outcomes. Despite recent advances in genomic and molecular profiling, conventional PrCa risk assessment tools heavily rely on a few clinical parameters, neglecting the prognostic potential of genomic biomarkers in the presence of clinical biomarkers. This study presents a computational pipeline to harmonize and evaluate the prognostic value of clinicogenomic profiles of patients in modelling progression free survival (PFS). PFS, …
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Theses and Dissertations
The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …
Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih
Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih
Electronic Theses and Dissertations
The objective of this study is to predict car prices using machine learning models and the DVM-CAR dataset, which includes over 1.4 million images and car specifi- cations from 899 car models. Key factors such as mileage, engine power, and year of registration were analyzed for their correlation with car prices. Extensive data cleaning was performed, including filling missing values, identifying outliers, and normalizing numerical variables. Discrete variables like car make and body type were encoded using one-hot encoding. Linear relationships were analyzed with Multiple Logistic Regression, and Random Forest models were used for nonlinear patterns. Model performance was evaluated …
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Honors College Theses
The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara
Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara
Journal of Environmental Science and Sustainable Development
The relocation of Indonesia's capital city is anticipated to promote inclusive economic growth while embracing cultural diversity. However, this transition may affect ultraviolet (UV) radiation exposure patterns. The study investigated variations in UV exposure in the IKN region, focusing on urban development factors such as land use and population density that affect public health, sun protection, and skin cancer prevention. The research hypothesized that UV radiation is significantly correlated with these factors. UV Index data from 2010-2023, a hierarchical clustering method, identifies complex data patterns without determining the number of clusters. XGBoost, a machine learning model, was used for handling …
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
CBER Conference
Data is the fundamental building block for advancements in artificial intelligence (AI), general AI (GAI), machine learning (ML), and large language models (LLMs). This study emphasizes the critical need for robust data infrastructure, arguing that without it, countries cannot fully benefit from technological advancements in various economic sectors. Governments possess vast repositories of both structured and unstructured data across multiple domains such as the judiciary, parliaments, and civil bureaucracy. However, these potential goldmines remain untapped due to inadequate data management capabilities and a lack of appreciation for the necessity of high-quality data. The research identifies key issues in public data …
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
Rose-Hulman Undergraduate Mathematics Journal
Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
All Dissertations
The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Data Science Undergraduate Honors Theses
Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …