Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (104)
- Artificial Intelligence and Robotics (80)
- Engineering (41)
- Medicine and Health Sciences (36)
- Statistics and Probability (36)
-
- Life Sciences (27)
- Social and Behavioral Sciences (23)
- Bioinformatics (19)
- Biomedical Informatics (19)
- Statistical Models (14)
- Applied Mathematics (13)
- Business (13)
- Databases and Information Systems (13)
- Mathematics (13)
- Electrical and Computer Engineering (12)
- Software Engineering (12)
- Applied Statistics (11)
- Other Computer Sciences (11)
- Theory and Algorithms (11)
- Numerical Analysis and Scientific Computing (9)
- Computer Engineering (8)
- Diseases (8)
- Earth Sciences (7)
- Longitudinal Data Analysis and Time Series (7)
- Medical Specialties (7)
- Numerical Analysis and Computation (7)
- Business Analytics (6)
- Business Intelligence (6)
- Institution
-
- Southern Methodist University (17)
- The Texas Medical Center Library (17)
- City University of New York (CUNY) (12)
- California Polytechnic State University, San Luis Obispo (11)
- Chapman University (9)
-
- California State University, San Bernardino (8)
- Technological University Dublin (8)
- Clemson University (7)
- Dartmouth College (7)
- Claremont Colleges (6)
- West Virginia University (6)
- Embry-Riddle Aeronautical University (5)
- Louisiana State University (5)
- Kennesaw State University (4)
- LSU New Orleans (4)
- San Jose State University (4)
- The University of Akron (4)
- Universitas Negeri Malang (4)
- CCT College Dublin (3)
- East Tennessee State University (3)
- Georgia Southern University (3)
- Mississippi State University (3)
- New Jersey Institute of Technology (3)
- Purdue University (3)
- The University of Southern Mississippi (3)
- University of Arkansas, Fayetteville (3)
- University of Connecticut (3)
- University of Montana (3)
- Washington University in St. Louis (3)
- Central Washington University (2)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (17)
- SMU Data Science Review (16)
- Master's Theses (9)
- Electronic Theses, Projects, and Dissertations (8)
- CMC Senior Theses (6)
-
- Dissertations, Theses, and Capstone Projects (6)
- Dissertations (5)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (5)
- LSU Doctoral Dissertations (5)
- All Dissertations (4)
- Computational and Data Sciences (PhD) Dissertations (4)
- Conference papers (4)
- Knowledge Engineering and Data Science (4)
- LSU New Orleans Theses and Dissertations (4)
- Williams Honors College, Honors Research Projects (4)
- All Theses (3)
- College of Engineering Summer Undergraduate Research Program (3)
- College of Graduate Studies: Theses & Dissertations (3)
- Computational and Data Sciences (MS) Theses (3)
- Dartmouth College Ph.D Dissertations (3)
- Dartmouth College Undergraduate Theses (3)
- Data Science Undergraduate Honors Theses (3)
- Discovery Undergraduate Interdisciplinary Research Internship (3)
- Dissertations and Theses (3)
- Doctoral Dissertations and Master's Theses (3)
- Electronic Theses and Dissertations (3)
- Graduate Student Theses, Dissertations, & Professional Papers (3)
- ICT (3)
- Theses and Dissertations (3)
- All Graduate Theses, Dissertations, and Other Capstone Projects (2)
- Publication Type
Articles 1 - 30 of 216
Full-Text Articles in Data Science
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Dissertations, Theses, and Capstone Projects
Jamaica Bay, located along the southeastern coast of New York City, acts as a biodiverse estuary of wetlands, meadows, and salt marsh islands. The purpose of this study is to analyze the water quality conditions of the region over time, comparing locations around the bay to identify hyperlocal features that influence larger trends in the hydrological system. Ten variables were used as water quality indicators, including total Kjeldahl nitrogen, salinity, pH, Secchi disk depth, and total phosphorus, among others, across five stations in the bay, between 1994 and 2024. After data cleaning and standardization methods were applied, principal component analysis …
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Dissertations, Theses, and Projects
The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Dissertations, Theses, and Capstone Projects
About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …
Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos
Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos
LSU Doctoral Dissertations
The adoption of conservation management practices is critical for improving soil health, enhancing nutrient use efficiency, and sustaining crop productivity in row crop systems in Louisiana. This study evaluated the role of conservation agronomic practices, soil biochemical indicators, and machine learning predictive models to improve soil nutrient dynamics, soil health indicators, microbial communities (MC), and crop productivity on a corn (Zea mays L.) research plot scale and in a commercial forty-hectare cotton (Gassypium hirsutum L.)-corn-soybean (Glycine max L.) rotation system in northeast Louisiana. The objectives of the study were to evaluate soil nutrient dynamics and MCs under …
Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega
Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega
Honors Theses
Running gait analysis plays a critical role in injury prevention and performance optimization, however, existing approaches often rely on specialized laboratory equipment or wearable sensors with limited interpretability. Recent advances in computer vision, particularly 2D human pose estimation, enable markerless motion analysis from standard video. However, progress remains constrained by the lack of publicly available datasets designed for running form analysis.
In this work, we introduce a preliminary dataset and benchmark for stride-level running gait analysis. The dataset consists of 73 treadmill running videos from 15 participants with varying experience levels, annotated with over 4,600 stride-level labels across multiple biomechanical …
Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones
Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones
Data Science Undergraduate Honors Theses
When companies acquire beverage brands, they typically value them based on total sales revenue. This traditional approach treats all sales equally over time, whether they are driven by genuine consumer demand or temporary discounts. This is important because while promotions can boost short-term sales, they tend to erode brand value over long periods of time. The measurement problem extends to acquisitions, where buyers lack the tools to distinguish real consumer demand from artificial promotional inflation.
This thesis develops a framework to separate genuine baseline demand from promotional dependence using Nielsen scanner data covering 189 beverage brands across 188,304 weekly observations …
Scenarioxp: A Complete Scenario-Based Testing Framework For The Exploration And Exploitation Of Autonomous Vehicle Validation Scenarios, Quentin Goss
Doctoral Dissertations and Master's Theses
Today is an age of exciting emerging technology where cutting-edge research in autonomous vehicles (AVs) reduces the active human participation in driving and extends awareness beyond human limitations of perception and reaction, improving driving safety and quality of the user experience as a result. The ever-increasing complexity of these autonomous systems poses many challenges towards the validation and verification (V\&V) of these complex systems under time and resource constraints, as the use of artificial intelligence and also the intricacy of the operating environment means that these systems are also black-box and non-deterministic. Scenario-based V\&V testing of such systems, which involves …
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
Dartmouth College Ph.D Dissertations
Multimodal large language models have achieved impressive performance on vision-language benchmarks by integrating visual encoders with large language models. Yet a critical gap persists between benchmark accuracy and genuine multimodal understanding: current evaluation frameworks assess performance by final answers alone, rewarding confident predictions while leaving systematic reasoning failures undetected.
This thesis addresses this gap through a unified framework that progresses from understanding to reasoning, using video as the most comprehensive multimodal testbed. Video inherently combines vision, audio, and language with temporal dynamics and massive token redundancy; techniques developed for video's comprehensive challenges transfer naturally to simpler multimodal tasks.
On understanding …
Spatial Temporal Modeling Of Infectious Disease Patterns In Texas, Robert E. Lashbrook
Spatial Temporal Modeling Of Infectious Disease Patterns In Texas, Robert E. Lashbrook
Earth & Environmental Sciences Theses
The Texas Department of State Health Services monitors numerous notifiable conditions statewide, including Campylobacter, Salmonella, Shiga toxin-producing Escherichia coli (STEC), Rabies, and West Nile virus (WNV). Given the substantial health, economic, and public health burden associated with these conditions, improving prediction is an important step toward reducing their overall impact. This study evaluated whether external demographic, social, climate, and environmental data could improve prediction of county-year disease activity across Texas. County level data was analyzed using supervised machine learning models, including linear regression, ridge regression, multilayer perceptron, random forest, XGBoost, as well as K-means clustering to identify broader …
Blens: Biomedical Literature Extraction And Scoring System, Tyler J. Simone
Blens: Biomedical Literature Extraction And Scoring System, Tyler J. Simone
Honors Theses and Capstones
Systematic reviews and meta-analyses represent the gold standard for evidence synthesis in healthcare, yet their manual execution remains labor-intensive, time-consuming, and vulnerable to human bias. With the exponential growth of biomedical literature, traditional literature screening and analysis has become increasingly unstable and noncomprehensive. This thesis presents the development and validation of an automate literature gathering and review system that integrates multiple scientific databases through a unified desktop application. The platform combines APIs from PubMed (NCBI Entrez), ClinicalTrials.gov, bioRxiv and medRxiv to enable simultaneous, standardized searching across peerreviewed and preprint sources. Built in Python with a PySide6 graphical interface, this standalone …
Network-Aware Airline-Specific Flight Delay Prediction Using Tree-Based Ensemble Models, Mary Dufie Afrane
Network-Aware Airline-Specific Flight Delay Prediction Using Tree-Based Ensemble Models, Mary Dufie Afrane
College of Graduate Studies: Theses & Dissertations
Flight delays pose persistent challenges to the efficiency and reliability of air transportation systems, affecting airlines, airports, regulators, and passengers alike. As traffic demand grows and operational environments become increasingly interconnected, accurately predicting both departure and arrival delays has become crucial for effective planning and mitigation. This study presents a network-aware, airline-specific framework for predicting flight delays in U.S. domestic air transportation systems using tree-based ensemble machine learning models. A large-scale dataset of 1.98 million flights, enriched with weather information, is used to develop predictive models for both departure and arrival delays. To capture the structural and operational complexity of …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
Topic Modeling The Cuny Graduate Center's Dissertations And Theses, Michael Mandiberg
Topic Modeling The Cuny Graduate Center's Dissertations And Theses, Michael Mandiberg
Open Educational Resources
This 4 week module is designed for Data Analysis, Data Visualization, and Digital Humanities courses at the MA/MS or advanced 400-level undergraduate level. It introduces students to textual analysis with topic modeling and requires a solid foundation in Python. The module uses Gensim and a Colab notebook to introduce a standard text analysis workflow used in Digital Humanities, archival research, and exploratory data analysis.
Students build a topic model describing 19,000 CUNY Graduate Center dissertations and theses. They work with an unexplored dataset to load and explore the data, prepare the corpus, train and evaluate a topic model, and interpret, …
Visualizing And Evaluating Binary Classifier Performance With Contingency Space, Colin D. Kehoe, Azim Ahmadzadeh
Visualizing And Evaluating Binary Classifier Performance With Contingency Space, Colin D. Kehoe, Azim Ahmadzadeh
Undergraduate Research Symposium
Traditional metrics for evaluating binary classifiers, such as Accuracy, F1 Score, and True Skill Statistic (TSS), often obscure the underlying tradeoffs between true positive and true negative performance—particularly in imbalanced or high-stakes domains. This poster introduces the Contingency Space, a two-dimensional representation of classifier behavior defined by true positive rate (TPR) and true negative rate (TNR). Within this space, scalar performance metrics become geometric surfaces, revealing how scores vary across the entire landscape of possible classifier outputs.
We present a Python package that implements this framework, enabling users to map model predictions into the Contingency Space, visualize metric surfaces …
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Undergraduate Research Symposium
Various computational models of first impressions have been developed to uncover the mechanisms driving these judgments. However, the implicit notion of a singular ``human'' often overlooks meaningful individual differences in beliefs, attitudes, and associations, as well as culturally grounded group-level constructs. In this paper, we extend Cultural Consensus Theory (CCT) to estimate culturally shared beliefs about faces by incorporating latent constructs structured around interpretable facial features extracted via computer vision algorithms. We apply our model to a large-scale dataset of people’s first impressions of faces. Our approach reveals a robust mapping between facial features and culturally constructed impressions, allowing us …
Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik
Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik
Honors Projects
Liver Transplantations are crucial treatment for end-stage liver disease. However, a persistent deficit of donor organs necessitates maximizing the utility of each available graft to minimize failure rates. We evaluated whether donor–recipient molecular immunogenicity metrics - Electrostatic and Hydrophobic Mismatch Scores (HMS/EMS) and eplet-based counts - improve post–liver-transplant survival prediction. The analytic cohort comprised adult, first time, single-organ deceased-donor transplants drawn from Scientific Registry of Transplant Recipients; follow-up was truncated at five years, and the endpoint was all-cause graft failure (earliest of graft failure or death; otherwise, censored). HLA variables were derived via high- resolution conversion and molecular mismatch computations …
Atlas Of Ai: Power, Politics And The Planetary Costs Of Artificial Intelligence - Book Review, Jelena Popov
Atlas Of Ai: Power, Politics And The Planetary Costs Of Artificial Intelligence - Book Review, Jelena Popov
Feminist Pedagogy
No abstract provided.
Sociohydrodynamics: Data-Driven Modeling Of Social Behavior, Daniel S. Seara, Jonathan Colen, Michel Fruchart, Yael Avni, David G. Martin, Vincenzo Vitelli
Sociohydrodynamics: Data-Driven Modeling Of Social Behavior, Daniel S. Seara, Jonathan Colen, Michel Fruchart, Yael Avni, David G. Martin, Vincenzo Vitelli
Data Science Faculty Publications
Living systems display complex behaviors driven by physical forces as well as decision-making. Hydrodynamic theories hold promise for simplified universal descriptions of socially generated collective behaviors. However, the construction of such theories is often divorced from the data they should describe. Here, we develop and apply a data-driven pipeline that links micromotives to macrobehavior by augmenting hydrodynamics with individual preferences that guide motion. We illustrate this pipeline on a case study of residential dynamics in the United States, for which census and sociological data are available. Guided by Census data, sociological surveys, and neural network analysis, we systematically assess standard …
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Discovery Undergraduate Interdisciplinary Research Internship
Understanding and accurately predicting crop yield is becoming increasingly important today in the face of global food security challenges, and thus, the availability of standardized data and scalable models is the need of the hour. To support this, researchers have developed CY-Bench (Crop Yield Benchmark), a comprehensive dataset that helps forecast maize and wheat yields on a global scale. This research project primarily involved working with the CY-Bench dataset aiming to improve crop yield prediction through machine learning. Initially, papers explaining the CY-Bench dataset and other papers for agriculture modeling were studied and analyzed in detail. The research then progressed …
Discovering And Designing Novel Perovskite Photovoltaic Materials Via Machine Learning, Junyeong Ahn
Discovering And Designing Novel Perovskite Photovoltaic Materials Via Machine Learning, Junyeong Ahn
Discovery Undergraduate Interdisciplinary Research Internship
Perovskite semiconductors are promising materials for high-efficiency photovoltaics due to their outstanding optoelectronic properties, emerging as a sustainable energy source through solar cell applications. Perovskites with the ABX₃ composition (A, B = metal or organic cations with varying oxidation states; X = chalcogen or halogen anions) have gained interest for their excellent phase stability and compositional tunability. However, combinatorial possibilities arising from the many choices of A, B, and X site species, and their respective mixing fractions, a large number of possible ABX₃ perovskites remain undiscovered. In this work, we used machine learning (ML) methods to design new stable and …
Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang
Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang
Lingnan Theses (MPhil & PhD)
Accurate forecasting of metro passenger flow is vital for efficient urban transportation management and optimal resource allocation in modern cities. Traditional ARIMA-based models effectively capture regular, cyclical patterns but struggle with sudden, nonlinear fluctuations caused by random events such as weather disruptions, special events, or service interruptions. Moreover, existing research predominantly focuses on individual stations, overlooking the complex cross-station interactions inherent in networked metro systems where passenger flows are interconnected across the entire network.
To address these critical limitations, we propose the Three-Stage Latent Dynamics Forecasting (T-LDF) Framework, a novel approach that systematically integrates temporal decomposition, latent dynamics extraction, and …
A Study Of Machine Learning Techniques In Solving Biochemical And Chemical Problems, Kenneth Micheal Plackowski
A Study Of Machine Learning Techniques In Solving Biochemical And Chemical Problems, Kenneth Micheal Plackowski
Chemistry and Chemical Biology ETDs
Data-driven approaches to solving problems in biology and chemistry require utilization of reliable techniques and machine learning algorithms are the modern reliable approach. This work presents three problems that involve use of supervised learning techniques when classification is the goal and unsupervised learning techniques when global data representation is the goal.
In the first problem, we demonstrate the use of unsupervised clustering techniques, self-organizing maps and K-means, to ascertain analyte detection capabilities of carbon nitride dots. In the second problem, we add scalability features to a functional group classification model applied to infrared data and evaluate its ability to inform …
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
Beyond: Undergraduate Research Journal
Flight time prediction plays a crucial role in modern air travel, benefiting airlines and passengers alike. Accurate predictions enable airlines to optimize schedules, allocate resources effectively, and ensure passenger safety and satisfaction. In recent years, machine learning models, such as neural networks and XGBoost, have gained popularity for predicting flight times. This study aims to compare the performance of neural network and XGBoost models in predicting flight times, considering factors such as weather conditions, air traffic control, and aircraft performance. The results indicate that both models are effective, with XGBoost achieving slightly higher accuracy. However, neural networks offer advantages in …
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Master's Theses
The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Dartmouth College Ph.D Dissertations
In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Capstone Projects
This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …
Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi
Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi
Open Educational Resources
This open-access machine learning course is a comprehensive 15-week curriculum developed and published on GitHub with full Google Colab compatibility. It combines theoretical concepts with hands-on Python coding, real-world datasets, and structured projects covering regression, classification, clustering, deep learning, transformers, and multimodal AI. The course is designed for students, educators, and researchers interested in applied machine learning, including biomedical applications. It includes explainable AI components and ethical discussions to align with modern AI standards. The course is maintained by BioMind AI Lab at CUNY.
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
Undergraduate Honors Theses
Machine learning is a method that employs statistical algorithms to identify patterns and make predictions from data. This study applies machine learning techniques to analyze data from Major League Baseball (MLB) teams between 1998 and 2024, with the goal of determining which factors strongly influence a team's likelihood of reaching the postseason and in accurately predicting the teams that do and do not qualify for the postseason. Data exploration and unsupervised machine learning methods such as clustering were used to identify underlying patterns in team performance metrics and determine potential significant contributors to team success. Many different supervised learning methods …
From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins
From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins
Electrical Engineering and Computer Science (MS) Theses
This study examines the relationship between climate conditions and residential water usage, focusing on how seasonal and environmental changes influence water consumption. Utilizing data from over 100,000 households across three micro-climate zones for over a five-year period, we apply statistical analysis and machine learning techniques to assess the impact of temperature, precipitation, evapotranspiration, and location on water usage. By integrating climate and billing data, this research provides a data-driven approach on water usage behaviors in Irvine, CA, in collaboration with Irvine Ranch Water District (IRWD).
Our analysis utilizes time series modeling, including a Seasonal Autoregressive Integrated Moving Average (SARIMA) and …
An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck
An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck
Electrical Engineering and Computer Science (MS) Theses
Closed-source large language models (LLMs) developed by large technology companies continue to grow in popularity. However, ethical conversations surrounding the safety of model outputs have been a prominent topic of discussion. This project aims to assess three leading closed-source LLMs: OpenAI’s ChatGPT, Google’s Gemini, and Anthropic’s Claude, to analyze how their outputs perform when treated as a subject of several psychological evaluation scales measuring biased behaviors against women. The Ambivalent Sexism Index, Modern Sexism Scale, and Belief in Sexism Shift evaluations were used to get descriptions of how the LLMs respond to traditional and modern prompts involving sexism and gender …