Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (1156)
- Medicine and Health Sciences (780)
- Life Sciences (765)
- Bioinformatics (568)
- Statistics and Probability (550)
-
- Biomedical Informatics (530)
- Engineering (527)
- Artificial Intelligence and Robotics (525)
- Social and Behavioral Sciences (519)
- Databases and Information Systems (212)
- Computer Engineering (208)
- Electrical and Computer Engineering (204)
- Applied Statistics (194)
- Medical Sciences (190)
- Business (189)
- Statistical Models (181)
- Applied Mathematics (175)
- Medical Specialties (173)
- Theory and Algorithms (149)
- Environmental Sciences (148)
- Mathematics (144)
- Other Computer Sciences (127)
- Data Storage Systems (123)
- Systems and Communications (120)
- Numerical Analysis and Scientific Computing (116)
- Public Health (116)
- Public Affairs, Public Policy and Public Administration (109)
- Statistical Methodology (109)
- Institution
-
- The Texas Medical Center Library (523)
- Old Dominion University (173)
- Southern Methodist University (144)
- Universitas Negeri Malang (113)
- City University of New York (CUNY) (101)
-
- CCT College Dublin (82)
- Chapman University (66)
- Kennesaw State University (63)
- University of Central Florida (62)
- Smith College (60)
- Air Force Institute of Technology (57)
- Embry-Riddle Aeronautical University (52)
- Singapore Management University (45)
- University of Arkansas, Fayetteville (45)
- Chinese Academy of Sciences (44)
- Purdue University (44)
- California Polytechnic State University, San Luis Obispo (39)
- Technological University Dublin (39)
- Illinois State University (38)
- University of Kentucky (38)
- University of Nebraska - Lincoln (38)
- New Jersey Institute of Technology (37)
- West Virginia University (37)
- Claremont Colleges (36)
- Virginia Commonwealth University (35)
- Clemson University (32)
- Dartmouth College (31)
- University of Texas at Arlington (27)
- East Tennessee State University (26)
- Minnesota State University, Mankato (26)
- Keyword
-
- Humans (278)
- Machine learning (241)
- Machine Learning (216)
- Deep learning (115)
- Computer Science (98)
-
- Deep Learning (93)
- Artificial Intelligence (65)
- Data science (58)
- Data Science (57)
- Natural Language Processing (56)
- COVID-19 (55)
- Artificial intelligence (53)
- Female (52)
- Male (50)
- Classification (49)
- Natural language processing (46)
- Animals (41)
- Data (41)
- Electronic Health Records (41)
- Neural Networks (40)
- Algorithms (38)
- Big data (37)
- Data mining (37)
- Statistics (36)
- Clustering (32)
- Computer science (31)
- Adult (30)
- NLP (30)
- Neural networks (30)
- AI (29)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (508)
- SMU Data Science Review (124)
- Knowledge Engineering and Data Science (113)
- Theses and Dissertations (111)
- ICT (82)
-
- Data Science and Data Mining (53)
- Dissertations (53)
- Statistical and Data Sciences: Faculty Publications (53)
- Electronic Theses and Dissertations (49)
- Dissertations, Theses, and Capstone Projects (45)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (44)
- Research Collection School Of Computing and Information Systems (37)
- Master's Theses (35)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (34)
- Data Science Undergraduate Honors Theses (31)
- Annual Symposium on Biomathematics and Ecology Education and Research (30)
- Computer Science Faculty Publications (30)
- Publications and Research (30)
- All Graduate Theses, Dissertations, and Other Capstone Projects (24)
- Computational and Data Sciences (PhD) Dissertations (24)
- Symposium of Student Scholars (24)
- All Dissertations (23)
- Articles (23)
- Electrical & Computer Engineering Faculty Publications (22)
- CBN Journal of Applied Statistics (JAS) (21)
- College of Graduate Studies: Theses & Dissertations (20)
- CMC Senior Theses (19)
- Theses (19)
- Electronic Theses, Projects, and Dissertations (18)
- Faculty Publications (18)
- Publication Type
- File Type
Articles 391 - 420 of 3232
Full-Text Articles in Data Science
Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha
Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha
All Faculty Open Access Publications
Volatility forecasting for financial institutions plays a pivotal role across a wide range of domains, such as risk management, option pricing, and market making. For instance, banks can incorporate volatility forecasts into stress testing frameworks to ensure they are holding sufficient capital during extreme market conditions. However, volatility forecasting is challenging because volatility can only be estimated, and different factors influence volatility, ranging from macroeconomic indicators to investor sentiments. While recent works show promising advances in machine learning and artificial intelligence for volatility forecasting, a comprehensive assessment of current statistical and learning-based methods is lacking. Thus, this paper aims to …
Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao
Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao
Faculty, Staff and Student Publications
Bone metastasis is a major cause of cancer death; however, the epigenetic determinants driving this process remain elusive. Here, we report that histone methyltransferase ASH1L is genetically amplified and is required for bone metastasis in men with prostate cancer. ASH1L rewires histone methylations and cooperates with HIF-1α to induce pro-metastatic transcriptome in invading cancer cells, resulting in monocyte differentiation into lipid-associated macrophage (LA-TAM) and enhancing their pro-tumoral phenotype in the metastatic bone niche. We identified IGF-2 as a direct target of ASH1L/HIF-1α and mediates LA-TAMs' differentiation and phenotypic changes by reprogramming oxidative phosphorylation. Pharmacologic inhibition of the ASH1L-HIF-1α-macrophages axis elicits …
Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene
Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene
Physics & Astronomy ETDs
The dominant Lambda Cold Dark Matter (LCDM) cosmological model, while remarkably successful, increasingly shows signs that it may not fully describe our Universe, as persistent tensions in expansion rates and structure formation remain unresolved. In this thesis, I challenge the LCDM paradigm using novel theoretical frameworks combined with rigorous numerical analyses. I demonstrate that the expansion-rate tension fundamentally reflects underlying distance disagreements, and that the Thomson scattering rate strongly restricts higher pre-recombination expansion rates without additional physics. Further, I present a novel cosmological model using a mirror dark sector and varying fundamental constants, revealing an observational degeneracy allowing significantly higher …
Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha
Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha
Computer Science and Engineering Theses and Dissertations
Quantum computing enables new approaches to data processing, especially in quantum machine learning. Unlike classical systems, quantum data must be synthesized through operations and can exist in superposition. Encoding choices affect efficiency, noise resilience, and trainability—key factors in quantum machine learning models. This dissertation enhances quantum data encodings by extending quantum read-only memory (QROM) beyond binary representations, improving efficiency and parallelism. It introduces new compilation methods for quantum random number generators (QRNGs), supporting non-parametric distributions for post-quantum cryptography. Additionally, it explores Cayley graph-based encodings to extract spectral features for quantum machine learning.
Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani
Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani
Computer Science and Engineering Theses and Dissertations
The rapid expansion of scientific literature has intensified the challenge of identifying relevant citations, particularly for newly published or under-cited papers. Traditional citation recommendation systems typically model static relationships or respond to past citation activity, offering limited predictive power for emerging works. In response, this thesis presents a temporal modeling framework for citation recommendation that anticipates future scholarly relevance by forecasting the latent representations of academic papers.
Building on prior work that utilized Temporal Graph Networks (TGNs) to model dynamic citation flows, we propose Graph-Time, a hybrid architecture that integrates a Graph Transformer with a GRU-based time series predictor. The …
A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer
A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer
Computer Science ETDs
Modern drug discovery and chemical biology research relies heavily on analyzing bioassay data. One of the many challenges in bioassay data analysis is identifying false trails, i.e., chemical compounds which initially appear to have desirable activity but are found to be problematic upon further investigation. Badapple (the BioAssay-Data Associative Promiscuity Pattern Learning Engine) was created over ten years ago to help researchers identify promiscuous compounds and thus avoid a common source of these false trails. Through an effort involving software engineering, cheminformatics, and biomedical data science we have developed Badapple 2.0, which incorporates updated assay records and expanded data semantics. …
Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja
Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja
Faculty, Staff and Student Publications
Background: Finite mixture models have been recently applied in time-to-event data to identify subgroups with distinct hazard functions, yet they often assume differing covariate effects on failure times across latent classes but homogeneous covariate distributions. This study aimed to develop a method for analyzing time-to-event data while accounting for unobserved heterogeneity within a mixture modeling framework.
Methods: A joint model was developed to incorporate latent survival trajectories and observed information for the joint analysis of time-to-event outcomes, correlated discrete and continuous covariates, and a latent class variable. It assumed covariate effects on survival times and covariate distributions vary across latent …
Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell
Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell
Capstone Projects
A common technique when investigating a disease is to profile gene expression, as this gives unique insights into the functions of a cell. Gene expression data gathered from single cell RNA sequencing can be encoded into a gene co-expression network, which is a graph of potential relationships between different genes. One method for interpreting data encoded as a graph is to use a graph neural network, or GNN. This project designs and implements a GNN architecture to accomplish classification tasks on graph data. Then, given a dataset of gene co-expression networks made from multiple single cell RNA sequencing studies, the …
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Capstone Projects
This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …
Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi
Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi
Open Educational Resources
This open-access machine learning course is a comprehensive 15-week curriculum developed and published on GitHub with full Google Colab compatibility. It combines theoretical concepts with hands-on Python coding, real-world datasets, and structured projects covering regression, classification, clustering, deep learning, transformers, and multimodal AI. The course is designed for students, educators, and researchers interested in applied machine learning, including biomedical applications. It includes explainable AI components and ethical discussions to align with modern AI standards. The course is maintained by BioMind AI Lab at CUNY.
Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon
Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon
Honors Theses
Effective grouping methods enhance classroom collaboration and allow for a student-centered teaching approach; however, traditional grouping methods are time-consuming, subjective, and can create inconsistent group dynamics. This project addresses these challenges by employing a data-driven approach to optimize student groups based on academic performance, behavior, attendance, language barriers, and teacher preferences. The minimum viable product is a web application with an algorithm-driven system to group students and a database storage for group results. During the initiation phase, a problem was defined with a proposed solution. During the planning phase, potential design choices and grouping methods were researched and assessed. During …
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Electronic Theses and Dissertations
The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).
The analysis focuses on six key predictors: Hot Carcass Weight, Back …
Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb
Data Science and Data Mining
Genome-Wide Association Studies (GWAS) are instrumental in identifying genetic variants linked to complex traits, providing valuable insights into trait heritability and biological mechanisms. This study applies GWAS to investigate flowering time in maize, a critical adaptive trait, using a diverse dataset of 5,000 recombinant inbred lines across eight environments. Traditional GWAS methods often encounter challenges in high-dimensional datasets due to the presence of multiple small-effect genetic loci. To address this, we compared two penalized regression methods—LASSO and Ridge regression—to perform variable selection and regression analysis within a GWAS framework. LASSO effectively reduced the number of predictors by selecting the most …
Data Science Job Salary Prediction Using Linear Regression, Dipok Deb
Data Science Job Salary Prediction Using Linear Regression, Dipok Deb
Data Science and Data Mining
In the evolving landscape of data science, accurate salary prediction plays a crucial role in shaping career expectations, informing educational strategies, and guiding organizational hiring decisions. This study investigates the key factors influencing entry-level data science salaries in the United States by applying a multiple linear regression model to a recent dataset spanning from 2020 to 2024. Through data preprocessing, transformation, and diagnostic evaluation, we identify how job roles, experience levels, employment types, work arrangements, residency status, and company size impact compensation. Despite challenges such as outliers, heteroscedasticity, and non-normal residuals, model refinements like the Box-Cox transformation and variable selection …
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Student Scholar Symposium Abstracts and Posters
For my Introduction to Statistics Class, I have been tasked with collecting unique, personal data to give insight into my daily routine. I decided to record nine different outcomes (two qualitative and seven quantitative). On February 6, 2025, I began with a blank Excel sheet, and so far, I have 57 full days of data collected. I will continue monitoring my findings for the remainder of the Spring 2025 Semester. Per my project instructions, I must include tables and graphs for my qualitative and quantitative outcomes. So far, I have collected daily quantitative data on my screen time (Instagram and …
Two-Sample Bi-Directional Causality Between Two Traits With Some Invalid Ivs In Both Directions Using Gwas Summary Statistics, Siyi Chen
School of Public Health Faculty Publications
Mendelian randomization (MR) is a widely used method for assessing causal relationships between risk factors and outcomes using genetic variants as instrumental variables (IVs). While traditional MR assumes uni-directional causality, bi-directional MR aims to identify the true causal direction. In uni-directional MR, invalid IVs due to pleiotropy can violate assumptions and introduce biases. In bi-directional MR, traditional MR can be performed separately for each direction, but the presence of invalid IVs poses even greater challenges. We introduce a new bi-directional MR method incorporating stepwise selection (Bidir-SW) designed to address these challenges. Our approach leverages public genome-wide association study (GWAS) datasets …
Handwritten Digit Recognition Using Machine Learning, Dipok Deb
Handwritten Digit Recognition Using Machine Learning, Dipok Deb
Data Science and Data Mining
Handwritten Digit Recognition (HDR) remains a fundamental benchmark in pattern recognition and machine learning due to its practical applications and inherent classification challenges posed by diverse handwriting styles. This study investigates and compares two classical statistical classifiers—Gaussian Naive Bayes (GNB) and Linear Discriminant Analysis (LDA)—to recognize the digits from the MNIST dataset. Both models assume underlying normality in feature distributions and offer computational efficiency, making them suitable for high-dimensional input such as image pixels. Using 60,000 training and 10,000 test samples, we evaluate model performance through accuracy, precision, recall, F1 score, and confusion matrices. The results reveal that while GNB …
Clustering Dataset Using K-Mean Clustering, Dipok Deb
Clustering Dataset Using K-Mean Clustering, Dipok Deb
Data Science and Data Mining
Clustering is a fundamental technique in unsupervised machine learning, widely applied in various domains such as pattern recognition, data segmentation, and anomaly detection. This study evaluates the performance of the K-Means clustering algorithm on multiple benchmark datasets, including low-dimensional, high-dimensional, and imbalanced datasets. The clustering results are assessed using four key evaluation metrics: Mean Squared Error (MSE), Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Silhouette Score. Experimental results demonstrate that K-Means performs effectively on datasets with well-separated clusters, particularly in high-dimensional spaces, where it achieves near-perfect clustering accuracy. However, its performance deteriorates in datasets with overlapping clusters and …
Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca
Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca
Theses and Dissertations
Missing data is pervasive in healthcare, where incomplete observations commonly arise from patient dropout, sensor failures, or privacy constraints. This research presents an investigation into handling such data, focusing on (1) Missingness-Aware Dynamic Ensemble Weighting (MDEW), (2) feature selection under varying missing rates, (3) autoencoder-based imputation (ODAE), and (4) a meta-feature analysis guiding pipeline selection. We evaluate our experiments on four diverse datasets, Cleveland Heart Disease, Diabetic Retinopathy, Breast Cancer Wisconsin, EEG Eye State. Our research shows that MDEW adaptively selects imputer classifier pipelines, outperforming single model and uniform averaging baselines at moderate to high missingness 10% to 50%. Filter …
Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala
Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala
Presentations
No abstract provided.
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
Undergraduate Honors Theses
Machine learning is a method that employs statistical algorithms to identify patterns and make predictions from data. This study applies machine learning techniques to analyze data from Major League Baseball (MLB) teams between 1998 and 2024, with the goal of determining which factors strongly influence a team's likelihood of reaching the postseason and in accurately predicting the teams that do and do not qualify for the postseason. Data exploration and unsupervised machine learning methods such as clustering were used to identify underlying patterns in team performance metrics and determine potential significant contributors to team success. Many different supervised learning methods …
The Little Diagram That Could: Geometric Properties And Statistical Applications Of Persistence Diagrams In Topological Data Analysis, Eugene Kler
McKelvey School of Engineering Graduate Student Theses & Dissertations
Topological Data Analysis (TDA) is a collection of techniques for data analysis that leverages topological invariants of spaces formed from data points. These methods excel at extracting useful information from noisy or sparse data, making them attractive to many mathematicians, statisticians, and scientists. In this thesis, we explore TDA on three fronts: algebraic foundations, statistical applications, and metric properties. Throughout, the central object of study is the Persistence Diagram (PD), a summary of the changes in homology that occur as one builds simplicial complexes from the data by increasing a parameter.
Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran
Data Science Undergraduate Honors Theses
Early and accurate wildfire detection is critical for minimizing environmental damage and ensuring a timely response. However, existing satellite-based wildfire datasets suffer from limitations such as coarse ground truth, poor spectral coverage, and class imbalance, which hinder progress in developing robust segmentation models. In this paper, we introduce Land8Fire, a new large-scale wildfire segmentation dataset composed of over 20,000 multispectral image patches derived from Landsat 8 and manually annotated for high-quality fire masks. Building on the ActiveFire dataset, Land8Fire improves ground truth reliability and offers predefined splits for consistent benchmarking. We evaluate a range of state-of-the-art convolutional and transformer-based models, …
Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba
Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba
Theses, Dissertations and Culminating Projects
Pharmaceutical treatments are essential for managing medical conditions, but drug-drug interactions (DDIs) pose significant risks to patient safety and healthcare outcomes. This research integrates Knowledge Graphs and Graph Neural Networks to predict DDIs by exploring complex drug relationships. We construct a comprehensive knowledge graph using DrugBank data (1,000 drugs, 155,774 interactions) enriched with molecular features from PubChem. Our methodology introduces a novel multi-modal approach by integrating transformer-based embeddings (ChemBERTa, SPECTER, and SBERT) to create 1152-dimensional feature vectors that capture structural, biomedical literature, and semantic properties of drugs. Formulating DDI prediction as a link prediction task, we compare three Graph Neural …
Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii
Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii
Theses, Dissertations and Culminating Projects
Environmental justice, as both a movement and a theoretical construct, continues to evolve in response to shifting societal, environmental, and technological conditions. This dissertation investigates the integration of big data, such as social media, remote sensing imagery, and internet search frequencies, into the identification, analysis, and remediation of environmental injustices. Framing environmental justice through the lenses of distributive and data justice, the project explores both the promises and pitfalls of using emergent data sources to enhance the spatial and temporal precision of environmental equity investigations. Through a combination of systematic literature review, spatial analysis, system dynamics simulation, and policy evaluation, …
From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins
From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins
Electrical Engineering and Computer Science (MS) Theses
This study examines the relationship between climate conditions and residential water usage, focusing on how seasonal and environmental changes influence water consumption. Utilizing data from over 100,000 households across three micro-climate zones for over a five-year period, we apply statistical analysis and machine learning techniques to assess the impact of temperature, precipitation, evapotranspiration, and location on water usage. By integrating climate and billing data, this research provides a data-driven approach on water usage behaviors in Irvine, CA, in collaboration with Irvine Ranch Water District (IRWD).
Our analysis utilizes time series modeling, including a Seasonal Autoregressive Integrated Moving Average (SARIMA) and …
An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck
An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck
Electrical Engineering and Computer Science (MS) Theses
Closed-source large language models (LLMs) developed by large technology companies continue to grow in popularity. However, ethical conversations surrounding the safety of model outputs have been a prominent topic of discussion. This project aims to assess three leading closed-source LLMs: OpenAI’s ChatGPT, Google’s Gemini, and Anthropic’s Claude, to analyze how their outputs perform when treated as a subject of several psychological evaluation scales measuring biased behaviors against women. The Ambivalent Sexism Index, Modern Sexism Scale, and Belief in Sexism Shift evaluations were used to get descriptions of how the LLMs respond to traditional and modern prompts involving sexism and gender …
A Narrative-Focused Machine Learning Approach To Predicting Feature Film Success, Arisa T. Trombley
A Narrative-Focused Machine Learning Approach To Predicting Feature Film Success, Arisa T. Trombley
Electrical Engineering and Computer Science (MS) Theses
For decades, the field of film production has been driven by marketability, and it has relied on gut feelings and subjectivity to produce feature films. The analysis of the relationship between a screenplay’s narrative and a film’s success has been widely overlooked due to the challenges involved in data acquisition and complexity. This study investigates the predictive power of narrative structure on film success and aims to build evidence for hypothesized narrative principles. The results suggest that narrative structural elements exhibit moderate predictive power, with strong support for the alignment of the 2nd act crucial moments and the 2nd act …
Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar
Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar
Data Science Undergraduate Honors Theses
As universities navigate financial constraints and resource allocation challenges, data driven financial analysis has become increasingly important. Universities employ various methods to assess financial efficiency, predict future expenditures, and optimize student credit hour distribution. However, the approaches to financial analysis vary widely, with some institutions leveraging advanced predictive modeling and business intelligence tools, while others rely on traditional budgeting techniques and manual forecasting.
This thesis examines how the University of Arkansas' (“Uark”) financial analysis methods compare to those of other institutions and alternative data-driven approaches. Using four years of financial and student credit hour data, this study evaluates cost trends …
Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore
Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore
Data Science Undergraduate Honors Theses
Fire station placement has a critical role in emergency response efficiency and community safety. Traditional optimization models focus on mainly the minimization of response times and the maximization of coverage. However, this approach may overlook potential socioeconomic disparities that can influence emergency demand. This study seeks to expand upon the existing project of zoning a fire station in Sugar Land, TX, by integrating spatial road network analysis and publicly available census data—including population density, median household income, and age-based vulnerability—into a Maximal Coverage Location Problem (MCLP) framework. Using a road network-based travel time with realistic constraints, the goal is to …