Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Old Dominion University (12)
- Claremont Colleges (11)
- East Tennessee State University (7)
- Southern Methodist University (5)
- University of New Mexico (5)
-
- Ursinus College (5)
- Chapman University (4)
- Florida Institute of Technology (4)
- Illinois State University (4)
- University of Missouri, St. Louis (4)
- West Virginia University (4)
- Kennesaw State University (3)
- Montclair State University (3)
- Murray State University (3)
- New Jersey Institute of Technology (3)
- University of Central Florida (3)
- Utah State University (3)
- Binghamton University (2)
- California Polytechnic State University, San Luis Obispo (2)
- City University of New York (CUNY) (2)
- Clemson University (2)
- Embry-Riddle Aeronautical University (2)
- Georgia Southern University (2)
- Marshall University (2)
- Purdue University (2)
- South Dakota State University (2)
- Southeastern University (2)
- University of Arkansas, Fayetteville (2)
- University of Connecticut (2)
- University of Kentucky (2)
- Keyword
-
- Machine Learning (13)
- Machine learning (9)
- Data Science (7)
- Deep learning (6)
- Data science (5)
-
- Graph Theory (5)
- Neural networks (5)
- Statistics (5)
- Algorithms (4)
- Artificial Intelligence (4)
- COVID-19 (4)
- Neural Networks (4)
- Clustered data (3)
- Deep Learning (3)
- Informative cluster size (3)
- Mathematics (3)
- Reinforcement Learning (3)
- Sentiment (3)
- Time series (3)
- ARIMA (2)
- Algebraic topology (2)
- Analytics (2)
- Artificial intelligence (2)
- Big data (2)
- Bioinformatics (2)
- CIE (2)
- Central Limit Theorem (2)
- Classification (2)
- Computer Vision (2)
- Computer vision (2)
- Publication Year
- Publication
-
- Mathematics & Statistics Faculty Publications (10)
- Electronic Theses and Dissertations (9)
- Theses and Dissertations (8)
- Mathematics, Computer Science & Statistics Presentations (5)
- SMU Data Science Review (4)
-
- Theses (4)
- Annual Symposium on Biomathematics and Ecology Education and Research (3)
- Data Science and Data Mining (3)
- Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- Journal of Humanistic Mathematics (3)
- All Dissertations (2)
- Branch Mathematics and Statistics Faculty and Staff Publications (2)
- CMC Senior Theses (2)
- CODEE Journal (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Data Science Undergraduate Honors Theses (2)
- Department of Computer Science Faculty Scholarship and Creative Works (2)
- Discovery Day - Daytona Beach (2)
- Honors College Theses (2)
- Honors Projects (2)
- Honors Theses and Capstones (2)
- MPP Published Research (2)
- Master's Theses (2)
- Mathematics & Statistics ETDs (2)
- Northeast Journal of Complex Systems (NEJCS) (2)
- Pitzer Senior Theses (2)
- SDSU Data Science Symposium (2)
- Scripps Senior Theses (2)
- Selected Honors Theses (2)
- Publication Type
- File Type
Articles 1 - 30 of 144
Full-Text Articles in Data Science
Toward Mapping Multiphase Multicomponent Mixtures With Neural Networks, Kristen L. Hallas, Melissa De Jesus, Christine J. Wu, Jianzhi Li, Jason Bernstein, Philip C. Myint
Toward Mapping Multiphase Multicomponent Mixtures With Neural Networks, Kristen L. Hallas, Melissa De Jesus, Christine J. Wu, Jianzhi Li, Jason Bernstein, Philip C. Myint
School of Mathematical & Statistical Sciences Faculty Publications
Equation of state (EOS) tables are commonly used in hydrodynamic simulations of high-pressure, high-temperature phenomena in fields like planetary science, astrophysics, and high-energy-density science. However, generating and storing EOS tables for multiphase, multicomponent mixtures over a wide range of pressures and temperatures is computationally infeasible due to their memory-intensive nature. To address this issue, we have developed a neural network-based machine learning model to predict new EOS tables for binary mixtures. In particular, a deep feedforward neural network trained on a set of ten EOS tables at particular mixture compositions is able to predict nine new (hold-out) EOS tables at …
Learning Motion Primitive Selection And Environment Abstraction, Edison Alberto Martinez Samaniego, Natalie Alexander, Kaelyn Weddle
Learning Motion Primitive Selection And Environment Abstraction, Edison Alberto Martinez Samaniego, Natalie Alexander, Kaelyn Weddle
Discovery Day - Daytona Beach
Learning Motion Primitive Selection and Environment Abstraction Advanced Air Mobility (AAM) is emerging as a transformative solution for short and medium range transportation; however, it introduces an operational model that differs significantly from conventional aviation. AAM vehicles are expected to operate closer to populated areas, with increased autonomy, in dense urban and suburban environments. These settings present constrained maneuvering conditions which highlights the importance of maintaining safe operation under degraded flight conditions. Abnormal conditions may endanger onboard passengers, people on the ground, and surrounding infrastructure, making rapid detection and mitigation essential to prevent loss of control. Recent research has explored …
Low-Complexity Polynomial Ring Learning For Quantum Space Assets, Lola Torres
Low-Complexity Polynomial Ring Learning For Quantum Space Assets, Lola Torres
Discovery Day - Daytona Beach
Secure communication for space-based systems requires cryptographic methods that remain both reliable and efficient under strict computational constraints. This work investigates a low-complexity polynomial ring learning algorithm designed for quantum space assets, including satellite–ground communication systems. The project focuses on post-quantum cryptographic principles, where encryption and decryption rely heavily on repeated polynomial operations; this can be computationally expensive with constrained platforms. This is addressed with reformulating polynomial multiplication as a structured linear transformation on coefficient vectors. By representing these operations as matrices with a cyclic structure, the structure allows the use of the discrete Fourier transform (DFT); this will simplify …
Unifying And Expanding Global And Local Variable Importance Methods For Explainable Machine Learning, Kelvyn K. Bladen
Unifying And Expanding Global And Local Variable Importance Methods For Explainable Machine Learning, Kelvyn K. Bladen
All Graduate Theses and Dissertations, Fall 2023 to Present
Machine learning methods are powerful analytical tools used across all scientific disciplines and many other fields of investigation for prediction and inference from diverse data sources. Despite their broad applicability, machine learning methods are often highly complex and difficult to interpret. Developing a greater understanding of which variables most influence a response is essential for increasing the interpretability of these models and supporting informed decision-making. This research focuses on improving how we evaluate the importance of these variables.
One common approach is to shuffle the values of a variable and see how much the model accuracy gets worse. Another approach …
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Dissertations
The environmental benefits of electric vehicle (EV) adoption depend on more than replacing internal combustion engine vehicles with electric powertrains. EV adoption reshapes electricity demand, interacts with regional generation mixes, and influences travel behavior and congestion, creating a coupled transportation-energy system in which vehicle and power-plant emissions must be evaluated together. This dissertation develops machine-learning frameworks for predicting energy consumption and emissions from vehicles and power grids under rising EV adoption. The first component forecasts grid emissions from EV charging. Using simulation data from NREL's Cambium database, a Prophet-based time-series framework predicts carbon dioxide, nitrous oxide, and methane emission rates …
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Mathematics Theses and Dissertations
This dissertation presents a computational framework for high-frequency options trading that combines Cross-Data-Type 1-D Convolutional Neural Networks (CDT-1D CNN) with Simpson-Sobolev regularization for directional prediction, and finite element methods (FEM) for realistic option pricing during backtesting. The core innovation lies in developing a mathematically rigorous regularization approach that maintains the adaptability of modern deep learning while enabling accurate evaluation through stochastic volatility models. The primary contribution is the Simpson-Sobolev regularization scheme, which extends traditional Sobolev regularization by incorporating Simpson’s rule for numerical integration. This approach achieves higher-order accuracy in approximating the Sobolev norms that control function smoothness. Simpson’s rule attains …
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson
CODEE Journal
In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.
This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Dissertations, Theses, and Capstone Projects
About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Northeast Journal of Complex Systems (NEJCS)
Financial markets play a critical role in resource allocation. Their performance depends on the decisions of millions of independent investors constantly reacting to one another. Their informational efficiency remains a subject of debate across economic systems. When informational efficiency is present at the weak form, historical price information should not consistently predict future returns. Several empirical tests of this hypothesis often focus on the behavior of aggregate market indices, and use individual efficiency proxies such as autocorrelation, GARCH-type volatility, or entropy-based measures to measure efficiency. This has often yielded mixed results, particularly in emerging markets. Here we show that testing …
Effectiveness Of The Guided Discovery Method In Teaching The Surface Area Of A Cylinder, Paul Ahortu
Effectiveness Of The Guided Discovery Method In Teaching The Surface Area Of A Cylinder, Paul Ahortu
2026 Symposium
This study investigates the impact of the guided discovery instructional method on students’ understanding of the surface area of a cylinder. A quasi-experimental pre-test–post-test design was conducted with 100 senior high school students in Cape Coast, Ghana, divided into experimental and comparison groups..
Results showed a substantial improvement in performance for students exposed to guided discovery, with mean scores increasing from 1.25 (pre-test) to 9.43 (post-test) and a large effect size (Cohen’s d = 2.70). Statistical analysis also revealed significant gender differences in achievement.
These findings indicate strong improvement following the guided discovery intervention and suggest its potential to enhance …
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
Theses and Dissertations
Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …
The Waldo Dataset, Mary E. Koone, Rosie Kallie, Vassilis Athisos, Laurel S. Stvan
The Waldo Dataset, Mary E. Koone, Rosie Kallie, Vassilis Athisos, Laurel S. Stvan
Computer Science and Engineering Datasets - Archive
Distinct from the task of predicting the author of a document (authorship attribution), we focus on addressing the issue of how to estimate the similarity between the written language styles of authors. To do so, we present a dataset of metadata derived by asking human annotators, who were presented with three documents, to identify which two were written by the same author and which was written by a different author. The dataset has over 400 such annotations, creating a companion to the Amazon Web Services (AWS) customer review dataset, laying the groundwork for crowdsourcing applications to other natural language processing …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Effective Wordle Heuristics, Ronald I. Greenberg
Effective Wordle Heuristics, Ronald I. Greenberg
Computer Science: Faculty Publications and Other Works
While previous researchers have performed an exhaustive search to determine an optimal Wordle strategy, that computation is very time consuming and produced a strategy using words that are unfamiliar to most people. With Wordle solutions being gradually eliminated (with a new puzzle each day and no reuse), an improved strategy could be generated each day, but the computation time makes a daily exhaustive search impractical. This paper shows that simple heuristics allow for fast generation of effective strategies and that little is lost by guessing only words that are possible solution words rather than more obscure words.
An Association Test For Ordinal Outcomes In Clustered Data With Informative Cluster Size, Hasika K. Wickrama Senevirathne, Sandipan Dutta
An Association Test For Ordinal Outcomes In Clustered Data With Informative Cluster Size, Hasika K. Wickrama Senevirathne, Sandipan Dutta
Mathematics & Statistics Faculty Publications
In cluster-correlated data, the number of observations in a cluster can be associated with the outcome from that cluster. This phenomenon is known as informative cluster size which can occur in cluster-randomized clinical trial data. Several studies have found that ignoring the issue of informative cluster size can produce biased results in the analysis of clustered data. Most of the existing methods for addressing informative cluster size are suited to continuous outcomes. However, ordinal outcomes and covariates are often encountered in clustered data obtained from large clinical studies. The existing methods for ordinal association testing in clustered data can produce …
Logistic-T Multinomial Mixture Model For Clustering For Microbiome Data, Wenshu Dai, Yuan Fang, Sanjeena Subedi
Logistic-T Multinomial Mixture Model For Clustering For Microbiome Data, Wenshu Dai, Yuan Fang, Sanjeena Subedi
Mathematics & Statistics Faculty Publications
The logistic-normal multinomial distribution has been used for modelling microbiome data obtained from high-throughput sequencing technologies, which are compositional in nature. A logistic-normal multinomial distribution is a hierarchical multinomial distribution that assumes the latent variable which are the additive log-ratio (ALR) transformed proportions in a multinomial distribution follows a Gaussian distribution. Model-based clustering algorithms have also been developed for clustering microbiome data based on the logistic-normal models. However, the Gaussian assumption may violated when the ALR transformed variable exhibit heavy-tailed distributions or has outliers. Our study introduces a novel mixture of logistic-t multinomial models that effectively address these challenges. Utilizing …
Integer-Valued Time Series Model Via Copula-Based Bivariate Skellam Distribution, Mohammed Alqawba, Norou Diawara, Mame Mor Sene
Integer-Valued Time Series Model Via Copula-Based Bivariate Skellam Distribution, Mohammed Alqawba, Norou Diawara, Mame Mor Sene
Mathematics & Statistics Faculty Publications
Time series analysis is crucial for modeling and forecasting diverse real-world phenomena. Traditional models typically assume continuous-valued data; however, many applications involve integer-valued series, often including negative integers. This paper introduces an approach that combines copula theory with the bivariate Skellam distribution to handle such integer-valued data effectively. Copulas are widely recognized for capturing complex dependencies among variables. By integrating copulas, our proposed method respects integer constraints while modeling positive, negative, and temporal dependencies accurately. Through simulation and an empirical study on a real-life example, we demonstrate that our class of models performs well. This approach has broad applicability in …
Identifying Relevant Covariates In Rna-Seq Analysis By Pseudo-Variable Augmentation, Yet Nguyen, Dan Nettleton
Identifying Relevant Covariates In Rna-Seq Analysis By Pseudo-Variable Augmentation, Yet Nguyen, Dan Nettleton
Mathematics & Statistics Faculty Publications
RNA-sequencing (RNA-seq) technology allows for the identification of differentially expressed genes, which are genes whose mean transcript abundance levels vary across conditions. In practice, RNA-seq datasets often include covariates that are of primary interest in addition to a set of covariates that are subject to selection. Some of these covariates may be relevant to gene expression levels, while others may be irrelevant. Ignoring relevant covariates or attempting to adjust for the effect of irrelevant covariates can compromise the identification of differentially expressed genes. To address this issue, we propose a variable selection method that uses pseudo-variables to control the expected …
Machine Learning-Based Intrusion Detection System For Iot Networks Using The Rt-Iot 2022 Dataset, Bukunmi Ebenezer Afolabi
Machine Learning-Based Intrusion Detection System For Iot Networks Using The Rt-Iot 2022 Dataset, Bukunmi Ebenezer Afolabi
Theses, Dissertations and Capstones
The rapid expansion of the Internet of Things (IoT) has transformed modern computing by enabling seamless connectivity among heterogeneous devices across diverse application domains. However, this increased interconnectivity has significantly enlarged the attack surface of IoT networks, exposing them to a wide range of sophisticated cyber threats. Conventional security mechanisms often lack the capability to detect emerging attacks in real time, thereby necessitating the development of intelligent Intrusion Detection Systems (IDS) capable of accurately identifying malicious network activities. This study developed and evaluated a machine learning-based intrusion detection framework for multiclass IoT attack detection using the RT-IoT2022 dataset. The dataset …
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Undergraduate Research Symposium
Various computational models of first impressions have been developed to uncover the mechanisms driving these judgments. However, the implicit notion of a singular ``human'' often overlooks meaningful individual differences in beliefs, attitudes, and associations, as well as culturally grounded group-level constructs. In this paper, we extend Cultural Consensus Theory (CCT) to estimate culturally shared beliefs about faces by incorporating latent constructs structured around interpretable facial features extracted via computer vision algorithms. We apply our model to a large-scale dataset of people’s first impressions of faces. Our approach reveals a robust mapping between facial features and culturally constructed impressions, allowing us …
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
Electronic Theses and Dissertations
This thesis develops a discrete stochastic linear systems interpretation of age–stage demographic evolution grounded in Leslie operators and realized in a discrete-event simulation implemented with salabim. The central claim is that one annual cycle of the simulation constitutes a cone-preserving, stochastic affine transformation on a high- dimensional population state vector indexed by age, sex, marital status, household type, employment, and education, and that the composition of yearly operators yields a random matrix product whose top Lyapunov exponent is the stochastic counterpart of the Perron–Frobenius growth rate (Caswell, 2001; Tuljapurkar, 1997)[1, 2]. The actuarial bridge is constructed by mapping simulated survival …
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
Electronic Theses and Dissertations
As single-cell RNA sequencing (scRNA-seq) data expands, robust methods for integrating diverse datasets are critical. This dissertation applies Persistent Homology (PH), a technique from Topological Data Analysis (TDA), to a collection of scRNA-seq datasets spanning eight tissue types to quantify how data integration affects topological features and biological interpretability. We assessed global topological structure using Betti curves, Euler characteristics, and persistence landscapes across raw, normalized, and integrated data representations. Our analysis revealed a performance inversion: while conventional methods excelled on unintegrated data, high-granularity topological methods, particularly those sensitive to global data structure, became superior after integration. This suggests a synergy …
An Income Subsystem As A Discrete Stochastic Leslie System: A Simulation-Based Approach, Fahd Nii Okantah Cobblah
An Income Subsystem As A Discrete Stochastic Leslie System: A Simulation-Based Approach, Fahd Nii Okantah Cobblah
Electronic Theses and Dissertations
This thesis formulates the household-income engine of an integrated population sim- ulator as a Discrete Stochastic Leslie System (DSLS). The nonnegative state vector nt ∈ Rk + aggregates income, savings, debt, employment, and transfers. (Here, the subscript + denotes the positive cone, i.e., vectors with nonnegative components). Annual evolution is linear in state, stochastic in coefficients: nt+1 = Ttnt + εt, with Tt : Rk + → Rk + cone-preserving. Exogenous macro drivers (inflation, employment, tax, salary inflation, mortgage) are forecast via ARIMA; forecasts multiply entries of Tt, preserving linearity in expectation while introducing realistic temporal correlation. The discrete-event implemented …
Modeling The Cancer Cell Growth Predictions Based On Classical Mathematical Models With Physics-Informed Neural Network, Widodo Samyono
Modeling The Cancer Cell Growth Predictions Based On Classical Mathematical Models With Physics-Informed Neural Network, Widodo Samyono
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Temporal Modeling And Forecasting Of Blood Glucose Dynamics In Individuals With Diabetes Mellitus, Mj Ruff
Temporal Modeling And Forecasting Of Blood Glucose Dynamics In Individuals With Diabetes Mellitus, Mj Ruff
University Honors Theses
People living with Diabetes Mellitus face significant health risks, including an increased likelihood of heart disease, stroke, and fluctuations in blood glucose levels. The unpredictable nature of glucose levels can lead to dangerous conditions such as ketoacidosis and hypoglycemia. This study employs advanced time series analysis tools to forecast the glucose levels for an individual diagnosed with Type 1 Diabetes Mellitus.
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Dartmouth College Ph.D Dissertations
In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …
A Novel Framework For Dynamic Graph Representation Learning With Mamba, Ashish Pandey
A Novel Framework For Dynamic Graph Representation Learning With Mamba, Ashish Pandey
Theses
Dynamic graph embedding is a key technique for modeling temporal dependencies in evolving networks. While transformer-based models perform well, their quadratic complexity limits scalability on long graph sequences. This thesis compares transformer approaches with the Mamba architecture-a linear-complexity state-space model—for temporal graph embedding.
Two frameworks are proposed: DG-Mamba and GDG-Mamba. DG-Mamba uses standard GCN-based spatial encoding, while GDG-Mamba incorporates domain-aware edge features using Graph Isomorphism Network with Edge Convolution (GraphGINE). Experiments on UCI, Reality Mining, Slashdot, Bitcoin-OTC, and SBM datasets show that Mamba-based models match or exceed transformer performance, especially on graphs with high temporal variability.
The thesis also applies …
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Student Scholar Symposium Abstracts and Posters
For my Introduction to Statistics Class, I have been tasked with collecting unique, personal data to give insight into my daily routine. I decided to record nine different outcomes (two qualitative and seven quantitative). On February 6, 2025, I began with a blank Excel sheet, and so far, I have 57 full days of data collected. I will continue monitoring my findings for the remainder of the Spring 2025 Semester. Per my project instructions, I must include tables and graphs for my qualitative and quantitative outcomes. So far, I have collected daily quantitative data on my screen time (Instagram and …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig
Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig
Honors College Theses
Analysis of a childhood game has led us to the problem of maximum independent sets in planar graphs. We wrote a graph creation utility using R to generate a random planar map and its dual graph. This utility then finds a graph’s maximal independent set using a variety of six algorithms. We investigate statistical connections between graph structure, colorability, and the maximal independent sets found using these algorithms over an incredibly large and procedurally generated dataset. We find one can always win the coloring game if the resultant graph is two-colorable. The algorithms perform statistically and practically significantly better on …