Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (1156)
- Medicine and Health Sciences (780)
- Life Sciences (765)
- Bioinformatics (568)
- Statistics and Probability (549)
-
- Biomedical Informatics (530)
- Engineering (527)
- Artificial Intelligence and Robotics (525)
- Social and Behavioral Sciences (519)
- Databases and Information Systems (212)
- Computer Engineering (208)
- Electrical and Computer Engineering (204)
- Applied Statistics (193)
- Medical Sciences (190)
- Business (189)
- Statistical Models (181)
- Applied Mathematics (175)
- Medical Specialties (173)
- Theory and Algorithms (149)
- Environmental Sciences (147)
- Mathematics (144)
- Other Computer Sciences (127)
- Data Storage Systems (123)
- Systems and Communications (120)
- Numerical Analysis and Scientific Computing (116)
- Public Health (116)
- Public Affairs, Public Policy and Public Administration (109)
- Statistical Methodology (109)
- Institution
-
- The Texas Medical Center Library (523)
- Old Dominion University (173)
- Southern Methodist University (144)
- Universitas Negeri Malang (113)
- City University of New York (CUNY) (100)
-
- CCT College Dublin (82)
- Chapman University (66)
- Kennesaw State University (63)
- University of Central Florida (62)
- Smith College (60)
- Air Force Institute of Technology (57)
- Embry-Riddle Aeronautical University (52)
- Singapore Management University (45)
- University of Arkansas, Fayetteville (45)
- Chinese Academy of Sciences (44)
- Purdue University (44)
- California Polytechnic State University, San Luis Obispo (39)
- Technological University Dublin (39)
- Illinois State University (38)
- University of Kentucky (38)
- University of Nebraska - Lincoln (38)
- New Jersey Institute of Technology (37)
- West Virginia University (37)
- Claremont Colleges (36)
- Virginia Commonwealth University (35)
- Clemson University (32)
- Dartmouth College (31)
- University of Texas at Arlington (27)
- East Tennessee State University (26)
- Minnesota State University, Mankato (26)
- Keyword
-
- Humans (278)
- Machine learning (241)
- Machine Learning (215)
- Deep learning (115)
- Computer Science (98)
-
- Deep Learning (93)
- Artificial Intelligence (65)
- Data science (58)
- Data Science (57)
- Natural Language Processing (56)
- COVID-19 (55)
- Artificial intelligence (53)
- Female (52)
- Male (50)
- Classification (49)
- Natural language processing (46)
- Animals (41)
- Data (41)
- Electronic Health Records (41)
- Neural Networks (40)
- Algorithms (38)
- Big data (37)
- Data mining (37)
- Statistics (36)
- Clustering (32)
- Computer science (31)
- Adult (30)
- NLP (30)
- Neural networks (30)
- AI (29)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (508)
- SMU Data Science Review (124)
- Knowledge Engineering and Data Science (113)
- Theses and Dissertations (111)
- ICT (82)
-
- Data Science and Data Mining (53)
- Dissertations (53)
- Statistical and Data Sciences: Faculty Publications (53)
- Electronic Theses and Dissertations (49)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (44)
- Dissertations, Theses, and Capstone Projects (44)
- Research Collection School Of Computing and Information Systems (37)
- Master's Theses (35)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (34)
- Data Science Undergraduate Honors Theses (31)
- Annual Symposium on Biomathematics and Ecology Education and Research (30)
- Computer Science Faculty Publications (30)
- Publications and Research (30)
- All Graduate Theses, Dissertations, and Other Capstone Projects (24)
- Computational and Data Sciences (PhD) Dissertations (24)
- Symposium of Student Scholars (24)
- All Dissertations (23)
- Articles (23)
- Electrical & Computer Engineering Faculty Publications (22)
- CBN Journal of Applied Statistics (JAS) (21)
- College of Graduate Studies: Theses & Dissertations (20)
- CMC Senior Theses (19)
- Theses (19)
- Electronic Theses, Projects, and Dissertations (18)
- Faculty Publications (18)
- Publication Type
- File Type
Articles 121 - 150 of 3231
Full-Text Articles in Data Science
Information Theory Analysis Of The Solar Wind Magnetic Structures For Space Weather Prediction, Katherine Holland
Information Theory Analysis Of The Solar Wind Magnetic Structures For Space Weather Prediction, Katherine Holland
Doctoral Dissertations and Master's Theses
Forecasting space weather at Earth is highly complicated, because of the limited measurements of the dynamic processes in the Sun that span multiple temporal, spatial, and energy-scales. The solar wind is a highly structured, multi-scale, evolving plasma and consists of coronal mass ejections (CMEs), stream interaction regions (SIRs), expanding flux tubes (Borovsky, 2008), and interplanetary magnetic field (IMF) discontinuities and fluctuations. The aim of this research is to improve our understanding of the evolution and dissipation of different scale-size solar wind magnetic structures as they move from the Sun-Earth Lagrange point 1 (L1) to Earth's bow shock and, ultimately, to …
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue
Articles
We introduce Saturated Hierarchical Atomic Incremental Learning (sHAIL), a learning paradigm in which complex tasks are approached through a sequence of simpler atomic subtasks, each mastered to saturation before progression. The central mechanism is a saturation criterion that detects when learning dynamics enter a plateau region, triggering consolidation and subsequent ascent to a higher level of task complexity. We develop a theoretical framework for sHAIL and show that it naturally gives rise to \emph{staircased convergence}: alternating phases of rapid improvement and genuine plateau. Within each level, classical convergence guarantees apply under standard smoothness conditions, while the hierarchical transitions are driven …
No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue
No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue
Articles
The rapid ascent of artificial intelligence (AI) is often portrayed as a revolution born from computer science and engineering. This narrative, however, obscures a fundamental truth: the theoretical and methodological core of AI is, and has always been, statistical. This paper systematically argues that the field of statistics provides the indispensable foundation for machine learning and modern AI. We deconstruct AI into nine foundational pillars—Inference, Density Estimation, Sequential Learning, Generalization, Representation Learning, Interpretability, Causality, Optimization, and Unification—demonstrating that each is built upon century-old statistical principles. From the inferential frameworks of hypothesis testing and estimation that underpin model evaluation, to the …
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental
Articles
Social insect colonies and ensemble machine learning methods represent two of the most successful examples of decentralized information processing in nature and computation respectively. Here we develop a rigorous mathematical framework demonstrating that ant colony decision-making and random forest learning are isomorphic under a common formalism of stochastic ensemble intelligence. We show that the mechanisms by which genetically identical ants achieve functional differentiation— through stochastic response to local cues and positive feedback—map precisely onto the bootstrap aggregation and random feature subsampling that decorrelate decision trees. Using tools from Bayesian inference, multi-armed bandit theory, and statistical learning theory, we prove that …
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
Articles
Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …
On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue
On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue
Articles
Data Science has ignited unprecedented academic, industrial, and pedagogical fervor, yet its status as a \textit{science} in the classical sense---comparable to physics or biology---remains profoundly unsettled. This article interrogates the epistemological foundations of Data Science by examining its hybrid theoretical lineage, from the Universal Approximation Theorem to the No-Free-Lunch Theorems, with special emphasis on the fundamental Bayesian optimality results for both regression and classification. We argue that Data Science is in a vigorous \textit{gestational period}, characterized not by an absence of principles but by a creative tension between empirical pragmatism and deep mathematical theory. The Cross-Validation score emerges as the …
On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue
On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue
Articles
Nature rarely reveals her secrets bluntly, yet in the Fibonacci sequence she grants us a glimpse of her quiet architecture of growth, harmony, and recursive stability \citep{Koshy2001Fibonacci, Livio2002GoldenRatio}. From spiral galaxies to the unfolding of leaves, this humble sequence reflects a universal grammar of balance. In this work, we introduce \emph{Fibonacci Ensembles}, a mathematically principled yet philosophically inspired framework for ensemble learning that complements and extends classical aggregation schemes such as bagging, boosting, and random forests \citep{Breiman1996Bagging, Breiman2001RandomForests, Friedman2001GBM, Zhou2012Ensemble, HastieTibshiraniFriedman2009ESL}. Two intertwined formulations unfold: (1) the use of normalized Fibonacci weights -- tempered through orthogonalization and Rao--Blackwell optimization -- …
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue
Articles
Ordinal data arise ubiquitously in survey research, psychology, medicine, economics, and recommender systems, yet kernel methods for such data typically rely on either nominal encodings or arbitrary numeric codings. The former discards order information; the lat- ter imposes a fictitious metric structure. This paper develops a principled framework for kernel design on ordinal scales and introduces a new class of Semantic–Aware Ordinal Ker- nels (SAOK) that simultaneously capture ordinal order and semantic proximity between categories. We begin by formalizing order–preserving embeddings of finite chains and characterizing a broad family of chain distances that are conditionally negative definite. Through Schoen- berg …
The Waldo Dataset, Mary E. Koone, Rosie Kallie, Vassilis Athisos, Laurel S. Stvan
The Waldo Dataset, Mary E. Koone, Rosie Kallie, Vassilis Athisos, Laurel S. Stvan
Computer Science and Engineering Datasets - Archive
Distinct from the task of predicting the author of a document (authorship attribution), we focus on addressing the issue of how to estimate the similarity between the written language styles of authors. To do so, we present a dataset of metadata derived by asking human annotators, who were presented with three documents, to identify which two were written by the same author and which was written by a different author. The dataset has over 400 such annotations, creating a companion to the Amazon Web Services (AWS) customer review dataset, laying the groundwork for crowdsourcing applications to other natural language processing …
Curriculum For A Two Semester Calculus Course Specializing In Life Science And Data Science, Patrick Mcclain
Curriculum For A Two Semester Calculus Course Specializing In Life Science And Data Science, Patrick Mcclain
LSU Master's Theses
Traditionally, introductory calculus has been designed for engineering and physics students, often leaving students majoring in data science or life sciences with a curriculum that lacks professional relevance and is overly reliant on problems that focus on computational fluency. This thesis proposes a two-semester sequence, called MATH 153X and MATH 154X, specifically tailored for the Louisiana State University (LSU) Dual Enrollment program and university-level data science and life sciences majors. By integrating modern computational tools—such as symbolic calculators and artificial intelligence (AI) tools—the proposed curriculum shifts the pedagogical focus from procedural symbolic manipulation toward conceptual literacy.
Through a series …
Human Subject Studies For The Alignment Of Llm-As-A-Judge Evaluation Metric For Science News, Gabriel Vega Osborne
Human Subject Studies For The Alignment Of Llm-As-A-Judge Evaluation Metric For Science News, Gabriel Vega Osborne
Knowledge and Creativity Expo
Science news has become an important vehicle to disseminate scientific breakthroughs, discoveries, and technological innovations. With the advancement of large language models and related AI models, it is possible to automatically generate science news from scientific papers, extending the reader population from domain scientists to a broader scope. However, how to evaluate the quality of the generated news warrants research. Traditional token based metrics have been shown to fail to evaluate the semantics and nuances of science news. Inspired by the fact that a major goal of science news is to educate readers with new knowledge, we thus propose knowledge …
Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun
Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun
Electronic Theses and Dissertations
The performance of deep neural networks (DNNs) is strongly influenced by the characteristics and quality of the underlying datasets. This Ph.D. dissertation addresses three pervasive data challenges-imbalance, quality degradation, and scarcity-that commonly hinder the effectiveness of DNNs in computer vision (CV) and natural language processing (NLP) applications.
Class imbalance remains one of the most frequent causes of degraded model generalization. While Focal Loss effectively mitigates inter-class imbalance by assigning higher weights to minority classes, it struggles with intra-class imbalance, particularly in video datasets where longer clips dominate feature representation. To address this, I implement and utilize …
What Would It Take To Compose The Ideal Library Dashboard As A 'Symphony' Of Library Data?, Evan Rusch, Heidi J. Southworth, Nat Gustafson-Sundell
What Would It Take To Compose The Ideal Library Dashboard As A 'Symphony' Of Library Data?, Evan Rusch, Heidi J. Southworth, Nat Gustafson-Sundell
Library Services Publications
Most areas of the library produce data informative about the scope and success of library services. Many areas use that data, but there might be a variety of approaches. In our library, we have developed online, interactive dashboards to understand the value of collections on our campus. Other library service areas might produce one-shot reports encapsulating their data, or they rely on analytics silos specific to their services. Our hope is to work toward a dashboard of all or most library services, possibly to include other learning services under the library roof. Instead of a ‘battle of the bands,’ we …
Kms-Net: Kolmogorov–Arnold-Based Multi-Scale Attention Network For Cardiac Segmentation, Abid Mehmood, Hassan Ali, David Noule Tolno, Sery Gahouidi Thierry S, Muhammad Saeed, Naeem Ahmed
Kms-Net: Kolmogorov–Arnold-Based Multi-Scale Attention Network For Cardiac Segmentation, Abid Mehmood, Hassan Ali, David Noule Tolno, Sery Gahouidi Thierry S, Muhammad Saeed, Naeem Ahmed
Research & Publications
Accurate segmentation of cardiac structures in 2D echocardiography is essential for diagnosing cardiovascular disease and computing clinical metrics such as chamber volumes and ejection fraction. Conventional U-Net architectures excel at extracting local spatial features but struggle with long-range dependencies inherent in noisy ultrasound images, while pure Transformer-based models capture global context at the expense of fine boundary detail. To address these limitations, we propose KMS-Net, a novel hybrid segmentation architecture that integrates Kolmogorov–Arnold Networks (KANs), a class of learnable, spline-based function approximators that replace fixed activation functions with trainable nonlinear mappings, alongside multi-scale attention mechanisms. Specifically, spline-based KAN layers (grid …
Terrorism By The Numbers: Event And Structural Determinants Of Attack Outcomes, Claire Lebakken
Terrorism By The Numbers: Event And Structural Determinants Of Attack Outcomes, Claire Lebakken
Student Research Symposium (SRS)
This project aims to identify which event-level and structural covariates are most predictive of terrorism outcomes. Using the Global Terrorism Database (1970–2020), we examine whether fatalities, injuries, attack type, target type, and actor type, combined with national-level conditions, can reliably predict outcomes such as lone actor versus group involvement, attack method, target selection, and property damage. Event-level data are merged with World Bank Development Indicators and Freedom House scores to incorporate economic and governance contexts. After cleaning the data, creating dummy variables, and log-transforming skewed measures (e.g., GDP per capita), we apply logistic and multinomial logistic regression models to test …
Associational Inference With Many Potential Covariates: Bayesian Information Criterion Elastic Net, Farideh Bagherzadeh Khiabani, I-Chan Huang, Jose Miguel Martinez Martinez, Shizue Izumi, Sedigheh Mirzaei, Tiange Zheng, Irina Dinu, Yutaka Yasui
Associational Inference With Many Potential Covariates: Bayesian Information Criterion Elastic Net, Farideh Bagherzadeh Khiabani, I-Chan Huang, Jose Miguel Martinez Martinez, Shizue Izumi, Sedigheh Mirzaei, Tiange Zheng, Irina Dinu, Yutaka Yasui
COBRA Preprint Series
Background: An emerging feature in modern biomedical research is collecting and analyzing numerous variables. In the presence of many potential covariates, inference becomes challenging requiring both distinguishing a set of covariates truly associated with an outcome and estimating their corresponding regression coefficients consistently. Traditional statistical inference typically focuses on estimating coefficients assuming a pre-specified set of covariates. Further, advanced machine/statistical learning methods performing both selection and estimation predominantly focus on outcome prediction rather than association inference.
Methods: Motivated by our epidemiological research on long-term childhood cancer survivors, where we aimed to investigate associations between a large pool of longitudinal symptom …
Data Centers In Mountain West Markets, 2026, Cason Noll, Krish Sharma, Maisoon Faris, Olivia K. Cheche, Caitlin J. Saladino, William E. Brown Jr.
Data Centers In Mountain West Markets, 2026, Cason Noll, Krish Sharma, Maisoon Faris, Olivia K. Cheche, Caitlin J. Saladino, William E. Brown Jr.
Transportation & Infrastructure
This fact sheet reports on the distribution and geographic concentration of data centers across the Mountain West states of Arizona, Colorado, Nevada, New Mexico and Utah as of March 6th, 2026. Using data from DataCenterMap, this fact sheet examines the number of data centers in each Mountain West state and further analyzes market-level distribution, defined as cities within each state where data centers are located. The data are used to compare state totals and to rank Mountain West markets from highest to lowest based on the number of data centers operating in that area.
Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D.
Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D.
SMU Data Science Review
Classical machine learning models and quantum kernel methods often struggle to capture quantum-coherent molecular features under the constraints of noisy intermediate-scale quantum (NISQ) hardware, limiting both predictive accuracy and scalability.
This paper introduces the Molecular Quantum Particle Algorithm (MQPA), a hybrid quantum–classical framework designed to achieve chemically accurate property prediction by integrating handcrafted molecular descriptors with parameterized quantum circuits. Molecular inputs, expressed as SMILES strings, are processed via RDKit and encoded through angle-based quantum gates with entangling layers in Qiskit [1]. Quantum parameters are optimized using simultaneous perturbation stochastic approximation (SPSA) [2], while classical regression layers leverage Adam [3] …
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
SMU Data Science Review
This study explores the feasibility of an AI-powered chatbot for HIPAA-aligned intake of emergency room patients seeking treatment for overdose and violence. The system utilizes AWS Amplify, an encrypted EC2 instance, and a secure S3 Bucket house on Amazon Web Services. Chat functionality is powered by a multi-agentic framework operating on Anthropic’s Claude Sonnet 4. Manual evaluation and exact match testing reveal the system reliably obtains and records relevant information during intake. Future work will focus on expanding accessibility by integrating voice functionality, obtaining HIPAA compliance certifications, and incorporating the chat system into existing healthcare networks.
Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed
Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed
SMU Data Science Review
Abstract. Large Language Models (LLMs) are being used increasingly within the healthcare industry to summarize complex clinical information, but their outputs can often reflect biases inherited from their training data. In healthcare, these biases are not just technical flaws, but they can lead to distorted and false information about vaccine safety, compromise patient trust, and lead to potential harmful outcomes. This study investigates bias found in LLM-generated outputs to question-answer pairs inspired by adverse vaccine reactions using COVID-19 data from the Vaccine Adverse Event Reporting System (VAERS) from 2020–2024. We examined whether training the LLMs on a known Bias Benchmark …
Nlp Bias And African American English, Kenya Roy, Faizan Javed
Nlp Bias And African American English, Kenya Roy, Faizan Javed
SMU Data Science Review
African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
SMU Data Science Review
Large-scale software systems produce vast volumes of logs and telemetry, making manual incident triage slow and error prone. This study presents an unsupervised anomaly detection pipeline that fuses logs, metrics, and traces through late fusion. Using Hybrid Ensemble modeling with Isolation Forest, and Long Short-Term Memory (LSTM) Deep Learning model, the system detects cross-service anomalies producing and assigning a composite triage score reflecting severity and impact. Ranked alerts are categorized into Critical, High, or Medium priorities for review. A retrieval-augmented generation (RAG) layer enriches results with contextual summaries for explainable triage. Evaluated on synthetic multi-service datasets, the pipeline …
A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia
A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia
SMU Data Science Review
Aspect-based sentiment analysis (ABSA) links opinions in text to specific product attributes (for example, battery life, screen quality, or delivery speed) rather than only assigning an overall star rating. This level of detail is important in domains such as e-commerce, where teams need to know which features customers praised and which they criticized. Traditional ABSA pipelines have relied on large language models (LLMs), which achieved high quality but were expensive to run and difficult to scale. This study evaluated whether small language models (SLMs) in the 1–3 billion parameter range could serve as a lower-cost alternative. We implemented a modular …
Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo
Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo
SMU Data Science Review
The United States has made it clear; it is imperative that the US wins the global AI race. This paper focuses on one of the most challenging puzzle pieces surfaced at the POWER Data Center conference (San Antonio, Sept. 30.); for Electric Reliability Council of Texas (ERCOT) the limiting factor is not generation alone but the need to balance generation and load to preserve grid reliability.
The regulatory landscape fundamentally changed with the passage of Texas Senate Bill 6 in June 2025, which mandates new large loads must "contribute to the recovery of the interconnecting electric utility’s costs" (Texas Legislature, …
Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn
Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn
SMU Data Science Review
The rapid integration of generative AI in finance introduces both opportunities and challenges, particularly when analyzing sensitive data such as Securities and Exchange Commission (SEC) filings. This study investigates the use of open-source Small Large Language Models (SLLMs), deployed locally through the Ollama and LangChain frameworks, combined with Retrieval-Augmented Generation (RAG) for extracting financial insights relevant to index performance and reporting quality. Two key objectives guide this work: (1) benchmarking multiple open-source SLLMs for sentiment analysis, multiple-choice reasoning, and financial question answering, and (2) assessing the feasibility of locally deployed SLLMs for domain-specific financial queries. A standardized set of 50 …
Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler
Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler
SMU Data Science Review
Electric Vehicles (EV) range anxiety remains one of the top barriers for broader adoption. Range anxiety can be attributed to battery pack age and degradation over time. This paper plans to explore how to address this issue by creating a machine learning model that can predict degradation based on usage, temperature, battery chemistry, charging habits and exploring whether other factors tie into range degradation. This research will be using real world charging data along with lab tested chemistry data to build a model that can be chemistry specific for degradation. This paper will help perspective used-EV buyers learn about battery …
“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington
“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington
Publications
Purpose: This study explores how middle-grade girls from predominantly underrepresented and underserved racially and ethnically minoritized (UUREM) backgrounds developed critical data literacy (CDL) through participation in a week-long residential STEM camp. Given the increasing importance of data science education in a data-driven world, this research examines how informal learning environments can support CDL development among youth from historically marginalized groups.
Design/Methodology/Approach: The study draws on qualitative interview data from eleven participants, and their group-produced artifacts to investigate how the camp experience supported engagement with data and the development of CDL. Interviews explored participants' experiences with data collection, organization, analysis, and …
Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift, Patrick D. Murphy
Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift, Patrick D. Murphy
Master's Theses
Semantic segmentation of eelgrass from drone imagery is crucial for coastal habitat monitoring, restoration, and management, as these habitats continue to see rapid changes due to climate change and human influence. However, the reliability of generalizing a deployed classification model relies on both high-accuracy segmentation as well as robust uncertainty quantification that holds up when conditions change over years or locations. Conformal prediction (CP) is a method that converts a classifier's output into prediction sets with a guaranteed average coverage level for in-distribution data. However, the “vanilla” conformal score can often under-cover in hard or out-of-distribution (OOD) regions under drift. …
A Unified Methodological Framework For Generating Digital Twins Of Multi Class Uncrewed Systems (Uxs), Sai Raghava Pathuri
A Unified Methodological Framework For Generating Digital Twins Of Multi Class Uncrewed Systems (Uxs), Sai Raghava Pathuri
Shelby Hall Graduate Research Forum Presentations
No abstract provided.