Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistics and Probability (37)
- Computer Sciences (34)
- Social and Behavioral Sciences (23)
- Artificial Intelligence and Robotics (18)
- Applied Statistics (16)
-
- Statistical Models (16)
- Business (14)
- Medicine and Health Sciences (14)
- Engineering (13)
- Applied Mathematics (10)
- Longitudinal Data Analysis and Time Series (10)
- Theory and Algorithms (7)
- Categorical Data Analysis (6)
- Statistical Methodology (6)
- Biostatistics (5)
- Electrical and Computer Engineering (5)
- Finance and Financial Management (5)
- Life Sciences (5)
- Mathematics (5)
- Other Computer Sciences (5)
- Diseases (4)
- Education (4)
- Environmental Sciences (4)
- Information Security (4)
- Law (4)
- Numerical Analysis and Computation (4)
- Other Public Health (4)
- Public Affairs, Public Policy and Public Administration (4)
- Keyword
-
- Machine Learning (17)
- NLP (15)
- Data Science (14)
- Machine learning (12)
- CNN (10)
-
- Deep Learning (9)
- Natural language processing (9)
- Statistics (8)
- Time series (8)
- Deep learning (7)
- Neural Networks (6)
- Random Forest (6)
- COVID-19 (5)
- Classification (5)
- LSTM (5)
- ARIMA (4)
- Clustering (4)
- Computer Science (4)
- Computer vision (4)
- GAN (4)
- LLM (4)
- LLMs (4)
- Natural Language Processing (4)
- XGBoost (4)
- AI (3)
- American Community Survey (3)
- BERT (3)
- Bias (3)
- Biostatistics (3)
- Data science (3)
- Publication
- Publication Type
Articles 1 - 30 of 144
Full-Text Articles in Data Science
Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan
Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan
SMU Journal of Undergraduate Research
Scientific communication is typically tied to promoting public engagement and interest in science, increasing scientific literacy, and playing an essential role in policymaking. The success of public communication of scientific findings is largely associated with secondary characteristics of research (e.g. the style of writing and presentation), rather than the primary content or research quality. But it is unclear to what extent the success of scientific literature intended for working scientists is influenced by those same secondary characteristics. Does the writing style of scientific articles impact their success in academic spheres? In this study, we explore the stylometric and formal characteristics …
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Mathematics Theses and Dissertations
This dissertation presents a computational framework for high-frequency options trading that combines Cross-Data-Type 1-D Convolutional Neural Networks (CDT-1D CNN) with Simpson-Sobolev regularization for directional prediction, and finite element methods (FEM) for realistic option pricing during backtesting. The core innovation lies in developing a mathematically rigorous regularization approach that maintains the adaptability of modern deep learning while enabling accurate evaluation through stochastic volatility models. The primary contribution is the Simpson-Sobolev regularization scheme, which extends traditional Sobolev regularization by incorporating Simpson’s rule for numerical integration. This approach achieves higher-order accuracy in approximating the Sobolev norms that control function smoothness. Simpson’s rule attains …
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
Civil and Environmental Engineering Theses and Dissertations
Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.
A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
The Digital Neuron: Neural Cellular Automata For Neural–Symbolic Translation, Nicole Assenza
The Digital Neuron: Neural Cellular Automata For Neural–Symbolic Translation, Nicole Assenza
SMU Data Science Review
A neural cellular automata (NCA) architecture, referred to as Pluto’s NCA, was developed to characterize bilateral communication and semantic reciprocity between symbolic representations and a spatially distributed update field. The architecture employs an encoder–automata–decoder pipeline that maps symbolic inputs into a multichannel state field and reconstructs them through agreement-driven attractor convergence within a stable semantic attractor landscape. System behavior was evaluated under controlled perturbations, including rhythmic desynchronization, graded ablations, correlated and independent noise, and percolation-based structural degradation. Quantities such as Agreement(t), internal coherence Aᵢ(t), the recovery time constant τ, and the critical percolation threshold pc were measured to assess stability, …
Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D.
Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D.
SMU Data Science Review
Classical machine learning models and quantum kernel methods often struggle to capture quantum-coherent molecular features under the constraints of noisy intermediate-scale quantum (NISQ) hardware, limiting both predictive accuracy and scalability.
This paper introduces the Molecular Quantum Particle Algorithm (MQPA), a hybrid quantum–classical framework designed to achieve chemically accurate property prediction by integrating handcrafted molecular descriptors with parameterized quantum circuits. Molecular inputs, expressed as SMILES strings, are processed via RDKit and encoded through angle-based quantum gates with entangling layers in Qiskit [1]. Quantum parameters are optimized using simultaneous perturbation stochastic approximation (SPSA) [2], while classical regression layers leverage Adam [3] …
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
SMU Data Science Review
This study explores the feasibility of an AI-powered chatbot for HIPAA-aligned intake of emergency room patients seeking treatment for overdose and violence. The system utilizes AWS Amplify, an encrypted EC2 instance, and a secure S3 Bucket house on Amazon Web Services. Chat functionality is powered by a multi-agentic framework operating on Anthropic’s Claude Sonnet 4. Manual evaluation and exact match testing reveal the system reliably obtains and records relevant information during intake. Future work will focus on expanding accessibility by integrating voice functionality, obtaining HIPAA compliance certifications, and incorporating the chat system into existing healthcare networks.
Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed
Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed
SMU Data Science Review
Abstract. Large Language Models (LLMs) are being used increasingly within the healthcare industry to summarize complex clinical information, but their outputs can often reflect biases inherited from their training data. In healthcare, these biases are not just technical flaws, but they can lead to distorted and false information about vaccine safety, compromise patient trust, and lead to potential harmful outcomes. This study investigates bias found in LLM-generated outputs to question-answer pairs inspired by adverse vaccine reactions using COVID-19 data from the Vaccine Adverse Event Reporting System (VAERS) from 2020–2024. We examined whether training the LLMs on a known Bias Benchmark …
Nlp Bias And African American English, Kenya Roy, Faizan Javed
Nlp Bias And African American English, Kenya Roy, Faizan Javed
SMU Data Science Review
African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
SMU Data Science Review
Large-scale software systems produce vast volumes of logs and telemetry, making manual incident triage slow and error prone. This study presents an unsupervised anomaly detection pipeline that fuses logs, metrics, and traces through late fusion. Using Hybrid Ensemble modeling with Isolation Forest, and Long Short-Term Memory (LSTM) Deep Learning model, the system detects cross-service anomalies producing and assigning a composite triage score reflecting severity and impact. Ranked alerts are categorized into Critical, High, or Medium priorities for review. A retrieval-augmented generation (RAG) layer enriches results with contextual summaries for explainable triage. Evaluated on synthetic multi-service datasets, the pipeline …
A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia
A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia
SMU Data Science Review
Aspect-based sentiment analysis (ABSA) links opinions in text to specific product attributes (for example, battery life, screen quality, or delivery speed) rather than only assigning an overall star rating. This level of detail is important in domains such as e-commerce, where teams need to know which features customers praised and which they criticized. Traditional ABSA pipelines have relied on large language models (LLMs), which achieved high quality but were expensive to run and difficult to scale. This study evaluated whether small language models (SLMs) in the 1–3 billion parameter range could serve as a lower-cost alternative. We implemented a modular …
Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo
Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo
SMU Data Science Review
The United States has made it clear; it is imperative that the US wins the global AI race. This paper focuses on one of the most challenging puzzle pieces surfaced at the POWER Data Center conference (San Antonio, Sept. 30.); for Electric Reliability Council of Texas (ERCOT) the limiting factor is not generation alone but the need to balance generation and load to preserve grid reliability.
The regulatory landscape fundamentally changed with the passage of Texas Senate Bill 6 in June 2025, which mandates new large loads must "contribute to the recovery of the interconnecting electric utility’s costs" (Texas Legislature, …
Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn
Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn
SMU Data Science Review
The rapid integration of generative AI in finance introduces both opportunities and challenges, particularly when analyzing sensitive data such as Securities and Exchange Commission (SEC) filings. This study investigates the use of open-source Small Large Language Models (SLLMs), deployed locally through the Ollama and LangChain frameworks, combined with Retrieval-Augmented Generation (RAG) for extracting financial insights relevant to index performance and reporting quality. Two key objectives guide this work: (1) benchmarking multiple open-source SLLMs for sentiment analysis, multiple-choice reasoning, and financial question answering, and (2) assessing the feasibility of locally deployed SLLMs for domain-specific financial queries. A standardized set of 50 …
Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler
Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler
SMU Data Science Review
Electric Vehicles (EV) range anxiety remains one of the top barriers for broader adoption. Range anxiety can be attributed to battery pack age and degradation over time. This paper plans to explore how to address this issue by creating a machine learning model that can predict degradation based on usage, temperature, battery chemistry, charging habits and exploring whether other factors tie into range degradation. This research will be using real world charging data along with lab tested chemistry data to build a model that can be chemistry specific for degradation. This paper will help perspective used-EV buyers learn about battery …
“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington
“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington
Publications
Purpose: This study explores how middle-grade girls from predominantly underrepresented and underserved racially and ethnically minoritized (UUREM) backgrounds developed critical data literacy (CDL) through participation in a week-long residential STEM camp. Given the increasing importance of data science education in a data-driven world, this research examines how informal learning environments can support CDL development among youth from historically marginalized groups.
Design/Methodology/Approach: The study draws on qualitative interview data from eleven participants, and their group-produced artifacts to investigate how the camp experience supported engagement with data and the development of CDL. Interviews explored participants' experiences with data collection, organization, analysis, and …
Mobile Computer Vision Application For Agricultural Disease Detection Of Pepper Diseases Using Two-Stage Deep Learning System, Carlos Jose Estevez, Mai Dang, Ryan Bass
Mobile Computer Vision Application For Agricultural Disease Detection Of Pepper Diseases Using Two-Stage Deep Learning System, Carlos Jose Estevez, Mai Dang, Ryan Bass
SMU Data Science Review
Plant diseases pose a significant threat to food security, particularly in developing countries where farmers often lack the resources and infrastructure for early detection. In nations like Mexico and the Dominican Republic, the spread of harmful plant diseases impacts key agricultural commodities, such as habanero peppers, leading to substantial yield losses. This study presents a computer vision system based on Convolutional Neural Networks (CNNs) and an object detection model (YOLO) to help farmers detect pepper diseases efficiently. The system uses a two-stage approach: YOLOv11n first detects pepper leaves in images, then a lightweight MobileNetV3Small model classifies whether the detected leaves …
Predicting Simulation Times For Multiphase Thermal-Hydraulic Models, Andrew Yule, Andrew Taylor
Predicting Simulation Times For Multiphase Thermal-Hydraulic Models, Andrew Yule, Andrew Taylor
SMU Data Science Review
Addressing the challenge of computationally intensive OLGA
simulations in the oil and gas industry, a machine learning framework is
developed for accurate runtime prediction. A specialized feature extraction
pipeline identifies key parameters—such as simulation time, time step,
number of branches, and section count—from OLGA input files that serve as
high-impact predictors. Multiple predictive models, including regression,
tree-based ensembles, and neural networks, are implemented to validate
accuracy and robustness. Results reveal that prioritizing simulations based on
predicted runtimes optimizes licensing resources and reduces operational
costs, making real-time scheduling more efficient. This research demonstrates
the effectiveness of data-driven runtime prediction in enhancing …
Ai-Powered Compliance: Accelerating Efficiency And Decision-Making For Compliance Related Inquiries., Amberly R. Rodriguez
Ai-Powered Compliance: Accelerating Efficiency And Decision-Making For Compliance Related Inquiries., Amberly R. Rodriguez
SMU Data Science Review
This research examines the potential of an AI-powered chatbot to streamline compliance workflows by reducing the time and effort required to locate and interpret complex compliance documents. The prototype integrates a centralized MySQL-based document repository, a contextual document querying engine, and a Streamlit web interface, enabling employees to retrieve accurate, document-backed answers within seconds. The system supports both stored and user-uploaded documents, with features such as automated summarization and source citations to enhance transparency and trust. Manual evaluation demonstrated notable gains in efficiency and accuracy compared to traditional search methods, with strong potential to improve adherence to compliance policies. Future …
Ai-Driven Optimization Of Wind Energy Distribution In Texas Using Multi-Agent Reinforcement Learning, Waleed Amer, Owolabi Oluwadamilola, Bassey Ogbonnaya
Ai-Driven Optimization Of Wind Energy Distribution In Texas Using Multi-Agent Reinforcement Learning, Waleed Amer, Owolabi Oluwadamilola, Bassey Ogbonnaya
SMU Data Science Review
Abstract. The integration of large-scale wind power into modern electrical grids presents persistent challenges due to variability, curtailment, and compliance with operational constraints. This study proposes a multi-agent reinforcement learning (MARL) framework for optimizing wind energy distribution within the Texas power grid. The system employs three specialized agents—managing wind curtailment, storage utilization, and load adjustments—to collaboratively balance supply and demand under dynamic grid conditions. Using historical operational data from the Electric Reliability Council of Texas (ERCOT), the framework was trained and evaluated on a range of scenarios encompassing both typical and extreme operating conditions. Results demonstrate substantial performance improvements compared …
A Comparative Time Series Analysis Of The Arima And Temporal Fusion Transformer (Tft) Models, Catherine Ticzon, Aaron Abromowitz, Bivin Sadler
A Comparative Time Series Analysis Of The Arima And Temporal Fusion Transformer (Tft) Models, Catherine Ticzon, Aaron Abromowitz, Bivin Sadler
SMU Data Science Review
Several new transformer-based time series models have been developed in the past five years and research has provided evidence of these models’ superior performance compared to classic statistical models such as ARIMA. While transformer-based models show impressive performance on baseline datasets, no research has been done on the robustness of these models on datasets with controlled modifications and in a replicable manner. In this paper, the Temporal Fusion Transformer (TFT) model was compared to the classical statistical model ARIMA on simulated data using multiple horizons. Data were simulated using a linear combination of exogenous variables; in total, 50 realizations of …
Analyzing The Global Happiness Index, Victoria Hernandez, Christy W. Wachira
Analyzing The Global Happiness Index, Victoria Hernandez, Christy W. Wachira
SMU Data Science Review
This study explores the Global Happiness Index using data compiled from the OECD and Our World in Data to identify key factors contributing to societal well-being. Six primary predictors were analyzed: GDP per capita, social support, healthy life expectancy, freedom to make life choices, generosity, and perceptions of corruption. Regression and clustering techniques were employed to uncover patterns among countries. By expanding the analytical scope beyond conventional economic and social indicators, this study helps identify new pathways for improving well-being across diverse cultural and economic landscapes. Additional variables such as perceived safety, political engagement, and values related to family and …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Mapping Food Justice: Urban Farms And The Examination Of Equitable Food Access, Aaron Avila, Mark Ayiah, Jake Stavely, Marc T. Sager, Maximilian K. Sherard, Anthony J. Petrosino
Mapping Food Justice: Urban Farms And The Examination Of Equitable Food Access, Aaron Avila, Mark Ayiah, Jake Stavely, Marc T. Sager, Maximilian K. Sherard, Anthony J. Petrosino
SMU Journal of Undergraduate Research
In this project, we worked alongside members from an urban farm in South Dallas to learn about issues related to food justice, urban farming, and food deserts. Using participatory design research methods, we created data visualizations showing how society can reduce inequities relating to food access produced in historically underserved neighborhoods. The research goals guiding this study are: a) to identify food deserts and urban farms in the Dallas-Fort Worth metropolitan region (DFW) and b) to determine which urban farms service the needs of these food deserts. To identify food deserts, we took two steps: First, we used open-access data …
Farming With Data: Tracing Critical Tensions Using Data Science For Food Justice, Marc Sager, Maximilan Sherard, Anthony Petrosino
Farming With Data: Tracing Critical Tensions Using Data Science For Food Justice, Marc Sager, Maximilan Sherard, Anthony Petrosino
Publications
In this manuscript, we explore the intersection of artificial intelligence (AI) and equitable learning in higher education, focusing on data science as a subset of AI and social justice as the core theme of equity. Our investigation sheds light on the nuanced tensions inherent in employing data science for social justice. Rooted in situated perspectives of learning and consequential learning, our study employs an instrumental case-study methodology and analysis techniques from interaction and conversation analysis. Collaborating with three undergraduate students and an urban farm, the students used data science practices to highlight inequities surrounding food justice and access to food. …
A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater
A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater
SMU Data Science Review
Paleography, the study of historical handwriting, is essential for preserving societal understanding of cultural, social, and legal frameworks from the past. Medieval manuscripts, often exhibiting refined craftsmanship, present unique challenges to modern readers due to differences in handwriting conventions and the absence of standardized punctuation and spaces. These texts hold valuable insights into the evolution of written communication, literacy, and language development. However, interpreting them requires specialized knowledge and technological solutions. Convolutional Neural Networks (CNNs) can be leveraged to classify scripts, an important step in Historical Document analysis. These models extract and analyze hierarchical features from images, addressing inconsistencies in …
Enhancing News Article Generation With Ai Tools, Adam Ercanbrack, Christopher Johnson, Max Pagan, Nibhrat Lohia
Enhancing News Article Generation With Ai Tools, Adam Ercanbrack, Christopher Johnson, Max Pagan, Nibhrat Lohia
SMU Data Science Review
Abstract. This paper aims to present a comparative analysis of a custom-built AI Journalist Assistant and ChatGPT 4.0 for news article generation. The goal is to evaluate the performance of each model based on accuracy, speed, ethical safeguards, and relevance, particularly in the context of journalism. While ChatGPT is widely used for general-purpose content creation, its reliance on older data and potential for plagiarism presents challenges in the fast-paced, high-stakes world of news reporting. To address these issues, we will design and implement an AI Assistant using Retrieval-Augmented Generation (RAG) techniques, focusing on real-time data access, bias reduction, and plagiarism …
Predictive Modeling Of Colorectal Cancer Risk: Leveraging Health, Demographic, And Socioeconomic Factors For Targeted Screening, Anish Bhandari, Shawn Deng, Michael Olheiser, Robert Slater
Predictive Modeling Of Colorectal Cancer Risk: Leveraging Health, Demographic, And Socioeconomic Factors For Targeted Screening, Anish Bhandari, Shawn Deng, Michael Olheiser, Robert Slater
SMU Data Science Review
Colorectal cancer (CRC) remains a significant public health concern, affecting millions in the United States and worldwide. This study investigates the risk factors associated with CRC using data from the Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial and aims to develop predictive models to identify high-risk individuals for targeted screening and increased awareness. The dataset integrates CRC incidence data from the National Cancer Institute with socioeconomic indicators from U.S. Census Bureau, linked by zip code. We employ Logistic Regression and Neural Network models to predict CRC risk, incorporating health, demographic, and socio-economic features. While the results suggest that …
Context-Switch Attacks: Understanding And Mitigating The Threat To Llm Applications, Sydney Holder, Bivin Sadler
Context-Switch Attacks: Understanding And Mitigating The Threat To Llm Applications, Sydney Holder, Bivin Sadler
SMU Data Science Review
Large Language Models (LLMs) are transforming conversational AI, yet their dependence on prompt-supplied context exposes them to context-switch attacks that covertly steer dialogue toward sensitive or malicious ends. A 70 one-sided conversation transcript evaluation set was constructed spanning various fraudulent scenarios. Each transcript embeds adversarial patterns drawn while preserving natural conversational flow. We introduce a hybrid defense that pairs a BERT-based semantic-drift detector (cosine-similarity threshold = 0.70) with a curated keyword and hack-phrase scanner to counter these threats. In aggregate, the system delivered 100 % recall, intercepting every simulated phishing or data-harvesting attempt. The keyword layer achieved perfect precision, generating …
The Prevalence And Impact Of Discourse In Social Media Networks: The 2024 Presidential Election, Kyle Kuberski, Xavier R. Mojica, Gwonchan J. Yoon, Brad Klein, Bivin P. Sadler
The Prevalence And Impact Of Discourse In Social Media Networks: The 2024 Presidential Election, Kyle Kuberski, Xavier R. Mojica, Gwonchan J. Yoon, Brad Klein, Bivin P. Sadler
SMU Data Science Review
Abstract. Social Media platforms serve as central hubs for global discourse, where political dialogue is widely shared and echoed. This exchange shapes civic participation and influences electoral outcomes, often with both intended and unintended consequences. In its inception, social media platforms served as message boards for the masses, yet manipulation and exploiting of systems via bot usage has made platforms susceptible to outside forces. Thus, false narratives and an artificial sense of consensus are endemically augmented. As for the 2016 and 2020 U.S. presidential elections, it is important to investigate the prevalence of evolved bot activity in both political and …