Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Journal

Discipline
Institution
Keyword
Publication Year
Publication
File Type

Articles 1 - 30 of 455

Full-Text Articles in Data Science

Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan Jul 2026

Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan

SMU Journal of Undergraduate Research

Scientific communication is typically tied to promoting public engagement and interest in science, increasing scientific literacy, and playing an essential role in policymaking. The success of public communication of scientific findings is largely associated with secondary characteristics of research (e.g. the style of writing and presentation), rather than the primary content or research quality. But it is unclear to what extent the success of scientific literature intended for working scientists is influenced by those same secondary characteristics. Does the writing style of scientific articles impact their success in academic spheres? In this study, we explore the stylometric and formal characteristics …


Unmasking Twitter Bots: An Applied Machine Learning Approach, Rayane El Raba’A, Layal Abu Daher Jun 2026

Unmasking Twitter Bots: An Applied Machine Learning Approach, Rayane El Raba’A, Layal Abu Daher

BAU Journal - Science and Technology

The rapid growth of social networks has led to increased challenges, such as fraud, cyberbullying, and the spread of automated accounts (bots). Detecting anomalies within these networks is essential to maintaining security and trust. This study explored machine learning algorithms: Random Forest, XGBoost, Support Vector Machine (SVM), and Logistic Regression for anomaly detection in social networks, specifically focusing on Twitter bot identification, By applying AI-driven data mining techniques to a dataset of 37,438 Twitter bot accounts dataset, the research evaluates the effectiveness of these models in detecting unusual patterns. XGBoost achieved the highest accuracy (84.9%), with an ROA_AUC of 0.87, …


Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer Jun 2026

Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer

CODEE Journal

Undergraduate instruction in ordinary differential equations (ODEs) is typically organized around the forward problem: finding solution trajectories when the governing equation and its parameters are known. In scientific practice, however, inverse problems are often more relevant, requiring unknown parameters to be inferred from noisy observations while assessing whether a proposed model is consistent with the data. We introduce PINNLab, an open-source MATLAB dashboard designed to help undergraduate students explore inverse modeling through physics-informed neural networks (PINNs). PINNLab presents PINNs as a complementary data-driven framework that connects differential equations, optimization, empirical data, and scientific machine learning. The instructional sequence is organized …


A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson Jun 2026

A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson

CODEE Journal

In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.

This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …


Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh May 2026

Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh

Spora: A Journal of Biomathematics

Ischemic hepatitis (IH) results from shock-related conditions that impair oxygenated blood flow to the liver, causing hepatocyte death. Diagnosis relies largely on clinical history due to the absence of specific diagnostic tests and limited ability to predict outcomes. This study applies machine learning methods to real-world IH patient data to improve outcome prediction. Biomedical indicators analyzed include creatinine, international normalized ratio (INR), aspartate aminotransferase (AST), alanine transaminase (ALT), and bilirubin. Data were collected from multiple U.S. centers through the Acute Liver Failure Study Group (ALFSG), a multicenter network focused on this rare condition. We implemented logistic regression, regression tree methods …


Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama May 2026

Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama

Northeast Journal of Complex Systems (NEJCS)

Financial markets play a critical role in resource allocation. Their performance depends on the decisions of millions of independent investors constantly reacting to one another. Their informational efficiency remains a subject of debate across economic systems. When informational efficiency is present at the weak form, historical price information should not consistently predict future returns. Several empirical tests of this hypothesis often focus on the behavior of aggregate market indices, and use individual efficiency proxies such as autocorrelation, GARCH-type volatility, or entropy-based measures to measure efficiency. This has often yielded mixed results, particularly in emerging markets. Here we show that testing …


From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach Apr 2026

From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach

The Transdisciplinary STEAM+ Journal

This paper presents the design, development, and analysis of my real-time audio visualization system created entirely in Python using PyQt5, called WaveCatcher. The system captures live audio input from a microphone and provides simultaneous visual feedback through multiple signal representations: including a time-domain waveform, a frequency-domain FFT spectrum, a scrolling spectrogram, harmonic peak visualization, dynamic range, amplitude envelope, spectral centroid, and spectral bandwidth. These features offer insight not just into the raw structure of sound, but into how humans perceive its qualities–like timbre! This terminology may seem intimidating—it certainly was when I first began learning it—but I’ll explain all of …


Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi Apr 2026

Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi

HCA Healthcare Journal of Medicine

The integration of artificial intelligence (AI) and machine learning (ML) into health care holds the potential to revolutionize patient care by enhancing clinical decision-making, improving diagnostic accuracy, and reducing costs. Despite this promise, adoption remains limited due to a range of technical, regulatory, educational, and cultural barriers. This paper examines these challenges and proposes strategies to support safe and effective implementation of AI in clinical practice.

Key barriers include the lack of model interpretability, often referred to as the "black box" problem, which undermines clinician trust and accountability in clinical settings, evolving regulatory frameworks and unresolved questions surrounding liability, and …


A Deep Learning-Based Approach For Bot Detection In Trending Hashtags On X, Mehboob Hussain, Muhammad Rizwan Rashid Rana, Muhammad Imran, Muhammad Shoaib, Muhammad Hasaan Mujtaba Apr 2026

A Deep Learning-Based Approach For Bot Detection In Trending Hashtags On X, Mehboob Hussain, Muhammad Rizwan Rashid Rana, Muhammad Imran, Muhammad Shoaib, Muhammad Hasaan Mujtaba

Makara Journal of Technology

The widespread presence of bots on social media platforms, such as X (formerly Twitter), poses a significant threat to the integrity of online information by facilitating the dissemination of misinformation and manipulating public discourse. This study proposes a robust deep learning-based framework, DeepBot, to detect bot participation in trending hashtags and discussions on X. The approach uses a dataset sourced from Kaggle, comprising user profile metadata, including follower count, tweet frequency, account verification status, and engagement metrics. The data were subjected to comprehensive preprocessing, including noise removal, part-of-speech (POS) tagging, and word embedding using the pre-trained GloVe model. RoBERTa is …


The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue Apr 2026

The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue

Chinese/English Journal of Educational Measurement and Evaluation | 教育测量与评估双语期刊

The Item Response Warehouse (IRW) is a repository of harmonized item response datasets designed to support secondary analysis and methodological research in psychological and educational measurement. This paper serves as a practical guide for researchers interested in using the IRW. We describe the structure of IRW datasets and the quantitative and qualitative metadata available for dataset selection, and we demonstrate how researchers can navigate the IRW website to explore and compare available tables. We further show how the IRW R and Python packages can be used to filter datasets programmatically, download response-level data, and generate standardized citations for reproducible research …


The Digital Neuron: Neural Cellular Automata For Neural–Symbolic Translation, Nicole Assenza Apr 2026

The Digital Neuron: Neural Cellular Automata For Neural–Symbolic Translation, Nicole Assenza

SMU Data Science Review

A neural cellular automata (NCA) architecture, referred to as Pluto’s NCA, was developed to characterize bilateral communication and semantic reciprocity between symbolic representations and a spatially distributed update field. The architecture employs an encoder–automata–decoder pipeline that maps symbolic inputs into a multichannel state field and reconstructs them through agreement-driven attractor convergence within a stable semantic attractor landscape. System behavior was evaluated under controlled perturbations, including rhythmic desynchronization, graded ablations, correlated and independent noise, and percolation-based structural degradation. Quantities such as Agreement(t), internal coherence Aᵢ(t), the recovery time constant τ, and the critical percolation threshold pc were measured to assess stability, …


Developing Strategies For Pce Outreach, Mariah Blankenbaker, Gordon Carlson, Daniel Adesoji, Levi Eck Apr 2026

Developing Strategies For Pce Outreach, Mariah Blankenbaker, Gordon Carlson, Daniel Adesoji, Levi Eck

SACAD: Scholarly Activities

In collaboration with Professional and Continuing Education (PCE), we produced projects to automate their internal tasks, as well as to promote their services. By taking advantage of software techniques, we bridged live-action footage with 2D and 3D computer visuals for promotional material. In addition, we researched the capabilities of creating a custom Generative Pre-trained Transformer (GPT) and trained it to analyze and interact with thousands of industry datapoints.


Deconstructing The Black Box: An Explainability Analysis Of Deep Learning Architectures In Cytopathology, Chase A. Garrett Apr 2026

Deconstructing The Black Box: An Explainability Analysis Of Deep Learning Architectures In Cytopathology, Chase A. Garrett

SACAD: Scholarly Activities

Deep learning shows strong potential in medical-image analysis, yet adoption in cyptopathology

remains limited. Cytopathology could benefit from deep learning applications by improving

diagnostic efficiency and accuracy. However deep learning comes with a notorious “black box”

that keeps the models from being transparent and trustworthy for widespread clinical adoption.

We conducted a comprehensive and comparative analysis of several deep learning architectures

for multi-class classification of acute leukemia types, ALL, AML, and normal healthy cells from

peripheral blood smear images. The models in this research include a Vision Transformer (ViT)

and a diverse selection of Convolutional Neural Network (CNN) models. The …


At-Home Computational And Data Literacy For Pre-K–5 Students: A Review Of The Literature, Marc Sager, Sarah Miller, Zarek Drozda Apr 2026

At-Home Computational And Data Literacy For Pre-K–5 Students: A Review Of The Literature, Marc Sager, Sarah Miller, Zarek Drozda

Journal of Educational Research and Practice

This systematic review investigates best practices for promoting data literacy development in pre-K–5 learners, with a specific focus on caregiver involvement in at-home learning environments. A comprehensive search and screening process identified 42 studies for inclusion. The studies were analyzed to determine effective strategies for fostering early data literacy. Key findings emphasize (1) early, developmentally appropriate engagement with attention to equity; (2) core learning outcomes in computational and data literacy; (3) learning approaches integrating real-world contexts and hands-on experiences; (4) balanced use of digital tools and unplugged activities; and (5) essential parental and caregiver involvement. The review highlights integrating data …


Complex Systems Mapping Of Fiscal Growth Dynamics At Strategic Maritime Chokepoints Using Time-Series Slopes, Rahul Balamurugan, Preethi Nanjundan, Avichal Sharma Apr 2026

Complex Systems Mapping Of Fiscal Growth Dynamics At Strategic Maritime Chokepoints Using Time-Series Slopes, Rahul Balamurugan, Preethi Nanjundan, Avichal Sharma

Northeast Journal of Complex Systems (NEJCS)

This study examines how maritime and trading states allocate public resources between defence, health, and economic growth around three strategic chokepoints the Strait of Malacca, the Strait of Hormuz, and the Suez Canal. The analysis extends the classic “guns versus butter” framing by treating defence and health spending as co-evolving components of an interconnected fiscal-growth system. Using World Development Indicators data (1999-2024), trend slopes are estimated for military spending (% of GDP), healthcare spending (% of GDP), and GDP growth (annual %). Two derived indicators are computed, a defence-to-health slope ratio (military slope/health slope) and a fiscal-balance proxy (health slope …


Regional Drought Modulation By Enso And Iod As Indicated By The Standardized Precipitation Index, Arpit Tiwari, Preethi Nanjundan, Tanu Sharma, Ravi Ranjan Kumar, Satyaban Bishoyi Ratna Apr 2026

Regional Drought Modulation By Enso And Iod As Indicated By The Standardized Precipitation Index, Arpit Tiwari, Preethi Nanjundan, Tanu Sharma, Ravi Ranjan Kumar, Satyaban Bishoyi Ratna

Northeast Journal of Complex Systems (NEJCS)

Understanding the modulation of drought by large-scale ocean–atmosphere teleconnections is crucial for strengthening drought prediction and resilience in India. This study investigates the influence of the El Niño–Southern Oscillation (ENSO) and the Indian Ocean Dipole (IOD) on meteorological drought characteristics across India from 1950 to 2024 using the Standardized Precipitation Index (SPI) at a 12-month timescale. Drought events were quantified in terms of frequency, duration, severity, and intensity and linked to ENSO–IOD variability through composite, correlation, and mediation analyses. Results reveal that El Niño events consistently correspond to widespread and severe droughts, particularly over central and southern India, with drought …


Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D. Mar 2026

Molecular Quantum Particle Algorithm (Mqpa): Hybrid Quantum-Classical Learning For Molecular Property Prediction, Jessica Mcphaul, Bivin Sadler Ph.D.

SMU Data Science Review

Classical machine learning models and quantum kernel methods often struggle to capture quantum-coherent molecular features under the constraints of noisy intermediate-scale quantum (NISQ) hardware, limiting both predictive accuracy and scalability.

This paper introduces the Molecular Quantum Particle Algorithm (MQPA), a hybrid quantum–classical framework designed to achieve chemically accurate property prediction by integrating handcrafted molecular descriptors with parameterized quantum circuits. Molecular inputs, expressed as SMILES strings, are processed via RDKit and encoded through angle-based quantum gates with entangling layers in Qiskit [1]. Quantum parameters are optimized using simultaneous perturbation stochastic approximation (SPSA) [2], while classical regression layers leverage Adam [3] …


Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun Mar 2026

Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun

SMU Data Science Review

This study explores the feasibility of an AI-powered chatbot for HIPAA-aligned intake of emergency room patients seeking treatment for overdose and violence. The system utilizes AWS Amplify, an encrypted EC2 instance, and a secure S3 Bucket house on Amazon Web Services. Chat functionality is powered by a multi-agentic framework operating on Anthropic’s Claude Sonnet 4. Manual evaluation and exact match testing reveal the system reliably obtains and records relevant information during intake. Future work will focus on expanding accessibility by integrating voice functionality, obtaining HIPAA compliance certifications, and incorporating the chat system into existing healthcare networks.


Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed Mar 2026

Bias Evaluation Of Healthcare Data With The Use Of Vbqa - A Vaers Inspired Bias Question Answer Dataset, Nolan Dulude, Renu Karthikeyan, Bivin Sadler, Faizan Javed

SMU Data Science Review

Abstract. Large Language Models (LLMs) are being used increasingly within the healthcare industry to summarize complex clinical information, but their outputs can often reflect biases inherited from their training data. In healthcare, these biases are not just technical flaws, but they can lead to distorted and false information about vaccine safety, compromise patient trust, and lead to potential harmful outcomes. This study investigates bias found in LLM-generated outputs to question-answer pairs inspired by adverse vaccine reactions using COVID-19 data from the Vaccine Adverse Event Reporting System (VAERS) from 2020–2024. We examined whether training the LLMs on a known Bias Benchmark …


Nlp Bias And African American English, Kenya Roy, Faizan Javed Mar 2026

Nlp Bias And African American English, Kenya Roy, Faizan Javed

SMU Data Science Review

African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …


Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh Mar 2026

Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh

SMU Data Science Review

Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …


Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal Mar 2026

Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal

SMU Data Science Review

Large-scale software systems produce vast volumes of logs and telemetry, making manual incident triage slow and error prone. This study presents an unsupervised anomaly detection pipeline that fuses logs, metrics, and traces through late fusion. Using Hybrid Ensemble modeling with Isolation Forest, and Long Short-Term Memory (LSTM) Deep Learning model, the system detects cross-service anomalies producing and assigning a composite triage score reflecting severity and impact. Ranked alerts are categorized into Critical, High, or Medium priorities for review. A retrieval-augmented generation (RAG) layer enriches results with contextual summaries for explainable triage. Evaluated on synthetic multi-service datasets, the pipeline …


A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia Mar 2026

A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia

SMU Data Science Review

Aspect-based sentiment analysis (ABSA) links opinions in text to specific product attributes (for example, battery life, screen quality, or delivery speed) rather than only assigning an overall star rating. This level of detail is important in domains such as e-commerce, where teams need to know which features customers praised and which they criticized. Traditional ABSA pipelines have relied on large language models (LLMs), which achieved high quality but were expensive to run and difficult to scale. This study evaluated whether small language models (SLMs) in the 1–3 billion parameter range could serve as a lower-cost alternative. We implemented a modular …


Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo Mar 2026

Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo

SMU Data Science Review

The United States has made it clear; it is imperative that the US wins the global AI race. This paper focuses on one of the most challenging puzzle pieces surfaced at the POWER Data Center conference (San Antonio, Sept. 30.); for Electric Reliability Council of Texas (ERCOT) the limiting factor is not generation alone but the need to balance generation and load to preserve grid reliability.

The regulatory landscape fundamentally changed with the passage of Texas Senate Bill 6 in June 2025, which mandates new large loads must "contribute to the recovery of the interconnecting electric utility’s costs" (Texas Legislature, …


Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn Mar 2026

Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn

SMU Data Science Review

The rapid integration of generative AI in finance introduces both opportunities and challenges, particularly when analyzing sensitive data such as Securities and Exchange Commission (SEC) filings. This study investigates the use of open-source Small Large Language Models (SLLMs), deployed locally through the Ollama and LangChain frameworks, combined with Retrieval-Augmented Generation (RAG) for extracting financial insights relevant to index performance and reporting quality. Two key objectives guide this work: (1) benchmarking multiple open-source SLLMs for sentiment analysis, multiple-choice reasoning, and financial question answering, and (2) assessing the feasibility of locally deployed SLLMs for domain-specific financial queries. A standardized set of 50 …


Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler Mar 2026

Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler

SMU Data Science Review

Electric Vehicles (EV) range anxiety remains one of the top barriers for broader adoption. Range anxiety can be attributed to battery pack age and degradation over time. This paper plans to explore how to address this issue by creating a machine learning model that can predict degradation based on usage, temperature, battery chemistry, charging habits and exploring whether other factors tie into range degradation. This research will be using real world charging data along with lab tested chemistry data to build a model that can be chemistry specific for degradation. This paper will help perspective used-EV buyers learn about battery …


Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda Feb 2026

Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda

Tanzania Journal of Science

The study successfully employed Monte Carlo (MC) simulation and a Machine Learning (ML) approach using a Random Forest Regression (RFR) model to develop optimized Multi-Detector CT (MDCT) protocols that significantly reduce radiation dose while maintaining diagnostic image quality. The MC engine accurately modeled X-ray spectra, and the RFR model demonstrated high predictive power for key metrics, achieving R2 scores of 0.97 for CTDIvol and over 0.92 for image quality metrics (Noise, CNR). Through multi-objective optimization guided by the RFR, the final protocol (Optimization-3) was found on the Pareto front, achieving a notable 35% dose reduction (from 15.5 mGy to 9.9 …


Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell Feb 2026

Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell

Proceedings from the Document Academy

Generative Artificial Intelligences (AIs) and current advanced large language models (LLMs) are algorithmically designed to generate text-based conversations as conversational agents (CAs), by replicating human language and conversational communication. Pairing human cognition with generative computationally coded cognition. We have never been here before: cerebral and artificial information collaborations and processing producing expressions that may or may not become visible as second-hand/secondary source documents.

Sensemaking or sense(un)making is a unique autonomous human drive cognitively, our information processing is sensemaking in action and expressions and articulations are evidence of the sensemaking cycle. Documentation [expressed or articulated through various mediums] are a product …


Correlation Analysis Of Factors Associated With Students’ Numeracy Skills Using Decision Tree Algorithms, Yenita Roza, Arisman Adnan, Zul Indra, Tuti Alawiyah Jan 2026

Correlation Analysis Of Factors Associated With Students’ Numeracy Skills Using Decision Tree Algorithms, Yenita Roza, Arisman Adnan, Zul Indra, Tuti Alawiyah

Numeracy

Numeracy is a critical competency for academic and everyday functioning. This study investigates the key factors associated with students’ numeracy skills by employing decision tree algorithms as a data mining technique. The dataset used in this study is educational assessment data from Indonesia. Utilizing a dataset comprising 6,953 entries and 60 variables from Education Report, the research adopts an exploratory approach involving data preprocessing, exploratory data analysis, and decision tree model construction. The findings reveal that students’ literacy skills serve as the most dominant predictor of numeracy proficiency, emerging as the root node in the decision tree structure. Additional associated …


Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng Jan 2026

Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng

Journal of Aviation Technology and Engineering

This study evaluates the effectiveness of log transformation in enhancing multiple regression models used to forecast air traffic movements (ATMs) in South Africa during the COVID-19 pandemic. Using 60 monthly observations from October 2016 to September 2021, the analysis incorporates variables such as revenue, lockdown levels, COVID-19 metrics, exchange rates, gross domestic product, and population. Two models are compared: one using raw ATMs and another with log-transformed ATMs as the dependent variable.

While the untransformed model shows stronger explanatory power (R² = 0.904, adjusted R² = 0.891) compared to the log-transformed model (R² = 0.772, adjusted R² = 0.741), the …