Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Data Science (8)
- Social and Behavioral Sciences (8)
- Applied Statistics (7)
- Computer Sciences (7)
- Engineering (5)
-
- Multivariate Analysis (5)
- Artificial Intelligence and Robotics (4)
- Mathematics (4)
- Sociology (4)
- Statistical Models (4)
- Arts and Humanities (3)
- Business (3)
- Databases and Information Systems (3)
- Medicine and Health Sciences (3)
- Quantitative, Qualitative, Comparative, and Historical Methodologies (3)
- Business Analytics (2)
- Communication (2)
- Computer Engineering (2)
- Data Storage Systems (2)
- Digital Humanities (2)
- Language Interpretation and Translation (2)
- Linguistics (2)
- Medical Specialties (2)
- Probability (2)
- Public Health (2)
- Software Engineering (2)
- Statistical Methodology (2)
- Institution
-
- Hunan Provincial Institute of Scientific and Technology Information (3)
- City University of New York (CUNY) (2)
- Murray State University (2)
- Arkansas Tech University (1)
- California Polytechnic State University, San Luis Obispo (1)
-
- California State University, Monterey Bay (1)
- Chapman University (1)
- Clemson University (1)
- Illinois State University (1)
- Missouri State University (1)
- Nova Southeastern University (1)
- Thomas Jefferson University (1)
- University of Arkansas, Fayetteville (1)
- University of Mississippi (1)
- University of Nebraska - Lincoln (1)
- Ursinus College (1)
- Utah State University (1)
- Keyword
-
- Data Visualization (2)
- Sports Analytics (2)
- Statistical Analysis (2)
- Accumulated local effects (1)
- Akaike Information Criterion (1)
-
- Algorithmic Design (1)
- Artificial Intelligence (1)
- Automated essay scoring (1)
- Autorotation (1)
- BERT (1)
- Bagging (1)
- Bayesian Information Criterion (1)
- Between the World and Me (1)
- Bias in AI (1)
- Boundary-aware (1)
- Climate Action (1)
- Climate Change (1)
- Clustering (1)
- Colorability (1)
- Community Engagement (1)
- Computer Vision (1)
- Confidence intervals (1)
- Count Min Sketch (1)
- Data analysis (1)
- Database (1)
- Deep learning (1)
- Discourse on Method (1)
- Disruptive innovation (1)
- Disruptive technology (1)
- Early recognition (1)
- Publication
-
- Journal of Scientific Information Research (3)
- Dissertations, Theses, and Capstone Projects (2)
- ATU Honors Projects (1)
- All Dissertations (1)
- All Graduate Reports and Creative Projects, Fall 2023 to Present (1)
-
- Capstone Projects and Master's Theses (1)
- Data Science Undergraduate Honors Theses (1)
- Department of Neurosurgery Faculty Papers (1)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (1)
- Essays in Developmental Psychology (1)
- GMAS Course Syllabi (1)
- Graduate Theses/Dissertations (1)
- Honors College Theses (1)
- Master's Theses (1)
- Mathematics, Computer Science & Statistics Presentations (1)
- Posters-at-the-Capitol (1)
- Student Scholar Symposium Abstracts and Posters (1)
- Theses and Dissertations (1)
- Publication Type
- File Type
Articles 1 - 21 of 21
Full-Text Articles in Categorical Data Analysis
An Assessment And Comparison Of Expert System Performance And Large Language Model Performance, Carter A. Lange
An Assessment And Comparison Of Expert System Performance And Large Language Model Performance, Carter A. Lange
ATU Honors Projects
This study compares the performance of knowledge-based expert systems (KBES) and large language models (LLMs) in narrow-domain tasks. Using Akinator as the representative KBES and ChatGPT as the representative LLM, fifty character-identification trials were conducted. Results show that both systems ultimately succeeded in identifying all characters, but their efficiency and accuracy differ. Akinator required fewer incorrect guesses and produced no identifiable total failures, or “errors,” while ChatGPT occasionally erred beyond possible continuation despite similar average guess counts. Statistical analysis revealed no significant difference in the number of questions required before success, but McNemar’s test indicated that ChatGPT made significantly more …
Bridging Machine Learning And Islamic Scholarship: A Study In Hadith Translation And Similarity Analysis, Asiyah R. Speight
Bridging Machine Learning And Islamic Scholarship: A Study In Hadith Translation And Similarity Analysis, Asiyah R. Speight
Student Scholar Symposium Abstracts and Posters
Translation of Islamic religious texts poses unique challenges requiring both linguistic and theological expertise. This study explores the application of neural machine translation (NMT) models to Arabic-English hadith translation while analyzing semantic similarity patterns across different human translations. Using the complete Sahih Bukhari corpus (7,550 hadiths) as the primary dataset, we adopt a dual approach combining transfer learning and comprehensive neural network analysis to demonstrate the critical impact of corpus size on model performance.
First, we fine-tune a pre-trained MarianMT Arabic-English translation model on the full Sahih Bukhari corpus, comparing models trained on 40 hadiths versus 7,550 hadiths. Performance is …
Engr 597: Special Projects In Engineering Science - Data Analytics, Ali Behnood Ph.D., P.E., M., Asce
Engr 597: Special Projects In Engineering Science - Data Analytics, Ali Behnood Ph.D., P.E., M., Asce
GMAS Course Syllabi
No abstract provided.
Online Prediction Of Streaming Data, Aleena Chanda
Online Prediction Of Streaming Data, Aleena Chanda
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
We present two new approaches for point prediction with streaming data based on a) the Count-Min sketch and b) Gaussian Process Priors with random bias. The methods are intended for the most general case where no true model can be usefully formulated for the data stream. In statistical contexts, this is often called the M open problem class. For the Count Min Sketch method we show that the predicted distribution function ^F converges to F under the assumption that the data consists of i.i.d samples from a fixed distribution function F. To implement the Gaussian Process Prior methods, we used …
Secondary Malignancies In Patients With Meningioma: A Surveillance, Epidemiology, And End Results Data Analysis, Maxwell W. Pickles, Thomas Z. Rohan, Shreya Vinjamuri, Nikolaos Mouchtouris, Roger Murayi, David P. Bray, James J. Evans
Secondary Malignancies In Patients With Meningioma: A Surveillance, Epidemiology, And End Results Data Analysis, Maxwell W. Pickles, Thomas Z. Rohan, Shreya Vinjamuri, Nikolaos Mouchtouris, Roger Murayi, David P. Bray, James J. Evans
Department of Neurosurgery Faculty Papers
BACKGROUND: The risk of secondary primary malignancies (SPMs) in meningioma patients is not well understood. In this unidirectional analysis, we evaluated the risk of SPMs occurring following a primary diagnosis of meningioma.
METHODS: The Surveillance, Epidemiology, and End Results (SEER-17) database (2000-2020) was used to identify 124,769 meningioma patients from a total of 9,208,295 cancer cases. Standardized incidence ratios (SIRs) were calculated using SEER's statistical analysis package to evaluate SPM risk. Basic demographic and treatment information was collected as well.
RESULTS: Of the 124,769 patients, 11,411 (9.2%) received diagnoses of an SPM, which correlates to a higher risk than the …
Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming Yu, Bin Deng, Zhengang Zhang
Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming Yu, Bin Deng, Zhengang Zhang
Journal of Scientific Information Research
[Purpose/significance] This study addresses the issue of inadequate perception of entity boundaries in traditional character-level modeling-based named entity recognition models by integrating syntax information containing entity boundary features into the task using a multi-head graph attention network with dense connections. This integration enhances the effectiveness of named entity recognition.
[Method/process] This study proposes a Syntax-enhanced Boundary-aware Named Entity Recognition Model (SynBNER), which utilizes BERT for text semantic representation and integrates syntax information using a dense-connected graph attention network. This integration incorporates implicit entity boundary information from syntax information into word representations, thereby enhancing the model's entity boundary perception capability.
[Result/conclusion] …
Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh
Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh
Master's Theses
Samara Seeds are a class of fruit most famously belonging to the Acer species and are characterized by their single-bladed geometry and their auto-rotation response during descent. This steady-state auto-rotation response is the subject of aerodynamic analysis which aim to quantify the performance. The period prior to the beginning of steady-state auto-rotation is classified as the transition regime and has not been the subject of intense scrutiny.
This thesis employs a data-driven approach to analyzing the kinematic and dynamic response of these seeds during both the transition and auto-rotation stages of flight to quantify the performance with respect to the …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Community Voices, Climate Action Choices: Working Towards A Resilient Monterey County, Lesley A. Solano Alonso
Community Voices, Climate Action Choices: Working Towards A Resilient Monterey County, Lesley A. Solano Alonso
Capstone Projects and Master's Theses
Vulnerable communities in Monterey County face disproportionate environmental and health impacts due to climate change, yet many residents remain unaware of the tools and resources available to support local action. This capstone project was implemented in partnership with Ecology Action (EA) and the Resilient Central Coast (RCC) campaign to increase awareness and engagement with the RCC platform. Serving diverse communities across Monterey County, the project included bilingual outreach efforts, community tabling, educational presentations, and a climate action survey. Over 650 residents were engaged directly, resulting in 99 new household sign-ups on the RCC website, a major milestone for the agency. …
Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig
Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig
Honors College Theses
Analysis of a childhood game has led us to the problem of maximum independent sets in planar graphs. We wrote a graph creation utility using R to generate a random planar map and its dual graph. This utility then finds a graph’s maximal independent set using a variety of six algorithms. We investigate statistical connections between graph structure, colorability, and the maximal independent sets found using these algorithms over an incredibly large and procedurally generated dataset. We find one can always win the coloring game if the resultant graph is two-colorable. The algorithms perform statistically and practically significantly better on …
Aleci: An R Package For Non-Parametric Confidence Intervals On Accumulated Local Effects Plots, Matthew R. Lister
Aleci: An R Package For Non-Parametric Confidence Intervals On Accumulated Local Effects Plots, Matthew R. Lister
All Graduate Reports and Creative Projects, Fall 2023 to Present
Machine learning models can take a collection of inputs and craft an output. The mathematical formulas these models use to calculate their outputs easily become too complex or time consuming for a human to analyze. Collectively, we refer to these as black box models. Accumulated local effects plots (ALE) are a method for adding interpretability and visibility into the effects that individual variables contribute to the predictions made by black box models. The method designed by D.W. Apley calculates equally spaced point estimates of the response value to construct a graph across the range of the variable of interest. AleCI …
Toward The Application Of Natural Language Processing In Electronic Health Record Analysis For Taxonomy Development, Latoya Mcdonald
Toward The Application Of Natural Language Processing In Electronic Health Record Analysis For Taxonomy Development, Latoya Mcdonald
All Dissertations
Electronic health records (EHRs) are pivotal resources for nurse practice because they increase the timeliness and reliability of patient information at the point of care and support access by multiple healthcare providers and the individual patients themselves. However, it is widely recognized that data extraction from EHRs is challenging due to the variability in the language used in clinical care notes and the lack of standardized terminology across healthcare systems. The broad objective of this dissertation is to develop taxonomy-based classification models for nursing care by applying feature engineering approaches to EHRs that include nursing care of ostomy patients following …
Representative Sampling, Jessica Choe, Lisa K. Lashley, Charles J. Golden
Representative Sampling, Jessica Choe, Lisa K. Lashley, Charles J. Golden
Essays in Developmental Psychology
The research process begins with an initial observation that scientists or researchers of various backgrounds want to engage in and understand, which prompts them to formulate theories and hypotheses, collect data, and use statistical procedures to organize, summarize, and interpret gathered data. Research conducted through questionnaires or surveys are often utilized to determine the nature of a population or the interests of members in a particular group.
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
Mathematics, Computer Science & Statistics Presentations
The purpose of this project was to perform a sentiment analysis of three texts used in Ursinus College's Common Intellectual Experience (CIE) course: Between the World and Me by Ta-Nehisi Coates, The New Jim Crow by Michelle Alexander and Discourse on Method by Rene Descartes. Word count and word cloud analysis were also performed on the texts as well as term frequency and bigram analysis.
Reshaping Future: Research Progress And Future Prospects Of Disruptive Technology Topics, Sanhong Deng, Jianming Guo, Yujie Shi
Reshaping Future: Research Progress And Future Prospects Of Disruptive Technology Topics, Sanhong Deng, Jianming Guo, Yujie Shi
Journal of Scientific Information Research
[Purpose/significance]Disruptive technology is regarded as a revolutionary force that "changes the rules of the game" and "reshapes the future pattern", and has gradually become a hot and difficult issue in interdisciplinary research. This paper summarizes the literature related to disruptive technologies, clarifies the concepts and characteristics of disruptive technologies, subdivides research topics and directions, summarizes research focuses and looks forward to future development trends, and provides reference for relevant personnel. [Method/ process]On the basis of sorting out the latest related research on disruptive technologies, this paper clarifies the research progress of disruptive technologies from three aspects: the concept and characteristics …
Visualizing The Rise Of American Women's Soccer, Hoi Wah Tang
Visualizing The Rise Of American Women's Soccer, Hoi Wah Tang
Dissertations, Theses, and Capstone Projects
The rise of American women’s soccer is a compelling narrative of determination, talent, and triumph on the global stage. From the inaugural FIFA Women’s World Cup in 1991 to the ongoing dominance of the U.S. Women’s National Team (USWNT), this journey encapsulates not only athletic excellence but also the broader struggles for equality, representation, and respect in sports.
This project aims to bring this story to life by presenting a rich tapestry of data visualizations, timelines, and multimedia elements that showcase key moments, players, and milestones. Through engaging graphics and interactive features, users will explore how the sport evolved, the …
Chatgpt Didn’T Write This: Evaluating The Impact Of Llms With A Case Study In Grading Cuny Language Immersion Program Student Essays, Benjamin Inbar
Chatgpt Didn’T Write This: Evaluating The Impact Of Llms With A Case Study In Grading Cuny Language Immersion Program Student Essays, Benjamin Inbar
Dissertations, Theses, and Capstone Projects
This study evaluates the capabilities and limitations of large language models (LLMs), specifically OpenAI’s ChatGPT-4o, in grading essays from students in the City University of New York’s Language Immersion Program. The program serves English language learners with diverse linguistic and demographic backgrounds, offering intensive language instruction to prepare students for academic success in college. Using a dataset of 30 pre- and post-program essays scored by program instructors and ChatGPT-4o under three paradigms, this research explores the alignment between human and AI-generated scores across five rubric-based competency areas. Findings reveal that ChatGPT-4o aligns moderately with human grading, with the strongest agreement …
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
Identifying Research Paths Based On Main Path Analysis: A Case Study Of Knowledge Graphs, Ruibin Wei, Yidan Wang, Yan Xu
Identifying Research Paths Based On Main Path Analysis: A Case Study Of Knowledge Graphs, Ruibin Wei, Yidan Wang, Yan Xu
Journal of Scientific Information Research
[Purpose/significance]The main path analysis of citation networks can be used to identify important literature in specific fields and can achieve the extraction of mainstream research threads. This paper will use the main path analysis method to analyze the research path of knowledge graphs and sort out the context of their research development. [Method/process]This paper firstly obtains research papers in the field of knowledge graphs from the Web of Science platform, then uses the HistCite software to generate a direct citation network of the literature, and then imports the data into Pajek to generate multiple main paths of the dataset, and …
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Theses and Dissertations
Employee turnover poses substantial challenges for technology firms, and understanding its key drivers through predictive modeling is essential for developing effective retention strategies. This study investigates factors influencing employee turnover in technology companies by implementing a predictive modeling approach on the IBM HR Analytics Employee Attrition dataset. The research aims were identifying key factors contributing to employee attrition, developing predictive models to forecast turnover risk, and analyzing interactions among significant predictors. By examining a range of features, the results highlight significant variables (Over Time, Monthly Income, Marital Status, etc.) of attrition and offer actionable insights for developing targeted employee retention …
Predicting Real Estate Prices Using Deep Learning Regression Models On Socio Spatial Data, Gentle Engworo
Predicting Real Estate Prices Using Deep Learning Regression Models On Socio Spatial Data, Gentle Engworo
Graduate Theses/Dissertations
ABSTRACT
Cities keep their own kind of ledger. Every block, bus stop, corner store, and year that slips by leaves a small entry about what homes are worth. That ledger is what we call socio-spatial data: simple facts about what a home is (its age), where it sits (latitude/longitude), how easy it is to get around (distance to the nearest MRT station), what’s nearby (number of convenience stores), and when it sold (transaction date). This thesis asks a practical question in that everyday language: given these common clues, can we predict home prices more accurately and explain why? Using 414 …