Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,358,113 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 48 of 155.

Semantic Structuring Of Digital Documents: Knowledge Graph Generation And Evaluation, Erik E. Luu 2024 Cal Poly

Semantic Structuring Of Digital Documents: Knowledge Graph Generation And Evaluation, Erik E. Luu

Master's Theses

In the era of total digitization of documents, navigating vast and heterogeneous data landscapes presents significant challenges for effective information retrieval, both for humans and digital agents. Traditional methods of knowledge organization often struggle to keep pace with evolving user demands, resulting in suboptimal outcomes such as information overload and disorganized data. This thesis presents a case study on a pipeline that leverages principles from cognitive science, graph theory, and semantic computing to generate semantically organized knowledge graphs. By evaluating a combination of different models, methodologies, and algorithms, the pipeline aims to enhance the organization and retrieval of digital documents. …


Cyberbullying Detection On Twitter Data Using Machine Learning Classifiers, Pradip Dhakal 2024 University of Central Florida

Cyberbullying Detection On Twitter Data Using Machine Learning Classifiers, Pradip Dhakal

Data Science and Data Mining

This study compares some of the popular machine learning techniques like Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbor, and Extreme Gradient Boosting to classify the tweets into three different categories: cyberbullying based on religion, cyberbullying based on ethnicity, or no cyberbullying. First, various data-cleaning approaches are used to clean the tweet data. After the data is clean and ready, the word embedding techniques, such as a bag of words and term frequency-Inverse document frequency, are used to convert the words into mathematical vectors. Finally, the model will be fitted using the combination of the above-mentioned word embedding techniques and machine …


Machine Learning-Based Design Of Doppler Tolerant Radar, Kyle Peter Wensell 2024 New Jersey Institute of Technology

Machine Learning-Based Design Of Doppler Tolerant Radar, Kyle Peter Wensell

Dissertations

In this work, machine learning theory is applied to the design of a radar detector in order to train a machine learning-based detector that is robust against Doppler shifts. The radar system is designed to work with data that would be otherwise intractable to conventional optimal detector design, such as transmitted noise waveforms and the effects of one-bit quantization at the receiver. The detection performance of the one-bit receiver is shown to match the performance of the derived square-law sign correlator detector. The resulting learning-based detector also introduces Doppler tolerance to the system, which allows for the successful detection of …


Financial Time Series Fusion, Completion, And Prediction With Deep Neural Networks, Dan Zhou 2024 New Jersey Institute of Technology

Financial Time Series Fusion, Completion, And Prediction With Deep Neural Networks, Dan Zhou

Dissertations

Time-series analysis is essential for a wide range of financial applications, including but not limited to bond valuation, firm earnings forecasts, firm fundamentals predictions, and firm characteristics imputations. Given its considerable value, the financial community has shown a strong interest in refining and advancing time-series analysis techniques. The study in this dissertation contributes to this field by employing advanced machine learning approaches, specifically graph neural networks, deep neural networks, and matrix/tensor methods. The primary objectives are twofold: first, to reveal complex correlations within financial time series to improve prediction accuracy, and second, to enhance the process of integrating and imputing …


Internet-Based Data Platforms Re-Define The Distributions Of Some Large Crabronid Wasps In Arkansas (Hymenoptera: Crabronidae), David E. Bowles 2024 Ozarks Biological TM

Internet-Based Data Platforms Re-Define The Distributions Of Some Large Crabronid Wasps In Arkansas (Hymenoptera: Crabronidae), David E. Bowles

Insecta Mundi

The geographic distributions of three large wasps, Sphecius speciosus (Drury), Stictia carolina Fabricius, and Stizus brevipennis Walsh (Hymenoptera: Crabronidae), occurring in Arkansas are defined using museum specimens and three internet-based data platforms. The internet-based data platforms generally provided more county location records than museum records. Using data from internet sources for easily identified species can better serve to illustrate the known distributions for some species thus making for a powerful tool elucidating distributional patterns and conservation planning.

ZooBank registration. urn:lsid:zoobank.org:pub:DCAE9192-1765-40CD-952B-0A094F413991


Try It Together - Qualitative Coding With Atlas.Ti, Danping DONG, Bryan LEOW 2024 Singapore Management University

Try It Together - Qualitative Coding With Atlas.Ti, Danping Dong, Bryan Leow

2024 AI for Research Week

This hands-on session introduces Atlas.ti, a well-established qualitative data analysis tool for analyzing your transcripts and textual data. The session will cover coding data, extracting insights, creating visualizations, and exploring the tool's latest AI features.


Try It Together: Transcribing Your Audio With Whisper Api, Bella RATMELIA 2024 Singapore Management University

Try It Together: Transcribing Your Audio With Whisper Api, Bella Ratmelia

2024 AI for Research Week

In this hands-on session, we will explore using the Whisper API to transcribe audio recordings from interviews, focus groups, and speeches. The session will delve into best practices and address common issues that may arise during the transcription process.


Exploring The Tradeoff Between Data Privacy And Utility With A Clinical Data Analysis Use Case, Eunyoung Im, Hyeoneui Kim, Hyungbok Lee, Xiaoqian Jiang, Ju Han Kim 2024 The Texas Medical Center Library

Exploring The Tradeoff Between Data Privacy And Utility With A Clinical Data Analysis Use Case, Eunyoung Im, Hyeoneui Kim, Hyungbok Lee, Xiaoqian Jiang, Ju Han Kim

Faculty, Staff and Student Publications

BACKGROUND: Securing adequate data privacy is critical for the productive utilization of data. De-identification, involving masking or replacing specific values in a dataset, could damage the dataset's utility. However, finding a reasonable balance between data privacy and utility is not straightforward. Nonetheless, few studies investigated how data de-identification efforts affect data analysis results. This study aimed to demonstrate the effect of different de-identification methods on a dataset's utility with a clinical analytic use case and assess the feasibility of finding a workable tradeoff between data privacy and utility.

METHODS: Predictive modeling of emergency department length of stay was used as …


Risk Factors For Pediatric Ischemic Stroke And Intracranial Hemorrhage: A National Electronic Health Record Based Study, Stuart Fraser, Samantha M Levy, Amee Moreno, Gen Zhu, Sean Savitz, Alicia Zha, Hulin Wu 2024 The Texas Medical Center Library

Risk Factors For Pediatric Ischemic Stroke And Intracranial Hemorrhage: A National Electronic Health Record Based Study, Stuart Fraser, Samantha M Levy, Amee Moreno, Gen Zhu, Sean Savitz, Alicia Zha, Hulin Wu

Faculty, Staff and Student Publications

BACKGROUND: Stroke is an important cause of morbidity in pediatrics. Large studies are needed to better understand the epidemiology, pathogenesis and risk factors associated with pediatric stroke. Large administrative datasets can provide information on risk factors in perinatal and childhood stroke at low cost. The aim of this hypothesis-generating study was to use a large administrative dataset to assess for prevalence and odds-ratios of rare exposures associated with pediatric stroke.

METHODS: The data for patients aged 0-18 with a diagnosis of either ischemic stroke or intracranial hemorrhage were extracted from the Cerner Health Facts EMR Database from 2000 to 2018. …


Predictive Analysis Of Local House Prices: Leveraging Machine Learning For Real Estate Valuation, Joey Hernandez, Danny Chang, Santiago Gutierrez, Paul Huggins 2024 Southern Methodist University

Predictive Analysis Of Local House Prices: Leveraging Machine Learning For Real Estate Valuation, Joey Hernandez, Danny Chang, Santiago Gutierrez, Paul Huggins

SMU Data Science Review

This paper presents a comprehensive study examining the real estate market potential in the dynamic urban landscapes of Frisco and Plano, Texas. Combining traditional real estate analysis with cutting-edge machine learning techniques, the study aims to predict home prices and assess investment feasibility. Leveraging these findings, the study proposes a strategic focus on predictive modeling and investment potential identification, emphasizing the continual refinement of machine learning models with updated data to accurately forecast changes in the real estate market. By harnessing the predictive power of these models, investors can identify high-growth areas and optimize their investment decisions, thus capitalizing on …


A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte 2024 Southern Methodist University

A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte

SMU Data Science Review

Current nonlinear time series methods such as neural networks forecast well. However, they act as a black box and are difficult to interpret, leaving the researchers and the audience with little insight into why the forecasts are the way they are. There is a need for a method that forecasts accurately while also being easy to interpret. This paper aims to develop a method to build an interpretable model for univariate and multivariate nonlinear time series data using wavelets and symbolic regression. The final method relies on multilayer perceptron (MLP) neural networks as a form of dimensionality reduction and the …


Intelligent Solutions For Retroactive Anomaly Detection And Resolution With Log File Systems, Derek G. Rogers, Chanvo Nguyen, Abhay Sharma 2024 Southern Methodist University

Intelligent Solutions For Retroactive Anomaly Detection And Resolution With Log File Systems, Derek G. Rogers, Chanvo Nguyen, Abhay Sharma

SMU Data Science Review

This paper explores the intricate challenges log files pose from data science and machine learning perspectives. Drawing inspiration from existing methods, LAnoBERT, PULL, LLMs, and the breadth of recent research, this paper aims to push the boundaries of machine learning for log file systems. Our study comprehensively examines the unique challenges presented in our problem setup, delineates the limitations of existing methods, and introduces innovative solutions. These contributions are organized to offer valuable insights, predictions, and actionable recommendations tailored for Microsoft's engineers working on log data analysis.


Baseball Decision-Making: Optimizing At-Bat Simulations, Varun Gopal, Krithika Kondakindi, Nibhrat Lohia, Morgan Williams 2024 Southern Methodist University

Baseball Decision-Making: Optimizing At-Bat Simulations, Varun Gopal, Krithika Kondakindi, Nibhrat Lohia, Morgan Williams

SMU Data Science Review

Pitch selection in baseball plays a crucial role, involving pitchers, catchers, and batters working together. This practice, dating back to early baseball, has seen teams try various methods to gain an advantage. This research aims to use reinforcement learning and pitch-by-pitch Statcast data to improve batting strategies. It also builds on previous statistical work (sabermetrics) to make better choices in pitch selection and plate discipline. The dataset used, including over 700,000 pitches for each full season and 200,000 pitches for the COVID-shortened 2020 season, encompasses a wealth of crucial metrics including pitch release point, velocity, and launch angle. This study …


Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler II, Alexandria Neff, Adam R. Ruthford 2024 Southern Methodist University

Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler Ii, Alexandria Neff, Adam R. Ruthford

SMU Data Science Review

This paper provides updated forecasts of energy demand in Texas and recognizes the impact of sustainable energy. It is important that the forecasts of the adoption of sustainable energy are reexamined after Winter Storm Uri crippled the Texas power grid and left many without power. This storm highlighted the issues the Texas power grid had and has continued to struggle with in supplying the state with energy. This paper will offer an overview of the relevant literature on the adoption of sustainable energy and relevant events that have occurred in the state of Texas that will give the reader the …


Multi-Class Emotion Classification With Xgboost Model Using Wearable Eeg Headband Data, James Khamthung, Nibhrat Lohia, Seement Srivastava 2024 Southern Methodist University

Multi-Class Emotion Classification With Xgboost Model Using Wearable Eeg Headband Data, James Khamthung, Nibhrat Lohia, Seement Srivastava

SMU Data Science Review

Electroencephalography (EEG) or brainwave signals serve as a valuable source for discerning human activities, thoughts, and emotions. This study explores the efficacy of EXtreme Gradient Boosting (XGBoost) models in sentiment classification using EEG signals, specifically those captured by the MUSE EEG headband. The MUSE device, equipped with four EEG electrodes (TP9, AF7, AF8, TP10), offers a cost-effective alternative to traditional EEG setups, which often utilize over 60 channels in laboratory-grade settings. Leveraging a dataset from previous MUSE research (Bird, J. et al., 2019), emotional states (positive, neutral, and negative) were observed in a male and a female participant, each for …


Building Effective Large Language Model Agents, Sydney Holder, Shreyash Taywade 2024 Southern Methodist University

Building Effective Large Language Model Agents, Sydney Holder, Shreyash Taywade

SMU Data Science Review

The advancement of large language models (LLMs) has significantly expanded the influence of artificial intelligence across various sectors. This paper explores building LLM agents to power applications and examines what is necessary to build an efficient and helpful AI assistant. The research investigates the core components necessary to create specialized agents, facilitate collaboration in problem-solving, and improve human task performance. The development and application of tools designed to augment the capabilities of LLM agents are also explored. The paper addresses the potential risks of the unknowns, such as hallucinations, which can compromise the success of agent-based solutions within LLM applications. …


Game Recommendation Analysis Using Steam Profiles And Reviews, Robert Blue, Luis Garcia, Jacob Turner 2024 Southern Methodist University

Game Recommendation Analysis Using Steam Profiles And Reviews, Robert Blue, Luis Garcia, Jacob Turner

SMU Data Science Review

Smaller game studios are at a disadvantage when it comes to getting their product noticed by users. This study aims to provide insights on how recommendation engines work so that these smaller studios can have their games noticed on Steam. Steam is one of the largest video game distribution services and they have a recommendation engine which promotes games to its user base. This study utilized user information such as number of games played, the type of games, and the hours played and created recommendation engines to identify the qualities in the game that are driving recommendations.


Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn 2024 Southern Methodist University

Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn

SMU Data Science Review

As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …


Detecting Drifts In Data Streams Using Kullback-Leibler (Kl) Divergence Measure For Data Engineering Applications, Jeomoan Francis Kurian, Mohamed Allali 2024 Chapman University

Detecting Drifts In Data Streams Using Kullback-Leibler (Kl) Divergence Measure For Data Engineering Applications, Jeomoan Francis Kurian, Mohamed Allali

Engineering Faculty Articles and Research

The exponential growth of data coupled with the widespread application of artificial intelligence(AI) presents organizations with challenges in upholding data accuracy, especially within data engineering functions. While the Extraction, Transformation, and Loading process addresses error-free data ingestion, validating the content within data streams remains a challenge. Prompt detection and remediation of data issues are crucial, especially in automated analytical environments driven by AI. To address these issues, this study focuses on detecting drifts in data distributions and divergence within data fields processed from different sample populations. Using a hypothetical banking scenario, we illustrate the impact of data drift on automated …


Academic Search And Discovery Tools In The Age Of Ai And Large Language Models: An Overview Of The Space, Aaron TAY 2024 Singapore Management University

Academic Search And Discovery Tools In The Age Of Ai And Large Language Models: An Overview Of The Space, Aaron Tay

2024 AI for Research Week

In the ever-evolving landscape of academic research, “AI tools” for literature search and synthesis are currently getting a lot of attention. These tools promise to ramp up productivity, enabling us to accomplish more in less time or absorb more knowledge without drowning in endless reading. With the sheer number of these systems increasing daily, it's natural to wonder: are they really worth our time and money? And if they are, how should we go about picking the right one from the multitude of options?

In this talk, I will share my views on how the space has developed over two …


Digital Commons powered by bepress