Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistics and Probability (37)
- Computer Sciences (34)
- Social and Behavioral Sciences (23)
- Artificial Intelligence and Robotics (18)
- Applied Statistics (16)
-
- Statistical Models (16)
- Business (14)
- Medicine and Health Sciences (14)
- Engineering (13)
- Applied Mathematics (10)
- Longitudinal Data Analysis and Time Series (10)
- Theory and Algorithms (7)
- Categorical Data Analysis (6)
- Statistical Methodology (6)
- Biostatistics (5)
- Electrical and Computer Engineering (5)
- Finance and Financial Management (5)
- Life Sciences (5)
- Mathematics (5)
- Other Computer Sciences (5)
- Diseases (4)
- Education (4)
- Environmental Sciences (4)
- Information Security (4)
- Law (4)
- Numerical Analysis and Computation (4)
- Other Public Health (4)
- Public Affairs, Public Policy and Public Administration (4)
- Keyword
-
- Machine Learning (17)
- NLP (15)
- Data Science (14)
- Machine learning (12)
- CNN (10)
-
- Deep Learning (9)
- Natural language processing (9)
- Statistics (8)
- Time series (8)
- Deep learning (7)
- Neural Networks (6)
- Random Forest (6)
- COVID-19 (5)
- Classification (5)
- LSTM (5)
- ARIMA (4)
- Clustering (4)
- Computer Science (4)
- Computer vision (4)
- GAN (4)
- LLM (4)
- LLMs (4)
- Natural Language Processing (4)
- XGBoost (4)
- AI (3)
- American Community Survey (3)
- BERT (3)
- Bias (3)
- Biostatistics (3)
- Data science (3)
- Publication
- Publication Type
Articles 121 - 144 of 144
Full-Text Articles in Data Science
Generating And Smoothing Handwriting With Long Short-Term Memory Networks, Muchigi Kimari, Edward Fry, Ikenna Nwaogu, Yumei Bennett, John Santerre
Generating And Smoothing Handwriting With Long Short-Term Memory Networks, Muchigi Kimari, Edward Fry, Ikenna Nwaogu, Yumei Bennett, John Santerre
SMU Data Science Review
This project explores the different neural network methods to generate synthetic handwriting text. The goal is to offer an AI tool that generates handwriting, while maintaining an individual’s style, to people suffering with Dysgraphia. As part of this project, an application development framework is setup on GitHub, in such a way that others can continue to explore and improve the AI tool.
A Machine Learning Method Of Determining Causal Inference Applied To Shifts In Voting Preferences Between 2012-2016, Jaclyn A. Coate, Reagan Meagher, Megan Riley, John Santerre
A Machine Learning Method Of Determining Causal Inference Applied To Shifts In Voting Preferences Between 2012-2016, Jaclyn A. Coate, Reagan Meagher, Megan Riley, John Santerre
SMU Data Science Review
This research investigates the application of machine learning techniques to assist in the execution of a synthetic control model. This model was performed to analyze counties within the United States that showed a voter shift from a majority of Democratic voter share to Republican between the 2012 and 2016 election cycles. The following study applies two steps of machine learning analysis. The first, which is the treatment discovery process, leverages a Random Forest to evaluate feature importance. The second step was the execution of the synthetic control model with two predictor variable lists. The first was the parametric method: …
Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman
Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman
SMU Data Science Review
Much progress has been made in text analysis, specifically within the statistical domain of Term Frequency (TF) and Inverse Document Frequency (IDF). However, there is much room for improvement especially within the area of discovering Emerging Trends. Emerging Trend Detection Systems (ETDS) depend on ingesting a collection of textual data and TF/IDF to identify new or up-trending topics within the Corpus. However, the tremendous rate of change and the amount of digital information presents a challenge that makes it almost impossible for a human expert to spot emerging trends without relying on an automated ETD system. Since the U.S. Government …
Flow-Based And Packet-Based Intrusion Detection Using Blstm, Brook Andreas, Jayaweera Dilruksha, Eric Mccandless
Flow-Based And Packet-Based Intrusion Detection Using Blstm, Brook Andreas, Jayaweera Dilruksha, Eric Mccandless
SMU Data Science Review
Abstract. Networks are always under the threat of malicious intrusions. Deep learning models are used to help identify and mitigate intrusions before damage can occur. Various types of deep learning models have been researched, built, and tested with the goal of improving intrusion detection and efficiencies. In this paper, a two-phase deep learning approach called a Hybrid Intrusion Detection System (HIDS) is proposed that uses Bi-Directional Long Short-Term Memory Neural Network (BLSTM) to assess both flow-based network data and packet-based data. This approach is unique because BLSTM is employed rather than a traditional Deep Neural Network (DNN) and two models …
Automated Machine Learning Framework For Demand Forecasting In Wholesale Beverage Alcohol Distribution, Jenna Ford, Christian Nava, Jonathan Tan, Bivin Sadler
Automated Machine Learning Framework For Demand Forecasting In Wholesale Beverage Alcohol Distribution, Jenna Ford, Christian Nava, Jonathan Tan, Bivin Sadler
SMU Data Science Review
This paper covers the development, testing, and implementation of an automatic framework for analyzing and forecasting demand for an alcoholic beverage distributor’s products at varying levels of granularity. Rather than look at macroscale geographic demand for a product from a distribution center, this framework will look at the localized customer level demand for that product before aggregating total demand. The approach will better capture individual behavior variations for each customer and allow for a more accurate estimation of the total monthly demand for that product. To best account for each product’s influencing factors, each product is analyzed separately per customer …
Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater
Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater
SMU Data Science Review
This paper proposes a method for the integration of natural language understanding in image classification to improve classification accuracy by making use of associated metadata. Traditionally, only image features have been used in the classification process; however, metadata accompanies images from many sources. This study implemented a multi-modal image classification model that combines convolutional methods with natural language understanding of descriptions, titles, and tags to improve image classification. The novelty of this approach was to learn from additional external features associated with the images using natural language understanding with transfer learning. It was found that the combination of ResNet-50 image …
Analysis Of The Commercial Real Estate Market In A Post Covid-19 World, Brandon Croom, Sean Kennedy, Sandesh Ojha, Justin Sparks
Analysis Of The Commercial Real Estate Market In A Post Covid-19 World, Brandon Croom, Sean Kennedy, Sandesh Ojha, Justin Sparks
SMU Data Science Review
The volatility in the commercial real estate market has been greatly influenced by the new societal practices brought about by the COVID-19 pandemic. The COVID-19 pandemic has added additional factors to already complex modeling to value and predict commercial real estate prices. Although multiple methodologies have been applied to commercial real estate valuation, these methods have not yet taken the COVID-19 pandemic factor into account. The main contribution of this article lies in developing an application for commercial real estate valuation which includes the COVID-19 pandemic factor. Thought this article a Hedonic model was developed to compare the impacts of …
Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels
Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels
SMU Data Science Review
Understanding diagnostic tests and examining important features of novel coronavirus (COVID-19) infection are essential steps for controlling the current pandemic of 2020. In this paper, we study the relationship between clinical diagnosis and analytical features of patient blood panels from the US, Mexico, and Brazil. Our analysis confirms that among adults, the risk of severe illness from COVID-19 increases with pre-existing conditions such as diabetes and immunosuppression. Although more than eight months into pandemic, more data have become available to indicate that more young adults were getting infected. In addition, we expand on the definition of COVID-19 test and discuss …
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Statistical Science Theses and Dissertations
This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.
In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …
Analysis Of Github Pull Requests, Canon Ellis
Analysis Of Github Pull Requests, Canon Ellis
Computer Science and Engineering Theses and Dissertations
The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud
SMU Data Science Review
A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …
Topic Modeling To Understand Technology Talent, Chad Madding, Allen Ansari, Chris Ballenger, Aswini Thota
Topic Modeling To Understand Technology Talent, Chad Madding, Allen Ansari, Chris Ballenger, Aswini Thota
SMU Data Science Review
Attracting technology talent in today’s hiring climate is more complicated than ever. Recruiting for technology talent in non-technology industries is even more challenging. This intense hiring landscape is motivating companies not only to attract the right talent but also to create a culture that can retain and grow that talent. In this paper, we developed algorithms and present insights that use data provided in reviews to glean information employers can use to address or even change their priorities to meet the demands of an ever-changing job market. The core of our research is to investigate and attribute the role of …
Cover Song Identification - A Novel Stem-Based Approach To Improve Song-To-Song Similarity Measurements, Lavonnia Newman, Dhyan Shah, Chandler Vaughn, Faizan Javed
Cover Song Identification - A Novel Stem-Based Approach To Improve Song-To-Song Similarity Measurements, Lavonnia Newman, Dhyan Shah, Chandler Vaughn, Faizan Javed
SMU Data Science Review
Music is incorporated into our daily lives whether intentional or unintentional. It evokes responses and behavior so much so there is an entire study dedicated to the psychology of music. Music creates the mood for dancing, exercising, creative thought or even relaxation. It is a powerful tool that can be used in various venues and through advertisements to influence and guide human reactions. Music is also often "borrowed" in the industry today. The practices of sampling and remixing music in the digital age have made cover song identification an active area of research. While most of this research is focused …
Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite
Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite
SMU Data Science Review
In this paper, modeling techniques for the forecasting of wind speed using historical values observed by Light Detection and Ranging (LIDAR) sensors in an offshore context are described. Both univariate time series and multivariate time series modeling techniques leveraging meteorological data collected simultaneously with the LIDAR data are evaluated for potential contributions to predictive ability. Accurate and timely ability to predict wind values is essential to the effective integration of wind power into existing power grid systems. It allows for both the management of rapid ramp-up / down of base production capacity due to highly variable wind power inputs and …
Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai
Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai
SMU Data Science Review
Social networks sometimes become a medium for threats, insults, and other types of cyberbullying. A large number of people are involved in online social networks. Hence, the protection of network users from anti-social behavior is a critical activity [19]. One of the significant tasks of such activity is the detection of toxic language. Abusive/Toxic language in user-generated online content has become an issue of increasing importance in recent years. Most current commercial methods use blacklists and regular expressions; however, these measures fall short when contending with more subtle, lesser-known examples of hate speech, profanity, or swearing[6]. Abusive language classification has …
Reducing Age Bias In Machine Learning: An Algorithmic Approach, Adriana Solange Garcia De Alford, Steven K. Hayden, Nicole Wittlin, Amy Atwood
Reducing Age Bias In Machine Learning: An Algorithmic Approach, Adriana Solange Garcia De Alford, Steven K. Hayden, Nicole Wittlin, Amy Atwood
SMU Data Science Review
In this paper, we study the prevalence of bias in machine learning; we explore the life cycle phases where bias is potentially introduced into a machine learning model; and lastly, we present how adversarial learning can be leveraged to measure unwanted bias and unfair behavior from a machine learning algorithm. This study focuses particularly on the topics of age bias in predicting employee attrition and presents a practical approach for how adversarial learning can be successful in mitigating age bias. To measure bias, we calculate group fairness metrics across five-year age groups and evaluate fairness between a baseline predictive model …
Forecasting Spare Parts Sporadic Demand Using Traditional Methods And Machine Learning - A Comparative Study, Bhuvana Adur Kannan, Ganesh Kodi, Oscar Padilla, Dough Gray, Barry C. Smith
Forecasting Spare Parts Sporadic Demand Using Traditional Methods And Machine Learning - A Comparative Study, Bhuvana Adur Kannan, Ganesh Kodi, Oscar Padilla, Dough Gray, Barry C. Smith
SMU Data Science Review
Sporadic demand presents a particular challenge to traditional time forecasting methods. In the past 50 years, there has been developments, such as, the Croston Model [3], which has improved forecast performance. With the rise of Machine Learning (ML) there is abundant research in the field of applying ML algorithms to predict sporadic demand [8][12][9]. However, most existing research has analyzed this problem from the demand side [17]. In this paper, we tackle this predictive analytics challenge from the supply side. We perform a comparative analysis utilizing a spare parts demand dataset from an Original Equipment Manufacturer (OEM). Since traditional measurements …
Floor Regularization And Investigation Of Transfer Learning Through Sharing Of Probability Distribution Parameters, Daniel Byrne, Stacey Smith, Joanna Duran, John Santerre
Floor Regularization And Investigation Of Transfer Learning Through Sharing Of Probability Distribution Parameters, Daniel Byrne, Stacey Smith, Joanna Duran, John Santerre
SMU Data Science Review
In this work we introduce a simple new regularization technique, aptly named Floor, which drops low weight connections on every forward pass whenever they fall below a specified event horizon threshold. We compare the results of this technique side by side on identical network architectures between regular Dropout and Floor algorithms. We report similar or improved regularization, with the Floor algorithm versus regular Dropout and/or in concert with regular Dropout.
In this paper we also describe our research into transfer learning by sharing of probability distribution parameters in which we investigated methods of transferring Gaussian prior parameters derived from the …
The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee
The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee
SMU Data Science Review
In this paper, we utilize next-generation sequencing (NGS) data from the LungMap project to identify and characterize the developmental RNA transcriptome in alveolar epithelial type II cells of embryonic mouse lungs of gestational ages embryonic days 16 (E16) and 18 (E18). Late gestation lung cellular maturation is necessary for survival at birth. Using R and the BioConductor packages for RNAseq analysis, we analyze changes in the mouse lung RNA transcriptome as this maturation process takes place. We particularly identify the cluster of genes whose expression changes markedly between immature (E16) and mature (E18) lungs which can be used to define …
Forecasting Power Consumption In Pennsylvania During The Covid-19 Pandemic: A Sarimax Model With External Covid-19 And Unemployment Variables, Jackson Au, Javier Saldaña Jr., Ben Spanswick, John Santerre
Forecasting Power Consumption In Pennsylvania During The Covid-19 Pandemic: A Sarimax Model With External Covid-19 And Unemployment Variables, Jackson Au, Javier Saldaña Jr., Ben Spanswick, John Santerre
SMU Data Science Review
In this paper, we present how electrical consumption can reveal insight into the novel COVID-19 pandemic spread. We analyze electrical power consumption provided by PPL Electric Utilities, Department of Labor’s unemployment claims, and the COVID-19 cases/deaths for the State of Pennsylvania to study the impact of the pandemic on the infrastructure. Using a SARIMA model as our benchmark and we analyzed the use of a SARIMAX model to forecast the power consumption in Pennsylvania 14 days ahead. Our work quantifies and illuminates the effect that the strict legislation passed to minimize the spread of COVID19 had a on power consumption. …
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
SMU Data Science Review
In this paper, we explore a representation methodology for the compression of DNA isolates. Using lossless string compression via tokenization of frequently repeated segments of DNA, we reduce the length of the isolates to be counted as k-mers for classification. With this new representation, we apply a previously established feature sampling method to dramatically reduce the feature space. In understanding the genetic diversity, we also look at conserving biological function across these spaces. Using a random forest model we were able to predict the resistance or susceptibility of bacteria with 85-90\% accuracy, with a 30-50\% reduction in overall isolate length, …
Spoken Language Recognition On Open-Source Datasets, Brady Arendale, Samira Zarandioon, Ryan Goodwin, Douglas Reynolds
Spoken Language Recognition On Open-Source Datasets, Brady Arendale, Samira Zarandioon, Ryan Goodwin, Douglas Reynolds
SMU Data Science Review
The field of speaker and language recognition is constantly being researched and developed, but much of this research is done on private or expensive datasets, making the field more inaccessible than many other areas of machine learning. In addition, many papers make performance claims without comparing their models to other recent research. With the recent development of public multilingual speech corpora such as Mozilla's Common Voice as well as several single-language corpora, we now have the resources to attempt to address both of these problems. We construct an eight-language dataset from Common Voice and a Google Bengali corpus as well …
Predicting Attrition - A Driver For Creating Value, Realizing Strategy, And Refining Key Hr Processes, Kevin Mendonsa, Maureen Stolberg, Vivek Viswanathan, Scott Crum
Predicting Attrition - A Driver For Creating Value, Realizing Strategy, And Refining Key Hr Processes, Kevin Mendonsa, Maureen Stolberg, Vivek Viswanathan, Scott Crum
SMU Data Science Review
Talent is the most important asset for every organization's success. While attrition (or churn) and turnover can refer to both employees and customers, this paper will focus on employee attrition only. Many organizations accept attrition as an inevitable cost of doing business and do nothing to adopt or implement mitigating strategies to combat it. World class companies on the other hand take deliberate measures to understand, control and mitigate attrition (turnover) at every stage. Unmitigated attrition can have a devastating effect on an organization's bottom line and market value. In addition, the “invisible" costs of low employee morale, reduced employee …
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Mathematics Theses and Dissertations
Cell assemblies, defined as groups of neurons forming temporal spike coordination, are thought to be fundamental units supporting major cognitive functions. However, detecting cell assemblies is challenging since they can occur at a range of time scales and with a range of precisions, from synchronous spikes to co-variations in firing rate. In this dissertation, we use a recently published cell assembly detection (CAD) algorithm that is capable of detecting assemblies at a range of time scales and precisions. We first showed that the CAD method can be applied to sparser spike train data than what have previously been reported. This …