Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 91 - 120 of 124

Full-Text Articles in Data Science

Covid-19 - A Graph Network Approach, Nibhrat Lohia, Rajesh Satluri, Suchismita Moharana, Venkat Kasarla Dec 2021

Covid-19 - A Graph Network Approach, Nibhrat Lohia, Rajesh Satluri, Suchismita Moharana, Venkat Kasarla

SMU Data Science Review

The effects of COVID-19 and its spreads are attributed to various factors. This study uses CDC open-source data on COVID-19 effected population with features ranging from location to ethnicity, to create a Knowledge Graph to measure the similarity between COVID-19 cases and estimate the risk for people likely affected by COVID-19. This data could be used to find correlations between distinct factors, like ethnicity and pre-existing health conditions, to find the vulnerability of a given COVID-19 patient. Using the Jaccard similarity coefficient, in the knowledge graph, we are able to identify and explore relationships between COVID-19 cases as well as …


Rocket Learn, Daanesh Ibrahim, Jules Stacy, David Stroud, Yusi Zhang Dec 2021

Rocket Learn, Daanesh Ibrahim, Jules Stacy, David Stroud, Yusi Zhang

SMU Data Science Review

Abstract. This paper covers the development, testing, and implementation of Reinforcement Learning methods designed to autonomously learn and optimize Rocket League play. This study aims to analyze and benchmark model frameworks commonly used in Reinforcement Learning applications. These models can be applied to tasks ranging in difficulty from simple to superhumanly complex, and this study will begin with and build upon simple models performing simple tasks. It will result in complex models performing difficult tasks. Models will be allowed to train autonomously on the game using mass parallelization to expedite training times with the goal of maximizing reward function scores. …


Pokégan: P2p (Pet To Pokémon) Stylizer, Michael B. Hedge, Morgan Nelson, Thomas Pengilly, Michael Weatherford Dec 2021

Pokégan: P2p (Pet To Pokémon) Stylizer, Michael B. Hedge, Morgan Nelson, Thomas Pengilly, Michael Weatherford

SMU Data Science Review

This paper covers the development, testing, and implementation of an automatic framework for converting common images of pets into a Pokémon cartoon with the style of a Pokémon trading card. The technique will first implement object detection for common animals to facilitate image segmentation and apply the appropriate style transfer model to ensure the most aesthetic stylization. It explores various methods to address artifacts in the results of common neural style transfer techniques using Generative Adversarial Networks (GANs). This research sets up a framework to create an app that converts user-submitted pet pictures to Pokémon styled images using the most …


Machine Learning Approach To Distinguish Ulcerative Colitis And Crohn’S Disease Using Smote (Synthetic Minority Oversampling Technique) Methods, Kris Ghimire, Walter Lai, Yasser Omar, Thad Schwebke, Jamie Vo Dec 2021

Machine Learning Approach To Distinguish Ulcerative Colitis And Crohn’S Disease Using Smote (Synthetic Minority Oversampling Technique) Methods, Kris Ghimire, Walter Lai, Yasser Omar, Thad Schwebke, Jamie Vo

SMU Data Science Review

Irritable Bowel Disease (IBD) affects a sizable portion of the US population, causing symptoms such as vomiting, abdominal pain, and diarrhea. Despite the disease’s prevalence, the precise cause is not fully understood. This study consists of endoscopic and histological data from patients diagnosed with IBD and a control population for reference. The machine learning models' focus is to classify patients into IBD types. Several models were analyzed, including decision trees, logistic regression, and k-nearest neighbors. In addition, various methods of SMOTE were applied to determine the most effective transformation and ensuring that the dataset is balanced. The best model with …


Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane Dec 2021

Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane

SMU Data Science Review

Traffic simulations are often used by city planners as a basis for predicting the impact of policies, plans, and operations. The complexities underpinning traffic simulations are often not described in detail yet can significantly impact the simulation outcome. Conflating underlying data for simulations is complex and hinders the interest in this type of exploration. This paper aims to elucidate critical features of traffic simulations that drive the generated metrics of the modeled urban environment. Specifically, this paper examines differences in two road graph networks for the metropolitan region of Houston, TX: a reduced network composed of 45,675 road links and …


Intelligent Investment Portfolio Management Using Time-Series Analytics And Deep Reinforcement Learning, Sachin Chavan, Pradeep Kumar, Tom Gianelle Dec 2021

Intelligent Investment Portfolio Management Using Time-Series Analytics And Deep Reinforcement Learning, Sachin Chavan, Pradeep Kumar, Tom Gianelle

SMU Data Science Review

Abstract. – With globalization, the capital markets have exploded in size and value, making them exceedingly difficult to predict. These days the public has access to real-time data of the market-leading to more participation. As a positive step, this might lead to better wealth distribution in the society, and it also adds to the random nature of the market, making it more unpredictable. The portfolio accounts consisting of stocks and bonds are considered serious investment assets. They can make or break a person’s future. It is also a way of shielding one against market risk or rising inflation. These accounts, …


Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez Dec 2021

Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez

SMU Data Science Review

Vacant lots have been associated with community violence for many years. Researchers have confirmed a positive correlation between vacant lots and vacant buildings with increased violence in urban and rural geographies. However, identifying vacant lots has been a challenge, and modeling methods were largely manual and time-intensive. This prevented cities and non-profit organizations from acting on the information since it was expensive and high-risk to develop remediation programs without clearly understanding where or how many vacant lots existed.

The primary objective of this study was to provide a predictive model that accelerates and improves the accuracy of prior land classification …


Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia Dec 2021

Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia

SMU Data Science Review

Across the United States, record numbers of wildfires are observed costing billions of dollars in property damage, polluting the environment, and putting lives at risk. The ability of emergency management professionals, city planners, and private entities such as insurance companies to determine if an area is at higher risk of a fire breaking out has never been greater. This paper proposes a novel methodology for identifying and characterizing zones with increased risks of forest fires. Methods involving machine learning techniques use the widely available and recorded data, thus making it possible to implement the tool quickly.


Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater Dec 2021

Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater

SMU Data Science Review

The prevalence of data has given consumers the power to make informed choices based off reviews, ratings, and descriptive statistics. However, when a local judge is coming up for re-election there is not any available data that aids voters in making data-driven decision on their vote. Currently court docket data is stored in text or PDFs with very little uniformity. Scaling the collection of this information could prove to be complicated and tiresome. There is a demand for an automated, intelligent system that can extract and organize useful information from the datasets. This paper covers the process of web scraping …


Using Machine Learning Methods To Predict The Movement Trajectories Of The Louisiana Black Bear, Daniel Clark, David Shaw, Armando Vela, Shane Weinstock, John Santerre, Joseph D. Clark May 2021

Using Machine Learning Methods To Predict The Movement Trajectories Of The Louisiana Black Bear, Daniel Clark, David Shaw, Armando Vela, Shane Weinstock, John Santerre, Joseph D. Clark

SMU Data Science Review

In 1992, the Louisiana black bear (Ursus americanus luteolus) was placed on the U.S. Endangered Species List. This was due to bear populations in Louisiana being small and isolated enough where their populations couldn’t intersect with other populations to grow. Interchange of individuals between subpopulations of bears in Louisiana is critical to maintain genetic diversity and avoid inbreeding effects. Utilizing GPS (Global Positioning System) data gathered from 31 radio-collared bears from 2010 through 2012, this research will investigate how bears traverse the landscape, which has implications for gene exchange. This paper will leverage machine learning tools to improve upon existing …


Analyzing Empirical Quality Metrics Of Deep Learning Models For Antimicrobial Resistance, Huy H. Nguyen, Sanjay Pillay, Allison Roderick, Hao Wang, John Santerre May 2021

Analyzing Empirical Quality Metrics Of Deep Learning Models For Antimicrobial Resistance, Huy H. Nguyen, Sanjay Pillay, Allison Roderick, Hao Wang, John Santerre

SMU Data Science Review

Antimicrobial Resistance (AMR) is a growing concern in the medical field. Over-prescription of antibiotics as well as bacterial mutations have caused some once lifesaving drugs to become ineffective against bacteria. However, the problem of AMR might be addressed using Machine Learning (ML) thanks to increased availability of genomic data and large computing resources. The Pathosystems Resource Integration Center (PATRIC) has genomic data of various bacterial genera with sample isolates that are either resistant or susceptible to certain antibiotics. Past research has used this database to use ML algorithms to model AMR with successful results, including accuracies over 80%. To better …


Analysis Of Individual Player Performances And Their Effect On Winning In College Soccer, Angelo Bravo, Thomas Karba, Sean Mcwhirter, Billy Nayden May 2021

Analysis Of Individual Player Performances And Their Effect On Winning In College Soccer, Angelo Bravo, Thomas Karba, Sean Mcwhirter, Billy Nayden

SMU Data Science Review

This study describes the process of modernizing the approach of the Southern Methodist University (SMU) Men's Soccer coaching staff through the use of location and tracking data from their matches in the 2019 season. This study utilizes a variety of modeling and analysis techniques to explore and categorize the data and use it to evaluate the types of plays that are most often correlated with victories. This study's contribution to college soccer analytics includes the implementation of a model to determine individual players' performance, the production of team-level metrics, and visualizations to increase the efficiency of the coaching staff's efforts. …


Machine Learning In The Health Industry: Predicting Congestive Heart Failure And Impactors, Alexandra Norman, James Harding, Daria Zhukova May 2021

Machine Learning In The Health Industry: Predicting Congestive Heart Failure And Impactors, Alexandra Norman, James Harding, Daria Zhukova

SMU Data Science Review

Cardiovascular diseases, Congestive Heart Failure in particular, are a leading cause of deaths worldwide. Congestive Heart Failure has high mortality and morbidity rates. The key to decreasing the morbidity and mortality rates associated with Congestive Heart Failure is determining a method to detect high-risk individuals prior to the development of this often-fatal disease. Providing high-risk individuals with advanced knowledge of risk factors that could potentially lead to Congestive Heart Failure, enhances the likelihood of preventing the disease through implementation of lifestyle changes for healthy living. When dealing with healthcare and patient data, there are restrictions that led to difficulties accessing …


Generating And Smoothing Handwriting With Long Short-Term Memory Networks, Muchigi Kimari, Edward Fry, Ikenna Nwaogu, Yumei Bennett, John Santerre May 2021

Generating And Smoothing Handwriting With Long Short-Term Memory Networks, Muchigi Kimari, Edward Fry, Ikenna Nwaogu, Yumei Bennett, John Santerre

SMU Data Science Review

This project explores the different neural network methods to generate synthetic handwriting text. The goal is to offer an AI tool that generates handwriting, while maintaining an individual’s style, to people suffering with Dysgraphia. As part of this project, an application development framework is setup on GitHub, in such a way that others can continue to explore and improve the AI tool.


A Machine Learning Method Of Determining Causal Inference Applied To Shifts In Voting Preferences Between 2012-2016, Jaclyn A. Coate, Reagan Meagher, Megan Riley, John Santerre May 2021

A Machine Learning Method Of Determining Causal Inference Applied To Shifts In Voting Preferences Between 2012-2016, Jaclyn A. Coate, Reagan Meagher, Megan Riley, John Santerre

SMU Data Science Review

This research investigates the application of machine learning techniques to assist in the execution of a synthetic control model. This model was performed to analyze counties within the United States that showed a voter shift from a majority of Democratic voter share to Republican between the 2012 and 2016 election cycles. The following study applies two steps of machine learning analysis. The first, which is the treatment discovery process, leverages a Random Forest to evaluate feature importance. The second step was the execution of the synthetic control model with two predictor variable lists. The first was the parametric method: …


Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman May 2021

Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman

SMU Data Science Review

Much progress has been made in text analysis, specifically within the statistical domain of Term Frequency (TF) and Inverse Document Frequency (IDF). However, there is much room for improvement especially within the area of discovering Emerging Trends. Emerging Trend Detection Systems (ETDS) depend on ingesting a collection of textual data and TF/IDF to identify new or up-trending topics within the Corpus. However, the tremendous rate of change and the amount of digital information presents a challenge that makes it almost impossible for a human expert to spot emerging trends without relying on an automated ETD system. Since the U.S. Government …


Flow-Based And Packet-Based Intrusion Detection Using Blstm, Brook Andreas, Jayaweera Dilruksha, Eric Mccandless Jan 2021

Flow-Based And Packet-Based Intrusion Detection Using Blstm, Brook Andreas, Jayaweera Dilruksha, Eric Mccandless

SMU Data Science Review

Abstract. Networks are always under the threat of malicious intrusions. Deep learning models are used to help identify and mitigate intrusions before damage can occur. Various types of deep learning models have been researched, built, and tested with the goal of improving intrusion detection and efficiencies. In this paper, a two-phase deep learning approach called a Hybrid Intrusion Detection System (HIDS) is proposed that uses Bi-Directional Long Short-Term Memory Neural Network (BLSTM) to assess both flow-based network data and packet-based data. This approach is unique because BLSTM is employed rather than a traditional Deep Neural Network (DNN) and two models …


Automated Machine Learning Framework For Demand Forecasting In Wholesale Beverage Alcohol Distribution, Jenna Ford, Christian Nava, Jonathan Tan, Bivin Sadler Jan 2021

Automated Machine Learning Framework For Demand Forecasting In Wholesale Beverage Alcohol Distribution, Jenna Ford, Christian Nava, Jonathan Tan, Bivin Sadler

SMU Data Science Review

This paper covers the development, testing, and implementation of an automatic framework for analyzing and forecasting demand for an alcoholic beverage distributor’s products at varying levels of granularity. Rather than look at macroscale geographic demand for a product from a distribution center, this framework will look at the localized customer level demand for that product before aggregating total demand. The approach will better capture individual behavior variations for each customer and allow for a more accurate estimation of the total monthly demand for that product. To best account for each product’s influencing factors, each product is analyzed separately per customer …


Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater Jan 2021

Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater

SMU Data Science Review

This paper proposes a method for the integration of natural language understanding in image classification to improve classification accuracy by making use of associated metadata. Traditionally, only image features have been used in the classification process; however, metadata accompanies images from many sources. This study implemented a multi-modal image classification model that combines convolutional methods with natural language understanding of descriptions, titles, and tags to improve image classification. The novelty of this approach was to learn from additional external features associated with the images using natural language understanding with transfer learning. It was found that the combination of ResNet-50 image …


Analysis Of The Commercial Real Estate Market In A Post Covid-19 World, Brandon Croom, Sean Kennedy, Sandesh Ojha, Justin Sparks Jan 2021

Analysis Of The Commercial Real Estate Market In A Post Covid-19 World, Brandon Croom, Sean Kennedy, Sandesh Ojha, Justin Sparks

SMU Data Science Review

The volatility in the commercial real estate market has been greatly influenced by the new societal practices brought about by the COVID-19 pandemic. The COVID-19 pandemic has added additional factors to already complex modeling to value and predict commercial real estate prices. Although multiple methodologies have been applied to commercial real estate valuation, these methods have not yet taken the COVID-19 pandemic factor into account. The main contribution of this article lies in developing an application for commercial real estate valuation which includes the COVID-19 pandemic factor. Thought this article a Hedonic model was developed to compare the impacts of …


Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels Jan 2021

Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels

SMU Data Science Review

Understanding diagnostic tests and examining important features of novel coronavirus (COVID-19) infection are essential steps for controlling the current pandemic of 2020. In this paper, we study the relationship between clinical diagnosis and analytical features of patient blood panels from the US, Mexico, and Brazil. Our analysis confirms that among adults, the risk of severe illness from COVID-19 increases with pre-existing conditions such as diabetes and immunosuppression. Although more than eight months into pandemic, more data have become available to indicate that more young adults were getting infected. In addition, we expand on the definition of COVID-19 test and discuss …


Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud Dec 2020

Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud

SMU Data Science Review

A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …


Topic Modeling To Understand Technology Talent, Chad Madding, Allen Ansari, Chris Ballenger, Aswini Thota Sep 2020

Topic Modeling To Understand Technology Talent, Chad Madding, Allen Ansari, Chris Ballenger, Aswini Thota

SMU Data Science Review

Attracting technology talent in today’s hiring climate is more complicated than ever. Recruiting for technology talent in non-technology industries is even more challenging. This intense hiring landscape is motivating companies not only to attract the right talent but also to create a culture that can retain and grow that talent. In this paper, we developed algorithms and present insights that use data provided in reviews to glean information employers can use to address or even change their priorities to meet the demands of an ever-changing job market. The core of our research is to investigate and attribute the role of …


Cover Song Identification - A Novel Stem-Based Approach To Improve Song-To-Song Similarity Measurements, Lavonnia Newman, Dhyan Shah, Chandler Vaughn, Faizan Javed Sep 2020

Cover Song Identification - A Novel Stem-Based Approach To Improve Song-To-Song Similarity Measurements, Lavonnia Newman, Dhyan Shah, Chandler Vaughn, Faizan Javed

SMU Data Science Review

Music is incorporated into our daily lives whether intentional or unintentional. It evokes responses and behavior so much so there is an entire study dedicated to the psychology of music. Music creates the mood for dancing, exercising, creative thought or even relaxation. It is a powerful tool that can be used in various venues and through advertisements to influence and guide human reactions. Music is also often "borrowed" in the industry today. The practices of sampling and remixing music in the digital age have made cover song identification an active area of research. While most of this research is focused …


Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite Sep 2020

Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite

SMU Data Science Review

In this paper, modeling techniques for the forecasting of wind speed using historical values observed by Light Detection and Ranging (LIDAR) sensors in an offshore context are described. Both univariate time series and multivariate time series modeling techniques leveraging meteorological data collected simultaneously with the LIDAR data are evaluated for potential contributions to predictive ability. Accurate and timely ability to predict wind values is essential to the effective integration of wind power into existing power grid systems. It allows for both the management of rapid ramp-up / down of base production capacity due to highly variable wind power inputs and …


Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai Sep 2020

Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai

SMU Data Science Review

Social networks sometimes become a medium for threats, insults, and other types of cyberbullying. A large number of people are involved in online social networks. Hence, the protection of network users from anti-social behavior is a critical activity [19]. One of the significant tasks of such activity is the detection of toxic language. Abusive/Toxic language in user-generated online content has become an issue of increasing importance in recent years. Most current commercial methods use blacklists and regular expressions; however, these measures fall short when contending with more subtle, lesser-known examples of hate speech, profanity, or swearing[6]. Abusive language classification has …


Reducing Age Bias In Machine Learning: An Algorithmic Approach, Adriana Solange Garcia De Alford, Steven K. Hayden, Nicole Wittlin, Amy Atwood Sep 2020

Reducing Age Bias In Machine Learning: An Algorithmic Approach, Adriana Solange Garcia De Alford, Steven K. Hayden, Nicole Wittlin, Amy Atwood

SMU Data Science Review

In this paper, we study the prevalence of bias in machine learning; we explore the life cycle phases where bias is potentially introduced into a machine learning model; and lastly, we present how adversarial learning can be leveraged to measure unwanted bias and unfair behavior from a machine learning algorithm. This study focuses particularly on the topics of age bias in predicting employee attrition and presents a practical approach for how adversarial learning can be successful in mitigating age bias. To measure bias, we calculate group fairness metrics across five-year age groups and evaluate fairness between a baseline predictive model …


Forecasting Spare Parts Sporadic Demand Using Traditional Methods And Machine Learning - A Comparative Study, Bhuvana Adur Kannan, Ganesh Kodi, Oscar Padilla, Dough Gray, Barry C. Smith Sep 2020

Forecasting Spare Parts Sporadic Demand Using Traditional Methods And Machine Learning - A Comparative Study, Bhuvana Adur Kannan, Ganesh Kodi, Oscar Padilla, Dough Gray, Barry C. Smith

SMU Data Science Review

Sporadic demand presents a particular challenge to traditional time forecasting methods. In the past 50 years, there has been developments, such as, the Croston Model [3], which has improved forecast performance. With the rise of Machine Learning (ML) there is abundant research in the field of applying ML algorithms to predict sporadic demand [8][12][9]. However, most existing research has analyzed this problem from the demand side [17]. In this paper, we tackle this predictive analytics challenge from the supply side. We perform a comparative analysis utilizing a spare parts demand dataset from an Original Equipment Manufacturer (OEM). Since traditional measurements …


Floor Regularization And Investigation Of Transfer Learning Through Sharing Of Probability Distribution Parameters, Daniel Byrne, Stacey Smith, Joanna Duran, John Santerre Sep 2020

Floor Regularization And Investigation Of Transfer Learning Through Sharing Of Probability Distribution Parameters, Daniel Byrne, Stacey Smith, Joanna Duran, John Santerre

SMU Data Science Review

In this work we introduce a simple new regularization technique, aptly named Floor, which drops low weight connections on every forward pass whenever they fall below a specified event horizon threshold. We compare the results of this technique side by side on identical network architectures between regular Dropout and Floor algorithms. We report similar or improved regularization, with the Floor algorithm versus regular Dropout and/or in concert with regular Dropout.

In this paper we also describe our research into transfer learning by sharing of probability distribution parameters in which we investigated methods of transferring Gaussian prior parameters derived from the …


The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee Aug 2020

The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee

SMU Data Science Review

In this paper, we utilize next-generation sequencing (NGS) data from the LungMap project to identify and characterize the developmental RNA transcriptome in alveolar epithelial type II cells of embryonic mouse lungs of gestational ages embryonic days 16 (E16) and 18 (E18). Late gestation lung cellular maturation is necessary for survival at birth. Using R and the BioConductor packages for RNAseq analysis, we analyze changes in the mouse lung RNA transcriptome as this maturation process takes place. We particularly identify the cluster of genes whose expression changes markedly between immature (E16) and mature (E18) lungs which can be used to define …