Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2021

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 331 - 360 of 368

Full-Text Articles in Data Science

Topic Modeling And Cultural Nature Of Citations, Marie Coraline Dumaz Jan 2021

Topic Modeling And Cultural Nature Of Citations, Marie Coraline Dumaz

Graduate Theses, Dissertations, and Problem Reports (ETD)

Ever since the beginning of research journals, the number of academic publications has been increasing steadily. Nowadays, especially, with the new importance of online open-access journals and databases, research papers are more easily available to read and share. It also becomes harder to keep up with novelties and grasp an idea of the general impact of a given researcher, institution, journal, or field. For this reason, different bibliometric indicators are now routinely used to classify and evaluate the impact or significance of individual researchers, conferences, journals, or entire scientific communities. In this thesis, we provide tools to study trends in …


Estimating The Azimuthal Mode Structure Of Ultra Low Frequency Waves And Its Effects On The Radial Diffusion Of Radiation Belt Electrons, Mohammad Barani Jan 2021

Estimating The Azimuthal Mode Structure Of Ultra Low Frequency Waves And Its Effects On The Radial Diffusion Of Radiation Belt Electrons, Mohammad Barani

Graduate Theses, Dissertations, and Problem Reports (ETD)

Characterizing the azimuthal mode number �� of Ultra Low Frequency (ULF) waves is critical to quantifying the radial diffusion of radiation belt electrons. A Wavelet cross-spectral technique is applied to the compressional ULF waves observed by multiple pairs of GOES and MMS satellites to estimate the mode structure of ULF waves. A more realistic distribution of mode numbers is achieved by inclusion of the modes corresponding to different wave propagation directions as well as at �� higher than fundamental mode number. For the event study of a geomagnetic storm using GOES data, ULF wave power is found to dominate at …


Analysis And Classification Of Software Fault-Proneness And Vulnerabilities, Mohammad Jamil Ahmad Jan 2021

Analysis And Classification Of Software Fault-Proneness And Vulnerabilities, Mohammad Jamil Ahmad

Graduate Theses, Dissertations, and Problem Reports (ETD)

Software bugs are expensive to fix and can lead to catastrophic consequences. Therefore, their analysis and the use of machine learning for prediction are of the utmost importance. Many prediction models have been proposed and different factors affecting the prediction performance have been extensively studied. This work addresses four topics in two areas in software engineering: software fault-proneness prediction and analysis and classification of security-related bug reports. The first topic focuses on the effect of the learning approach (i.e., the way software fault-proneness prediction models are trained and tested) on the performance of software fault-proneness prediction which lacks extensive research …


A Hybrid Gene Selection Strategy Based On Fisher And Ant Colony Optimization Algorithm For Breast Cancer Classification, Mohammed Hamim, Ismail El Moudden, Mohan D. Pant, Hicham Moutachaouik, Mustapha Hain Jan 2021

A Hybrid Gene Selection Strategy Based On Fisher And Ant Colony Optimization Algorithm For Breast Cancer Classification, Mohammed Hamim, Ismail El Moudden, Mohan D. Pant, Hicham Moutachaouik, Mustapha Hain

EVMS School of Health Professions Faculty Publications

Breast cancer poses the greatest threat to human life and especially to women's life. Despite the progress made in data mining technology in recent years, the ability to predict and diagnose such fatal diseases based on gene expression data still reveals a limited prediction performance, which may not be surprising since most of the genes in expression data are believed to be irrelevant or redundant. The dimensionality reduction process may be considered as a crucial step to analyze gene expression data, as it can reduce the high dimensionality of the breast cancer datasets, which may result into a better prediction …


A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui Jan 2021

A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui

Faculty, Staff and Student Publications

Uncovering and fixing errors in biomedical terminologies is essential so that they provide accurate knowledge to downstream applications that rely on them. Non-lattice-based methods have been applied to identify various kinds of inconsistencies in different biomedical terminologies. In previous work, we have introduced two inference-based approaches that were applied in an exhaustive manner to audit hierarchical relations in the Gene Ontology: (1) Lexical-based inference framework, and (2) Subsumption-based sub-term inference framework. However, it is unclear how effective these exhaustive approaches perform compared with their corresponding non-lattice-based approaches. Therefore, in this paper, we implement the non-lattice versions of these two exhaustive …


A Review And Evaluation Of Techniques For Improved Feature Detection In Mass Spectrometry Data, Annika R. Tostengard, Rob Smith Jan 2021

A Review And Evaluation Of Techniques For Improved Feature Detection In Mass Spectrometry Data, Annika R. Tostengard, Rob Smith

Graduate Student Theses, Dissertations, & Professional Papers

Mass spectrometry (MS) is used in analysis of chemical samples to identify the molecules present and their quantities. This analytical technique has applications in many fields, from pharmacology to space exploration. Its impacts on medicine are particularly significant, since MS aids in the identification of molecules associated with disease; for instance, in proteomics, MS allows researchers to identify proteins that are associated with autoimmune disorders, cancers, and other conditions. Since the applications are so wide-ranging and the tool is ubiquitous across so many fields, it is critical that the analytical methods used to collect data are sound.

Data analysis in …


Inference Of Surface Velocities From Oblique Time Lapse Photos And Terrestrial Based Lidar At The Helheim Glacier, Franklyn T. Dunbar Ii Jan 2021

Inference Of Surface Velocities From Oblique Time Lapse Photos And Terrestrial Based Lidar At The Helheim Glacier, Franklyn T. Dunbar Ii

Graduate Student Theses, Dissertations, & Professional Papers

Using time dependent observations derived from terrestrial LiDAR and oblique
time-lapse imagery, we demonstrate that a Bayesian approach to glacial motion es-
timation provides a concise way to incorporate multiple data products into a single
motion estimation procedure effectively producing surface velocity estimates with
an associated uncertainty. This approach brings both improved computational effi-
ciency, and greater scalability across observational time-frames when compared to
existing methods. To gauge efficacy, we apply these methods to a set of observa-
tions from the Helheim Glacier, a critical actor in contemporary mass loss trends
observed in the Greenland Ice Sheet. We find that …


Clustering Data To Classify Hearthstone Decks, Tim Inzitari Jan 2021

Clustering Data To Classify Hearthstone Decks, Tim Inzitari

Williams Honors College, Honors Research Projects

The esports game of "Hearthstone" is a collectible card game with a competitive format that has every team submit 4 decks of 30 cards each. Using K-Means clustering an adaptable way to group data for classifying can be made that works well in every update of the game. This system will take in a list of decks and cluster them to easily classify large amounts of information in a timely fashion. This system will be able to be used by the Universities esports department for years to come to aid the preparation of "Hearthstone" matches. This model uses qualities about …


The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr Jan 2021

The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr

Statistical and Data Sciences: Faculty Publications

We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.


Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse, Omar El-Gayar, Abdullah Wahbeh, Tareq Nasralah, Ahmed El Noshokaty, Mohammad A. Al-Ramahi Jan 2021

Mental Health And The Covid-19 Pandemic: Analysis Of Twitter Discourse, Omar El-Gayar, Abdullah Wahbeh, Tareq Nasralah, Ahmed El Noshokaty, Mohammad A. Al-Ramahi

Computer Information Systems Faculty Publications (Archived)

This study analyzed Twitter discourse to understand the association of the COVID-19 pandemic with mental health. The study compared tweets’ volume over time, tweets’ volume per mental health category, emotions, and the top hashtags on mental health before and after November 2019, the month on which the first COVID-19 case was reported. We analyzed a total of 273 million English tweets on mental health collected from 56 million unique users. Results and analysis showed a significant shift in trend for the volume of tweets on mental health over time. There was also a notable increase in the volume of tweets …


Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv Jan 2021

Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv

Theses and Dissertations

With data becoming a new form of currency, its analysis has become a top priority in both academia and industry, furthering advancements in high-performance computing and machine learning. However, these large, real-world datasets come with additional complications such as noise and class overlap. Problems are magnified when with multi-class data is presented, especially since many of the popular algorithms were originally designed for binary data. Another challenge arises when the number of examples are not evenly distributed across all classes in a dataset. This often causes classifiers to favor the majority class over the minority classes, leading to undesirable results …


An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill Jan 2021

An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill

Articles

This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble's annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to …


Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems, Rui Wang, Yusheng Xu Jan 2021

Representer Theorems In Banach Spaces: Minimum Norm Interpolation, Regularized Learning And Semi-Discrete Inverse Problems, Rui Wang, Yusheng Xu

Mathematics & Statistics Faculty Publications

Learning a function from a finite number of sampled data points (measurements) is a fundamental problem in science and engineering. This is often formulated as a minimum norm interpolation (MNI) problem, a regularized learning problem or, in general, a semi discrete inverse problem (SDIP), in either Hilbert spaces or Banach spaces. The goal of this paper is to systematically study solutions of these problems in Banach spaces. We aim at obtaining explicit representer theorems for their solutions, on which convenient solution methods can then be developed. For the MNI problem, the explicit representer theorems enable us to express the infimum …


Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries, Chibuzo Iloabuchi Jan 2021

Association Of Incident Cancer To Low-Value Care And Healthcare Cost Burden Among Elderly Medicare Beneficiaries, Chibuzo Iloabuchi

Graduate Theses, Dissertations, and Problem Reports (ETD)

In the United States (US), 25% of healthcare spending is considered wasteful because it is spent reimbursing low-value care. Low-value care is the utilization of healthcare services, medical tests, and procedures that have unclear or no clinical benefit to patients but still exposes them to risk. World-wide, low-value care imposes a significant economic burden on patients, payers, governments, and society. Cancer care among older adults > 65 years is one of the biggest drivers of healthcare expenditure in the US and accounts for nearly 40% of all spending, and low-value care among cancer patients is prevalent and contributes to the financial …


A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments, Roger Sayre, Kevin Butler, Keith Van Graafeiland, Sean Breyer, Dawn Wright, Charlie Frye, Deniz Karagulle, Madeline Martin, Jill Cress, Tom Allen, Rebecca J. Allee, Rost Parsons, Bjorn Nyberg, Mark J. Costello, Peter Harris, Frank E. Muller-Karger Jan 2021

A Global Ecological Classification Of Coastal Segment Units To Complement Marine Biodiversity Observation Network Assessments, Roger Sayre, Kevin Butler, Keith Van Graafeiland, Sean Breyer, Dawn Wright, Charlie Frye, Deniz Karagulle, Madeline Martin, Jill Cress, Tom Allen, Rebecca J. Allee, Rost Parsons, Bjorn Nyberg, Mark J. Costello, Peter Harris, Frank E. Muller-Karger

Political Science & Geography Faculty Publications

A new data layer provides Coastal and Marine Ecological Classification Standard (CMECS) labels for global coastal segments at 1 km or shorter resolution. These characteristics are summarized for six US Marine Biodiversity Observation Network (MBON) sites and one MBON Pole to Pole of the Americas site in Argentina. The global coastlines CMECS classifications were produced from a partitioning of a 30 m Landsat-derived shoreline vector that was segmented into 4 million 1 km or shorter segments. Each segment was attributed with values from 10 variables that represent the ecological settings in which the coastline occurs, including properties of the adjacent …


Plant Species Identification In The Wild Based On Images Of Organs, Meghana Kovur Jan 2021

Plant Species Identification In The Wild Based On Images Of Organs, Meghana Kovur

Graduate Theses, Dissertations, and Problem Reports (ETD)

Image-based plant species identification in the wild is a difficult problem for several reasons. First, the input data is subject to a very high degree of variability because it is captured under fully unconstrained conditions. The same plant species may look very different in different images, while different species can often appear very similar, challenging even the recognition skills of human experts in the field. The large intra-class and small inter-class image variability makes this a fine-grained visual classification problem. One way to cope with this variability and to reduce image background noise is to predict species based on the …


Ensemble Encoder-Decoder Models For Predicting Land Transformation, Pariya Pourmohammadi Jan 2021

Ensemble Encoder-Decoder Models For Predicting Land Transformation, Pariya Pourmohammadi

Graduate Theses, Dissertations, and Problem Reports (ETD)

In studying dynamic and complex processes which are influenced by a system of inter-connected driving variables, it is crucial to apply models that can learn the complexity of the interactions. Land transformation is one of such complex processes, prediction of which can help to mitigate severe climate situations and improve the resiliency of communities. In this study, a multi-spectral set of data cubes is used to capture various characteristics of a geographic region. Based on the data cube, a feature space is constructed using socio-economic attributes, terrain characteristics, and landscape traits of the study region. Two-dimensional and three-dimensional convolutional neural …


Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang Jan 2021

Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang

Graduate Theses, Dissertations, and Problem Reports (ETD)

Automated single-pulse search approaches are necessary as ever-increasing amount of observed data makes the manual inspection impractical. Detecting radio pulsars using single-pulse searches, however, is a challenging problem for machine learning because pul- sar signals often vary significantly in brightness, width, and shape and are only detected in a small fraction of observed data.

The research work presented in this dissertation is focused on development of ma- chine learning algorithms and approaches for single-pulse searches in the time domain. Specifically, (1) We developed a two-stage single-pulse search approach, named Single- Pulse Event Group IDentification (SPEGID), which automatically identifies and clas- …


Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington Jan 2021

Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington

Graduate Theses, Dissertations, and Problem Reports (ETD)

The necessary materials for most human activities are water and energy. Integrated analysis to accurately forecast water and energy consumption enables the implementation of efficient short and long-term resource management planning as well as expanding policy and research possibilities for the supportive infrastructure. However, the integral relationship between water and energy (water-energy nexus) poses a difficult problem for modeling. The accessibility and physical overlay of data sets related to water-energy nexus is another main issue for a reliable water-energy consumption forecast. The framework of urban metabolism (UM) uses several types of data to build a global view and highlight issues …


Why We Need Better Corporate Governance Data, Jens Frankenreiter, Cathy Hwang, Yaron Nili, Eric L. Talley Jan 2021

Why We Need Better Corporate Governance Data, Jens Frankenreiter, Cathy Hwang, Yaron Nili, Eric L. Talley

Scholarship@WashULaw

Three decades of finance, economics, and legal studies in corporate governance have been built substantially on data sets with nearly unknown provenance. A new paper sets to correct this fatal flaw of contemporary corporate governance research by debuting a brand new resource—the Cleaning Corporate Governance database.


Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee Jan 2021

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee

Senior Independent Study Theses

According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.

We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …


Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang Jan 2021

Identifying Sleep-Related Factors Associated With Cognitive Function In A Hispanics/Latinos Cohort: A Dual Random Forest Approach, Li Xiaojin, Cui Licong, Wang Fei, Paul E Schulz, Guo-Qiang Zhang

Faculty, Staff and Student Publications

Disordered sleep is associated with poor cognitive function and cognitive decline. However, little is known regarding the association of sleep-related factors with cognitive function in underrepresented cohorts such as the Hispanic/Latino population. Leveraging the National Sleep Research Resource, one of the most comprehensive collections of sleep studies, we identified a Hispanic/Latino cohort of 1,031 lower cognitive function cases and 2,062 normal controls. We developed a novel dual random forest (DRF) approach to discriminate cases against controls for estimating the potential impact of sleep-related variables related to the decline of cognitive function. Several important sleep-related factors were identified which may be …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke Jan 2021

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue Jan 2021

Super-Resolution Imaging Of Remote Sensed Brightness Temperature Using A Convolutional Neural Network, Kellen A. Donahue

Graduate Student Theses, Dissertations, & Professional Papers

Steady improvements to the instruments used in remote sensing has led to much higher resolution data, often contemporaneous with lower resolution instruments that continue to collect data. There is a clear opportunity to reconcile recent high resolution satellite data with the lower resolution data of the past. Super-resolution (SR) imaging is a technique that increases the spatial resolution of image data by training statistical methods on simultaneously occurring lower and higher resolution data sets. The special sensor microwave/imager (SSMI) and advanced microwave scanning radiometer (AMSR2) brightness temperature data products are well suited to super-resolution imaging, and SR can be used …


Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle Jan 2021

Interactive Visual Self-Service Data Classification Approach To Democratize Machine Learning, Sridevi Narayana Wagle

All Master's Theses

Machine learning algorithms often produce models considered as complex black-box models by both end users and developers. Such algorithms fail to explain the model in terms of the domain they are designed for. The proposed Iterative Visual Logical Classifier (IVLC) is an interpretable machine learning algorithm that allows end users to design a model and classify data with more confidence and without having to compromise on the accuracy. Such technique is especially helpful when dealing with sensitive and crucial data like cancer data in the medical domain with high cost of errors. With the help of the proposed interactive and …


K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant Jan 2021

K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant

Theses and Dissertations

Traditional density-based clustering approaches rely on a distance-based parameter to define data connectivity and density. However, an appropriate value of this parameter can be difficult to determine as it is highly dependent on the underlying distribution of the data. In particular, distribution parameters affect the scale of inter-group distances (e.g., variance); this dependence leads to a well-known inability to simultaneously detect clusters at varying levels of density. In this work, connectivity and density are defined according to the rank-order induced by the distance metric (i.e., invariant to the expected scale of the distances). Connectivity by k-nearest neighbors and density by …


Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini Jan 2021

Continual Learning For Multi-Label Drifting Data Streams Using Homogeneous Ensemble Of Self-Adjusting Nearest Neighbors, Gavin Alberghini

Theses and Dissertations

Multi-label data streams are sequences of multi-label instances arriving over time to a multi-label classifier. The properties of the data stream may continuously change due to concept drift. Therefore, algorithms must adapt constantly to the new data distributions. In this paper we propose a novel ensemble method for multi-label drifting streams named Homogeneous Ensemble of Self-Adjusting Nearest Neighbors (HESAkNN). It leverages a self-adjusting kNN as a base classifier with the advantages of ensembles to adapt to concept drift in the multi-label environment. To promote diverse knowledge within the ensemble, each base classifier is given a unique subset of features and …


Research Data Curation And Management Bibliography, Charles W. Bailey Jr. Jan 2021

Research Data Curation And Management Bibliography, Charles W. Bailey Jr.

Copyright, Fair Use, Scholarly Communication, etc.

Preface

The Research Data Curation and Management Bibliography includes over 800 selected English-language articles and books that are useful in understanding the curation of digital research data in academic and other research institutions.

The "digital curation" concept is still evolving. In "Digital Curation and Trusted Repositories: Steps toward Success," Christopher A. Lee and Helen R. Tibbo define digital curation as follows:

Digital curation involves selection and appraisal by creators and archivists; evolving provision of intellectual access; redundant storage; data transformations; and, for some materials, a commitment to long-term preservation. Digital curation is stewardship that provides for the reproducibility and re-use …


Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam Jan 2021

Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic triggered a large-scale work-from-home trend globally in recent months. In this paper, we study the phenomenon of “work-from-home” (WFH) by performing social listening. We propose an analytics pipeline designed to crawl social media data and perform text mining analyzes on textual data from tweets scrapped based on hashtags related to WFH in COVID-19 situation. We apply text mining and NLP techniques to analyze the tweets for extracting the WFH themes and sentiments (positive and negative). Our Twitter theme analysis adds further value by summarizing the common key topics, allowing employers to gain more insights on areas of …


Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang Jan 2021

Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang

Theses and Dissertations (Comprehensive)

Terrorism becomes more rampant in recent years because of separatism and extreme nationalism, which brings a serious threat to the national security of many countries in the world. The analysis of spatial and temporal patterns of terror data is significant in containing terrorism. This thesis focuses on building and applying a temporal point process called self-exciting point process to fit the terror data from 1970 to 2018 of 10 countries. The data come from the Global Terrorism database. Further, an application in predicting the number of terror events based on the self-exciting model is another main innovative idea, in which …