Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (52)
- Applied Statistics (46)
- Data Science (37)
- Computer Sciences (30)
- Statistical Methodology (27)
-
- Longitudinal Data Analysis and Time Series (22)
- Categorical Data Analysis (21)
- Biostatistics (20)
- Social and Behavioral Sciences (17)
- Engineering (15)
- Artificial Intelligence and Robotics (13)
- Business (13)
- Theory and Algorithms (11)
- Medicine and Health Sciences (10)
- Multivariate Analysis (10)
- Mathematics (9)
- Numerical Analysis and Scientific Computing (8)
- Other Computer Sciences (8)
- Other Statistics and Probability (8)
- Survival Analysis (8)
- Applied Mathematics (7)
- Probability (7)
- Statistical Theory (7)
- Analysis (5)
- Business Analytics (5)
- Computer Engineering (5)
- Public Affairs, Public Policy and Public Administration (5)
- Technology and Innovation (5)
- Keyword
-
- Statistics (37)
- Biostatistics (11)
- Machine Learning (11)
- Deep Learning (10)
- LSTM (7)
-
- Machine learning (7)
- CNN (6)
- Classification (6)
- Data Science (5)
- NLP (5)
- Time Series (5)
- Time series (5)
- ARIMA (4)
- Computer Science (4)
- Data science (4)
- Forecasting (4)
- Logistic regression (4)
- Regression (4)
- Analysis (3)
- Deep learning (3)
- Neural Network (3)
- Neural network (3)
- Random Forest (3)
- Random forest (3)
- Sports (3)
- Transfer Learning (3)
- ARMA (2)
- Algorithm (2)
- Applied statistics (2)
- Bayesian (2)
- Publication Year
- Publication
- Publication Type
Articles 61 - 90 of 119
Full-Text Articles in Statistics and Probability
Statistical Models And Analysis Of Univariate And Multivariate Degradation Data, Lochana Palayangoda
Statistical Models And Analysis Of Univariate And Multivariate Degradation Data, Lochana Palayangoda
Statistical Science Theses and Dissertations
For degradation data in reliability analysis, estimation of the first-passage time (FPT) distribution to a threshold provides valuable information on reliability characteristics. Recently, Balakrishnan and Qin (2019; Applied Stochastic Models in Business and Industry, 35:571-590) studied a nonparametric method to approximate the FPT distribution of such degradation processes if the underlying process type is unknown. In this thesis, we propose improved techniques based on saddlepoint approximation, which enhance upon their suggested methods. Numerical examples and Monte Carlo simulation studies are used to illustrate the advantages of the proposed techniques. Limitations of the improved techniques are discussed and some possible solutions …
Forecasting San Francisco Bay Area Rapid Transit (Bart) Ridership, Swee K. Chew, Alec Lepe, Aaron Tomkins, Peter Scheirer
Forecasting San Francisco Bay Area Rapid Transit (Bart) Ridership, Swee K. Chew, Alec Lepe, Aaron Tomkins, Peter Scheirer
SMU Data Science Review
In this paper, we present a forecasting analysis of the San Francisco Bay Area Rapid Transit (BART) ridership data utilizing a number of different time series methods. BART is a major public transportation system in the Bay Area and it relies heavily on its riders' fares; having models that generate accurate ridership numbers better enables the agency to project revenue and help manage future expenses. For our time series modeling, we utilized autoregressive integrated moving average (ARIMA), deep neural networks (DNN), state space models, and long short-term memory (LSTM) to predict monthly ridership. As there is such a wide range …
Universal Vector Neural Machine Translation With Effective Attention, Joshua Yi, Satish Mylapore, Ryan Paul, Robert Slater
Universal Vector Neural Machine Translation With Effective Attention, Joshua Yi, Satish Mylapore, Ryan Paul, Robert Slater
SMU Data Science Review
Neural Machine Translation (NMT) leverages one or more trained neural networks for the translation of phrases. Sutskever intro- duced a sequence to sequence based encoder decoder model which be- came the standard for NMT based systems. Attention mechanisms were later introduced to address the issues with the translation of long sen- tences and improving overall accuracy. In this paper, we propose two improvements to the encoder decoder based NMT approach. Most trans- lation models are trained as one model for one translation. We introduce a neutral/universal model representation that can be used to predict more than one language depending on …
Demand Forecasting In Wholesale Alcohol Distribution: An Ensemble Approach, Tanvi Arora, Rajat Chandna, Stacy Conant, Bivin Sadler, Robert Slater
Demand Forecasting In Wholesale Alcohol Distribution: An Ensemble Approach, Tanvi Arora, Rajat Chandna, Stacy Conant, Bivin Sadler, Robert Slater
SMU Data Science Review
In this paper, historical data from a wholesale alcoholic beverage distributor was used to forecast sales demand. Demand forecasting is a vital part of the sale and distribution of many goods. Accurate forecasting can be used to optimize inventory, improve cash ow, and enhance customer service. However, demand forecasting is a challenging task due to the many unknowns that can impact sales, such as the weather and the state of the economy. While many studies focus effort on modeling consumer demand and endpoint retail sales, this study focused on demand forecasting from the distributor perspective. An ensemble approach was applied …
Demand Forecasting For Alcoholic Beverage Distribution, Lei Jiang, Kristen M. Rollins, Meredith Ludlow, Bivin Sadler
Demand Forecasting For Alcoholic Beverage Distribution, Lei Jiang, Kristen M. Rollins, Meredith Ludlow, Bivin Sadler
SMU Data Science Review
Forecasting demand is one of the biggest challenges in any business, and the ability to make such predictions is an invaluable resource to a company. While difficult, predicting demand for products should be increasingly accessible due to the volume of data collected in businesses and the continuing advancements of machine learning models. This paper presents forecasting models for two vodka products for an alcoholic beverage distributing company located in the United States with the purpose of improving the company’s ability to forecast demand for those products. The results contain exploratory data analysis to determine the most important variables impacting demand, …
Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock
Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock
SMU Journal of Undergraduate Research
Professional development in its most traditional form is a classroom setting with a lecturer and an overwhelming amount of information. It is no surprise, then, that informal professional development away from institutions and on the teacher's own terms is a growing phenomenon due to an increased presence of educators on social media. These communities of educators use hashtags to broadcast to each other, with general hashtags such as #edchat having the broadest audience. However, many math educators usethe hashtags #ITeachMath and #MTBoS, communities I was interested in learning more about. I built a python script that used Tweepy to connect …
Quantitative Model For Setting Manufacturer's Suggested Retail Price, Peter Byrd, Jonathan Knowles, Dmitry Andreev, Jacob Turner, Brian Mente, Laroux Wallace
Quantitative Model For Setting Manufacturer's Suggested Retail Price, Peter Byrd, Jonathan Knowles, Dmitry Andreev, Jacob Turner, Brian Mente, Laroux Wallace
SMU Data Science Review
In this paper, we present a quantitative approach to model the manufacturer’s suggested retail price (MSRP) for children’s doll- houses and establish relationships among key features that contribute most to establishing MSRP. Determination of the MSRP is a critical step in how consumers respond with their wallets when purchasing an item. KidKraft, a global leader in toys and juvenile products, sets MSRP subjectively using product experts. The process is arduous and time consuming requiring the focus of specialized resources and knowledge of the interaction between key attributes and their impact on consumer value. An accurate prediction of MSRP during the …
Mapping Relationships And Positions Of Objects In Images Using Mask And Bounding Box Data, Jaime M. Villanueva Jr, Anantharam Subramanian, Vishal Ahir, Andrew Pollock
Mapping Relationships And Positions Of Objects In Images Using Mask And Bounding Box Data, Jaime M. Villanueva Jr, Anantharam Subramanian, Vishal Ahir, Andrew Pollock
SMU Data Science Review
In this paper we present novel methods for automatically annotating images with relationship and position tags that are derived using mask and bounding box data. A Mask Region-based Convolutional Neural Network (Mask R-CNN) is used as the foundation for the ob- ject detection process. The relationships are found by manipulating the bounding box and mask segmentation outputs of a Mask R-CNN. The absolute positions, the positions of the objects relative to the image, and the relative positions, the positions of objects relative to the other objects, are then associated with the images as annotations that are out- put in order …
Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang
Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang
Statistical Science Theses and Dissertations
This dissertation contains two topics: (1) A Comparative Study of Statistical Methods for Quantifying and Testing Between-study Heterogeneity in Meta-analysis with Focus on Rare Binary Events; (2) Estimation of Variances in Cluster Randomized Designs Using Ranked Set Sampling.
Meta-analysis, the statistical procedure for combining results from multiple studies, has been widely used in medical research to evaluate intervention efficacy and safety. In many practical situations, the variation of treatment effects among the collected studies, often measured by the heterogeneity parameter, may exist and can greatly affect the inference about effect sizes. Comparative studies have been done for only one or …
Identifying Customer Churn In After-Market Operations Using Machine Learning Algorithms, Vitaly Briker, Richard Farrow, William Trevino, Brent Allen
Identifying Customer Churn In After-Market Operations Using Machine Learning Algorithms, Vitaly Briker, Richard Farrow, William Trevino, Brent Allen
SMU Data Science Review
This paper presents a comparative study on machine learning methods as they are applied to product associations, future purchase predictions, and predictions of customer churn in aftermarket operations. Association rules are used help to identify patterns across products and find correlations in customer purchase behaviour. Studying customer behaviour as it pertains to Recency, Frequency, and Monetary Value (RFM) helps inform customer segmentation and identifies customers with propensity to churn. Lastly, Flowserve’s customer purchase history enables the establishment of churn thresholds for each customer group and assists in constructing a model to predict future churners. The aim of this model is …
Personalized Detection Of Anxiety Provoking News Events Using Semantic Network Analysis, Jacquelyn Cheun Phd, Luay Dajani, Quentin B. Thomas
Personalized Detection Of Anxiety Provoking News Events Using Semantic Network Analysis, Jacquelyn Cheun Phd, Luay Dajani, Quentin B. Thomas
SMU Data Science Review
In the age of hyper-connectivity, 24/7 news cycles, and instant news alerts via social media, mental health researchers don't have a way to automatically detect news content which is associated with triggering anxiety or depression in mental health patients. Using the Associated Press news wire, a semantic network was built with 1,056 news articles containing over 500,000 connections across multiple topics to provide a personalized algorithm which detects problematic news content for a given reader. We make use of Semantic Network Analysis to surface the relationship between news article text and anxiety in readers who struggle with mental health disorders. …
Achieving Optimal Horizontal Drill Operations, Daniel J. Serna, James Vasquez, Donald Markley
Achieving Optimal Horizontal Drill Operations, Daniel J. Serna, James Vasquez, Donald Markley
SMU Data Science Review
In this paper, we present a novel method of predicting the onset of a slide event in horizontal drilling operations. Horizontal drilling operations attempt to create a well through a subsurface as quickly as possible by rotating a drill through the subsurface. A slide event occurs when the drill begins to inefficiently rotate through the subsurface, resulting in a significantly reduced rate of penetration. Slide events can be prevented, or significantly reduced in their impact, when their onset is accurately predicted. We present a method of accurately predicting the onset of slide events with a time-series based predictive model that …
Predicting Wind Turbine Blade Erosion Using Machine Learning, Casey Martinez, Festus Asare Yeboah, Scott Herford, Matt Brzezinski, Viswanath Puttagunta
Predicting Wind Turbine Blade Erosion Using Machine Learning, Casey Martinez, Festus Asare Yeboah, Scott Herford, Matt Brzezinski, Viswanath Puttagunta
SMU Data Science Review
Using time-series data and turbine blade inspection assessments, we present a classification model in order to predict remaining turbine blade life in wind turbines. Capturing the kinetic energy of wind requires complex mechanical systems, which require sophisticated maintenance and planning strategies. There are many traditional approaches to monitoring the internal gearbox and generator, but the condition of turbine blades can be difficult to measure and access. Accurate and cost- effective estimates of turbine blade life cycles will drive optimal investments in repairs and improve overall performance. These measures will drive down costs as well as provide cheap and clean electricity …
Machine Learning In Support Of Electric Distribution Asset Failure Prediction, Robert D. Flamenbaum, Thomas Pompo, Christopher Havenstein, Jade Thiemsuwan
Machine Learning In Support Of Electric Distribution Asset Failure Prediction, Robert D. Flamenbaum, Thomas Pompo, Christopher Havenstein, Jade Thiemsuwan
SMU Data Science Review
In this paper, we present novel approaches to predicting as- set failure in the electric distribution system. Failures in overhead power lines and their associated equipment in particular, pose significant finan- cial and environmental threats to electric utilities. Electric device failure furthermore poses a burden on customers and can pose serious risk to life and livelihood. Working with asset data acquired from an electric utility in Southern California, and incorporating environmental and geospatial data from around the region, we applied a Random Forest methodology to predict which overhead distribution lines are most vulnerable to fail- ure. Our results provide evidence …
Identifying Undervalued Players In Fantasy Football, Christopher D. Morgan, Caroll Rodriguez, Korey Macvittie, Robert Slater, Daniel W. Engels
Identifying Undervalued Players In Fantasy Football, Christopher D. Morgan, Caroll Rodriguez, Korey Macvittie, Robert Slater, Daniel W. Engels
SMU Data Science Review
In this paper we present a model to predict player performance in fantasy football. In particular, identifying high-performance players can prove to be a difficult problem, as there are on occasion players capable of high performance whose past metrics give no indication of this capacity. These "sleepers"' are often undervalued, and the acquisition of such players can have notable impact on a fantasy football team's overall performance. We constructed a regression model that accounts for players' past performance and athletic metrics to predict their future performance. The model we built performs favorably in predicting athlete performance in relation to other …
Machine Learning Predicts Aperiodic Laboratory Earthquakes, Olha Tanyuk, Daniel Davieau, Charles South, Daniel W. Engels
Machine Learning Predicts Aperiodic Laboratory Earthquakes, Olha Tanyuk, Daniel Davieau, Charles South, Daniel W. Engels
SMU Data Science Review
In this paper we find a pattern of aperiodic seismic signals that precede earthquakes at any time in a laboratory earthquake’s cycle using a small window of time. We use a data set that comes from a classic laboratory experiment having several stick-slip displacements (earthquakes), a type of experiment which has been studied as a simulation of seismologic faults for decades. This data exhibits similar behavior to natural earthquakes, so the same approach may work in predicting the timing of them. Here we show that by applying random forest machine learning technique to the acoustic signal emitted by a laboratory …
Longitudinal Analysis With Modes Of Operation For Aes, Dana Geislinger, Cory Thigpen, Daniel W. Engels
Longitudinal Analysis With Modes Of Operation For Aes, Dana Geislinger, Cory Thigpen, Daniel W. Engels
SMU Data Science Review
In this paper, we present an empirical evaluation of the randomness of the ciphertext blocks generated by the Advanced Encryption Standard (AES) cipher in Counter (CTR) mode and in Cipher Block Chaining (CBC) mode. Vulnerabilities have been found in the AES cipher that may lead to a reduction in the randomness of the generated ciphertext blocks that can result in a practical attack on the cipher. We evaluate the randomness of the AES ciphertext using the standard key length and NIST randomness tests. We evaluate the randomness through a longitudinal analysis on 200 billion ciphertext blocks using logistic regression and …
Sample Size Calculation Of Clinical Trials With Correlated Outcomes, Dateng Li
Sample Size Calculation Of Clinical Trials With Correlated Outcomes, Dateng Li
Statistical Science Theses and Dissertations
In this thesis, we investigate sample size calculation for three kinds of clinical trials: (1). Randomized controlled trials (RCTs) with longitudinal count outcomes; (2). Cluster randomized trials (CRTs) with count outcomes; (3). CRTs with multiple binary co-primary endpoints.
Advances In Measurement Error Modeling, Linh Nghiem
Advances In Measurement Error Modeling, Linh Nghiem
Statistical Science Theses and Dissertations
Measurement error in observations is widely known to cause bias and a loss of power when fitting statistical models, particularly when studying distribution shape or the relationship between an outcome and a variable of interest. Most existing correction methods in the literature require strong assumptions about the distribution of the measurement error, or rely on ancillary data which is not always available. This limits the applicability of these methods in many situations. Furthermore, new correction approaches are also needed for high-dimensional settings, where the presence of measurement error in the covariates adds another level of complexity to the desirable structure …
Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams
Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams
Statistical Science Theses and Dissertations
Much recent research has focused on methods for combining a probability sample with a non-probability sample to improve estimation by making use of information from both sources. If units exist in both samples, it becomes necessary to link the information from the two samples for these units. Record linkage is a technique to link records from two lists that refer to the same unit but lack a unique identifier across both lists. Record linkage assigns a probability to each potential pair of records from the lists so that principled matching decisions can be made. Because record linkage is a probabilistic …
Self-Driving Cars: Evaluation Of Deep Learning Techniques For Object Detection In Different Driving Conditions, Ramesh Simhambhatla, Kevin Okiah, Shravan Kuchkula, Robert Slater
Self-Driving Cars: Evaluation Of Deep Learning Techniques For Object Detection In Different Driving Conditions, Ramesh Simhambhatla, Kevin Okiah, Shravan Kuchkula, Robert Slater
SMU Data Science Review
Deep Learning has revolutionized Computer Vision, and it is the core technology behind capabilities of a self-driving car. Convolutional Neural Networks (CNNs) are at the heart of this deep learning revolution for improving the task of object detection. A number of successful object detection systems have been proposed in recent years that are based on CNNs. In this paper, an empirical evaluation of three recent meta-architectures: SSD (Single Shot multi-box Detector), R-CNN (Region-based CNN) and R-FCN (Region-based Fully Convolutional Networks) was conducted to measure how fast and accurate they are in identifying objects on the road, such as vehicles, pedestrians, …
Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, Kj Price, Rosalie Multari
Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, Kj Price, Rosalie Multari
SMU Data Science Review
The challenges of learning a new language can be reduced with real-time feedback on pronunciation and language usage. Today there are readily available technologies which provide such feedback on spoken languages, by translating the voice of the learner into written text. For someone seeking to learn American Sign Language (ASL), there is however no such feedback application available. A learner of American Sign Language might reference websites or books to obtain an image of a hand sign for a word. This process is like looking up a word in a dictionary, and if the person wanted to know if they …
Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith
Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith
SMU Data Science Review
In this paper, we compare demand forecasting methods used by the supply chain department at Bilports to open-source forecasting methods. The design and implementation of the open-source forecasting system also attempts to use several external datasets such as consumer sentiment, housing permit starts, and weather to improve prediction quality. Additionally, the performance of the forecast is evaluated by the reduction of shipment lead times from China, the company’s primary vendor. The objective of our paper is to improve Bilports’s forecasting capabilities. The primary motivation of this paper is to increase forecasting accuracy and identify the weaknesses of the methods used …
Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler
Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler
SMU Data Science Review
In this paper, we present an analysis of survey data with the goal of determining if the KadAfrica training program, a social organization in Uganda, has a significant effect on the lives of the girls who participate in the program. This is done through an observational study of girl’s responses to several pre-program and post-program questions. These questions include topics such as the girl’s access to hygiene materials and their personal views on family finances. In addition to providing an analysis of historical data, we established a data platform in which future data can be stored and analyzed in an …
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
SMU Data Science Review
In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …
Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater
Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater
SMU Data Science Review
Satellite Imagery is one of the most widely used sources to analyze geographic features and environments in the world. The data gathered from satellites are used to quantify many vital problems facing our society, such as the impact of natural disasters, shore erosion, rising water levels, and urban growth rates. In this paper, we construct machine learning and deep learning algorithms for repairing anomalies in the Landsat satellite imagery data which arise for various reasons ranging from cloud obstruction to satellite malfunctions. The accuracy of GIS data is crucial to ensuring the models produced from such data are as close …
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
SMU Data Science Review
In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
SMU Data Science Review
Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …
Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros
Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros
SMU Data Science Review
Abstract. Current techniques for calculating and generating models used for analyzing the Earth’s magnetic field are laborious and time-consuming. We assert that machine learning can have a significant impact on building magnetic field models more quickly and on various levels of complexity, specifically as it pertains to data cleansing and sorting. Our approach to this problem uses a reverse iterative multi-phase process for data cleansing, in which, initially, the CHAOS-6 model data is examined to determine if machine learning can be used to differentiate between useful data components for spherical harmonics, versus data noise. During this phase, six different machine …
Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock
Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock
SMU Data Science Review
Through microblogging applications, such as Twitter, people actively document their lives even in times of natural disasters such as hurricanes and earthquakes. While first responders and crisis-teams are able to help people who call 911, or arrive at a designated shelter, there are vast amounts of information being exchanged online via Twitter that provide real-time, location-based alerts that are going unnoticed. To effectively use this information, the Tweets must be verified for authenticity and categorized to ensure that the proper authorities can be alerted. In this paper, we create a Crisis Message Corpus from geotagged Tweets occurring during 7 hurricanes …