Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,823 Full-Text Articles 23,921 Authors 9,922,835 Downloads 283 Institutions

All Articles in Statistics and Probability

Faceted Search

12,823 full-text articles. Page 241 of 487.

Twitter And Disasters: A Social Resilience Fingerprint, Benjamin A. Rachunok, Jackson B. Bennett, Roshanak Nateghi 2019 Purdue University

Twitter And Disasters: A Social Resilience Fingerprint, Benjamin A. Rachunok, Jackson B. Bennett, Roshanak Nateghi

Purdue University Libraries Open Access Publishing Fund

Understanding the resilience of a community facing a crisis event is critical to improving its adaptive capacity. Community resilience has been conceptualized as a function of the resilience of components of a community such as ecological, infrastructure, economic, and social systems, etc. In this paper, we introduce the concept of a “resilience fingerprint” and propose a multi-dimensional method for analyzing components of community resilience by leveraging existing definitions of community resilience with data from the social network Twitter. Twitter data from 14 events are analyzed and their resulting resilience fingerprints computed. We compare the fingerprints between events and show that …


Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, KJ Price, Rosalie Multari 2019 Southern Methodist University

Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, Kj Price, Rosalie Multari

SMU Data Science Review

The challenges of learning a new language can be reduced with real-time feedback on pronunciation and language usage. Today there are readily available technologies which provide such feedback on spoken languages, by translating the voice of the learner into written text. For someone seeking to learn American Sign Language (ASL), there is however no such feedback application available. A learner of American Sign Language might reference websites or books to obtain an image of a hand sign for a word. This process is like looking up a word in a dictionary, and if the person wanted to know if they …


What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren 2019 University of Northern Colorado

What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren

The Qualitative Report

Statistical and research consulting is defined as the collaboration of a statistician or methodologist with another professional for devising solutions to research problems. An in-depth, interview qualitative approach was taken to answer the research question of what makes a good research consultant. The authors interviewed four faculty members in the field of statistics and research methods and two experienced graduate student consultants. In-depth, face-to-face interviews revealed common themes regarding consultancy skills, resourcefulness, communication and interpersonal skills. The participants discussed how to improve consulting sessions and deal with clients with different statistics levels and backgrounds. Participants felt there was no difference …


Combining Chicken Retina Rna-Seq Data Across Studies To Strengthen Biomarker Detection, Sarah Szvetecz 2019 James Madison University

Combining Chicken Retina Rna-Seq Data Across Studies To Strengthen Biomarker Detection, Sarah Szvetecz

Senior Honors Projects, 2010-2019

Various studies have identified the chicken embryo (Gallus gallus) as a useful model to study the retinogenesis process in humans. This project uses data from two specific RNA sequencing (RNA-seq) studies to investigate retina developmental biology. These studies are done in two different labs using different protocols, as such they cannot be compared directly. Study 1 contains chicken retina samples from embryonic day 3, 5 and 8; while study 2 has retina samples from embryonic day 8, 16, and 18 of developmental age. We apply a normalization method on both studies to account for differences in the two …


The Reproducibility Crisis In Scientific Research, Sarah Eline 2019 James Madison University

The Reproducibility Crisis In Scientific Research, Sarah Eline

Senior Honors Projects, 2010-2019

Following the push for evidence based practice, came a huge proliferation of research journals and journal articles. With this increase in quantity came an increased concern about the quality of these articles being published, which led to a multifield investigation regarding the reproducibility of scientific research. With studies in the fields of psychology and biomedicine only reaching approximately a 30% reproducibility rate, a conversation has been sparked that spans across every field of research. Upon further investigation, various causes for this reproducibility crisis have surfaced which include, lack of data sharing/ transparency, statistical errors, funding corruption, and the culture surrounding …


Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith 2019 Southern Methodist University

Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith

SMU Data Science Review

In this paper, we compare demand forecasting methods used by the supply chain department at Bilports to open-source forecasting methods. The design and implementation of the open-source forecasting system also attempts to use several external datasets such as consumer sentiment, housing permit starts, and weather to improve prediction quality. Additionally, the performance of the forecast is evaluated by the reduction of shipment lead times from China, the company’s primary vendor. The objective of our paper is to improve Bilports’s forecasting capabilities. The primary motivation of this paper is to increase forecasting accuracy and identify the weaknesses of the methods used …


Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler 2019 Southern Methodist University

Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler

SMU Data Science Review

In this paper, we present an analysis of survey data with the goal of determining if the KadAfrica training program, a social organization in Uganda, has a significant effect on the lives of the girls who participate in the program. This is done through an observational study of girl’s responses to several pre-program and post-program questions. These questions include topics such as the girl’s access to hygiene materials and their personal views on family finances. In addition to providing an analysis of historical data, we established a data platform in which future data can be stored and analyzed in an …


Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley 2019 Southern Methodist University

Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley

SMU Data Science Review

In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …


Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater 2019 SMU

Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater

SMU Data Science Review

Satellite Imagery is one of the most widely used sources to analyze geographic features and environments in the world. The data gathered from satellites are used to quantify many vital problems facing our society, such as the impact of natural disasters, shore erosion, rising water levels, and urban growth rates. In this paper, we construct machine learning and deep learning algorithms for repairing anomalies in the Landsat satellite imagery data which arise for various reasons ranging from cloud obstruction to satellite malfunctions. The accuracy of GIS data is crucial to ensuring the models produced from such data are as close …


Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza III, Jose Quinonez, Misael Santana, Nibhrat Lohia 2019 Southern Methodist University

Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia

SMU Data Science Review

In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …


Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi 2019 Southern Methodist University

Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi

SMU Data Science Review

Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …


Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros 2019 Southern Methodist University

Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros

SMU Data Science Review

Abstract. Current techniques for calculating and generating models used for analyzing the Earth’s magnetic field are laborious and time-consuming. We assert that machine learning can have a significant impact on building magnetic field models more quickly and on various levels of complexity, specifically as it pertains to data cleansing and sorting. Our approach to this problem uses a reverse iterative multi-phase process for data cleansing, in which, initially, the CHAOS-6 model data is examined to determine if machine learning can be used to differentiate between useful data components for spherical harmonics, versus data noise. During this phase, six different machine …


Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock 2019 Southern Methodist University

Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock

SMU Data Science Review

Through microblogging applications, such as Twitter, people actively document their lives even in times of natural disasters such as hurricanes and earthquakes. While first responders and crisis-teams are able to help people who call 911, or arrive at a designated shelter, there are vast amounts of information being exchanged online via Twitter that provide real-time, location-based alerts that are going unnoticed. To effectively use this information, the Tweets must be verified for authenticity and categorized to ensure that the proper authorities can be alerted. In this paper, we create a Crisis Message Corpus from geotagged Tweets occurring during 7 hurricanes …


Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam, David Churchman, Abigail Morton Garland 2019 Southern Methodist University

Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam, David Churchman, Abigail Morton Garland

SMU Data Science Review

The state of Texas requires all public high school students to take End of Course (EOC) exams. The results of these exams are made nominally public, but in a shape and format that precludes ready analysis. To the extent possible, principles of tidy data will be applied to clean and analyze the publicly released data file for the 2014 English II EOC exam, providing insights into the EOC program and a case for better public data from the Texas Education Administration (TEA).


A Mathematical Investigation On Tumor-Immune Dynamics: The Impact Of Vaccines On The Immune Response, Jonathan Quinonez, Neethi Dasu, Mahboobi Qureshi 2019 Larkin Hospital (Florida)

A Mathematical Investigation On Tumor-Immune Dynamics: The Impact Of Vaccines On The Immune Response, Jonathan Quinonez, Neethi Dasu, Mahboobi Qureshi

Rowan-Virtua Research Day

Mathematical models analyzing tumor-immune interactions provide a framework by which to address specific scenarios in regard to tumor-immune dynamics. Important aspects of tumor-immune surveillance to consider is the elimination of tumor cells from a host’s cell-mediated immunity as well as the implications of vaccines derived from synthetic antigen. In present studies, our mathematical model examined the role of synthetic antigen to the strength of the immune system. The constructed model takes into account accepted knowledge of immune function as well as prior work done by de Pillis et al. All equations describing tumor-immune growth, antigen presentation, immune response, and interaction …


Development Of A School Boredom Proneness Scale For Children, Taylor Carrington 2019 James Madison University

Development Of A School Boredom Proneness Scale For Children, Taylor Carrington

Educational Specialist, 2009-2019

One common phrase heard from students is, “I’m bored.” However, there is no real understanding of what this actually means. In this study, elementary-age students were asked to respond to a newly developed School Boredom Proneness Scale (SBPS) including questions relating to a five-factor model of boredom. Students were also asked to rate how often they become bored at school and how bored they seem compared to classmates. In addition to student responses, parents and teachers were asked to rate how bored they thought the student was, and teachers were additionally asked to rate students’ level of work completion. The …


Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang 2019 University of Rhode Island

Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang

Senior Honors Projects

Inspired by the landmark paper “Emergence of Scaling in Random Networks” by Barabási and Albert, the field of network science has focused heavily on the power law distribution in recent years. This distribution has been used to model everything from the popularity of sites on the World Wide Web to the number of citations received on a scientific paper. The feature of this distribution is highlighted by the fact that many nodes (websites or papers) have few connections (internet links or citations) while few “hubs” are connected to many nodes. These properties lead to two very important observed effects: the …


Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark 2019 Utah State University

Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Filtered historical simulation with an underlying GARCH process can be used as a valuable tool in VaR analysis, as it derives risk estimates that are sensitive to the distributional properties of the historical data of the produced predictive density. I examine the applications to risk analysis that filtered historical simulation can provide, as well as an interpretation of the predictive density as a poor man’s Bayesian posterior distribution. The predictive density allows us to make associated probabilistic statements regarding the results for VaR analysis, giving greater measurement of risk and the ability to maintain the optimal level of risk per …


Dynamic Attribute-Level Best Worst Discrete Choice Experiments, Amanda Working, Mohammed Alqawba, Norou Diawara 2019 Old Dominion University

Dynamic Attribute-Level Best Worst Discrete Choice Experiments, Amanda Working, Mohammed Alqawba, Norou Diawara

Mathematics & Statistics Faculty Publications

Dynamic modelling of decision maker choice behavior of best and worst in discrete choice experiments (DCEs) has numerous applications. Such models are proposed under utility function of decision maker and are used in many areas including social sciences, health economics, transportation research, and health systems research. After reviewing references on the study of such experiments, we present example in DCE with emphasis on time dependent best-worst choice and discrimination between choice attributes. Numerical examples of the dynamic DCEs are simulated, and the associated expected utilities over time of the choice models are derived using Markov decision processes. The estimates are …


Simulation As A Predictor In Probability, Xiaona Zhou 2019 CUNY New York City College of Technology

Simulation As A Predictor In Probability, Xiaona Zhou

Publications and Research

In this study, we simulate bivariate normal data. We gain intuition about the bivariate normal distribution by comparing the generated data to the associated bivariate normal density surface. We also get results about covariance and correlation. We will use tools from linear algebra to discuss transformations of random normal vectors, and the use of contours.


Digital Commons powered by bepress