Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 4801 - 4830 of 12823

Full-Text Articles in Statistics and Probability

Twitter And Disasters: A Social Resilience Fingerprint, Benjamin A. Rachunok, Jackson B. Bennett, Roshanak Nateghi May 2019

Twitter And Disasters: A Social Resilience Fingerprint, Benjamin A. Rachunok, Jackson B. Bennett, Roshanak Nateghi

Purdue University Libraries Open Access Publishing Fund

Understanding the resilience of a community facing a crisis event is critical to improving its adaptive capacity. Community resilience has been conceptualized as a function of the resilience of components of a community such as ecological, infrastructure, economic, and social systems, etc. In this paper, we introduce the concept of a “resilience fingerprint” and propose a multi-dimensional method for analyzing components of community resilience by leveraging existing definitions of community resilience with data from the social network Twitter. Twitter data from 14 events are analyzed and their resulting resilience fingerprints computed. We compare the fingerprints between events and show that …


Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, Kj Price, Rosalie Multari May 2019

Asl Reverse Dictionary - Asl Translation Using Deep Learning, Ann Nelson, Kj Price, Rosalie Multari

SMU Data Science Review

The challenges of learning a new language can be reduced with real-time feedback on pronunciation and language usage. Today there are readily available technologies which provide such feedback on spoken languages, by translating the voice of the learner into written text. For someone seeking to learn American Sign Language (ASL), there is however no such feedback application available. A learner of American Sign Language might reference websites or books to obtain an image of a hand sign for a word. This process is like looking up a word in a dictionary, and if the person wanted to know if they …


What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren May 2019

What Makes A Good Research Consultant?, Justin Harding, Samantha Estrada, Michael Floren

The Qualitative Report

Statistical and research consulting is defined as the collaboration of a statistician or methodologist with another professional for devising solutions to research problems. An in-depth, interview qualitative approach was taken to answer the research question of what makes a good research consultant. The authors interviewed four faculty members in the field of statistics and research methods and two experienced graduate student consultants. In-depth, face-to-face interviews revealed common themes regarding consultancy skills, resourcefulness, communication and interpersonal skills. The participants discussed how to improve consulting sessions and deal with clients with different statistics levels and backgrounds. Participants felt there was no difference …


Combining Chicken Retina Rna-Seq Data Across Studies To Strengthen Biomarker Detection, Sarah Szvetecz May 2019

Combining Chicken Retina Rna-Seq Data Across Studies To Strengthen Biomarker Detection, Sarah Szvetecz

Senior Honors Projects, 2010-2019

Various studies have identified the chicken embryo (Gallus gallus) as a useful model to study the retinogenesis process in humans. This project uses data from two specific RNA sequencing (RNA-seq) studies to investigate retina developmental biology. These studies are done in two different labs using different protocols, as such they cannot be compared directly. Study 1 contains chicken retina samples from embryonic day 3, 5 and 8; while study 2 has retina samples from embryonic day 8, 16, and 18 of developmental age. We apply a normalization method on both studies to account for differences in the two …


The Reproducibility Crisis In Scientific Research, Sarah Eline May 2019

The Reproducibility Crisis In Scientific Research, Sarah Eline

Senior Honors Projects, 2010-2019

Following the push for evidence based practice, came a huge proliferation of research journals and journal articles. With this increase in quantity came an increased concern about the quality of these articles being published, which led to a multifield investigation regarding the reproducibility of scientific research. With studies in the fields of psychology and biomedicine only reaching approximately a 30% reproducibility rate, a conversation has been sparked that spans across every field of research. Upon further investigation, various causes for this reproducibility crisis have surfaced which include, lack of data sharing/ transparency, statistical errors, funding corruption, and the culture surrounding …


Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith May 2019

Demand Forecasting: An Open-Source Approach, Murtada Shubbar, Jared Smith

SMU Data Science Review

In this paper, we compare demand forecasting methods used by the supply chain department at Bilports to open-source forecasting methods. The design and implementation of the open-source forecasting system also attempts to use several external datasets such as consumer sentiment, housing permit starts, and weather to improve prediction quality. Additionally, the performance of the forecast is evaluated by the reduction of shipment lead times from China, the company’s primary vendor. The objective of our paper is to improve Bilports’s forecasting capabilities. The primary motivation of this paper is to increase forecasting accuracy and identify the weaknesses of the methods used …


Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler May 2019

Kadafrica: Survey Analysis To Support Research For Smallholder Farmers, Gregory Asamoah, Robert Gill, Frank Sclafani, Bivin Sadler

SMU Data Science Review

In this paper, we present an analysis of survey data with the goal of determining if the KadAfrica training program, a social organization in Uganda, has a significant effect on the lives of the girls who participate in the program. This is done through an observational study of girl’s responses to several pre-program and post-program questions. These questions include topics such as the girl’s access to hygiene materials and their personal views on family finances. In addition to providing an analysis of historical data, we established a data platform in which future data can be stored and analyzed in an …


Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley May 2019

Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley

SMU Data Science Review

In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …


Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater May 2019

Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater

SMU Data Science Review

Satellite Imagery is one of the most widely used sources to analyze geographic features and environments in the world. The data gathered from satellites are used to quantify many vital problems facing our society, such as the impact of natural disasters, shore erosion, rising water levels, and urban growth rates. In this paper, we construct machine learning and deep learning algorithms for repairing anomalies in the Landsat satellite imagery data which arise for various reasons ranging from cloud obstruction to satellite malfunctions. The accuracy of GIS data is crucial to ensuring the models produced from such data are as close …


Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia May 2019

Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia

SMU Data Science Review

In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …


Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi May 2019

Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi

SMU Data Science Review

Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …


Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros May 2019

Machine Learning Vs Conventional Analysis Techniques For The Earth’S Magnetic Field Study, Sheri Loftin, Sarah J. Fite, Laura V. Bishop, Stavros Kotsiaros

SMU Data Science Review

Abstract. Current techniques for calculating and generating models used for analyzing the Earth’s magnetic field are laborious and time-consuming. We assert that machine learning can have a significant impact on building magnetic field models more quickly and on various levels of complexity, specifically as it pertains to data cleansing and sorting. Our approach to this problem uses a reverse iterative multi-phase process for data cleansing, in which, initially, the CHAOS-6 model data is examined to determine if machine learning can be used to differentiate between useful data components for spherical harmonics, versus data noise. During this phase, six different machine …


Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock May 2019

Leveraging Natural Language Processing Applications And Microblogging Platform For Increased Transparency In Crisis Areas, Ernesto Carrera-Ruvalcaba, Johnson Ekedum, Austin Hancock, Ben Brock

SMU Data Science Review

Through microblogging applications, such as Twitter, people actively document their lives even in times of natural disasters such as hurricanes and earthquakes. While first responders and crisis-teams are able to help people who call 911, or arrive at a designated shelter, there are vast amounts of information being exchanged online via Twitter that provide real-time, location-based alerts that are going unnoticed. To effectively use this information, the Tweets must be verified for authenticity and categorized to ensure that the proper authorities can be alerted. In this paper, we create a Crisis Message Corpus from geotagged Tweets occurring during 7 hurricanes …


Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam, David Churchman, Abigail Morton Garland May 2019

Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam, David Churchman, Abigail Morton Garland

SMU Data Science Review

The state of Texas requires all public high school students to take End of Course (EOC) exams. The results of these exams are made nominally public, but in a shape and format that precludes ready analysis. To the extent possible, principles of tidy data will be applied to clean and analyze the publicly released data file for the 2014 English II EOC exam, providing insights into the EOC program and a case for better public data from the Texas Education Administration (TEA).


A Mathematical Investigation On Tumor-Immune Dynamics: The Impact Of Vaccines On The Immune Response, Jonathan Quinonez, Neethi Dasu, Mahboobi Qureshi May 2019

A Mathematical Investigation On Tumor-Immune Dynamics: The Impact Of Vaccines On The Immune Response, Jonathan Quinonez, Neethi Dasu, Mahboobi Qureshi

Rowan-Virtua Research Day

Mathematical models analyzing tumor-immune interactions provide a framework by which to address specific scenarios in regard to tumor-immune dynamics. Important aspects of tumor-immune surveillance to consider is the elimination of tumor cells from a host’s cell-mediated immunity as well as the implications of vaccines derived from synthetic antigen. In present studies, our mathematical model examined the role of synthetic antigen to the strength of the immune system. The constructed model takes into account accepted knowledge of immune function as well as prior work done by de Pillis et al. All equations describing tumor-immune growth, antigen presentation, immune response, and interaction …


Development Of A School Boredom Proneness Scale For Children, Taylor Carrington May 2019

Development Of A School Boredom Proneness Scale For Children, Taylor Carrington

Educational Specialist, 2009-2019

One common phrase heard from students is, “I’m bored.” However, there is no real understanding of what this actually means. In this study, elementary-age students were asked to respond to a newly developed School Boredom Proneness Scale (SBPS) including questions relating to a five-factor model of boredom. Students were also asked to rate how often they become bored at school and how bored they seem compared to classmates. In addition to student responses, parents and teachers were asked to rate how bored they thought the student was, and teachers were additionally asked to rate students’ level of work completion. The …


Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang May 2019

Do Metabolic Networks Follow A Power Law? A Psamm Analysis, Ryan Geib, Lubos Thoma, Ying Zhang

Senior Honors Projects

Inspired by the landmark paper “Emergence of Scaling in Random Networks” by Barabási and Albert, the field of network science has focused heavily on the power law distribution in recent years. This distribution has been used to model everything from the popularity of sites on the World Wide Web to the number of citations received on a scientific paper. The feature of this distribution is highlighted by the fact that many nodes (websites or papers) have few connections (internet links or citations) while few “hubs” are connected to many nodes. These properties lead to two very important observed effects: the …


Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark May 2019

Predictive Distributions Via Filtered Historical Simulation For Financial Risk Management, Tyson Clark

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Filtered historical simulation with an underlying GARCH process can be used as a valuable tool in VaR analysis, as it derives risk estimates that are sensitive to the distributional properties of the historical data of the produced predictive density. I examine the applications to risk analysis that filtered historical simulation can provide, as well as an interpretation of the predictive density as a poor man’s Bayesian posterior distribution. The predictive density allows us to make associated probabilistic statements regarding the results for VaR analysis, giving greater measurement of risk and the ability to maintain the optimal level of risk per …


Dynamic Attribute-Level Best Worst Discrete Choice Experiments, Amanda Working, Mohammed Alqawba, Norou Diawara May 2019

Dynamic Attribute-Level Best Worst Discrete Choice Experiments, Amanda Working, Mohammed Alqawba, Norou Diawara

Mathematics & Statistics Faculty Publications

Dynamic modelling of decision maker choice behavior of best and worst in discrete choice experiments (DCEs) has numerous applications. Such models are proposed under utility function of decision maker and are used in many areas including social sciences, health economics, transportation research, and health systems research. After reviewing references on the study of such experiments, we present example in DCE with emphasis on time dependent best-worst choice and discrimination between choice attributes. Numerical examples of the dynamic DCEs are simulated, and the associated expected utilities over time of the choice models are derived using Markov decision processes. The estimates are …


Simulation As A Predictor In Probability, Xiaona Zhou May 2019

Simulation As A Predictor In Probability, Xiaona Zhou

Publications and Research

In this study, we simulate bivariate normal data. We gain intuition about the bivariate normal distribution by comparing the generated data to the associated bivariate normal density surface. We also get results about covariance and correlation. We will use tools from linear algebra to discuss transformations of random normal vectors, and the use of contours.


Tdp-43 Proteinopathy In Aging: Associations With Risk-Associated Gene Variants And With Brain Parenchymal Thyroid Hormone Levels, Peter T. Nelson, Zsombor Gal, Wang-Xia Wang, Dana M. Niedowicz, Sergey C. Artiushin, Samuel Wycoff, Angela Wei, Gregory A. Jicha, David W. Fardo May 2019

Tdp-43 Proteinopathy In Aging: Associations With Risk-Associated Gene Variants And With Brain Parenchymal Thyroid Hormone Levels, Peter T. Nelson, Zsombor Gal, Wang-Xia Wang, Dana M. Niedowicz, Sergey C. Artiushin, Samuel Wycoff, Angela Wei, Gregory A. Jicha, David W. Fardo

Pathology and Laboratory Medicine Faculty Publications

TDP-43 proteinopathy is very prevalent among the elderly (affecting at least 25% of individuals over 85 years of age) and is associated with substantial cognitive impairment. Risk factors implicated in age-related TDP-43 proteinopathy include commonly inherited gene variants, comorbid Alzheimer's disease pathology, and thyroid hormone dysfunction. To test parameters that are associated with aging-related TDP-43 pathology, we performed exploratory analyses of pathologic, genetic, and biochemical data derived from research volunteers in the University of Kentucky Alzheimer's Disease Center autopsy cohort (n = 136 subjects). Digital pathologic methods were used to discriminate and quantify both neuritic and intracytoplasmic TDP-43 pathology …


Understanding Water Consumption And Energy Trends In New York City, Wen Yong Huang, Johann Thiel May 2019

Understanding Water Consumption And Energy Trends In New York City, Wen Yong Huang, Johann Thiel

Publications and Research

In this study, we will be using the NYC Open Data website to examine publicly available data sets on water and energy consumption in New York City. In particular, we will use various scientific programming and machine learning modules in Python to analyze and visualize trends in water and energy usage within the five boroughs.


Sampling Studies For Longitudinal Functional Data, Toni Jassel May 2019

Sampling Studies For Longitudinal Functional Data, Toni Jassel

Theses, Dissertations and Culminating Projects

We study the data setting consisting of functional data sets repeatedly observed over time. The focus is on the dynamic prediction of the future trajectory for a subject. Regression methods based on dynamic functional models are used for dynamic prediction of individual trajectories. We propose strategies for the selection of the study sampling design in the context of longitudinal functional data. An application to simulated child growth data is presented. The height-for-age z-score (HAZ) was the response variable in the functional dynamic models for prediction. The intent was to recommend four months for removal in our initial historic data set. …


Statistical Modeling Of Count Data With Over-Dispersion Or Zero-Inflation Problems, Chengxin Zhang May 2019

Statistical Modeling Of Count Data With Over-Dispersion Or Zero-Inflation Problems, Chengxin Zhang

Theses, Dissertations and Culminating Projects

In this study, we will analyze a supply retailing company’s data to model the relationship between their customer’s past purchase behavior to predict their future online purchase behavior. The data was divided into time periods from 2016: P1-P6(January 31st to July 30th) and P7(July 31st to August 27th ). Based on customer’s past purchase information from the P1-P6 period, such as money spent, number of cart additions, transactions type, number of unique purchase dates, number of unique purchase skus, number of page views, number browse dates, company size, and number of products purchased, we aim to find if these information …


Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett May 2019

Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett

Statistical and Data Sciences: Faculty Publications

This study considered the influence of misperceptions of typical versus self-identified important peers' heavy drinking on personal heavy drinking intentions and frequency utilizing data from a complete social network of college students. The study sample included data from 1,313 students (44% male, 57% White, 15% Hispanic/Latinx) collected during the fall and spring semesters of their freshman year. Students provided perceived heavy drinking frequency for a typical student peer and up to 10 identified important peers. Personal past-month heavy drinking frequency was assessed for all participants at both time points. By comparing actual with perceived heavy drinking frequencies, measures of misperceptions …


What Can We Do? Puzzling Over The Interpretation Of Heredity And Variation From Galton To Genetic Engineering, Peter J. Taylor May 2019

What Can We Do? Puzzling Over The Interpretation Of Heredity And Variation From Galton To Genetic Engineering, Peter J. Taylor

Working Papers on Science in a Changing World

First six chapters of a book motivated as follows: When I had mentioned to colleagues that I was exploring some significant issues overlooked by both sides in nature-nurture debates, the typical response was “we know, of course, that nature and nurture are intertwined”; they never asked “which nature-nurture science are you referring to?” It occurred to me that, in the long history of nature-nurture debates, opposing sides had always assumed or implied that these different scientific approaches were speaking to the same issues. If that were the case, then the challenge—something I was already puzzling over—was how best to draw …


A Bayesian Framework For Estimating Seismic Wave Arrival Time, Hua Zhong May 2019

A Bayesian Framework For Estimating Seismic Wave Arrival Time, Hua Zhong

Graduate Theses and Dissertations

Because earthquakes have a large impact on human society, statistical methods for better studying earthquakes are required. One characteristic of earthquakes is the arrival time of seismic waves at a seismic signal sensor. Once we can estimate the earthquake arrival time accurately, the earthquake location can be triangulated, and assistance can be sent to that area correctly. This study presents a Bayesian framework to predict the arrival time of seismic waves with associated uncertainty. We use a change point framework to model the different conditions before and after the seismic wave arrives. To evaluate the performance of the model, we …


Telomeres, Nutrition And Mortality: Risk Factors For The Rate Of Telomere Length Decline And The Associations Between Telomere Length, Nutrition And Mortality, Saruna Ghimire May 2019

Telomeres, Nutrition And Mortality: Risk Factors For The Rate Of Telomere Length Decline And The Associations Between Telomere Length, Nutrition And Mortality, Saruna Ghimire

UNLV Theses, Dissertations, Professional Papers, and Capstones

Introduction: Telomeres are nucleoprotein structures located at the ends of eukaryotic chromosomes, thought to protect the DNA from damage. As a person experiences stressors, harmful exposures, and other diseases throughout their life, telomeres are thought to become damaged and their length shortened, decreasing their ability to protect the DNA. Nutrition is an important aspect of healthy aging. Preservation of telomere length (TL) is thought to be one of the mechanisms by which good nutrition can delay or prevent the development of chronic disease and death. Recent evidence of preservation of TL with good nutrition is promising. Thus, the aim of …


K-Tuple Sampling From Partially Rank-Ordered Sets, Marvin Javier May 2019

K-Tuple Sampling From Partially Rank-Ordered Sets, Marvin Javier

UNLV Theses, Dissertations, Professional Papers, and Capstones

With the introduction of Ranked Set Sampling (RSS), McIntyre (1952) demonstrated that using ranking information to select units for measurement can lead to estimators with reduced variance when compared to their counterparts based on a simple random sample of the same size. This is done by selecting a set of units, and without direct measurement, ranking the units in the set before identifying one unit for measurement. This ranking of the units can be done through judgement ranking (such as visual assessment), or by using a correlated auxiliary variable.

In its original form, RSS does not allow for ties when …


Health Disparities Among Sexual And Gender Minorities, Jennifer Keeley May 2019

Health Disparities Among Sexual And Gender Minorities, Jennifer Keeley

UNLV Theses, Dissertations, Professional Papers, and Capstones

Decades of research has shown that sexual and gender minorities (SGMs) experience adverse health and mental health outcomes to a greater extent than their heterosexual peers. The need to better understand and eliminate health disparities in the SGM population was recognized by the National Institute on Minority Health and Health Disparities (NIMHD) at NIH. The Secretary of Health at the Department of Health and Human Services approved the designation of the SGM population as a health disparities population in 2016 and called for SGM studies to examine the health needs of the SGM population across SGM subgroups via large representative …