Leveraging Reviews To Improve User Experience,
2019
Southern Methodist University
Leveraging Reviews To Improve User Experience, Anthony Schams, Iram Bakhtiar, Cristina Stanley
SMU Data Science Review
In this paper, we will explore and present a method of finding characteristics of a restaurant using its reviews through machine learning algorithms. We begin by building models to predict the ratings of individual reviews using text and categorical features. This is to examine the efficacy of the algorithms to the task. Both XGBoost and logistic regression will be examined. With these models, our goal is then to identify key phrases in reviews that are correlated with positive and negative experience. Our analysis makes use of review data publicly made available by Yelp. Key bigrams extracted were non-specific to the …
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem,
2019
Southern Methodist University
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
SMU Data Science Review
In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …
Machine Learning Pipeline For Exoplanet Classification,
2019
Southern Methodist University
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
SMU Data Science Review
Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …
Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam,
2019
Southern Methodist University
Tidying And Analysis Of The 2014 Texas English Ii End-Of-Course Exam, David Churchman, Abigail Morton Garland
SMU Data Science Review
The state of Texas requires all public high school students to take End of Course (EOC) exams. The results of these exams are made nominally public, but in a shape and format that precludes ready analysis. To the extent possible, principles of tidy data will be applied to clean and analyze the publicly released data file for the 2014 English II EOC exam, providing insights into the EOC program and a case for better public data from the Texas Education Administration (TEA).
Sampling Studies For Longitudinal Functional Data,
2019
Montclair State University
Sampling Studies For Longitudinal Functional Data, Toni Jassel
Theses, Dissertations and Culminating Projects
We study the data setting consisting of functional data sets repeatedly observed over time. The focus is on the dynamic prediction of the future trajectory for a subject. Regression methods based on dynamic functional models are used for dynamic prediction of individual trajectories. We propose strategies for the selection of the study sampling design in the context of longitudinal functional data. An application to simulated child growth data is presented. The height-for-age z-score (HAZ) was the response variable in the functional dynamic models for prediction. The intent was to recommend four months for removal in our initial historic data set. …
Statistical Modeling Of Count Data With Over-Dispersion Or Zero-Inflation Problems,
2019
Montclair State University
Statistical Modeling Of Count Data With Over-Dispersion Or Zero-Inflation Problems, Chengxin Zhang
Theses, Dissertations and Culminating Projects
In this study, we will analyze a supply retailing company’s data to model the relationship between their customer’s past purchase behavior to predict their future online purchase behavior. The data was divided into time periods from 2016: P1-P6(January 31st to July 30th) and P7(July 31st to August 27th ). Based on customer’s past purchase information from the P1-P6 period, such as money spent, number of cart additions, transactions type, number of unique purchase dates, number of unique purchase skus, number of page views, number browse dates, company size, and number of products purchased, we aim to find if these information …
A Bayesian Framework For Estimating Seismic Wave Arrival Time,
2019
University of Arkansas, Fayetteville
A Bayesian Framework For Estimating Seismic Wave Arrival Time, Hua Zhong
Graduate Theses and Dissertations
Because earthquakes have a large impact on human society, statistical methods for better studying earthquakes are required. One characteristic of earthquakes is the arrival time of seismic waves at a seismic signal sensor. Once we can estimate the earthquake arrival time accurately, the earthquake location can be triangulated, and assistance can be sent to that area correctly. This study presents a Bayesian framework to predict the arrival time of seismic waves with associated uncertainty. We use a change point framework to model the different conditions before and after the seismic wave arrives. To evaluate the performance of the model, we …
Comparing Elo, Glicko, Irt, And Bayesian Irt Statistical Models For Educational And Gaming Data,
2019
University of Arkansas, Fayetteville
Comparing Elo, Glicko, Irt, And Bayesian Irt Statistical Models For Educational And Gaming Data, Breanna Morrison
Graduate Theses and Dissertations
Statistical models used for estimating skill or ability levels often vary by field, however their underlying mathematical models can be very similar. Differences in the underlying models can be due to the need to accommodate data with different underlying formats and structure. As the models from varying fields increase in complexity, their ability to be applied to different types of data may have the ability to increase. Models that are applied to educational or psychological data have advanced to accommodate a wide range of data formats, including increased estimation accuracy with sparsely populated data matrices. Conversely, the field of online …
Raman And Surface Enhanced Raman Spectroscopy For Forensic Analysis: Case Studies On The Identification Of Illicit Substances And Artist Pigments,
2019
CUNY Graduate Center
Raman And Surface Enhanced Raman Spectroscopy For Forensic Analysis: Case Studies On The Identification Of Illicit Substances And Artist Pigments, Abed Haddad
Dissertations, Theses, and Capstone Projects
Raman spectroscopy is an effective tool for detecting trace amounts of material by fingerprint-like vibrational spectra. At times, the weak intensity of Raman scattering can make it difficult to distinguish trace materials. This shortcoming is addressed by surface‐enhanced Raman spectroscopy (SERS), which produces strong signal enhancements when target compounds are near metal nanoparticles. For the first part of this thesis, the identification of fentanyl and carfentanil, main culprits in the opioid epidemic, was done using normal Raman and the SERS spectroscopy. As an aid in the assignment of the spectral lines, a computational model was built using Density Functional Theory …
A Comparison Of Standard Denoising Methods For Peptide Identification,
2019
East Tennessee State University
A Comparison Of Standard Denoising Methods For Peptide Identification, Skylar Carpenter
Electronic Theses and Dissertations
Peptide identification using tandem mass spectrometry depends on matching the observed spectrum with the theoretical spectrum. The raw data from tandem mass spectrometry, however, is often not optimal because it may contain noise or measurement errors. Denoising this data can improve alignment between observed and theoretical spectra and reduce the number of peaks. The method used by Lewis et. al (2018) uses a combined constant and moving threshold to denoise spectra. We compare the effects of using the standard preprocessing methods baseline removal, wavelet smoothing, and binning on spectra with Lewis et. al’s threshold method. We consider individual methods and …
Advanced Statistics In Arkansas Sports Reporting,
2019
University of Arkansas, Fayetteville
Advanced Statistics In Arkansas Sports Reporting, Andrew Lee Epperson
Graduate Theses and Dissertations
This study seeks to analyze how Arkansas’ sports journalists are adapting to the recent surge in available advanced statistics that are being used by certain national news organizations. Using in-depth qualitative research that includes in-depth interviews with a number of individuals in the print, broadcast, and athletics side of sports coverage, we discover how journalists and coaches use these next-generation analytics, what they fundamentally mean for the evolution of each respective path, and why so few Arkansas reporters and writers use them at the time of this paper’s defense. We see how budgets and deadlines restrict the use of these …
A Hidden Markov Factor Analysis Framework For Seizure Detection In Epilepsy Patients,
2019
University of Arkansas, Fayetteville
A Hidden Markov Factor Analysis Framework For Seizure Detection In Epilepsy Patients, Mahboubeh Madadi
Graduate Theses and Dissertations
Approximately 1% of the world population suffers from epilepsy. Continuous long-term electroencephalographic (EEG) monitoring is the gold-standard for recording epileptic seizures and assisting in the diagnosis and treatment of patients with epilepsy. Detection of seizure from the recorded EEG is a laborious, time consuming and expensive task. In this study, we propose an automated seizure detection framework to assist electroencephalographers and physicians with identification of seizures in recorded EEG signals. In addition, an automated seizure detection algorithm can be used for treatment through automatic intervention during the seizure activity and on time triggering of the injection of a radiotracer to …
Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis,
2019
Boise State University
Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis, Samuel Anyaso-Samuel
Boise State University Theses and Dissertations
The statistical process control (SPC) chart is an effective tool for the analysis, interpretation, and visualization of data from sequential processes. Commonly used SPC charts such as the Shewhart, CUSUM and EWMA charts are widely implemented in detecting distributional shifts in various processes. With recent scientific and technological advancements, massive amounts of data continue to be generated by production, medical, agricultural and many other industrial processes. Conventional SPC charts have significant drawbacks in monitoring such processes, specifically when the velocity of the data flow is greater than the run time of the monitoring procedure. In the literature, dynamic sampling control …
Deep Learning, Medical Physics And Cargo Cult Science.,
2019
University of San Francisco
Deep Learning, Medical Physics And Cargo Cult Science., Miguel Romero Phd, Gilmer Valdes Phd, Timothy Solberg Phd, Yannet Interian Phd
Creative Activity and Research Day - CARD
Deep learning algorithms have become widely popular, with considerable success in fields where datasets have hundreds of thousands or million points. As deep learning is increasingly applied to the fields of medical physics and radiation oncology, a reasonable question follows: are these techniques the best approach, given the unique conditions in our field? In this study, we investigate the dependence of dataset size on the performance of deep learning algorithms compared with more traditional radiomics-based methods.
Deep Neural Network Architectures For Music Genre Classification,
2019
University of San Francisco
Deep Neural Network Architectures For Music Genre Classification, Kai Middlebrook, Shyam Sudhakaran, Kunal Sonar, David Guy Brizan
Creative Activity and Research Day - CARD
With the recent advancements in technology, many tasks in fields such as computer vision, natural language processing, and signal processing have been solved using deep learning architectures. In the audio domain, these architectures have been used to learn musical features of songs to predict: moods, genres, and instruments. In the case of genre classification, deep learning models were applied to popular datasets--which are explicitly chosen to represent their genres--and achieved state-of-the-art results. However, these results have not been reproduced on less refined datasets. To this end, we introduce an un-curated dataset which contains genre labels and 30-second audio previews for …
The Andersen Likelihood Ratio Test With A Random Split Criterion Lacks Power,
2019
University College of Teacher Education Styria
The Andersen Likelihood Ratio Test With A Random Split Criterion Lacks Power, Georg Krammer
Journal of Modern Applied Statistical Methods
The Andersen LRT uses sample characteristics as split criteria to evaluate Rasch model fit, or theory driven hypothesis testing for a test. The power and Type I error of a random split criterion was evaluated with a simulation study. Results consistently show a random split criterion lacks power.
Weighted Version Of Generalized Inverse Weibull Distribution,
2019
University of Kashmir, Srinagar, India
Weighted Version Of Generalized Inverse Weibull Distribution, Sofi Mudiasir, S. P. Ahmad
Journal of Modern Applied Statistical Methods
Weighted distributions are used in many fields, such as medicine, ecology, and reliability. A weighted version of the generalized inverse Weibull distribution, known as weighted generalized inverse Weibull distribution (WGIWD), is proposed. Basic properties including mode, moments, moment generating function, skewness, kurtosis, and Shannon’s entropy are studied. The usefulness of the new model was demonstrated by applying it to a real-life data set. The WGIWD fits better than its submodels, such as length biased generalized inverse Weibull (LGIW), generalized inverse Weibull (GIW), inverse Weibull (IW) and inverse exponential (IE) distributions.
Calibration Of Measurements,
2019
University of British Columbia
Calibration Of Measurements, Edward Kroc, Bruno D. Zumbo
Journal of Modern Applied Statistical Methods
Traditional notions of measurement error typically rely on a strong mean-zero assumption on the expectation of the errors conditional on an unobservable “true score” (classical measurement error) or on the data themselves (Berkson measurement error). Weakly calibrated measurements for an unobservable true quantity are defined based on a weaker mean-zero assumption, giving rise to a measurement model of differential error. Applications show it retains many attractive features of estimation and inference when performing a naive data analysis (i.e. when performing an analysis on the error-prone measurements themselves), and other interesting properties not present in the classical or Berkson cases. Applied …
Estimation Of Mean With Two-Parameter Ratio-Product-Ratio Estimator In Double Sampling Using Ancillary Information Under Non-Response,
2019
Vikram University, Ujjain, India
Estimation Of Mean With Two-Parameter Ratio-Product-Ratio Estimator In Double Sampling Using Ancillary Information Under Non-Response, Surya K. Pal, Housila P. Singh
Journal of Modern Applied Statistical Methods
Ratio-product-ratio estimators with two parameters in double sampling under non-response are considered along with their properties. Practical conditions are obtained in which the suggested estimators are more proficient than other existing estimators. An example is given.
Capturing Heterogeneity Of Covariate Effects In Hidden Subpopulations In The Presence Of Censoring And Large Number Of Covariates,
2019
University of Nevada, Las Vegas
Capturing Heterogeneity Of Covariate Effects In Hidden Subpopulations In The Presence Of Censoring And Large Number Of Covariates, Farhad Shokoohi, Abbas Khalili, Masoud Asgharian, Shili Lin
Mathematical Sciences Faculty Research
The advent of modern technology has led to a surge of high-dimensional data in biology and health sciences such as genomics, epigenomics and medicine. The high-grade serous ovarian cancer (HGS-OvCa) data reported by The Cancer Genome Atlas (TCGA) Research Network is one example. The TCGA and other research groups have analyzed several aspects of these data. Here we study the relationship between Disease Free Time (DFT) after surgery among ovarian cancer patients and their DNA methylation profiles of genomic features. Such studies pose additional challenges beyond the typical big data problem due to population substructure and censoring. Despite the availability …
