Robust Ancova, Curvature, And The Curse Of Dimensionality,
2019
University of Southern California
Robust Ancova, Curvature, And The Curse Of Dimensionality, Rand Wilcox
Journal of Modern Applied Statistical Methods
There is a substantial collection of robust analysis of covariance (ANCOVA) methods that effectively deals with non-normality, unequal population slope parameters, outliers, and heteroscedasticity. Some are based on the usual linear model and others are based on smoothers (nonparametric regression estimators). However, extant results are limited to one or two covariates. A minor goal here is to extend a recently-proposed method, based on the usual linear model, to situations where there are up to six covariates. The usual linear model might provide a poor approximation of the true regression surface. The main goal is to suggest a method, based on …
A Strategy For Using Bias And Rmse As Outcomes In Monte Carlo Studies In Statistics,
2019
University of Minnesota - Twin Cities
A Strategy For Using Bias And Rmse As Outcomes In Monte Carlo Studies In Statistics, Michael Harwell
Journal of Modern Applied Statistical Methods
To help ensure important patterns of bias and accuracy are detected in Monte Carlo studies in statistics this paper proposes conditioning bias and root mean square error (RMSE) measures on estimated Type I and Type II error rates. A small Monte Carlo study is used to illustrate this argument.
A Robust Nonparametric Measure Of Effect Size Based On An Analog Of Cohen's D, Plus Inferences About The Median Of The Typical Difference,
2019
University of Southern California
A Robust Nonparametric Measure Of Effect Size Based On An Analog Of Cohen's D, Plus Inferences About The Median Of The Typical Difference, Rand Wilcox
Journal of Modern Applied Statistical Methods
The paper describes a nonparametric analog of Cohen's d, Q. It is established that a confidence interval for Q can be computed via a method for computing a confidence interval for the median of D = X1 − X2, which in turn is related to making inferences about P(X1 < X2).
Should We Give Up On Causality?,
2019
The Ohio State University
Should We Give Up On Causality?, Tom Knapp
Journal of Modern Applied Statistical Methods
No abstract provided.
Striving For Simple But Effective Advice For Comparing The Central Tendency Of Two Populations,
2019
University of St Andrews
Striving For Simple But Effective Advice For Comparing The Central Tendency Of Two Populations, Graeme Ruxton, Markus Neuhäuser
Journal of Modern Applied Statistical Methods
Nguyen et al. (2016) offered advice to researchers in the commonly-encountered situation where they are interested in testing for a difference in central tendency between two populations. Their data and the available literature support very simple advice that strikes the best balance between ease of implementation, power and reliability. Specifically, apply Satterthwaite’s test, with preliminary ranking of the data if a strong deviation from normality is expected, or is suggested by visual inspection of the data. This simple guideline will serve well except when dealing with small samples of discrete data, when more sophisticated treatment may be required.
Logistic Regression: An Inferential Method For Identifying The Best Predictors,
2019
University of Southern California
Logistic Regression: An Inferential Method For Identifying The Best Predictors, Rand Wilcox
Journal of Modern Applied Statistical Methods
When dealing with a logistic regression model, there is a simple method for estimating the strength of the association between the jth covariate and the dependent variable when all covariates are entered into the model. There is the issue of determining whether the jth independent variable has a stronger or weaker association than the kth independent variable. This note describes a method for dealing with this issue that was found to perform reasonably well in simulations.
Data Analytics Pipeline For Rna Structure Analysis Via Shape,
2019
University of Nebraska at Omaha
Data Analytics Pipeline For Rna Structure Analysis Via Shape, Quinn Nelson
UNO Student Research and Creative Activity Fair
Coxsackievirus B3 (CVB3) is a cardiovirulent enterovirus from the family Picornaviridae. The RNA genome houses an internal ribosome entry site (IRES) in the 5’ untranslated region (5’UTR) that enables cap-independent translation. Ample evidence suggests that the structure of the 5’UTR is a critical element for virulence. We probe RNA structure in solution using base-specific modifying agents such as dimethyl sulfate as well as backbone targeting agents such as N-methylisatoic anhydride used in Selective 2’-Hydroxyl Acylation Analyzed by Primer Extension (SHAPE). We have developed a pipeline that merges and evaluates base-specific and SHAPE data together with statistical analyses that provides confidence …
Sustainable Energy Governance In South Tyrol (Italy): A Probabilistic Bipartite Network Model,
2019
EURAC Research, Italy
Sustainable Energy Governance In South Tyrol (Italy): A Probabilistic Bipartite Network Model, Jessica Belest, Laura Secco, Elena Pisani, Alberto Caimo
Articles
At the national scale, almost all of the European countries have already achieved energy transition targets, while at the regional and local scales, there is still some potential to further push sustainable energy transitions. Regions and localities have the support of political, social, and economic actors who make decisions for meeting existing social, environmental and economic needs recognising local specificities.
These actors compose the sustainable energy governance that is fundamental to effectively plan and manage energy resources. In collaborative relationships, these actors share, save, and protect several kinds of resources, thereby making energy transitions deeper and more effective.
This research …
Session: 4 Multilinear Subspace Learning And Its Applications To Machine Learning,
2019
SDSMT
Session: 4 Multilinear Subspace Learning And Its Applications To Machine Learning, Randy Hoover, Kyle Caudle Dr., Karen Braman Dr.
SDSU Data Science Symposium
Multi-dimensional data analysis has seen increased interest in recent years. With more and more data arriving as 2-dimensional arrays (images) as opposed to 1-dimensioanl arrays (signals), new methods for dimensionality reduction, data analysis, and machine learning have been pursued. Most notably have been the Canonical Decompositions/Parallel Factors (commonly referred to as CP) and Tucker decompositions (commonly regarded as a high order SVD: HOSVD). In the current research we present an alternate method for computing singular value and eigenvalue decompositions on multi-way data through an algebra of circulants and illustrate their application to two well-known machine learning methods: Multi-Linear Principal Component …
Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning,
2019
Sanford Health
Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning, Arielle Selya, Eric L. Johnson
SDSU Data Science Symposium
Diabetes poses a variety of medical complications to patients, resulting in a high rate of unplanned medical visits, which are costly to patients and healthcare providers alike. However, unplanned medical visits by their nature are very difficult to predict. The current project draws upon electronic health records (EMR’s) of adult patients with diabetes who received care at Sanford Health between 2014 and 2017. Various machine learning methods were used to predict which patients have had an unplanned medical visit based on a variety of EMR variables (age, BMI, blood pressure, # of prescriptions, # of diagnoses on problem list, A1C, …
Nonparametric Depth And Quantile Regression For Functional Data,
2019
Indian Statistical Institute, Kolkata
Nonparametric Depth And Quantile Regression For Functional Data, Joydeep Chowdhury, Probal Chaudhuri
Journal Articles
We investigate nonparametric regression methods based on spatial depth and quantiles when the response and the covariate are both functions. As in classical quantile regression for finite dimensional data, regression techniques developed here provide insight into the influence of the functional covariate on different parts, like the center as well as the tails, of the conditional distribution of the functional response. Depth and quantile based nonparametric regression methods are useful to detect heteroscedasticity in functional regression. We derive the asymptotic behavior of the nonparametric depth and quantile regression estimates, which depend on the small ball probabilities in the covariate space. …
Pedestrian Safety -- Fundamental To A Walkable City,
2019
Southern Methodist University
Pedestrian Safety -- Fundamental To A Walkable City, Joshua Herrera, Patrick Mcdevitt, Preeti Swaminathan, Raghuram Srinivas
SMU Data Science Review
In this paper, we present a method to identify urban areas with a higher likelihood of pedestrian safety related events. Pedestrian safety related events are pedestrian-vehicle interactions that result in fatalities, injuries, accidents without injury, or near--misses between pedestrians and vehicles. To develop a solution to this problem of identifying likely event locations, we assemble data, primarily from the City of Cincinnati and Hamilton County, that include safety reports from a five year period, geographic information for these events, citizen survey of pedestrian reported concerns, non-emergency requests for service for any cause in the city, property values and public transportation …
Improving Vix Futures Forecasts Using Machine Learning Methods,
2019
Southern Methodist University
Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater
SMU Data Science Review
The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …
Ample Provision: A Preliminary Study Relating Budget Composition And High School Graduation Rates In Select Washington State Public School Districts,
2019
Central Washington University
Ample Provision: A Preliminary Study Relating Budget Composition And High School Graduation Rates In Select Washington State Public School Districts, Gregory P. Gadow
All Undergraduate Projects
How to allocate scarce resources for an optimal outcome is of keen interest to those who set the budgets in public education. Simply throwing money at schools is not enough; it is important that money is spent where it will do the most good. This study considers Washington State public school districts and examines how the share of per-student expenditures in seven budget categories relates to on-time high school graduation rates. It is an investigative study, exploring whether there is enough evidence to merit further, more in-depth research. Using budget and graduation information from academic years 1997-98 through 2016-17 for …
Comparative Analysis Of Students’ Performance Between Online And On Campus In An Introductory Statistics Course,
2019
Georgia College and State University
Comparative Analysis Of Students’ Performance Between Online And On Campus In An Introductory Statistics Course, Kendal Mcdonald
The Corinthian
In this research, we compare students’ performance in an online and on-campus introductory statistics and probability course at Georgia College. MyStatLab is the learning management system used in both the online and on-campus courses for homework and quizzes. The online data is produced by five summer courses between Summer 2014 to Summer 2017 and the on-campus data is produced from nine on-campus courses from Spring 2014, Spring 2016, and Spring 2017. For homework, the research compares the scores made between online and on-campus. For quizzes, we test if there is a difference between the scores and the number of attempts …
Step Away From Stepwise,
2019
Pomona College
Step Away From Stepwise, Gary N. Smith
Pomona Economics
Stepwise regression is a popular data-mining tool that uses statistical significance to select the explanatory variables to be used in a multiple-regression model. A fundamental problem with stepwise regression is that some real explanatory variables that have causal effects on the dependent variable may happen to not be statistically significant, while nuisance variables may be coincidentally significant. As a result, the model may fit the data well in-sample, but do poorly out-of-sample. Many Big-Data researchers believe that, the larger the number of possible explanatory variables, the more useful is stepwise regression for selecting explanatory variables. The reality is that stepwise …
Controlling For Confounding Via Propensity Score Methods Can Result In Biased Estimation Of The Conditional Auc: A Simulation Study,
2019
Old Dominion University
Controlling For Confounding Via Propensity Score Methods Can Result In Biased Estimation Of The Conditional Auc: A Simulation Study, Hadiza I. Galadima, Donna K. Mcclish
Community & Environmental Health Faculty Publications
In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not random in observational studies, comparisons of outcomes between exposed and nonexposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of conditional odds ratio and hazard ratio. However, research is lacking on the performance of propensity score methods for covariate adjustment when estimating the …
The Dark Sky Character Of Archaeological Landscapes: Cultural Meaning And Conservation Strategies,
2019
Technological University Dublin
The Dark Sky Character Of Archaeological Landscapes: Cultural Meaning And Conservation Strategies, Frank Prendergast
Book/Book Chapter
This paper presents the first ever study of light pollution at selected Irish prehistoric archaeological landscapes. The concepts of cosmology and landscape are first briefly described and followed by a summary of early human settlement of the island. Building on this, the extant corpus of early prehistoric megalithic burial tombs is illustrated to show their contrasting distribution patterns and typology. Analysis of tomb locations using nearest-neighbour statistical methods reveals evidence of intentional clustering. Further geo-statistical analysis identifies the geographical locations and the density ranking of these nucleated clusters - a feature especially evident in the passage tomb tradition on this …
Be Wary Of Black-Box Trading Algorithms,
2019
Pomona College
Be Wary Of Black-Box Trading Algorithms, Gary N. Smith
Pomona Economics
Black-box algorithms now account for nearly a third of all U. S. stock trades. It is a mistake to think that these algorithms possess superhuman intelligence. In reality, computers do not have the common sense and wisdom that humans have accumulated by living. Trading algorithms are particularly dangerous because they are so efficient at discovering statistical patterns—but so utterly useless in judging whether the discovered patterns are meaningful.
The Scaling Limit Of The Membrane Model,
2019
Delft University of Technology
The Scaling Limit Of The Membrane Model, Alessandra Cipriani, Biltu Dan, Rajat Subhra Hazra
Journal Articles
On the integer lattice, we consider the discrete membrane model, a random interface in which the field has Laplacian interaction. We prove that, under appropriate rescaling, the discrete membrane model converges to the continuum membrane model in d ≥ 2. Namely, it is shown that the scaling limit in d = 2, 3 is a Holder continuous random field, while in d ≥ 4 the membrane model converges to a random distribution. As a by-product of the proof in d = 2, 3, we obtain the scaling limit of the maximum. This work complements the analogous results of Caravenna and …
