Failure Of Care Acquisition: Identifying Risk Factors In American Health Disparities,
2017
DePauw University
Failure Of Care Acquisition: Identifying Risk Factors In American Health Disparities, Nicholas Downing, Mamunur Rashid
Student Research
We examined the effects of various demographic and socioeconomic risk factors that influence an adult's decision not to obtain medical care in the United States utilizing data from the 2015 National Health Interview Survey (NHIS). Bivariate analysis and multivariate logistic regression revealed that family income, insurance status and whether one worries about paying medical bills make individuals nearly 80% less likely to obtain care than their counterparts. This study provides evidence that certain risk factors, especially those directly related to one's socioeconomic status, may put individuals at greater risk for failure to obtain care. Interventions in policy may be needed …
An Investigation Of The Accuracy Of Parallel Analysis For Determining The Number Of Factors In A Factor Analysis,
2017
Western Kentucky University
An Investigation Of The Accuracy Of Parallel Analysis For Determining The Number Of Factors In A Factor Analysis, Mandy Matsumoto
Mahurin Honors College Capstone Experience/Thesis Projects
Exploratory factor analysis is an analytic technique used to determine the number of factors in a set of data (usually items on a questionnaire) for which the factor structure has not been previously analyzed. Parallel analysis (PA) is a technique used to determine the number of factors in a factor analysis. There are a number of factors that affect the results of a PA: the choice of the eigenvalue percentile, the strength of the factor loadings, the number of variables, and the sample size of the study. Although PA is the most accurate method to date to determine which factors …
Marketing The Mountain State: A Large N Study Of User Engagement On Twitter,
2017
Illinois State University
Marketing The Mountain State: A Large N Study Of User Engagement On Twitter, Kirk Richardson
Capstone Projects – Politics and Government
Much of the evolving research on the use of social media in destination marketing emphasizes how information diffusion influences the reputational image of place. The present study uses Twitter data to focus on the relative differences in user engagement across discrete account types. Specifically, this is done to examine how the official destination marketing organization of Montana—the Montana Office of Tourism (MTOT)—performs relative to other account types. Several regression analyses conducted on Twitter data associated with an ongoing MTOT place branding campaign reveal that tweets sent from ‘official’ accounts are more likely to be retweeted, and are estimated to receive …
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data,
2017
East Tennessee State University
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch
Electronic Theses and Dissertations
Missing data is one of the challenges we are facing today in modeling valid statistical models. It reduces the representativeness of the data samples. Hence, population estimates, and model parameters estimated from such data are likely to be biased.
However, the missing data problem is an area under study, and alternative better statistical procedures have been presented to mitigate its shortcomings. In this paper, we review causes of missing data, and various methods of handling missing data. Our main focus is evaluating various multiple imputation (MI) methods from the multiple imputation of chained equation (MICE) package in the statistical software …
Detecting And Evaluating Therapy Induced Changes In Radiomics Features Measured From Non-Small Cell Lung Cancer To Predict Patient Outcomes,
2017
The University of Texas MD Anderson Cancer Center UTHealth Graduate School of Biomedical Sciences
Detecting And Evaluating Therapy Induced Changes In Radiomics Features Measured From Non-Small Cell Lung Cancer To Predict Patient Outcomes, Xenia J. Fave
Dissertations and Theses (Open Access)
The purpose of this study was to investigate whether radiomics features measured from weekly 4-dimensional computed tomography (4DCT) images of non-small cell lung cancers (NSCLC) change during treatment and if those changes are prognostic for patient outcomes or dependent on treatment modality. Radiomics features are quantitative metrics designed to evaluate tumor heterogeneity from routine medical imaging. Features that are prognostic for patient outcome could be used to monitor tumor response and identify high-risk patients for adaptive treatment. This would be especially valuable for NSCLC due to the high prevalence and mortality of this disease.
A novel process was designed to …
Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association,
2017
University of Missouri-St. Louis
Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane
Theses
Alzheimer Disease (AD) is difficult to diagnose by using genetic testing or other traditional methods. Unlike diseases with simple genetic risk components, there exists no single marker determining as to whether someone will develop AD. Furthermore, AD is highly heterogeneous and different subgroups of individuals develop the disease due to differing factors. Traditional diagnostic methods using perceivable cognitive deficiencies are often too little too late due to the brain having suffered damage from decades of disease progression. In order to observe AD at early stages prior to the observation of cognitive deficiencies, biomarkers with greater accuracy are required. By using …
Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation,
2017
Murray State University
Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr
Scholars Week
Timing methods and performance metrics are important in the heavily industrialized world we live in. Industrial plants use metrics to measure quality of production, help make decisions, and drive the strategy of the organization. However, there are many factors to be considered when measuring performance based on a metric; of which we will be analyzing the importance of product variation. We will be analyzing assembly line timings, whilst controlling for product variance, to show the importance differences between products makes in one’s ability to predict performance. In addition, we will be analyzing the current “statistical” methods used by an industrial …
Disability In Long-Term Care Residents Explained By Prevalent Geriatric Syndromes, Not Long-Term Care Home Characteristics: A Cross-Sectional Study,
2017
University of Toronto
Disability In Long-Term Care Residents Explained By Prevalent Geriatric Syndromes, Not Long-Term Care Home Characteristics: A Cross-Sectional Study, Natasha E. Lane, Walter P. Wodchis, Cynthia M. Boyd, Thérèse A. Stukel
Dartmouth Scholarship
Self-care disability is dependence on others to conduct activities of daily living, such as bathing, eating and dressing. Among long-term care residents, self-care disability lowers quality of life and increases health care costs. Understanding the correlates of self-care disability in this population is critical to guide clinical care and ongoing research in Geriatrics. This study examines which resident geriatric syndromes and chronic conditions are associated with residents’ self-care disability and whether these relationships vary across strata of age, sex and cognitive status. It also describes the proportion of variance in residents’ self-care disability that is explained by residents’ geriatric syndromes …
Studying The Optimal Scheduling For Controlling Prostate Cancer Under Intermittent Androgen Suppression,
2017
Center for Applied Mathematics and Statistics and the Department of Mathematical Sciences, New Jersey Institute of Technology
Studying The Optimal Scheduling For Controlling Prostate Cancer Under Intermittent Androgen Suppression, Sunil K. Dhar, Hans R. Chaudhry, Bruce G. Bukiet, Zhiming Ji, Nan Gao, Thomas W. Findley
Harvard University Biostatistics Working Paper Series
This retrospective study shows that the majority of patients’ correlations between PSA and Testosterone during the on-treatment period is at least 0.90. Model-based duration calculations to control PSA levels during off-treatment are provided. There are two pairs of models. In one pair, the Generalized Linear Model and Mixed Model are both used to analyze the variability of PSA at the individual patient level by using the variable “Patient ID” as a repeated measure. In the second pair, Patient ID is not used as a repeated measure but additional baseline variables are included to analyze the variability of PSA.
Gamma/Hadron Separation For The Hawc Observatory,
2017
Michigan Technological University
Gamma/Hadron Separation For The Hawc Observatory, Michael J. Gerhardt
Dissertations, Master's Theses and Master's Reports
The High-Altitude Water Cherenkov (HAWC) Observatory is a gamma-ray observatory sensitive to gamma rays from 100 GeV to 100 TeV with an instantaneous field of view of ~2 sr. It is located on the Sierra Negra plateau in Mexico at an elevation of 4,100 m and began full operation in March 2015. The purpose of the detector is to study relativistic particles that are produced by interstellar and intergalactic objects such as: pulsars, supernova remnants, molecular clouds, black holes and more. To achieve optimal angular resolution, energy reconstruction and cosmic ray background suppression for the extensive air showers detected by …
Informational Index And Its Applications In High Dimensional Data,
2017
University of Kentucky
Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan
Theses and Dissertations--Statistics
We introduce a new class of measures for testing independence between two random vectors, which uses expected difference of conditional and marginal characteristic functions. By choosing a particular weight function in the class, we propose a new index for measuring independence and study its property. Two empirical versions are developed, their properties, asymptotics, connection with existing measures and applications are discussed. Implementation and Monte Carlo results are also presented.
We propose a two-stage sufficient variable selections method based on the new index to deal with large p small n data. The method does not require model specification and especially focuses …
A Traders Guide To The Predictive Universe- A Model For Predicting Oil Price Targets And Trading On Them,
2016
Washington University in St. Louis
A Traders Guide To The Predictive Universe- A Model For Predicting Oil Price Targets And Trading On Them, Jimmie Harold Lenz
Doctor of Business Administration Dissertations
At heart every trader loves volatility; this is where return on investment comes from, this is what drives the proverbial “positive alpha.” As a trader, understanding the probabilities related to the volatility of prices is key, however if you could also predict future prices with reliability the world would be your oyster. To this end, I have achieved three goals with this dissertation, to develop a model to predict future short term prices (direction and magnitude), to effectively test this by generating consistent profits utilizing a trading model developed for this purpose, and to write a paper that anyone with …
Tutorial For Using The Center For High Performance Computing At The University Of Utah And An Example Using Random Forest,
2016
Utah State University
Tutorial For Using The Center For High Performance Computing At The University Of Utah And An Example Using Random Forest, Stephen Barton
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
Random Forests are very memory intensive machine learning algorithms and most computers would fail at building models from datasets with millions of observations. Using the Center for High Performance Computing (CHPC) at the University of Utah and an airline on-time arrival dataset with 7 million observations from the U.S. Department of Transportation Bureau of Transportation Statistics we built 316 models by adjusting the depth of the trees and randomness of each forest and compared the accuracy and time each took. Using this dataset we discovered that substantial restrictions to the size of trees, observations allowed for each tree, and variables …
Advanced Data Analysis - Lecture Notes,
2016
University of New Mexico
Advanced Data Analysis - Lecture Notes, Erik B. Erhardt, Edward J. Bedrick, Ronald M. Schrader
Open Textbooks
Lecture notes for Advanced Data Analysis (ADA1 Stat 427/527 and ADA2 Stat 428/528), Department of Mathematics and Statistics, University of New Mexico, Fall 2016-Spring 2017. Additional material including RMarkdown templates for in-class and homework exercises, datasets, R code, and video lectures are available on the course websites: https://statacumen.com/teaching/ada1 and https://statacumen.com/teaching/ada2 .
Contents
I ADA1: Software
- 0 Introduction to R, Rstudio, and ggplot
II ADA1: Summaries and displays, and one-, two-, and many-way tests of means
- 1 Summarizing and Displaying Data
- 2 Estimation in One-Sample Problems
- 3 Two-Sample Inferences
- 4 Checking Assumptions
- 5 One-Way Analysis of Variance
III ADA1: Nonparametric, categorical, …
Implementing Some Basic Simuation Designs Using The Simsem Package In R,
2016
CUNY John Jay College
Implementing Some Basic Simuation Designs Using The Simsem Package In R, Keith A. Markus
Open Educational Resources
The purpose of this tutorial is to provide a very basic introduction to implementing three simple research designs using the simsem package in R. R is an open source statistical computing environment (R Core Team, 2015). For more information about R, see the R Project homepage (https://www.r-project.org/) and the Comprehensive R Archive Network (CRAN) web page (https://cran.r-project.org/). The lavaan package provides functions for fitting and evaluating structural equation models (Rosseel, 2012). For further information about the lavaan package including tutorials, see the lavaan Project web page (http://lavaan.ugent.be/). The simsem package (Pornprasertmanit, Miller & Schoemann, 2016) provides functions to facilitate structural …
The Influence Of The Electric Supply Industry On Economic Growth In Less Developed Countries,
2016
University of Southern Mississippi
The Influence Of The Electric Supply Industry On Economic Growth In Less Developed Countries, Edward Richard Bee
Dissertations
This study measures the impact that electrical outages have on manufacturing production in 135 less developed countries using stochastic frontier analysis and data from World Bank’s Investment Climate surveys. Outages of electricity, for firms with and without backup power sources, are the most frequently cited constraint on manufacturing growth in these surveys.
Outages are shown to reduce output below the production frontier by almost five percent in Africa and by a lower percentage in South Asia, Southeast Asia and the Middle East and North Africa. Production response to outages is quadratic in form. Outages also increase labor cost, reduce exports …
Quantifying Transit Access In New York City: Formulating An Accessibility Index For Analyzing Spatial And Social Patterns Of Public Transportation,
2016
CUNY Hunter College
Quantifying Transit Access In New York City: Formulating An Accessibility Index For Analyzing Spatial And Social Patterns Of Public Transportation, Maxwell S. Siegel
Theses and Dissertations
This paper aims to analyze accessibility within New York City’s transportation system through creating unique accessibility indices. Indices are detailed and implemented using GIS, analyzing the distribution of transit need and access. Regression analyses are performed highlighting relationships between demographics and accessibility and recommendations for transit expansion are presented.
Multivariate Thinking In An Intro Stats Course – Is It Possible?,
2016
Embry-Riddle Aeronautical University
Multivariate Thinking In An Intro Stats Course – Is It Possible?, Beverly Wood
Publications
Many of our students have an intuitive sense that there is more to the story than univariate or bivariate data can tell us. We can acknowledge and encourage that habit of digging deeper by demonstrating some ways to look at additional variables. Simpson’s paradox and side-by-side scatter plots are ways to provide a glimpse of more complex analysis that are accessible to students in an introductory course with or without strong quantitative skills.
Bivariate Negative Binomial Hurdle With Random Spatial Effects,
2016
Western Michigan University
Bivariate Negative Binomial Hurdle With Random Spatial Effects, Robert Mcnutt
Dissertations
Count data with excess zeros widely occur in ecology, epidemiology, marketing, and many other disciplines. Mixture distributions consisting of a point mass at zero and a separate discrete distribution are often employed in regression models to account for excessive zero observations in the data. While Poisson models are very popular for count data, Negative Binomial models provide greater flexibility due to their ability to account for overdispersion.
This research focuses on developing a method for analyzing bivariate count data with excess zeros collected over a lattice. A bivariate Zero-Inflated Negative Binomial Hurdle (ZINBH) regression model with spatial random effects is …
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization,
2016
Fox Chase Cancer Center
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
COBRA Preprint Series
Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …
