Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (18)
- Social and Behavioral Sciences (13)
- Statistical Methodology (13)
- Education (10)
- Design of Experiments and Sample Surveys (9)
-
- Computer Sciences (8)
- Mathematics (8)
- Applied Statistics (7)
- Data Science (7)
- Educational Assessment, Evaluation, and Research (7)
- Applied Mathematics (6)
- Environmental Sciences (6)
- Life Sciences (6)
- Engineering (5)
- Medicine and Health Sciences (5)
- Social Statistics (5)
- Biostatistics (4)
- Categorical Data Analysis (4)
- Multivariate Analysis (4)
- Public Affairs, Public Policy and Public Administration (4)
- Statistical Theory (4)
- Artificial Intelligence and Robotics (3)
- Business (3)
- Longitudinal Data Analysis and Time Series (3)
- Probability (3)
- Biometry (2)
- Chemical Engineering (2)
- Earth Sciences (2)
- Institution
- Keyword
-
- Machine learning (6)
- Survival analysis (4)
- Evaluation (3)
- Odds ratio (3)
- Archimedean copula (2)
-
- Bayesian (2)
- Classification (2)
- Deep learning (2)
- Design parameters (2)
- Equivalence test (2)
- Experimental design (2)
- Monte Carlo simulation (2)
- Poisson regression (2)
- Statistics (2)
- Stochastic block models (2)
- Variable selection (2)
- Zero-inflated (2)
- Adaptive test (1)
- Additive outliers (1)
- Adjusting for covariates (1)
- Alcoholism (1)
- Amazon product review data (1)
- Artificial intelligence (1)
- Asymmetric (1)
- Asymptotic normality (1)
- Atrial fibrillation in Hispanics (1)
- Attitudes about disability (1)
- Autoregressive (1)
- BRBM GLM (1)
- Bathymetry (1)
Articles 31 - 60 of 121
Full-Text Articles in Statistics and Probability
Ensemble Data Fitting For Bathymetric Models Informed By Nominal Data, Samantha Zambo
Ensemble Data Fitting For Bathymetric Models Informed By Nominal Data, Samantha Zambo
Dissertations
Due to the difficulty and expense of collecting bathymetric data, modeling is the primary tool to produce detailed maps of the ocean floor. Current modeling practices typically utilize only one interpolator; the industry standard is splines-in-tension.
In this dissertation we introduce a new nominal-informed ensemble interpolator designed to improve modeling accuracy in regions of sparse data. The method is guided by a priori domain knowledge provided by artificially intelligent classifiers. We recast such geomorphological classifications, such as ‘seamount’ or ‘ridge’, as nominal data which we utilize as foundational shapes in an expanded ordinary least squares regression-based algorithm. To our knowledge …
Asymmetric Multivariate Archimedean Copula Models And Semi-Competing Risks Data Analysis, Ziyan Guo
Asymmetric Multivariate Archimedean Copula Models And Semi-Competing Risks Data Analysis, Ziyan Guo
Dissertations
Many multivariate models have been proposed and developed to model high dimensional data when the dimension of a data set is greater than 2 (d ≥ 3). The existing multivariate models often force the “exchangeable” structure for part or the whole model, are not very flexible which tends to be of limited use in practice. There is a demand for developing and studying multivariate models with any pre-specified bivariate margins.
Suppose there exists such a class of flexible models with any pre-specified bivariate margins. Given a multivariate data, what is the distribution function and how to easily estimate the parameters …
A Management Strategy Evaluation Of The Impacts Of Interspecific Competition And Recreational Fishery Dynamics On Vermilion Snapper (Rhomboplites Aurorubens) In The Gulf Of Mexico, Megumi C. Oshima
Dissertations
In the Gulf of Mexico (GOM), Vermilion Snapper (Rhomboplites auroruben), are believed to compete with Red Snapper directly for prey and habitat. The two species share similar diets and have significant spatial overlap in the Gulf. Red Snapper are thought to be the dominate competitor, forcing Vermilion Snapper to feed on less nutritious prey when local resources are depleted. In addition to ecological pressures, GOM Vermilion Snapper support substantial commercial and recreational fisheries. Over the past decade, recreational landings have steadily increased, reaching a historical high in 2018. One cause may be stricter regulations for similar target species such as …
On Simes’S Second Conjecture: An Extended Single-Step Simes Test Procedure For Multiple Testing, Matthew G. Hudson
On Simes’S Second Conjecture: An Extended Single-Step Simes Test Procedure For Multiple Testing, Matthew G. Hudson
Dissertations
One of the major concerns with multiple tests of significance is controlling the family wise error rate. Various methods have been developed to ensure that the false positive rate be maintained at some prespecified level. One of the most well know being the Bonferroni procedure. Simes presented an improved Bonferroni procedure for testing the global hypothesis that is more powerful and less conservative, especially with positively correlated tests. While Simes’s procedure is more powerful, it does not allow for making inferences on the individual hypotheses. However, the Simes procedure has since become the foundation of many p-value based multiple testing …
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Dissertations
In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.
First, to improve the prediction accuracy of learning …
Statistical Properties And Applications Of Press Statistic, Ida Marie Alcantara
Statistical Properties And Applications Of Press Statistic, Ida Marie Alcantara
Dissertations
The most popularly used statistic R2 has a fundamental weakness in model building: it favors adding more predictors to the model because R2 can only increase. In effect, the additional predictors start fitting the noise in data. Other criterion in selecting a regression model such as R2 adj , AIC, SBC, and Mallow’s Cp does not guarantee the model selected will also make better prediction of future values. To avoid this, data scientists withhold a percentage of the data for validation purposes. The PRESS statistic does something similar by withholding each observation in calculating its own …
Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan
Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan
Dissertations
Spatial and temporal dependencies are ubiquitous properties of data in numerous domains. The popularity of spatial and temporal data mining has thus grown with the increasing prevalence of massive data. The presence of spatial and temporal attributes not only provides complementary useful perspectives, but also poses new challenges to the representation and integration into the learning procedure. In this dissertation, the involved spatial and temporal dependencies are explored with three genres: sample-wise, feature-wise, and target-wise. A family of novel methodologies is developed accordingly for the dependency representation in respective scenarios.
First, dependencies among discrete, continuous and repeated observations are studied …
Statistical Models For Correlated Data, Xiaomeng Niu
Statistical Models For Correlated Data, Xiaomeng Niu
Dissertations
Correlated data arise frequently in many studies where multiple response variables or repeatedly measured responses within subjects are correlated. My dissertation topic lies broadly in developing various statistical methodologies for correlated types of data such as longitudinal data, clustered data, and multivariate data.
Multiple response variables might be relevant within subjects. A univariate procedure fitting each response separately does not take into account the correlation among responses. To improve estimation efficiency for the regression parameter, this study proposes two estimation procedures by accommodating correlations among the response variables. The proposed procedures do not require knowledge of the true correlation structure …
Denoising Large Neuroimage Mri Data Using Spatial Random Effect Models, Leonard Chukuma Johnson
Denoising Large Neuroimage Mri Data Using Spatial Random Effect Models, Leonard Chukuma Johnson
Dissertations
Spatial smoothing in Magnetic Resonance image (MRI) involves applying a filter to remove high frequency information and consequently improves signal-to-noise ratio that can greatly aid neurosurgeons in pre-surgical planning stages of tumor resection. This immensely reduces the time spent on Electrical stimulation mapping (ESM) prior to surgery. MRI's three-dimensional data provides voxel intensities with complex spatial relationship. The standard de facto spatial smoothing method, Gaussian Kernel smoothing, is satisfactory since a uniform smoothing is done for the whole brain. Secondly, the kernel smoothing technique assumes normality for the voxel intensity, but there is ample evidence in current research that indicates …
Statistical Properties Of Population Stability Index, Bilal Yurdakul
Statistical Properties Of Population Stability Index, Bilal Yurdakul
Dissertations
Population stability is an important concept in model management. It is crucial to monitor whether the current population has changed from the population used during development of a model. For example, has the distribution of credit scores changed, and is the existing credit score model still valid? Population change may occur for many reasons–change in the economic environment, strategic change in the business, policy changes within the company, or changes in regulatory environment.
The population stability index (PSI) is a statistic that measures how much a variable has shifted over time, and is used to monitor applicability of a statistical …
Statistical And Clinical Equivalence Of Measurements, Puntipa Wanitjirattikal
Statistical And Clinical Equivalence Of Measurements, Puntipa Wanitjirattikal
Dissertations
This study proposes a test for statistical equivalence of two measurements. Typically, a new measurement process Υ is compared to an existing or standard measurement process Χ. We are assuming that Χ and Υ are measurements on the same scale. The paired t-test may be used to check for significant difference between (Χ, Υ) pairs. However, the paired t-test is intended to detect shift-type relationships of the form Υ=Χ+δ1 and may have low power for scale-type relations of the form Υ=γΧ.
We propose a test that has reasonable power to …
Diagnostics For Choosing Between Stratified Logrank And Stratified Wilcoxon, Jhoanne Marsh C. Gatpatan
Diagnostics For Choosing Between Stratified Logrank And Stratified Wilcoxon, Jhoanne Marsh C. Gatpatan
Dissertations
Martinez and Naranjo (2010) proposed a pretest for choosing between Logrank or Wilcoxon test in a two - sample case. However, in the presence of covariates, comparing two populations without adjusting for covariates would yield misleading results. In this study, we propose several pretests that will help the analyst decide to use stratified Logrank or stratified Wilcoxon tests in comparing two survival curves after covariates have been taken into account. Power performance of each adaptive test was done through simulations under PH and non-PH cases.
Spatial Analysis Of Time Between Two Consecutive Dental And Two Consecutive Well-Child Visits For Foster Care Youth, Chenyang Shi
Spatial Analysis Of Time Between Two Consecutive Dental And Two Consecutive Well-Child Visits For Foster Care Youth, Chenyang Shi
Dissertations
Foster care youth is a medically vulnerable population. Poor dental health and irregular well-child visits may cause serious health-related issues, such as mental disorder, nutrition imbalance, tooth damage, etc. Michigan requires all youth in foster care to receive annual dental and well-child visits. Usually, the study of foster care well-child and dental visits include two parts: time between two consecutive visits (gap time) and number of visits. For this study, a longitudinal-spatial model that has the flexibility to analyze the well-child/dental gap times and number of visits was developed. The longitudinal data (2009-2012) on Michigan foster care youth from 10 …
Development Of Traditional And Rank-Based Algorithms For Linear Models With Autoregressive Errors And Multivariate Logistic Regression With Spatial Random Effects, Shaofeng Zhang
Dissertations
Linear models are the most commonly used statistical methods in many disciplines. One of the model assumptions is that the error terms (residuals) are independent and identically distributed. This assumption is often violated and autoregressive error terms are often encountered by researchers. The most popular technique to deal with linear models with autoregressive errors is perhaps the autoregressive integrated moving average (ARIMA). Another common approach is generalized least squares, such as Cochrane-Orcutt estimation and Prais-Winsten estimation. However, these usually have poor behaviors when fitting small samples. To address this problem, a double bootstrap method was proposed by McKnight et al. …
Subgroup Analysis And Growth Curve Models For Longitudinal Data, Nichole Andrews
Subgroup Analysis And Growth Curve Models For Longitudinal Data, Nichole Andrews
Dissertations
In clinical trials and biomedical studies, treatments are compared to determine which one is effective against illness. Growth curve analysis can be beneficial in longitudinal biomedical studies, as we can evaluate the treatment effect on the response over time. The generalized growth curve model using polynomial regression is proposed for longitudinal data. An optimal degree for the polynomial is obtained using the BIQIF, an adaptation of the Bayesian information criterion. Quadratic inference functions are used to estimate the parameters of the model, which takes into account the fact that repeated measurements from the same subject are more likely to be …
The Influence Of The Electric Supply Industry On Economic Growth In Less Developed Countries, Edward Richard Bee
The Influence Of The Electric Supply Industry On Economic Growth In Less Developed Countries, Edward Richard Bee
Dissertations
This study measures the impact that electrical outages have on manufacturing production in 135 less developed countries using stochastic frontier analysis and data from World Bank’s Investment Climate surveys. Outages of electricity, for firms with and without backup power sources, are the most frequently cited constraint on manufacturing growth in these surveys.
Outages are shown to reduce output below the production frontier by almost five percent in Africa and by a lower percentage in South Asia, Southeast Asia and the Middle East and North Africa. Production response to outages is quadratic in form. Outages also increase labor cost, reduce exports …
Some Nonparametric Ordered Restricted Inference Problems In The Context Of A Statistical Education Study, Bradford M. Dykes
Some Nonparametric Ordered Restricted Inference Problems In The Context Of A Statistical Education Study, Bradford M. Dykes
Dissertations
Over the past 10 years, the Department of Statistics at Western Michigan University has developed a question generating system that can be used for creating multiple forms of exams, quizzes and homework for online and face-to-face use. This system can also be used to provide students with a form of instantaneous feedback. With the goal of analyzing how different levels of feedback in an online learning environment impacts students' performance on assignments, this study presents data collected on two semesters of students enrolled in three different meeting types (strictly online, typical face-to-face, and honors face-to-face) of an introductory Statistics course. …
Statistical Methodology For Data With Multiple Limits Of Detection, Robert M. Flikkema
Statistical Methodology For Data With Multiple Limits Of Detection, Robert M. Flikkema
Dissertations
Limitations of instruments used to collect continuous data sometimes lead to obtaining observations lower than a limit of detection. These observations are known as nondetects. They could be zeroes, or positive numbers, but they are too small to be recorded by a measuring device. Nondetects frequently occur in environmental data. Trace amounts of chemicals can exist in soil or groundwater and are undetectable by a machine reading. These observations pose a problem to researchers since the true values are unknown.
Simulations in the literature have led to inconsistent conclusions regarding what estimation technique to use with nondetect data when estimating …
Empirical Evaluation Of Different Features Of Design In Confirmatory Factor Analysis, Deyab Almaleki
Empirical Evaluation Of Different Features Of Design In Confirmatory Factor Analysis, Deyab Almaleki
Dissertations
Factor analysis (FA) is the study of variance within a group. Within-subject variance (WSV) is affected by multiple features in a study context, such as: the study experimental design (ED) and sampling design (SD), thus anything that influences or changes variance may affect the conclusions related to FA.
The aim of this study was to provide empirical evaluation of the influence of different aspects of ED and SD on WSV in the context of FA in terms of model precision and model estimate stability. Four Monte Carlo population correlation matrices were hypothesized based on different communality magnitudes (high, moderate, low, …
Bivariate Negative Binomial Hurdle With Random Spatial Effects, Robert Mcnutt
Bivariate Negative Binomial Hurdle With Random Spatial Effects, Robert Mcnutt
Dissertations
Count data with excess zeros widely occur in ecology, epidemiology, marketing, and many other disciplines. Mixture distributions consisting of a point mass at zero and a separate discrete distribution are often employed in regression models to account for excessive zero observations in the data. While Poisson models are very popular for count data, Negative Binomial models provide greater flexibility due to their ability to account for overdispersion.
This research focuses on developing a method for analyzing bivariate count data with excess zeros collected over a lattice. A bivariate Zero-Inflated Negative Binomial Hurdle (ZINBH) regression model with spatial random effects is …
Bayesian Rank Based Methods For Linear And Generalized Linear Models, James Kodzo Dzikunu
Bayesian Rank Based Methods For Linear And Generalized Linear Models, James Kodzo Dzikunu
Dissertations
A Bayesian Rank Based Method for linear models is developed in this research. The estimation of the regression coefficients is based on the full conditional distributions utilizing a rank based initial fit. The data likelihood is based on the asymptotic distribution of the gradient function and the asymptotic linearity of this rank-based procedure. Prior distributions are put on regression coefficient(s) and scale parameter(s). The effects of different priors on this scale parameter(s) are studied. Using these full conditional distributions, the estimates are obtained by a Markov Chain Monte-Carlo (MCMC) procedure. The results of our simulation studies show that these Bayesian …
Macrobenthic Communities In The Northern Gulf Of Mexico Hypoxic Zone: Testing The Pearson-Rosenberg Model, Shivakumar Shivarudrappa
Macrobenthic Communities In The Northern Gulf Of Mexico Hypoxic Zone: Testing The Pearson-Rosenberg Model, Shivakumar Shivarudrappa
Dissertations
The Pearson and Rosenberg (P-R) conceptual model of macrobenthic succession was used to assess the impact of hypoxia (dissolved oxygen [DO] ≤ 2 mg/L) on the macrobenthic community on the continental shelf of northern Gulf of Mexico for the first time. The model uses a stress-response relationship between environmental parameters and the macrobenthic community to determine the ecological condition of the benthic habitat. The ecological significance of dissolved oxygen in a benthic habitat is well understood. In addition, the annual recurrence of bottom-water hypoxia on the Louisiana/Texas shelf during summer months is well documented.
The P-R model illustrates the decreasing …
Poisson Versus Negative Binomial Regression In The Analysis Of Count Data, Barbie Ann L. Bugna
Poisson Versus Negative Binomial Regression In The Analysis Of Count Data, Barbie Ann L. Bugna
Dissertations
Commonly used tests for treatment effect in kx2 frequency data are Poisson regression, negative binomial regression, and Cochran-Mantel-Haentzel. In practice, Poisson regression or CMH is used as default, and NB regression is used only when there is reason to believe the data has overdispersion beyond what is expected of Poisson counts.
We show that the Poisson regression is sensitive to the Poisson assumption, and does not maintain its size in the presence of overdispersion. In particular, it tends to interpret overdispersion as significant treatment effect. Thus there is a need for a reliable pretest for the Poisson assumption. A commonly …
Rank Based Procedures For Ordered Alternative Models, Yuanyuan Shao
Rank Based Procedures For Ordered Alternative Models, Yuanyuan Shao
Dissertations
The ordered alternatives in a one-way layout with k ordered treatment levels are appropriate for many applications, especially in psychology and medicine. There is extensive literature in this area, and many parametric and nonparametric approaches have been introduced. Abelson-Tukey (AT) test is a frequently used parametric method. Its coefficients provide an ideal way of combining means for the purpose of detecting a monotonic relationship between the independent and dependent variables. The AT method, though, is not robust. Furthermore, our initial empirical studies show that it is not more powerful than the Jonckheere-Terpstra (JT) and the Hettmansperger- Norton (HN) nonparametric tests …
Failing To Replicate: Hypothesis Testing As A Crucial Key To Make Direct Replications More Credible And Predictable, Pedro Fernando Mateu Bullón
Failing To Replicate: Hypothesis Testing As A Crucial Key To Make Direct Replications More Credible And Predictable, Pedro Fernando Mateu Bullón
Dissertations
Theory cannot be fully validated unless the original results have been replicated, resulting in conclusion consistency. Replications are the strongest source to verify research findings and knowledge claims. Sciences such as medicine, chemistry, physics, genetics, and biology, are considered successful because their knowledge claims are buttressed by a large set of replications of original studies. Unfortunately in the social sciences many attempts to replicate fail and thus there is a continuing need for replication studies to confirm facts, expand knowledge to gain new understanding, and verify hypotheses. Two plausible explanations for the failure to replicate in the social sciences could …
Three Essays On Panel Data Estimation, Alexander Houser
Three Essays On Panel Data Estimation, Alexander Houser
Dissertations
This work discusses various aspects of panel data estimation. In chapter one, an algorithm for semiparametric random effects estimation is proposed. The performance of bootstrap-based confidence intervals for the proposed estimators are examined and found reasonable. The algorithm is also applied to a set of U.S. state level medical expenditure data to estimate the medical Engel curve. In the second chapter, the predictive performance of various parametric and semiparametric panel data estimators is compared on the same dataset of U.S. state level medical expenditures as well as out of sample forecast performance and bootstrap bias-corrected mean square errors of the …
Lnference On Differences In K Means For Data With Excess Zeros And Detection Limits, Haolai Jiang
Lnference On Differences In K Means For Data With Excess Zeros And Detection Limits, Haolai Jiang
Dissertations
Many data have excess zeros or unobservable values falling below detection limit. For example, data on hospitalization costs incurred by members of a health insurance plan will have zeros for the percentage who did not get sick. Benzene exposure measurements on petroleum re nery workers have some exposures fall below the limit of detection. Traditional methods of inference like one-way ANOVA are not appropriate to analyze such data since the point mass at zero violates typical distribution assumptions.
For testing for equality of means of k distributions, we will propose a likelihood ratio test that accounts for excess zeros or …
Comparison Of Hazard, Odds And Risk Ratio In The Two-Sample Survival Problem, Benedict P. Dormitorio
Comparison Of Hazard, Odds And Risk Ratio In The Two-Sample Survival Problem, Benedict P. Dormitorio
Dissertations
Cox proportional hazards is the standard method for analyzing treatment efficacy when time-to-event data is available. In the absence of time-to-event, investigators may use logistic regression which only requires relative frequencies of events, or Poisson regression which requires only interval-summarized frequency tables of time-to-event. When event frequencies are used instead of time-to-events, does it always result in a loss in power?
We investigate the relative performance of the three methods. In particular, we compare the power of tests based on the respective effect-size estimates (1)hazard ratio (HR), (2)odds ratio (OR), and (3)risk ratio (RR). We use a variety of survival …
A Metaevaluation Of Energy Efficiency Evaluations, Brandy Brown
A Metaevaluation Of Energy Efficiency Evaluations, Brandy Brown
Dissertations
This study systematically reviews the methodological characteristics of energy efficiency evaluations and uses metaevaluation to assess its quality. Metaevaluation is used to systematically assess the quality of evaluation products, confirm that evaluations deliver sound findings and conclusions, are useful to the client, are credible, are ethically conducted, and are done as cost-effective as possible. The results of this study show that the ability to accurately assess evaluation for methodological quality using evaluations reports as a primary data source depends on the presence of detailed descriptions of evaluation methods. Furthermore, the study suggests that methodological variations of energy efficiency evaluations coalesce …
Multivariate Autoregressive Time Series Using Schweppe Weighted Wilcoxon Estimates, Jaime Burgos
Multivariate Autoregressive Time Series Using Schweppe Weighted Wilcoxon Estimates, Jaime Burgos
Dissertations
The increasing needs of forecasting techniques has led to the popularity of the vector autoregressive model in multivariate time series analysis, which has become of typical use across different fields due to its simplicity in application. The traditional method for estimating the model parameters is the least squares minimization, due to the linear nature of the model and its similarity with multivariate linear regression. However, since least squares estimates are sensitive to outliers, more robust techniques have become of interest. This manuscript investigates a robust alternative by obtaining the estimates using a weighted Wilcoxon dispersion with Schweppe-type weights. The first …