Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Changsha University of Science and Technology (570)
- COBRA (362)
- University of Kentucky (43)
- Central Bank of Nigeria (28)
- Southern Methodist University (27)
-
- Virginia Commonwealth University (23)
- Kennesaw State University (20)
- University of Denver (18)
- City University of New York (CUNY) (17)
- University of Arkansas, Fayetteville (16)
- University of Nebraska - Lincoln (16)
- Georgia Southern University (15)
- California Polytechnic State University, San Luis Obispo (14)
- Michigan Technological University (14)
- University of Louisville (13)
- University of Nevada, Las Vegas (13)
- Old Dominion University (12)
- Air Force Institute of Technology (11)
- Western Michigan University (11)
- Claremont Colleges (9)
- East Tennessee State University (9)
- The Texas Medical Center Library (9)
- The University of Akron (9)
- Bethel University (8)
- Clemson University (8)
- Rochester Institute of Technology (8)
- South Dakota State University (8)
- University of Central Florida (8)
- University of North Florida (8)
- West Virginia University (8)
- Keyword
-
- Road engineering (53)
- Statistics (50)
- Bridge engineering (31)
- Numerical simulation (31)
- Machine learning (18)
-
- Regression (18)
- Cable-stayed bridge (16)
- Causal inference (16)
- Prediction (16)
- Asphalt pavement (15)
- Simulation (15)
- Morgridge College of Education (14)
- Research Methods and Information Science (14)
- Research Methods and Statistics (14)
- Bootstrap (13)
- Missing data (13)
- Tunnel engineering (13)
- Machine Learning (12)
- Mechanical property (12)
- Model selection (12)
- Subgrade engineering (12)
- Classification (11)
- Genetics (11)
- Suspension bridge (11)
- Multiple testing (10)
- Psychology (10)
- Concrete (9)
- Cross-validation (9)
- Logistic regression (9)
- Stability (9)
- Publication Year
- Publication
-
- Journal of China & Foreign Highway (570)
- U.C. Berkeley Division of Biostatistics Working Paper Series (114)
- Harvard University Biostatistics Working Paper Series (74)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (59)
- UW Biostatistics Working Paper Series (56)
-
- Electronic Theses and Dissertations (41)
- Theses and Dissertations--Statistics (38)
- Theses and Dissertations (34)
- CBN Journal of Applied Statistics (JAS) (28)
- The University of Michigan Department of Biostatistics Working Paper Series (28)
- COBRA Preprint Series (24)
- Statistical Science Theses and Dissertations (16)
- Symposium of Student Scholars (15)
- College of Graduate Studies: Theses & Dissertations (14)
- Dissertations, Master's Theses and Master's Reports (14)
- Dissertations (13)
- Graduate Theses and Dissertations (13)
- Articles (11)
- SMU Data Science Review (10)
- Master's Theses (9)
- Williams Honors College, Honors Research Projects (9)
- All Dissertations (8)
- Dissertations and Theses (Open Access) (8)
- Psychology Student Works (8)
- UNF Graduate Theses and Dissertations (8)
- SDSU Data Science Symposium (7)
- Data Science and Data Mining (6)
- Publications and Research (6)
- Department of Statistics: Dissertations, Theses, and Student Research (5)
- Dissertations and Theses (5)
- Publication Type
- File Type
Articles 901 - 930 of 1562
Full-Text Articles in Statistics and Probability
Evaluation Of The Utility Of Informative Priors In Bayesian Structural Equation Modeling With Small Samples, Hao Ma
Education Policy and Leadership Theses and Dissertations
The estimation of parameters in structural equation modeling (SEM) has been primarily based on the maximum likelihood estimator (MLE) and relies on large sample asymptotic theory. Consequently, the results of the SEM analyses with small samples may not be as satisfactory as expected. In contrast, informative priors typically do not require a large sample, and they may be helpful for improving the quality of estimates in the SEM models with small samples. However, the role of informative priors in the Bayesian SEM has not been thoroughly studied to date. Given the limited body of evidence, specifying effective informative priors remains …
“The Prediction Of Fantasy Football”, Chelsea Robinson
“The Prediction Of Fantasy Football”, Chelsea Robinson
Mathematics Senior Capstone Papers
In this paper, we consider the game fantasy football, which allows people to simulate being a National Football League team owner. Imaginary owners select from the best players in the NFL and compete on weekly basis based upon player performances on the field. Fantasy football has become popular over the years. In 2011, according to the Fantasy Sports Trade Association there were 35 million people that played fantasy sports online in the United States and Canada. The most major companies that use fantasy football are Yahoo, ESPN, and NFL, even though there are more platforms. Many people use these platforms …
Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder
Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder
Dissertations, 2020-current
The Rasch model is commonly used to calibrate multiple choice items. However, the sample sizes needed to estimate the Rasch model can be difficult to attain (e.g., consider a small testing company trying to pretest new items). With small sample sizes, auxiliary information besides the item responses may improve estimation of the item parameters. The purpose of this study was to determine if incorporating item property information (i.e., characteristics of the items related to item difficulty) in a random effects linear logistic test model (RE-LLTM) would improve estimation of item difficulty. A simulation study was conducted that varied sample size, …
Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig
Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig
Masters Theses, 2020-current
In the absence of random assignment, researchers must consider the impact of selection bias – pre-existing covariate differences between groups due to differences among those entering into treatment and those otherwise unable to participate. Propensity score matching (PSM) and generalized boosted modeling (GBM) are two quasi-experimental pre-processing methods that strive to reduce the impact of selection bias before analyzing a treatment effect. PSM and GBM both examine a treatment and comparison group and either match or weight members of those groups to create new, balanced groups. The new, balanced groups theoretically can then be used as a proxy for the …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Introduction To Research Statistical Analysis: An Overview Of The Basics, Christian Vandever
Introduction To Research Statistical Analysis: An Overview Of The Basics, Christian Vandever
HCA Healthcare Journal of Medicine
This article covers many statistical ideas essential to research statistical analysis. Sample size is explained through the concepts of statistical significance level and power. Variable types and definitions are included to clarify necessities for how the analysis will be interpreted. Categorical and quantitative variable types are defined, as well as response and predictor variables. Statistical tests described include t-tests, ANOVA and chi-square tests. Multiple regression is also explored for both logistic and linear regression. Finally, the most common statistics produced by these methods are explored.
Demand Forecasting In Wholesale Alcohol Distribution: An Ensemble Approach, Tanvi Arora, Rajat Chandna, Stacy Conant, Bivin Sadler, Robert Slater
Demand Forecasting In Wholesale Alcohol Distribution: An Ensemble Approach, Tanvi Arora, Rajat Chandna, Stacy Conant, Bivin Sadler, Robert Slater
SMU Data Science Review
In this paper, historical data from a wholesale alcoholic beverage distributor was used to forecast sales demand. Demand forecasting is a vital part of the sale and distribution of many goods. Accurate forecasting can be used to optimize inventory, improve cash ow, and enhance customer service. However, demand forecasting is a challenging task due to the many unknowns that can impact sales, such as the weather and the state of the economy. While many studies focus effort on modeling consumer demand and endpoint retail sales, this study focused on demand forecasting from the distributor perspective. An ensemble approach was applied …
Data-Driven Investment Decisions In P2p Lending: Strategies Of Integrating Credit Scoring And Profit Scoring, Yan Wang
Doctor of Data Science and Analytics Dissertations
In this dissertation, we develop and discuss several loan evaluation methods to guide the investment decisions for peer-to-peer (P2P) lending. In evaluating loans, credit scoring and profit scoring are the two widely utilized approaches. Credit scoring aims at minimizing the risk while profit scoring aims at maximizing the profit. This dissertation addresses the strengths and weaknesses of each scoring method by integrating them in various ways in order to provide the optimal investment suggestions for different investors. Before developing the methods for loan evaluation at the individual level, we applied the state-of-the-art method called the Long Short Term Memory (LSTM) …
A Monte Carlo Analysis Of Standard Error-Based Methods For Computing Confidence Intervals, Elayna Wichert
A Monte Carlo Analysis Of Standard Error-Based Methods For Computing Confidence Intervals, Elayna Wichert
Masters Theses & Specialist Projects
The objective of this study is to empirically test existing techniques to calculate the likely range of values for a Classical Test Theory true score given an observed score. The traditional method for forming these confidence intervals has used the standard error of measurement (SEM) as the basis for this confidence interval. An alternate equation, the standard error of estimate (SEE), has been recommended in place of the SEM for this purpose, yet it remains overlooked in the field of psychometrics. It is important that the correct equation be used in various applications in personnel psychology. Monte Carlo analyses were …
Psychometric Properties Of A Year-End Form Four Chemistry Paper 1 In Selected Schools In Petaling Utama District, Mohamed Razali Rose Enne Emellia
Psychometric Properties Of A Year-End Form Four Chemistry Paper 1 In Selected Schools In Petaling Utama District, Mohamed Razali Rose Enne Emellia
Student Works (2020-2029)
The low-performance of science stream students in Chemistry especially in high-stakes testing, has alarmed the stakeholders. This problem has led to the need of analyzing the instruments used to measure the achievement of students. The present study examined the psychometric properties of multiple-choice questions in the Year-end Form Four Chemistry Paper 1 using the Rasch Model in the Klang Valley. A quantitative research design has been used for this study. The sample comprised of 435 Form Four Pure Science students from four randomly selected secondary schools in the Petaling Utama district. The Rasch analysis was conducted in two stages. The …
Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter
Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter
SDSU Data Science Symposium
Lifetime percentile is an important indicator of product reliability. However, the sampling distribution of a percentile estimator for any lifetime distribution is not a bell shaped one. As a result, the well-known Shewhart-type control chart cannot be applied to monitor the product lifetime percentiles. In this presentation, Bootstrap control charts based on maximum likelihood estimator (MLE) are proposed for monitoring Pareto percentiles. An intensive simulation study is conducted to compare the performance among the proposed MLE Bootstrap control chart and Shewhart-type control chart.
An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone
An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone
Published and Grey Literature from PhD Candidates
Data mining techniques have numerous applications in bankcard response modeling. Logistic regression has been used as the standard modeling tool in the financial industry because of its almost always desirable performance and its interpretability. In this paper, we propose a hybrid bankcard response model, which integrates decision tree-based chi-square automatic interaction detection (CHAID) into logistic regression. In the first stage of the hybrid model, CHAID analysis is used to detect the possible potential variable interactions. Then in the second stage, these potential interactions are served as the additional input variables in logistic regression. The motivation of the proposed hybrid model …
A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone
A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone
Published and Grey Literature from PhD Candidates
We propose a two-stage hybrid approach with neural networks as the new feature construction algorithms for bankcard response classifications. The hybrid model uses a very simple neural network structure as the new feature construction tool in the first stage, then the newly created features are used as the additional input variables in logistic regression in the second stage. The model is compared with the traditional one-stage model in credit customer response classification. It is observed that the proposed two-stage model outperforms the one-stage model in terms of accuracy, the area under the ROC curve, and KS statistic. By creating new …
A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni
A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni
Published and Grey Literature from PhD Candidates
This paper aims to explore models based on the extreme gradient boosting (XGBoost) approach for business risk classification. Feature selection (FS) algorithms and hyper-parameter optimizations are simultaneously considered during model training. The five most commonly used FS methods including weight by Gini, weight by Chi-square, hierarchical variable clustering, weight by correlation, and weight by information are applied to alleviate the effect of redundant features. Two hyper-parameter optimization approaches, random search (RS) and Bayesian tree-structuredParzen Estimator (TPE), are applied in XGBoost. The effect of different FS and hyper-parameter optimization methods on the model performance are investigated by the Wilcoxon Signed Rank …
Is The Reliability Of Objective Originality Scores Confounded By Elaboration?, Shannon Marie Maio
Is The Reliability Of Objective Originality Scores Confounded By Elaboration?, Shannon Marie Maio
Electronic Theses and Dissertations
The increased use of text-mining models as a scoring mechanism for divergent thinking (DT) tasks has sparked concerns about the ways in which automated Originality scores may be influenced by other dimensions of DT, especially Elaboration. The debate centers around the question of whether too much variance in automated Originality scores is accounted for by the number of words a participant uses in a response (i.e., Elaboration), and, thus, how the influence of Elaboration can affect the reliability of Originality scores. Here, a partial correlation analysis, in conjunction with text-mining and psychometric modeling, is conducted to test the degree to …
Comprehensive Research Synthesis: An Approach To Mixed Methods Research Syntheses, Lilian Linialy Chimuma
Comprehensive Research Synthesis: An Approach To Mixed Methods Research Syntheses, Lilian Linialy Chimuma
Electronic Theses and Dissertations
Mixed methods research synthesis (MMRS) is an emerging application of both mixed-methods research (MMR) and review research. MMRS promises to comprehensively address intricate contemporary research and evaluation questions given diverse evidence sources (across quantitative, qualitative, and MMR primary studies). The significance of concurrently addressing methodological issues for new research developments is widely noted in the literature. Current efforts attempt to streamline methodological practices along with application of the MMRS approach. Researchers have proposed conceptual frameworks to guide the application and practice of MMRS studies. Despite these efforts, complications and disagreements persist. In response to these concerns, this study developed a …
Shrinkage Priors For Isotonic Probability Vectors And Binary Data Modeling, Philip S. Boonstra, Daniel R. Owen, Jian Kang
Shrinkage Priors For Isotonic Probability Vectors And Binary Data Modeling, Philip S. Boonstra, Daniel R. Owen, Jian Kang
The University of Michigan Department of Biostatistics Working Paper Series
This paper outlines a new class of shrinkage priors for Bayesian isotonic regression modeling a binary outcome against a predictor, where the probability of the outcome is assumed to be monotonically non-decreasing with the predictor. The predictor is categorized into a large number of groups, and the set of differences between outcome probabilities in consecutive categories is equipped with a multivariate prior having support over the set of simplexes. The Dirichlet distribution, which can be derived from a normalized cumulative sum of gamma-distributed random variables, is a natural choice of prior, but using mathematical and simulation-based arguments, we show that …
A Bivariate Life Distribution And Notions Of Negative Dependence, Prajamitra Bhuyan, Shyamal Ghosh, Priyanka Majumder, Murari Mitra
A Bivariate Life Distribution And Notions Of Negative Dependence, Prajamitra Bhuyan, Shyamal Ghosh, Priyanka Majumder, Murari Mitra
Journal Articles
Bivariate life distributions used to model negative dependence typically possess certain limitations; in particular, the correlation coefficient takes values in a restricted subrange of (Formula presented.). We construct a new bivariate life distribution to remedy this. Properties of the proposed distribution are studied. It is shown that the distribution satisfies most of the popular notions of negative dependence prevalent in the literature. Stress–strength reliability bounds are obtained, and parameter estimation methodology has been discussed. Performance of the estimators are compared through a simulation study.
Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu
Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu
College of Graduate Studies: Theses & Dissertations
The misclassification simulation extrapolation (MC-SIMEX) method proposed by Küchenho et al. is a general method of handling categorical data with measurement error. It consists of two steps, the simulation and extrapolation steps. In the simulation step, it simulates observations with varying degrees of measurement error. Then parameter estimators for varying degrees of measurement error are obtained based on these observations. In the extrapolation step, it uses a parametric extrapolation function to obtain the parameter estimators for data with no measurement error. However, as shown in many studies, the parameter estimators are still biased as a result of the parametric extrapolation …
Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil
Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil
College of Graduate Studies: Theses & Dissertations
Modern research strategies rely predominantly on three steps, data collection, data analysis, and inference. In research, if the data is not collected as designed, researchers may face challenges of having incomplete data, especially when it is non-ignorable. These situations affect the subsequent steps of evaluation and make them difficult to perform. Inference with incomplete data is a challenging task in data analysis and clinical trials when missing data related to the condition under the study. Moreover, results obtained from incomplete data are prone to biases. Parameter estimation with non-ignorable missing data is even more challenging to handle and extract useful …
Webexpo : Vers Une Meilleure Interprétation Des Mesures D'Exposition Professionnelle Aux Substances Chimiques Sur Les Lieux De Travail, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc
Webexpo : Vers Une Meilleure Interprétation Des Mesures D'Exposition Professionnelle Aux Substances Chimiques Sur Les Lieux De Travail, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc
Rapports de recherche scientifique
Une grande partie de l'activité des hygiénistes du travail consiste à mesurer les niveaux d'exposition professionnelle des travailleurs. La plupart des rapports d'évaluation de l'exposition font état d'une importante variabilité spatiale et temporelle quant à l'intensité de l'exposition, laquelle fluctue souvent du simple au décuple en dépit de conditions apparemment similaires. Il en résulte depuis toujours un défi de taille en ce qui concerne l'interprétation des niveaux mesurés par rapport aux valeurs limites d’exposition professionnelle (VLEP). Il existe désormais un cadre consensuel – issu d'une élaboration progressive au cours des deux dernières décennies – concernant l'analyse des niveaux d’exposition par …
Webexpo: Towards A Better Interpretation Of Measurements Of Occupational Exposure To Chemicals In The Workplace, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc
Webexpo: Towards A Better Interpretation Of Measurements Of Occupational Exposure To Chemicals In The Workplace, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc
Rapports de recherche scientifique
Une grande partie de l'activité des hygiénistes du travail consiste à mesurer les niveaux d'exposition professionnelle des travailleurs. La plupart des rapports d'évaluation de l'exposition font état d'une importante variabilité spatiale et temporelle quant à l'intensité de l'exposition, laquelle fluctue souvent du simple au décuple en dépit de conditions apparemment similaires. Il en résulte depuis toujours un défi de taille en ce qui concerne l'interprétation des niveaux mesurés par rapport aux valeurs limites d’exposition professionnelle (VLEP). Il existe désormais un cadre consensuel – issu d'une élaboration progressive au cours des deux dernières décennies – concernant l'analyse des niveaux d’exposition par …
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Theses and Dissertations
Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …
Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii
Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii
Theses and Dissertations
Although advances in modern computational algorithms have provided researchers the ability to work problems which were once too computationally complex to solve, problems with high computation or large parameter spaces still remain. Problems such as those involving Time Series can be such problems. Chapter 1 looks at the the use of Exponentially Weighted Moving Averages developed by \citep{holt2004forecasting, winters1960forecasting} which were thought to provide sufficient solutions to these Time Series. A discussion is provided which illustrates the shortcomings of the EWMA and how its infinite number of possible starting values provides the modeler with an endless number of possible solutions …
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Theses and Dissertations--Statistics
Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Theses and Dissertations--Statistics
The Bayesian adjustment for confounding (BAC) is a Bayesian model averaging method to select and adjust for confounding factors when evaluating the average causal effect of an exposure on a certain outcome. We extend the BAC method to time-to-event outcomes. Specifically, the posterior distribution of the exposure effect on a time-to-event outcome is calculated as a weighted average of posterior distributions from a number of candidate proportional hazards models, weighing each model by its ability to adjust for confounding factors. The Bayesian Information Criterion based on the partial likelihood is used to compare different models and approximate the Bayes factor. …
Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui
Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui
Theses and Dissertations--Statistics
In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.
In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …
Measuring Variability In Model Performance Measures, Matthew Rutledge
Measuring Variability In Model Performance Measures, Matthew Rutledge
Theses and Dissertations--Statistics
As data become increasingly available, statisticians are confronted with both larger sample sizes and larger numbers of predictors. While both of these factors are beneficial in building better predictive models and allowing for better inference, models can become difficult to interpret and often include variables of little practical significance. This dissertation provides methods that assist model builders to better understand and select from a collection of candidate models. We study the asymptotic distribution of AIC and propose a graphical tool to assist practitioners in comparing and contrasting candidate models. Real-world examples show how this graphic might be used and a …
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Theses and Dissertations--Statistics
Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …
Moment Kernels For T-Central Subspace, Weihang Ren
Moment Kernels For T-Central Subspace, Weihang Ren
Theses and Dissertations--Statistics
The T-central subspace allows one to perform sufficient dimension reduction for any statistical functional of interest. We propose a general estimator using a third moment kernel to estimate the T-central subspace. In particular, in this dissertation we develop sufficient dimension reduction methods for the central mean subspace via the regression mean function and central subspace via Fourier transform, central quantile subspace via quantile estimator and central expectile subsapce via expectile estima- tor. Theoretical results are established and simulation studies show the advantages of our proposed methods.