Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (61)
- Applied Statistics (58)
- Statistical Models (48)
- Statistical Methodology (41)
- Social and Behavioral Sciences (39)
-
- Mathematics (31)
- Other Statistics and Probability (30)
- Life Sciences (26)
- Statistical Theory (26)
- Education (25)
- Multivariate Analysis (23)
- Applied Mathematics (20)
- Data Science (18)
- Computer Sciences (17)
- Longitudinal Data Analysis and Time Series (17)
- Sociology (16)
- Bioinformatics (15)
- Medicine and Health Sciences (15)
- Educational Assessment, Evaluation, and Research (12)
- Probability (12)
- Categorical Data Analysis (11)
- Quantitative, Qualitative, Comparative, and Historical Methodologies (11)
- Business (10)
- Psychology (9)
- Survival Analysis (9)
- Engineering (8)
- Geography (8)
- Artificial Intelligence and Robotics (7)
- Institution
- Keyword
-
- Morgridge College of Education (46)
- Research Methods and Information Science (43)
- Research Methods and Statistics (42)
- Statistics (14)
- Causal inference (8)
-
- Machine learning (8)
- Bayesian (6)
- Simulation (6)
- Propensity score (5)
- Deep learning (4)
- Education (4)
- Machine Learning (4)
- Missing data (4)
- Mixed data (4)
- Regression (4)
- Variable selection (4)
- Average treatment effect (3)
- Bioinformatics (3)
- Bootstrap (3)
- Breast cancer (3)
- Confirmatory factor analysis (3)
- Depression (3)
- Diabetes (3)
- Feature Selection (3)
- Grounded theory (3)
- Mathematics (3)
- Meta-analysis (3)
- Metabolomics (3)
- Mixed methods (3)
- Negative binomial distribution (3)
Articles 211 - 240 of 261
Full-Text Articles in Statistics and Probability
Comparison Of Two Parameter Estimation Techniques For Stochastic Models, Thomas C. Robacker
Comparison Of Two Parameter Estimation Techniques For Stochastic Models, Thomas C. Robacker
Electronic Theses and Dissertations
Parameter estimation techniques have been successfully and extensively applied to deterministic models based on ordinary differential equations but are in early development for stochastic models. In this thesis, we first investigate using parameter estimation techniques for a deterministic model to approximate parameters in a corresponding stochastic model. The basis behind this approach lies in the Kurtz limit theorem which implies that for large populations, the realizations of the stochastic model converge to the deterministic model. We show for two example models that this approach often fails to estimate parameters well when the population size is small. We then develop a …
Parental Involvement, Students' Self-Engagement, And Academic Achievement: A Structural Equation Model, Cuirong Wu
Parental Involvement, Students' Self-Engagement, And Academic Achievement: A Structural Equation Model, Cuirong Wu
Electronic Theses and Dissertations
In the ever-evolving landscape of China's education system, the gap between the rural and urban was always an important issue, not only the multifaceted interplay of parental involvement, student self-engagement, and academic performance. This study utilized a comprehensive national survey dataset through a thorough understanding of cultural nuances and educational intricacies specific to the Chinese context. Its aim was not only to decipher the underlying constructs of parental involvement and student self-engagement but also to investigate how these factors impacted academic achievement among students. The research unfolded in several stages using a representative sample of 10750 students from 112 schools …
Novel Applications Of And Extensions To Linear Regression Methods For The Biomedical And Materials Sciences., Joe Bible
Electronic Theses and Dissertations
In this work we present three topics, each of which centered on either the application or modification of various linear regression methods. Our work with respect to the “Materials Genome” project while undermined by oversimplification and data integrity issues in its early stages, provides a sound platform from which the project can proceed successfully. Building upon a growing body of knowledge around the use of Weighted Generalized Estimating Equations (WGEE), our second investigation proposes an extension to that framework intended to address the inherent bias present in the analysis of clustered longitudinal data with potentially informative cluster sizes and temporal …
Optcluster : An R Package For Determining The Optimal Clustering Algorithm And Optimal Number Of Clusters., Michael N. Sekula
Optcluster : An R Package For Determining The Optimal Clustering Algorithm And Optimal Number Of Clusters., Michael N. Sekula
Electronic Theses and Dissertations
Determining the best clustering algorithm and ideal number of clusters for a particular dataset is a fundamental difficulty in unsupervised clustering analysis. In biological research, data generated from Next Generation Sequencing technology and microarray gene expression data are becoming more and more common, so new tools and resources are needed to group such high dimensional data using clustering analysis. Different clustering algorithms can group data very differently. Therefore, there is a need to determine the best groupings in a given dataset using the most suitable clustering algorithm for that data. This paper presents the R package optCluster as an efficient …
Summary Of Survival Analysis With Sas Procedures., Derek Duane Childers 1990-
Summary Of Survival Analysis With Sas Procedures., Derek Duane Childers 1990-
Electronic Theses and Dissertations
The research conducted for this thesis was performed to summarize some of the most commonly used survival analysis techniques as well as to create one macro that will provide the solutions for these techniques. Some of the techniques that this thesis focuses on are survival and hazard functions, mean and median survival times, life table, log rank test, proportional hazards/model building, and competing risk. To further analyze these survival analysis techniques I will use the Bone Marrow Transplantation for Leukemia dataset. This trial consists of either acute myelocytic leukemia (AML 99 patients) or acute lymphoblastic leukemia (ALL 38 patients). There …
A Simulation-Based Task Analysis Using Agent-Based, Discrete Event And System Dynamics Simulation, Anastasia Angelopoulou
A Simulation-Based Task Analysis Using Agent-Based, Discrete Event And System Dynamics Simulation, Anastasia Angelopoulou
Electronic Theses and Dissertations
Recent advances in technology have increased the need for using simulation models to analyze tasks and obtain human performance data. A variety of task analysis approaches and tools have been proposed and developed over the years. Over 100 task analysis methods have been reported in the literature. However, most of the developed methods and tools allow for representation of the static aspects of the tasks performed by expert system-driven human operators, neglecting aspects of the work environment, i.e. physical layout, and dynamic aspects of the task. The use of simulation can help face the new challenges in the field of …
Mahalanobis Kernel-Based Support Vector Data Description For Detection Of Large Shifts In Mean Vector, Vu Nguyen
Electronic Theses and Dissertations
Statistical process control (SPC) applies the science of statistics to various process control in order to provide higher-quality products and better services. The K chart is one among the many important tools that SPC offers. Creation of the K chart is based on Support Vector Data Description (SVDD), a popular data classifier method inspired by Support Vector Machine (SVM). As any methods associated with SVM, SVDD benefits from a wide variety of choices of kernel, which determines the effectiveness of the whole model. Among the most popular choices is the Euclidean distance-based Gaussian kernel, which enables SVDD to obtain a …
Assessing The Social And Ecological Factors That Influence Childhood Overweight And Obesity, Katie Callahan
Assessing The Social And Ecological Factors That Influence Childhood Overweight And Obesity, Katie Callahan
Electronic Theses and Dissertations
The prevalence of childhood overweight and obesity is increasing at an alarming rate in the United States. Currently more than 1 in 3 children aged 2-19 are overweight or obese. This is of major concern because childhood overweight and obesity leads to chronic conditions such as type II diabetes and tracks into adulthood, where more severe adverse health outcomes arise. In this study I used the premise of the social ecological model (SEM) to analyze the common levels that a child is exposed to daily; the intrapersonal level, the interpersonal level, the school level, and the community level to better …
Penalized Regressions For Variable Selection Model, Single Index Model And An Analysis Of Mass Spectrometry Data., Yubing Wan
Electronic Theses and Dissertations
The focus of this dissertation is to develop statistical methods, under the framework of penalized regressions, to handle three different problems. The first research topic is to address missing data problem for variable selection models including elastic net (ENet) method and sparse partial least squares (SPLS). I proposed a multiple imputation (MI) based weighted ENet (MI-WENet) method based on the stacked MI data and a weighting scheme for each observation. Numerical simulations were implemented to examine the performance of the MIWENet method, and compare it with competing alternatives. I then applied the MI-WENet method to examine the predictors for the …
Analyses Of 2002-2013 China’S Stock Market Using The Shared Frailty Model, Chao Tang
Analyses Of 2002-2013 China’S Stock Market Using The Shared Frailty Model, Chao Tang
Electronic Theses and Dissertations
This thesis adopts a survival model to analyze China’s stock market. The data used are the capitalization-weighted stock market index (CSI 300) and the 300 stocks for creating the index. We define the recurrent events using the daily return of the selected stocks and the index. A shared frailty model which incorporates the random effects is then used for analyses since the survival times of individual stocks are correlated. Maximization of penalized likelihood is presented to estimate the parameters in the model. The covariates are selected using the Akaike information criterion (AIC) and the variance inflation factor (VIF) to avoid …
Some New Probability Distributions Based On Random Extrema And Permutation Patterns, Jie Hao
Some New Probability Distributions Based On Random Extrema And Permutation Patterns, Jie Hao
Electronic Theses and Dissertations
In this paper, we study a new family of random variables, that arise as the distribution of extrema of a random number N of independent and identically distributed random variables X1,X2, ..., XN, where each Xi has a common continuous distribution with support on [0,1]. The general scheme is first outlined, and SUG and CSUG models are introduced in detail where Xi is distributed as U[0,1]. Some features of the proposed distributions can be studied via its mean, variance, moments and moment-generating function. Moreover, we make some other choices for …
Comparison Of Different Methods For Estimating Log-Normal Means, Qi Tang
Comparison Of Different Methods For Estimating Log-Normal Means, Qi Tang
Electronic Theses and Dissertations
The log-normal distribution is a popular model in many areas, especially in biostatistics and survival analysis where the data tend to be right skewed. In our research, a total of ten different estimators of log-normal means are compared theoretically. Simulations are done using different values of parameters and sample size. As a result of comparison, ``A degree of freedom adjusted" maximum likelihood estimator and Bayesian estimator under quadratic loss are the best when using the mean square error (MSE) as a criterion. The ten estimators are applied to a real dataset, an environmental study from Naval Construction Battalion Center (NCBC), …
Are Highly Dispersed Variables More Extreme? The Case Of Distributions With Compact Support, Benedict E. Adjogah
Are Highly Dispersed Variables More Extreme? The Case Of Distributions With Compact Support, Benedict E. Adjogah
Electronic Theses and Dissertations
We consider discrete and continuous symmetric random variables X taking values in [0; 1], and thus having expected value 1/2. The main thrust of this investigation is to study the correlation between the variance, Var(X) of X and the value of the expected maximum E(Mn) = E(X1,...,Xn) of n independent and identically distributed random variables X1,X2,...,Xn, each distributed as X. Many special cases are studied, some leading to very interesting alternating sums, and some progress is made towards a general theory.
Patient Rule Induction Method For Subgroup Identification Given Censored Data., Patrick James Trainor
Patient Rule Induction Method For Subgroup Identification Given Censored Data., Patrick James Trainor
Electronic Theses and Dissertations
The identification of subgroups in clinical studies is an important aspect of personalized medicine. In order to develop tailored therapeutics, the factors that characterize subgroups with differential prognosis, response to treatment, and incidence of adverse events or toxicities must be elucidated. We present a generalization of a statistical learning algorithm, Patient Rule Induction Method (PRIM), that is well suited for this task given a right-censored time-to-event outcome measure. This algorithm works to recursively partition a covariate space into mutually exclusive boxes that can be utilized to define subgroups. Conceptually the algorithm is similar to classification and regression trees but rather …
Statistical Methods For Assessing Treatment Effects For Observational Studies., Kristopher C. Gardner 1984-
Statistical Methods For Assessing Treatment Effects For Observational Studies., Kristopher C. Gardner 1984-
Electronic Theses and Dissertations
Though randomized clinical (RCTs) trials are the gold standard for comparing treatments, they are often infeasible or exclude clinically important subjects, or generally represent an idealized medical setting rather than real practice. Observational data provide an opportunity to study practice-based evidence, but also present challenges for analysis. Traditional statistical methods which are suitable for RCTs may be inadequate for the observational studies. In this project, four of the most popular statistical methods for observational studies: ANCOVA, propensity score matching, regression with the propensity score as a covariate, and instrumental variables (IV) are investigated through application to MarketScan insurance claims data. …
Time Series Decomposition Using Singular Spectrum Analysis, Cheng Deng
Time Series Decomposition Using Singular Spectrum Analysis, Cheng Deng
Electronic Theses and Dissertations
Singular Spectrum Analysis (SSA) is a method for decomposing and forecasting time series that recently has had major developments but it is not yet routinely included in introductory time series courses. An international conference on the topic was held in Beijing in 2012. The basic SSA method decomposes a time series into trend, seasonal component and noise. However there are other more advanced extensions and applications of the method such as change-point detection or the treatment of multivariate time series. The purpose of this work is to understand the basic SSA method through its application to the monthly average sea …
Statistical Analysis Of Depression And Social Support Change In Arab Immigrant Women In Usa, Hazhar Blbas
Statistical Analysis Of Depression And Social Support Change In Arab Immigrant Women In Usa, Hazhar Blbas
Electronic Theses and Dissertations
Arab Muslim immigrant women encounter many stressors and are at risk for depression. Social supports from husbands, family and friends are generally considered mitigating resources for depression. However, changes in social support over time and the effects of such supports on depression at a future time period have not been fully addressed in the literature This thesis investigated the relationship between demographic characteristics, changes in social support, and depression in Arab Muslim immigrant women to the USA. A sample of 454 married Arab Muslim immigrant women provided demographic data, scores on social support variables and depression at three time periods …
How Many Are Out There? A Novel Approach For Open And Closed Systems, Zia Rehman
How Many Are Out There? A Novel Approach For Open And Closed Systems, Zia Rehman
Electronic Theses and Dissertations
We propose a ratio estimator to determine population estimates using capture-recapture sampling. It's different than traditional approaches in the following ways: (1) Ordering of recaptures: Currently data sets do not take into account the "ordering" of the recaptures, although this crucial information is available to them at no cost. (2) Dependence of trials and cluster sampling: Our model explicitly considers trials to be dependent and improves existing literature which assumes independence. (3) Rate of convergence: The percentage sampled has an inverse relationship with population size, for a chosen degree of accuracy. (4) Asymptotic Attainment of Minimum Variance (Open Systems: (=population …
Compound Identification Using Penalized Linear Regression., Ruiqi Liu
Compound Identification Using Penalized Linear Regression., Ruiqi Liu
Electronic Theses and Dissertations
In this study, we propose a new method for compound identification using penalized linear regression. Compound identification is often achieved by matching the experimental mass spectra to the mass spectra stored in a reference library based on mass spectral similarity. In the context of the linear regression, the response variable is an experimental mass spectrum (i.e., query) and all the compounds in the reference library are the independent variables. However, the number of compounds in the reference library is much larger than the range of m/z values so that the data become high dimensional data with suffering from singularity. For …
Level Crossing Times In Mathematical Finance, Ofosuhene Osei
Level Crossing Times In Mathematical Finance, Ofosuhene Osei
Electronic Theses and Dissertations
Level crossing times and their applications in finance are of importance, given certain threshold levels that represent the "desirable" or "sell" values of a stock. In this thesis, we make use of Wald's lemmas and various deep results from renewal theory, in the context of finance, in modelling the growth of a portfolio of stocks. Several models are employed .
Sparse Ridge Fusion For Linear Regression, Nozad Mahmood
Sparse Ridge Fusion For Linear Regression, Nozad Mahmood
Electronic Theses and Dissertations
For a linear regression, the traditional technique deals with a case where the number of observations n more than the number of predictor variables p (n > p). In the case n < p, the classical method fails to estimate the coefficients. A solution of the problem is the case of correlated predictors is provided in this thesis. A new regularization and variable selection is proposed under the name of Sparse Ridge Fusion (SRF). In the case of highly correlated predictor, the simulated examples and a real data show that the SRF always outperforms the lasso, eleastic net, and the S-Lasso, and the results show that the SRF selects more predictor variables than the sample size n while the maximum selected variables by lasso is n size.
Two Methodologies: How Well Can Universities Predict Retention, Tiffany Lynette Gregory
Two Methodologies: How Well Can Universities Predict Retention, Tiffany Lynette Gregory
Electronic Theses and Dissertations
Student retention has been a long standing focus in higher education research with one of the earliest work dating back to 1937. Many researchers have proposed factors that affect a student's decision to depart from the university without successfully completing a degree. It is important to not only research different attributes and characteristics that affect student departure but it is also important to study different statistical methodologies. With the advancement in technology, new methodologies such as the Classification and Regression Tree (CART) have proven to yield significant results in a variety of research fields. As these new statistical methodologies emerge, …
Comparison Of Time Series And Functional Data Analysis For The Study Of Seasonality., Jake Allen
Comparison Of Time Series And Functional Data Analysis For The Study Of Seasonality., Jake Allen
Electronic Theses and Dissertations
Classical time series analysis has well known methods for the study of seasonality. A more recent method of functional data analysis has proposed phase-plane plots for the representation of each year of a time series. However, the study of seasonality within functional data analysis has not been explored extensively. Time series analysis is first introduced, followed by phase-plane plot analysis, and then compared by looking at the insight that both methods offer particularly with respect to the seasonal behavior of a variable. Also, the possible combination of both approaches is explored, specifically with the analysis of the phase-plane plots. The …
Solving The Differential Equation For The Probit Function Using A Variant Of The Carleman Embedding Technique., Kelechukwu Iroajanma Alu
Solving The Differential Equation For The Probit Function Using A Variant Of The Carleman Embedding Technique., Kelechukwu Iroajanma Alu
Electronic Theses and Dissertations
The probit function is the inverse of the cumulative distribution function associated with the standard normal distribution. It is of great utility in statistical modelling. The Carleman embedding technique has been shown to be effective in solving first order and, less efficiently, second order nonlinear differential equations. In this thesis, we show that solutions to the second order nonlinear differential equation for the probit function can be approximated efficiently using a variant of the Carleman embedding technique.
Power Analysis For Alternative Tests For The Equality Of Means., Haiyin Li
Power Analysis For Alternative Tests For The Equality Of Means., Haiyin Li
Electronic Theses and Dissertations
The two sample t-test is the test usually taught in introductory statistics courses to test for the equality of means of two populations. However, the t-test is not the only test available to compare the means of two populations. The randomization test is being incorporated into some introductory courses. There is also the bootstrap test. It is also not uncommon to decide the equality of the means based on confidence intervals for the means of these two populations. Are all those methods equally powerful? Can the idea of non-overlapping t confidence intervals be extended to bootstrap confidence intervals? The powers …
A Comparative Study Between The Standards Of Learning And In-Class Grades., Randetta Lynn Fuller
A Comparative Study Between The Standards Of Learning And In-Class Grades., Randetta Lynn Fuller
Electronic Theses and Dissertations
We examined the Standards of Learning mathematics scores and in-class grades for a rural Virginia county public school system. We looked at third, fourth, fifth, sixth, and seventh grades as well as Algebra I, Algebra II, and Geometry classes. The purpose of this was to determine whether or not there is a strong correlation between the Standards of Learning and the students' in-class grades. Had a strong enough correlation between the Standards of Learning and in-class grades been found we would have used only the in-class grades to predict the Standard of Learning test scores. However, we found that the …
Spousal Concordance In Academic Achievements And Intelligence And Family-Based Association Studies Identified Novel Loci Associated With Intelligence., Yue Pan
Electronic Theses and Dissertations
Assortative Mating, the tendency for mate selection to occur on the basis of similar traits, plays an essential role in understanding the genetic variation on academic achievements and intelligence (IQ). It is an important mechanism explaining spousal concordance. We used principal component analysis (PCA) for spousal correlation. There is a significant positive correlation between spouses by the new variable PC1 (correlation coefficient=0.515, p<0.0001). We further research the genetic factor that affects IQ by using the same data. We performed a low density genome-wide association (GWA) analysis with a family-based association test to identify genetic variants that associated with intelligence as measured by WAIS full-score IQ (FSIQ). NTM at 11q25 (rs411280, p=0.000764) and NR3C2 at 4q31.23 (rs3846329, p=0.000675) were 2 novel genes that haven't been associated with IQ from other studies. This study may serve as a resource for replication in other populations and a foundation for future investigations.
Early Stopping Of A Neural Network Via The Receiver Operating Curve., Daoping Yu
Early Stopping Of A Neural Network Via The Receiver Operating Curve., Daoping Yu
Electronic Theses and Dissertations
This thesis presents the area under the ROC (Receiver Operating Characteristics) curve, or abbreviated AUC, as an alternate measure for evaluating the predictive performance of ANNs (Artificial Neural Networks) classifiers. Conventionally, neural networks are trained to have total error converge to zero which may give rise to over-fitting problems. To ensure that they do not over fit the training data and then fail to generalize well in new data, it appears effective to stop training as early as possible once getting AUC sufficiently large via integrating ROC/AUC analysis into the training process. In order to reduce learning costs involving the …
Item Order Effects On Attitude Measures, Pei-Hua Chen
Item Order Effects On Attitude Measures, Pei-Hua Chen
Electronic Theses and Dissertations
The purpose of this dissertation was to examine the effects of altered item order on attitude measures for both computerized adaptive and conventional survey formats. Based on items modified from a dissertation/thesis completion survey (Green & Kluever, 1997) with three scales, three survey versions were generated with items ordered by difficulty as hard-to-easy (H-E), easy-to-hard (E-H), and five medium trait level items presented first followed by randomly ordered items (M-R) for conventional survey format. Significant differences in item difficulty and item discrimination were found for two of the three scales. Differences in scale reliability were detected for the procrastination and …
Estimating The Difference Of Percentiles From Two Independent Populations., Romual Eloge Tchouta
Estimating The Difference Of Percentiles From Two Independent Populations., Romual Eloge Tchouta
Electronic Theses and Dissertations
We first consider confidence intervals for a normal percentile, an exponential percentile and a uniform percentile. Then we develop confidence intervals for a difference of percentiles from two independent normal populations, two independent exponential populations and two independent uniform populations. In our study, we mainly focus on the maximum likelihood to develop our confidence intervals. The efficiency of this method is examined via coverage rates obtained in a simulation study done with the statistical software R.