Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (162)
- Applied Statistics (91)
- Engineering (86)
- Statistical Models (59)
- Operations Research, Systems Engineering and Industrial Engineering (47)
-
- Design of Experiments and Sample Surveys (41)
- Operational Research (35)
- Statistical Methodology (34)
- Social and Behavioral Sciences (32)
- Life Sciences (29)
- Applied Mathematics (22)
- Medicine and Health Sciences (22)
- Business (20)
- Longitudinal Data Analysis and Time Series (19)
- Data Science (17)
- Multivariate Analysis (17)
- Computer Sciences (16)
- Mathematics (14)
- Survival Analysis (14)
- Environmental Sciences (13)
- Electrical and Computer Engineering (12)
- Aviation (11)
- Psychology (11)
- Oceanography and Atmospheric Sciences and Meteorology (10)
- Arts and Humanities (9)
- Bioinformatics (9)
- Categorical Data Analysis (9)
- Genetics and Genomics (9)
- Institution
- Keyword
-
- Bayesian (21)
- Statistics (20)
- Physical Sciences and Mathematics, Statistics and Probability (12)
- Machine learning (11)
- Simulation (11)
-
- Biostatistics (9)
- Classification (9)
- Survival analysis (8)
- Design of experiments (7)
- Semiparametric (7)
- Bayesian Hierarchical Models (6)
- Bayesian methods (6)
- Goodness-of-fit tests (6)
- Longitudinal data (6)
- Markov chain Monte Carlo (6)
- Maximum likelihood (6)
- Physical Sciences and Mathematics, Statistics (6)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Variable selection (6)
- #antcenter (5)
- Clinical trial (5)
- Clustering (5)
- EM algorithm (5)
- Experimental design (5)
- Gene expression (5)
- Gibbs sampling (5)
- MCMC (5)
- Neural networks (5)
- Optimization (5)
- Parameter estimation (5)
- Publication Year
Articles 181 - 210 of 565
Full-Text Articles in Statistics and Probability
Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg
Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg
Theses and Dissertations
Dropout is a common issue in randomized controlled clinical trials and can negatively impact the internal validity of a study and potentially bias the treatment effect. When subjects discontinue study participation, they are not being given the opportunity to gain from the investigational therapy as if they had remained in the study, defeating one of the main purposes of clinical trials, providing treatment. Specifically in unblinded studies, such as the wait-list control (WLC) design, dropout is often due to group membership. Subjects allocated to the control group often dropout at higher rates than in the treatment group. Adaptive designs have …
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Theses and Dissertations
Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …
Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii
Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii
Theses and Dissertations
Although advances in modern computational algorithms have provided researchers the ability to work problems which were once too computationally complex to solve, problems with high computation or large parameter spaces still remain. Problems such as those involving Time Series can be such problems. Chapter 1 looks at the the use of Exponentially Weighted Moving Averages developed by \citep{holt2004forecasting, winters1960forecasting} which were thought to provide sufficient solutions to these Time Series. A discussion is provided which illustrates the shortcomings of the EWMA and how its infinite number of possible starting values provides the modeler with an endless number of possible solutions …
Utilizing Design Structure For Improving Design Selection And Analysis, Ahlam Ali Alzharani
Utilizing Design Structure For Improving Design Selection And Analysis, Ahlam Ali Alzharani
Theses and Dissertations
Recent work has shown that the structure for design plays a role in the simplicity or complexity of data analysis. To increase the knowledge of research in these areas, this dissertation aims to utilize design structure for improving design selection and analysis. In this regard, minimal dependent sets and block diagonal structure are both important concepts that are relevant to the orthogonality of the columns of a design. We are interested in finding ways to improve the data analysis especially for active effect detection by utilizing minimal dependent sets and block diagonal structure for design.
We introduce a new classification …
An Empirical Comparison Of Machine Learning Models For Classification, Nubaira Rizvi
An Empirical Comparison Of Machine Learning Models For Classification, Nubaira Rizvi
Theses and Dissertations
Classification problems are tackled across various industries throughout multiple disciplines. A model used for classification attempts to predict the class of an outcome variable based on some predictors. There are number of classification models available. But as the underlying population distribution of the predictors is always unknown it is difficult to know which model fits the situation best. Several studies have been done on which supervised model performs better given specific datasets. But little work has been done to compare the models’ performance for predicting one or more outcomes under multivariate settings.
This study compares the performance of seven popular …
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Theses and Dissertations
Humans are exposed to multiple chemicals every day. Epidemiological studies have shown that chemical mixtures are associated with cancers, allergies, neurodevelopmental disorders, and other adverse health effects. To assess these associations, investigators are increasingly using chemical mixture approaches like weighted quantile sum (WQS) regression. In these studies, the research objectives are to determine whether a mixture of correlated chemicals is associated with an adverse health outcome and to identify the important chemicals. However, as experimental equipment measures each exposure to a chemical-specific detection limit, the exposures are unknown between zero and the detection limit. Indeed, the number of exposures below …
Natural Lead-In Approaches To Response-Adaptive Allocation In Clinical Trials, Erin E. Donahue
Natural Lead-In Approaches To Response-Adaptive Allocation In Clinical Trials, Erin E. Donahue
Theses and Dissertations
Response-adaptive (RA) allocation designs can be implemented in clinical trials to skew the allocation of incoming subjects toward the better performing treatment group based on the previously accrued subjects' responses. These designs alleviate potential ethical concerns of equally allocating subjects in a trial when one treatment arm is inferior. The RA design can be generalized to include covariate information in the covariate-adjusted response-adaptive (CARA) design, which aims to maximize treatment successes conditional on a set of patient characteristics. While RA and CARA designs can improve the treatment of patients, they have unstable estimators and increased variability in early stages of …
Phenotype Extraction: Estimation And Biometrical Genetic Analysis Of Individual Dynamics, Kevin L. Mckee
Phenotype Extraction: Estimation And Biometrical Genetic Analysis Of Individual Dynamics, Kevin L. Mckee
Theses and Dissertations
Within-person data can exhibit a virtually limitless variety of statistical patterns, but it can be difficult to distinguish meaningful features from statistical artifacts. Studies of complex traits have previously used genetic signals like twin-based heritability to distinguish between the two. This dissertation is a collection of studies applying state-space modeling to conceptualize and estimate novel phenotypic constructs for use in psychiatric research and further biometrical genetic analysis. The aims are to: (1) relate control theoretic concepts to health-related phenotypes; (2) design statistical models that formally define those phenotypes; (3) estimate individual phenotypic values from time series data; (4) consider hierarchical …
Estimating Response Status In Sequential Multiple Assignment (Smar)-Like Trials, Keighly Bradbrook
Estimating Response Status In Sequential Multiple Assignment (Smar)-Like Trials, Keighly Bradbrook
Theses and Dissertations
Sequential, multiple assignment, randomized trials (SMARTs) allow investigators to develop and compare experimental treatment regimens in which individuals are successively randomized to different treatments based on some set of predetermined rules. The rules used to make decisions on when and how to switch treatments are based on a chosen set of tailoring variables. Although not always true, intermediate response is commonly used as the primary tailoring variable as it is often predictive of future treatment success. As such, successful implementation depends on identifying patients who respond to treatment, though in some situations such mechanisms may not exist. Further, patient-level covariates …
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
Theses and Dissertations
Diversity of intrinsic neural attributes and network connections is known to exist in many areas of the brain and is thought to significantly affect neural coding. Recent theoretical and experimental work has argued that in uncoupled networks, coding is most accurate at intermediate levels of heterogeneity. I explore this phenomenon through two distinct approaches: a theoretical mathematical modeling approach and a data-driven statistical modeling approach.
Through the mathematical approach, I examine firing rate heterogeneity in a feedforward network of stochastic neural oscillators utilizing a high-dimensional model. The firing rate heterogeneity stems from two sources: intrinsic (different individual cells) and network …
The Epsilon-Skew Rayleigh Distribution, By John Greene, John M. Greene
The Epsilon-Skew Rayleigh Distribution, By John Greene, John M. Greene
Theses and Dissertations
In this dissertation, a new family of skew distributions is introduced and developed, the Epsilon Skew Rayleigh. The members of this family are bimodal skewed distributions with location, scale and skewness parameters. There exist two unimodal parameter cases. The distribution can be skewed or symmetric. This distribution family has many applications including population demographics, signal dynamics, ocean wave heights and hardware failure rates. The effects of the parameters are described and developed. We derive the moment generating and maximum likelihood functions, as well as the expected value, median, modes, variance, skewness and kurtosis. The properties of a random variable with …
Time Series Analysis Of Weather Data In South Carolina, Geophrey Odero
Time Series Analysis Of Weather Data In South Carolina, Geophrey Odero
Theses and Dissertations
This thesis discusses time series analysis of weather data in South Carolina for the last fifteen years (January 2003 to December 2017) for Columbia, Greenville and North Myrtle Beach. The first part presents a brief overview of different variables that are used in the analysis. That is, temperature, dew point, humidity and sea level pressure. A short discussion of time series data is also introduced. The second part is about modeling the variables. The models of choice are presented, fitted and model diagnostics is carried out. In the third part, we discuss background on climates of the cities and model …
Statistical L-Moment And L-Moment Ratio Estimation And Their Applicability In Network Analysis, Timothy S. Anderson
Statistical L-Moment And L-Moment Ratio Estimation And Their Applicability In Network Analysis, Timothy S. Anderson
Theses and Dissertations
This research centers on finding the statistical moments, network measures, and statistical tests that are most sensitive to various node degradations for the Barabási-Albert, Erdös-Rényi, and Watts-Strogratz network models. Thirty-five different graph structures were simulated for each of the random graph generation algorithms, and sensitivity analysis was undertaken on three different network measures: degree, betweenness, and closeness. In an effort to find the statistical moments that are the most sensitive to degradation within each network, four traditional moments: mean, variance, skewness, and kurtosis as well as three non-traditional moments: L-variance, L-skewness, and L-kurtosis were examined. Each of these moments were …
Sample Size Requirements And Considerations For Models To Assess Human-Machine System Performance, Jennifer S. G. Lopez
Sample Size Requirements And Considerations For Models To Assess Human-Machine System Performance, Jennifer S. G. Lopez
Theses and Dissertations
Hierarchical Linear Models (HLMs), also known as multi-level models, are an extension of multiple regression analysis and can aid in the understanding of human and machine workloads of a system. These models allow for prediction and testing in systems with hierarchies of two or more levels. The complex interrelated variability of these multi-level models exists in operational settings, such as the Air Force Distributed Common Ground System Full Motion Video (AF DCGS FMV) community which is composed of individuals (Level-1), groups (Level-2), units (Level-3), and organizations (Level-4). Through the development of sample size requirements and considerations for multi-level models, this …
Garch Modeling Of Value At Risk And Expected Shortfall Using Bayesian Model Averaging, Ismail Kheir
Garch Modeling Of Value At Risk And Expected Shortfall Using Bayesian Model Averaging, Ismail Kheir
Theses and Dissertations
This thesis conducts Value at Risk (VaR) and Expected Shortfall (ES) estimation using GARCH modeling and Bayesian Model Averaging (BMA). BMA considers multiple models weighted by some information criterion. Through BMA, this thesis finds that VaR and ES estimates can be improved through enhanced modeling of the data generation process.
Investigations On Multiple Interval Estimators, Taeho Kim
Investigations On Multiple Interval Estimators, Taeho Kim
Theses and Dissertations
Multiple interval estimation for a set of parameters is investigated. To begin, a strategy of optimization for a multiple interval estimator (MIE) is introduced. This approach allocates distinct optimized levels to individual interval estimators so that the global expected content can be minimized while the global coverage probability is still maintained at a global level. This optimal allocation is achieved by a decision theoretic procedure which consists of two global risk functions. The major part of this manuscript is devoted to two multiple interval estimation procedures. Both procedures adopt prior information added to the classical setting, but these procedures do …
Statistical Analysis Of Interval-Censored Data Subject To Additional Complications, Qiang Zheng
Statistical Analysis Of Interval-Censored Data Subject To Additional Complications, Qiang Zheng
Theses and Dissertations
Survival analysis is an important branch of statistics that studies time to event data (or survival data), in which the response variable is time to a certain event of interest. The most prominent feature of survival data is that the response is not exactly observed due to limits of the study design or nature of the event of interest. Interval-censored data are a common type of survival data and occur frequently in real life studies where subjects are examined at periodical follow ups. The response time is usually not observed, but the status of the event of interest is known …
Extension Of Risk-Based Measure Of Time-Varying Prognostic Discrimination For Survival Models, Shujie Chen
Extension Of Risk-Based Measure Of Time-Varying Prognostic Discrimination For Survival Models, Shujie Chen
Theses and Dissertations
The Cox proportional hazards (PH) model and time dependent PH model are the most popular survival models in survival analysis. The hazard discrimination summary HDS(t) proposed by Liang and Heagerty [2017] is used to evaluate the mean hazard difference between cases and controls at time t. Liang and Heagerty [2017] evaluated the discrimination performance under the PH model and time dependent PH model with right censoring.
In this thesis, first, we further investigate their method via comprehensive simulations including 1) We extend the simulation in Liang and Heagerty [2017] under the PH model by adding more scenarios such as different …
Estimation Problems For Pooled Data, Xichen Mou
Estimation Problems For Pooled Data, Xichen Mou
Theses and Dissertations
In epidemiological applications, individual specimens (e.g., blood, urine, etc.) are often pooled together to detect the presence of disease or to measure the concentration level of a specific biomarker. Due to the advantage of cost efficiency, pooled data are also seen in diverse areas such as genetics, animal ecology, and environmental science. With pooled data, individual observations are masked and new statistical methods are needed to estimate characteristics such as disease prevalence, the underlying density function of a biomarker, etc. We focus on three estimation problems for pooled data. Chapters 2 and 3 propose nonparametric estimators for the density function …
Multivariate Probit Models For Interval-Censored Failure Time Data, Yifan Zhang
Multivariate Probit Models For Interval-Censored Failure Time Data, Yifan Zhang
Theses and Dissertations
Survival analysis is an important branch of statistics that analyzes the time to event data. The events of interest can be death, disease occurrence, the failure of a machine part, etc.. One important feature of this type of data is censoring: information on time to event is not observed exactly due to loss to follow-up or non-occurrence of interested event before the trial ends. Censored data are commonly observed in clinical trials and epidemiological studies, since monitoring a person’s health over time after treatment is often required in medical or health studies. In this dissertation we focus on studying multivariate …
Cocyclic Hadamard Matrices: An Efficient Search Based Algorithm, Jonathan S. Turner
Cocyclic Hadamard Matrices: An Efficient Search Based Algorithm, Jonathan S. Turner
Theses and Dissertations
This dissertation serves as the culmination of three papers. “Counting the decimation classes of binary vectors with relatively prime fixed-density" presents the first non-exhaustive decimation class counting algorithm. “A Novel Approach to Relatively Prime Fixed Density Bracelet Generation in Constant Amortized Time" presents a novel lexicon for binary vectors based upon the Discrete Fourier Transform, and develops a bracelet generation method based upon the same. “A Novel Legendre Pair Generation Algorithm" expands upon the bracelet generation algorithm and includes additional constraints imposed by Legendre Pairs. It further presents an efficient sorting and comparison algorithm based upon symmetric functions, as well …
Mle And Bayesian Methods To Analyze Data With Missing Values Below The Limit Of Detection, Xinxin Hu
Mle And Bayesian Methods To Analyze Data With Missing Values Below The Limit Of Detection, Xinxin Hu
Theses and Dissertations
As pesticides are widely used in agriculture, more and more people who work at places like farm are exposed to the pesticides. According to enviroment re- searches [Villarejo; 2003; Reigart and Roberts; 1999], being exposed to some kind of pesticides like Organophosphorus (OP) insecticides has significantly effected the health of farmworkers and their family. The actual level of pesticides can be detected with some limitation for now. However, it is hard to detect when the level is below the limit of detection (LOD). Therefore, the goal of our research is to propose several different methods to analyze data …
Inflated Standard Errors Of Mcmc Estimates In Irt, Dongho Shin
Inflated Standard Errors Of Mcmc Estimates In Irt, Dongho Shin
Theses and Dissertations
Two widely used algorithms for estimating item response theory (IRT) parameters are Markov chain Monte Carlo (MCMC) and the EM algorithm. In general, the MCMC algorithm has advantages over the EM algorithm - for example, the MCMC algorithm allows one to estimate the desired posterior distribution and also works more straightforwardly with complex IRT models. This ease of use, allows one to implement the MCMC algorithm without carefully consideration. Previous studies, Hendrix (2011) and Lee (2016), noted that the estimated standard errors from the MCMC algorithm are larger than those from the EM algorithm. Therefore, this study investigate the reason …
Randomization Analysis Driven Software, Steph-Yves Louis
Randomization Analysis Driven Software, Steph-Yves Louis
Theses and Dissertations
The application of a method of randomization for a clinical trial frequently summarizes to using Simple Randomization. Even though the latter method provides favorable characteristics, if the collected sample is not large enough, it still presents the highest chance of imbalance both marginally in the treatment groups and locally in terms of the covariates. Methods of Permuted Block Randomization, Urn Randomization, Stratified Permuted Block Randomization, and Minimization represent popular alternative methods that one should consider depending on the goal of the study. A comparison of the previously mentioned methods is carried to evaluate their performance with samples that are not …
Spatio-Temporal Analysis Of Precipitation And Flood Data From South Carolina, Haigang Liu
Spatio-Temporal Analysis Of Precipitation And Flood Data From South Carolina, Haigang Liu
Theses and Dissertations
Spatio-temporal data are everywhere: we encounter them on TV, in newspapers, on computer screens, on tablets, and on plain paper maps. As a result, researchers in di- verse areas are increasingly faced with the task of modeling geographically-referenced and temporally-correlated data. In this dissertation, we propose two different spa- tiotemporal models to capture the behavior of rainfall and flood data in the state of South Carolina.
Both models are built using a Bayesian hierarchical framework, which involves specifying the true underlying process in the first level and the spatio-temporal ran- dom effect in the second level of the hierarchy. The …
Cluster Analysis Of Mixed-Mode Data, Yawei Liang
Cluster Analysis Of Mixed-Mode Data, Yawei Liang
Theses and Dissertations
In the modern world, data have become increasingly more complex and often contain different types of features. Two very common types of features are continuous and discrete variables. Clustering mixed-mode data, which include both continuous and discrete variables, can be done in various ways. Furthermore, a continuous variable can take any value between its minimum and maximum. Types of continuous vari- ables include bounded or unbounded normal variables, uniform variables, circular variables, etc. Discrete variables include types other than continuous variables, such as binary variables, categorical (nominal) variables, Poisson variables, etc. Difficulties in clustering mixed-mode data include handling the association …
Regression For Pooled Testing Data With Biomedical Applications, Juexin Lin
Regression For Pooled Testing Data With Biomedical Applications, Juexin Lin
Theses and Dissertations
Since first introduced by Dorfman in 1943, pooled testing has been widely used as a cost and time effective testing protocol in the variety of applications. This dis- sertation consists of three projects that reveal the use of pooling techniques in the disease prevention from the perspective of regression. For disease monitoring and control, individual covariates information are often of practical interest and yield meaningful interpretations. It is natural to model the outcome of interest, which can be either a disease status (binary) or a biomarker concentration index (continuous), with individual-specific covariates through a regression analysis. Chapter 2 focuses on …
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Theses and Dissertations
In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Theses and Dissertations
Response-Adaptive (RA) designs are used to adaptively allocate patients in clinical trials. These methods have been generalized to include Covariate-Adjusted Response-Adaptive (CARA) designs, which adjust treatment assignments for a set of covariates while maintaining features of the RA designs. Challenges may arise in multi-center trials if differential treatment responses and/or effects among sites exist. We propose Site-Adjusted Response-Adaptive (SARA) approaches to account for inter-center variability in treatment response and/or effectiveness, including either a fixed site effect or both random site and treatment-by-site interaction effects to calculate conditional probabilities. These success probabilities are used to update assignment probabilities for allocating patients …
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Theses and Dissertations
As researchers increasingly use web-based surveys, the ease of dropping out in the online setting is a growing issue in ensuring data quality. One theory is that dropout or attrition occurs in phases that can be generalized to phases of high dropout and phases of stable use. In order to detect these phases, several methods are explored. First, existing methods and user-specified thresholds are applied to survey data where significant changes in the dropout rate between two questions is interpreted as the start or end of a high dropout phase. Next, survey dropout is considered as a time-to-event outcome and …