Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 361 - 390 of 565

Full-Text Articles in Statistics and Probability

The Effect Of Smoking On Tuberculosis Incidence In Burdened Countries, Natalie Noel Ellison Mar 2012

The Effect Of Smoking On Tuberculosis Incidence In Burdened Countries, Natalie Noel Ellison

Theses and Dissertations

It is estimated that one third of the world's population is infected with tuberculosis. Though once thought a "dead" disease, tuberculosis is very much alive. The rise of drug resistant strains of tuberculosis, and TB-HIV coinfection have made tuberculosis an even greater worldwide threat. While HIV, poverty, and public health infrastructure are historically assumed to affect the burden of tuberculosis, recent research has been done to implicate smoking in this list. This analysis involves combining data from multiple sources in order determine if smoking is a statistically significant factor in predicting the number of incident tuberculosis cases in a country. …


Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel Jan 2012

Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel

Theses and Dissertations

High-throughput genomic datasets obtained from microarray or sequencing studies have revolutionized the field of molecular biology over the last decade. The complexity of these new technologies also poses new challenges to statisticians to separate biological relevant information from technical noise. Two methods are introduced that address important issues with normalization of array comparative genomic hybridization (aCGH) microarrays and the analysis of RNA sequencing (RNA-Seq) studies. Many studies investigating copy number aberrations at the DNA level for cancer and genetic studies use comparative genomic hybridization (CGH) on oligo arrays. However, aCGH data often suffer from low signal to noise ratios resulting …


Screening Designs That Minimize Model Dependence, Kenneth P. Fairchild Dec 2011

Screening Designs That Minimize Model Dependence, Kenneth P. Fairchild

Theses and Dissertations

When approaching a new research problem, we often use screening designs to determine which factors are worth exploring in more detail. Before exploring a problem, we don't know which factors are important. When examining a large number of factors, it is likely that only a handful are significant and that even fewer two-factor interactions will be significant. If there are important interactions, it is likely that they are connected with the handful of significant main effects. Since we don't know beforehand which factors are significant, we want to choose a design that gives us the highest probability a priori of …


Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson Jul 2011

Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson

Theses and Dissertations

Walmart and other “big box” stores seek to expand in rural markets, possibly due to cheap land and lack of zoning laws. In August 2000, Walmart opened a store in Ephraim, a small rural town in central Utah. It is of interest to understand how Walmart's entrance into the local market changes the sales tax revenue base for Ephraim and for the surrounding municipalities. It is thought that small “Mom and Pop” stores go out of business because they cannot compete with Walmart's prices, leading to a decrease in variety, selection, convenience, and most importantly, sales tax revenue base in …


An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey Jul 2011

An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey

Theses and Dissertations

Bayesian statistical methods have long been computationally out of reach because the analysis often requires integration of high-dimensional functions. Recent advancements in computational tools to apply Markov Chain Monte Carlo (MCMC) methods are making Bayesian data analysis accessible for all statisticians. Two such computer tools are Win-BUGS and SASR 9.2's PROC MCMC. Bayesian methodology will be introduced through discussion of fourteen statistical examples with code and computer output to demonstrate the power of these computational tools in a wide variety of settings.


Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler Jun 2011

Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler

Theses and Dissertations

University students often complete evaluations of their courses and instructors. The evaluation tool typically contains questions about the course and the instructor on an ordinal Likert scale. We assess instructor effectiveness while adjusting for known confounders. We present a probit regression model with a latent variable to measure the instructor effectiveness accounting for student specific covariates, such as student grade in the course, high school and university GPA, and ACT score.


A Bayesian Approach To Missile Reliability, Taylor Hardison Redd Jun 2011

A Bayesian Approach To Missile Reliability, Taylor Hardison Redd

Theses and Dissertations

Each year, billions of dollars are spent on missiles and munitions by the United States government. It is therefore vital to have a dependable method to estimate the reliability of these missiles. It is important to take into account the age of the missile, the reliability of different components of the missile, and the impact of different launch phases on missile reliability. Additionally, it is of importance to estimate the missile performance under a variety of test conditions, or modalities. Bayesian logistic regression is utilized to accurately make these estimates. This project presents both previously proposed methods and ways to …


Adaptive Threat Detector Testing Using Bayesian Gaussian Process Models, Bradley Thomas Ferguson May 2011

Adaptive Threat Detector Testing Using Bayesian Gaussian Process Models, Bradley Thomas Ferguson

Theses and Dissertations

Detection of biological and chemical threats is an important consideration in the modern national defense policy. Much of the testing and evaluation of threat detection technologies is performed without appropriate uncertainty quantification. This paper proposes an approach to analyzing the effect of threat concentration on the probability of detecting chemical and biological threats. The approach uses a probit semi-parametric formulation between threat concentration level and the probability of instrument detection. It also utilizes a bayesian adaptive design to determine at which threat concentrations the tests should be performed. The approach offers unique advantages, namely, the flexibility to model non-monotone curves …


Overcoming Pose Limitations Of A Skin-Cued Histograms Of Oriented Gradients Dismount Detector Through Contextual Use Of Skin Islands And Multiple Support Vector Machines, Jonathon R. Climer Mar 2011

Overcoming Pose Limitations Of A Skin-Cued Histograms Of Oriented Gradients Dismount Detector Through Contextual Use Of Skin Islands And Multiple Support Vector Machines, Jonathon R. Climer

Theses and Dissertations

This thesis provides a novel visualization method to analyze the impact that articulations in dismount pose and camera aspect angle have on histograms of oriented gradients (HOG) features and eventual detections. Insights from these relationships are used to identify limitations in a state of the art skin cued HOG dismount detector's ability to detect poses not in a standard upright stances. Improvements to detector performance are made by further leveraging available skin information, reducing false detections by an additional order of magnitude. In addition, a method is outlined for training supplemental support vector machines (SVMs) from computer generated data, for …


Variable Selection And Parameter Estimation Using A Continuous And Differentiable Approximation To The L0 Penalty Function, Douglas Nielsen Vanderwerken Mar 2011

Variable Selection And Parameter Estimation Using A Continuous And Differentiable Approximation To The L0 Penalty Function, Douglas Nielsen Vanderwerken

Theses and Dissertations

L0 penalized likelihood procedures like Mallows' Cp, AIC, and BIC directly penalize for the number of variables included in a regression model. This is a straightforward approach to the problem of overfitting, and these methods are now part of every statistician's repertoire. However, these procedures have been shown to sometimes result in unstable parameter estimates as a result on the L0 penalty's discontinuity at zero. One proposed alternative, seamless-L0 (SELO), utilizes a continuous penalty function that mimics L0 and allows for stable estimates. Like other similar methods (e.g. LASSO and SCAD), SELO produces sparse solutions because the penalty function is …


Hierarchical Bayesian Methods For Evaluation Of Traffic Project Efficacy, Andrew Nolan Olsen Mar 2011

Hierarchical Bayesian Methods For Evaluation Of Traffic Project Efficacy, Andrew Nolan Olsen

Theses and Dissertations

A main objective of Departments of Transportation is to improve the safety of the roadways over which they have jurisdiction. Safety projects, such as cable barriers and raised medians, are utilized to reduce both crash frequency and crash severity. The efficacy of these projects must be evaluated in order to use resources in the best way possible. Five models are proposed for the evaluation of traffic projects: (1) a Bayesian Poisson regression model; (2) a hierarchical Poisson regression model building on model (1) by adding hyperpriors; (3) a similar model correcting for overdispersion; (4) a dynamic linear model; and (5) …


Parameter Estimation For The Two-Parameter Weibull Distribution, Mark A. Nielsen Mar 2011

Parameter Estimation For The Two-Parameter Weibull Distribution, Mark A. Nielsen

Theses and Dissertations

The Weibull distribution, an extreme value distribution, is frequently used to model survival, reliability, wind speed, and other data. One reason for this is its flexibility; it can mimic various distributions like the exponential or normal. The two-parameter Weibull has a shape (γ) and scale (β) parameter. Parameter estimation has been an ongoing search to find efficient, unbiased, and minimal variance estimators. Through data analysis and simulation studies, the following three methods of estimation will be discussed and compared: maximum likelihood estimation (MLE), method of moments estimation (MME), and median rank regression (MRR). The analysis of wind speed data from …


Utilizing Universal Probability Of Expression Code (Upc) To Identify Disrupted Pathways In Cancer Samples, Michelle Rachel Withers Mar 2011

Utilizing Universal Probability Of Expression Code (Upc) To Identify Disrupted Pathways In Cancer Samples, Michelle Rachel Withers

Theses and Dissertations

Understanding the role of deregulated biological pathways in cancer samples has the potential to improve cancer treatment, making it more effective by selecting treatments that reverse the biological cause of the cancer. One of the challenges with pathway analysis is identifying a deregulated pathway in a given sample. This project develops the Universal Probability of Expression Code (UPC), a profile of a single deregulated biological path- way, and projects it into a cancer cell to determine if it is present. One of the benefits of this method is that rather than use information from a single over-expressed gene, it pro- …


Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini Nov 2010

Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini

Theses and Dissertations

The role of abnormal DNA methylation in the progression of disease is a growing area of research that relies upon the establishment of sound statistical methods. The common method for declaring there is differential methylation between two groups at a given CpG site, as summarized by the difference between proportions methylated db=b1-b2, has been through use of a Filtered Two Sample t-test, using the recommended filter of 0.17 (Bibikova et al., 2006b). In this dissertation, we performed a re-analysis of the data used in recommending the threshold by fitting a mixed-effects ANOVA model. It was determined that the 0.17 filter …


Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham Nov 2010

Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham

Theses and Dissertations

Over the past few decades, Cluster Randomized Trials (CRT) have become a design of choice in many research areas. One of the most critical issues in planning a CRT is to ensure that the study design is sensitive enough to capture the intervention effect. The assessment of power and sample size in such studies is often faced with many challenges due to several methodological difficulties. While studies on power and sample size for cluster designs with one and two levels are abundant, the evaluation of required sample size for three-level designs has been generally overlooked. First, the nesting effect introduces …


Stereotype Logit Models For High Dimensional Data, Andre Williams Oct 2010

Stereotype Logit Models For High Dimensional Data, Andre Williams

Theses and Dissertations

Gene expression studies are of growing importance in the field of medicine. In fact, subtypes within the same disease have been shown to have differing gene expression profiles (Golub et al., 1999). Often, researchers are interested in differentiating a disease by a categorical classification indicative of disease progression. For example, it may be of interest to identify genes that are associated with progression and to accurately predict the state of progression using gene expression data. One challenge when modeling microarray gene expression data is that there are more genes (variables) than there are observations. In addition, the genes usually demonstrate …


Assessment Of Acgh Clustering Methodologies, Serena F. Baker Oct 2010

Assessment Of Acgh Clustering Methodologies, Serena F. Baker

Theses and Dissertations

Array comparative genomic hybridization (aCGH) is a technique for identifying duplications and deletions of DNA at specific locations across a genome. Potential objectives of aCGH analysis are the identification of (1) altered regions for a given subject, (2) altered regions across a set of individuals, and (3) clinically relevant clusters of hybridizations. aCGH analysis can be particularly useful when it identifies previously unknown clusters with clinical relevance. This project focuses on the assessment of existing aCGH clustering methodologies. Three methodologies are considered: hierarchical clustering, weighted clustering of called aCGH data, and clustering based on probabilistic recurrent regions of alteration within …


Statistical Image Recovery From Laser Speckle Patterns With Polarization Diversity, Donald B. Dixon Sep 2010

Statistical Image Recovery From Laser Speckle Patterns With Polarization Diversity, Donald B. Dixon

Theses and Dissertations

This research extends the theory and understanding of the laser speckle imaging technique. This non-traditional imaging technique may be employed to improve space situational awareness and image deep space objects from a ground-based sensor system. The use of this technique is motivated by the ability to overcome aperture size limitations and the distortion effects from Earth’s atmosphere. Laser speckle imaging is a lensless, coherent method for forming two-dimensional images from their autocorrelation functions. Phase retrieval from autocorrelation data is an ill-posed problem where multiple solutions exist. This research introduces polarization diversity as a method for obtaining additional information so the …


Application Of Convex Methods To Identification Of Fuzzy Subpopulations, Ryan Lee Eliason Sep 2010

Application Of Convex Methods To Identification Of Fuzzy Subpopulations, Ryan Lee Eliason

Theses and Dissertations

In large observational studies, data are often highly multivariate with many discrete and continuous variables measured on each observational unit. One often derives subpopulations to facilitate analysis. Traditional approaches suggest modeling such subpopulations with a compilation of interaction effects. However, when many interaction effects define each subpopulation, it becomes easier to model membership in a subpopulation rather than numerous interactions. In many cases, subjects are not complete members of a subpopulation but rather partial members of multiple subpopulations. Grade of Membership scores preserve the integrity of this partial membership. By generalizing an analytic chemistry concept related to chromatography-mass spectrometry, we …


An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates Jun 2010

An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates

Theses and Dissertations

The analysis of weighted co-expression gene sets is gaining momentum in systems biology. In addition to substantial research directed toward inferring co-expression networks on the basis of microarray/high-throughput sequencing data, inferential methods are being developed to compare gene networks across one or more phenotypes. Common gene set hypothesis testing procedures are mostly confined to comparing average gene/node transcription levels between one or more groups and make limited use of additional network features, e.g., edges induced by significant partial correlations. Ignoring the gene set architecture disregards relevant network topological comparisons and can result in familiar n<


An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall May 2010

An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall

Theses and Dissertations

Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …


Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed May 2010

Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed

Theses and Dissertations

The practice of sequential testing is followed by the evaluation of accuracy, but often not by the evaluation of cost. This research described and compared three sequential testing strategies: believe the negative (BN), believe the positive (BP) and believe the extreme (BE), the latter being a less-examined strategy. All three strategies were used to combine results of two medical tests to diagnose a disease or medical condition. Descriptions of these strategies were provided in terms of accuracy (using the maximum receiver operating curve or MROC) and cost of testing (defined as the proportion of subjects who need 2 tests to …


A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber Apr 2010

A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber

Theses and Dissertations

Most studies on maturation and body composition using the Fels Longitudinal data mention peak height velocity (PHV) as an important outcome measure. The PHV is often derived from growth models such as the triple logistic model fitted to the stature (height) data. The age at PHV is sometimes ordinalized to designate an individual as an early, average or late maturer. In theory, age at PHV is the age at which the rate of growth reaches the maximum. Theoretically, for a well behaved growth function, this could be obtained by setting the second derivative of the growth function to zero and …


Parameter Estimation In Linear-Linear Segmented Regression, Erika Lyn Hernandez Apr 2010

Parameter Estimation In Linear-Linear Segmented Regression, Erika Lyn Hernandez

Theses and Dissertations

Segmented regression is a type of nonlinear regression that allows differing functional forms to be fit over different ranges of the explanatory variable. This paper considers the simple segmented regression case of two linear segments that are constrained to meet, often called the linear-linear model. Parameter estimation in the case where the joinpoint between the regimes is unknown can be tricky. Using a simulation study, four estimators for the parameters of the linear-linear model are evaluated. The bias and mean squared error of the estimators are considered under differing parameter combinations and sample sizes. Parameters estimated in the model are …


Cluster And Classification Analysis Of Fossil Invertebrates Within The Bird Spring Formation, Arrow Canyon, Nevada: Implications For Relative Rise And Fall Of Sea-Level, Scott L. Morris Apr 2010

Cluster And Classification Analysis Of Fossil Invertebrates Within The Bird Spring Formation, Arrow Canyon, Nevada: Implications For Relative Rise And Fall Of Sea-Level, Scott L. Morris

Theses and Dissertations

Carbonate strata preserve indicators of local marine environments through time. Such indicators often include microfossils that have relatively unique conditions under which they can survive, including light, nutrients, salinity, and especially water temperature. As such, microfossils are environmental proxies. When these microfossils are preserved in the rock record, they constitute key components of depositional facies. Spence et al. (2004, 2007) has proposed several approaches for determining the facies of a given stratigraphic succession based upon these proxies. Cluster analysis can be used to determine microfossil groups that represent specific environmental conditions. Identifying which microfossil groups exist through time can indicate …


Extensions Of Nearest Shrunken Centroid Method For Classification, Tomohiko Funai Mar 2010

Extensions Of Nearest Shrunken Centroid Method For Classification, Tomohiko Funai

Theses and Dissertations

Stylometry assumes that the essence of the individual style of an author can be captured using a number of quantitative criteria, such as the relative frequencies of noncontextual words (e.g., or, the, and, etc.). Several statistical methodologies have been developed for authorship analysis. Jockers et al. (2009) utilize Nearest Shrunken Centroid (NSC) classification, a promising classification methodology in DNA microarray analysis for authorship analysis of the Book of Mormon. Schaalje et al. (2010) develop an extended NSC classification to remedy the problem of a missing author. Dabney (2005) and Koppel et al. (2009) suggest other modifications of NSC. This paper …


Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly Mar 2010

Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly

Theses and Dissertations

In an ever-increasing data rich environment, actionable information must be extracted, filtered, and correlated from massive amounts of disparate often free text sources. The usefulness of the retrieved information depends on how we accomplish these steps and present the most relevant information to the analyst. One method for extracting information from free text is Latent Dirichlet Allocation (LDA), a document categorization technique to classify documents into cohesive topics. Although LDA accounts for some implicit relationships such as synonymy (same meaning) it often ignores other semantic relationships such as polysemy (different meanings), hyponym (subordinate), meronym (part of), and troponomys (manner). To …


Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway Mar 2010

Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway

Theses and Dissertations

The normal distribution is a commonly seen distribution in nature, education, and business. Data that are mounded or bell shaped are easily found across various fields of study. Although there is high utility with the normal distribution; often the full range can not be observed. The truncated normal distribution accounts for the inability to observe the full range and allows for inferring back to the original population. Depending on the amount of truncation, the truncated normal has several distinct shapes. A simulation study evaluating the performance of the maximum likelihood estimators and method of moment estimators is conducted and a …


Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi Mar 2010

Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi

Theses and Dissertations

In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …


An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger Dec 2009

An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger

Theses and Dissertations

Clinical drug trials are costly and time-consuming. Bayesian methods alleviate the inefficiencies in the testing process while providing user-friendly probabilistic inference and predictions from the sampled posterior distributions, saving resources, time, and money. We propose a dynamic linear model to estimate the mean response at each dose level, borrowing strength across dose levels. Our model permits nonmonotonicity of the dose-response relationship, facilitating precise modeling of a wider array of dose-response relationships (including the possibility of toxicity). In addition, we incorporate an adaptive approach to the design of the clinical trial, which allows for interim decisions and assignment to doses based …