Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Brigham Young University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 144

Full-Text Articles in Statistics and Probability

Development Of A Reliable Confidence Interval For The Kappa Statistic, With Application To The Semiconductor Industry, Rachelle Curtis, Dr. G. Bruce Schallje Sep 2013

Development Of A Reliable Confidence Interval For The Kappa Statistic, With Application To The Semiconductor Industry, Rachelle Curtis, Dr. G. Bruce Schallje

Journal of Undergraduate Research

In the Semiconductor industry, it is necessary to compare the performance of different test-tapes to assure that each produces similar quality work. Test tapes sort each die, or computer chip, into one of a fixed number of bins, depending on how well it functions. When validating a new test tape, several die are tested using both the old and the new test tape and the bin assignments are compared. Kappa, a statistical measure of agreement, compares the number of matching bin assignments to the number expected if the test tapes are independent. Unfortunately, the closer Kappa is to one (perfect …


A Statistical Report Of Byu Premedical Students From 2002-2007, Daniel Chan, David Kaiser Sep 2013

A Statistical Report Of Byu Premedical Students From 2002-2007, Daniel Chan, David Kaiser

Journal of Undergraduate Research

In 2006 the average number of BYU medical school applications for any one applicant was 17. With the chance of matriculation per application only 3.8%, yet the cost of each application running as high as $700 each, it is important for BYU premedical students to make informed decisions when choosing to which of 142 medical schools to apply.


Transcription Factor Binding Site Identification Using Mathematical Algorithms, Colin Rogerson, Dr. W. Evan Johnson Sep 2013

Transcription Factor Binding Site Identification Using Mathematical Algorithms, Colin Rogerson, Dr. W. Evan Johnson

Journal of Undergraduate Research

The purpose of our research project was, originally, first to identify the likely binding sites of the Estrogen Receptor transcription factor, and second to identify likely co-factors that interact with Estrogen Receptor in the binding process. We planned to do this using a computational algorithm which scanned sequences of DNA, and by utilizing the position weight matrices of many transcription factors, statistically identify likely binding sites and co-factors.


An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger, Dr. C. Shane Reese Sep 2013

An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger, Dr. C. Shane Reese

Journal of Undergraduate Research

The Food and Drug Administration is responsible for testing and screening clinical drugs before market entry. The screening process involves different phases which determine a drug’s potency, recommended dosage, and potential side effects. Current statistical designs used to model the efficacy of a drug at varying dosage levels are robust yet inefficient.


Investigation Of Three Methods Of Variable Selection: Linear Model Selection, Cart And Treed Regression, Todd R. Nelson, Dr. Scott Grimshaw Aug 2013

Investigation Of Three Methods Of Variable Selection: Linear Model Selection, Cart And Treed Regression, Todd R. Nelson, Dr. Scott Grimshaw

Journal of Undergraduate Research

My project for the Research and Creative Activities Scholarship Award looks at the effectiveness and accuracy of a popular new method of variable selection. The method under investigation uses tree-based models to select important variables and then uses classical least squares estimation to fit a regression model with the variables contained in the tree. This research tests the strengths and weaknesses of this approach to modeling as compared with more classical methods.


Modeling Life Expectancy Using Weibull Results Computer Simulation, Scott Morris, Dr. Scott Grimshaw Aug 2013

Modeling Life Expectancy Using Weibull Results Computer Simulation, Scott Morris, Dr. Scott Grimshaw

Journal of Undergraduate Research

The purpose of our project was to model a person’s life expectancy using lifetimes from their ancestry. Typical mortality models are based on data from large populations, and are not necessarily representative of a specific individual. Our primary objective was to model my personal life expectancy based on the men in my own ancestry by looking at their ages at death. The results from the study are unique to my own life, but the methodology applies to anyone who gathers their own ancestral data. Furthermore, these results can be useful in the medical field and in fnancial and retirement planning.


Just A Few More Moments: The G-And-H Distribution, Patrick A. Turley Aug 2013

Just A Few More Moments: The G-And-H Distribution, Patrick A. Turley

Journal of Undergraduate Research

Every mathematical model in the social sciences has its assumptions that may or may not be justified. These may be assumptions related to the behavior of the subjects of study (e.g. that people act in their best interests) or assumptions related to the model itself (e.g. that one variable will affect another linearly). A standard assumption in most regression models is that the error of the model is normally distributed, an assumption that doesn’t hold in most cases. In this research, the more flexible g-and-h distribution was examined as an alternative to the normal distribution.


Comparison Of Bootstrap Confidence Interval Methods In A Real Application: The Kappa Statistic In Industry, Krik O. Monson, Dr. Bruce G. Schaalje Aug 2013

Comparison Of Bootstrap Confidence Interval Methods In A Real Application: The Kappa Statistic In Industry, Krik O. Monson, Dr. Bruce G. Schaalje

Journal of Undergraduate Research

The kappa statistic is defined as the proportion of agreement between two instruments after chance agreement is removed from consideration (Cohen 1960). It is very useful when assessing the similarities of two instruments or procedures. Industries often use an automated system for testing their products to ensure a level of quality before packaging their goods. Since these systems parts must be periodically replaced, it is necessary to determine if the new components have the same level of accuracy or performance as the old ones. The kappa statistic would be very useful in this situation if a more reliable method for …


The Influence Of Rna Editing On Mirna Processing And Gene Regulation, Brent Shepherd, Dr. Evan Johnson Aug 2013

The Influence Of Rna Editing On Mirna Processing And Gene Regulation, Brent Shepherd, Dr. Evan Johnson

Journal of Undergraduate Research

The emergence of high-throughput DNA sequencing within the past decade has redefined the world of molecular research. Mapping a human genome, a project that cost the United States government almost $3 billion and took 13 years to complete, is now speculated to soon be less than a week-long procedure costing $1000 or less1,2. With high-throughput DNA sequencing machines, we can rapidly process genomic data for relatively cheap amounts of money. As a result, genomic information is becoming more widely available for many different species. In addition to DNA, high-throughput machines can also be used to sequence other nucleotide-based molecules, including …


Reverse Mortgages: A Source Of Stability In A Volatile Interest Market?, Grant Hodgson, Dr. H. Dennis Tolley Aug 2013

Reverse Mortgages: A Source Of Stability In A Volatile Interest Market?, Grant Hodgson, Dr. H. Dennis Tolley

Journal of Undergraduate Research

In reverse mortgages, homeowners over the age of 62 take a loan out based on the equity of their home. While reverse mortgages can be beneficial for both the mortgager and mortgagee, there is also a great deal of risk involved because it is possible for the value of the loan to become larger than the equity on the home, depending on interest rates and life time of the mortgage. It is important to have a model that accurately predicts the risk associated with the mortgage. Recognizing and appropriately managing risk is becoming increasingly important as the “baby boom” generation …


Identification Of Important Pathways In Cancerogenesis, Michelle Withers, Dr. W. Evan Johnson Aug 2013

Identification Of Important Pathways In Cancerogenesis, Michelle Withers, Dr. W. Evan Johnson

Journal of Undergraduate Research

The human body interacts at a cellular level through a series of events known as biological pathways. These pathways consist of communications at the molecular, cellular, and genetic levels to maintain a healthy, functional body. When the pathways are disrupted, new pathways are formed that can lead to disease. Currently, researchers are trying to understand the role of disrupted and normal pathways in cancer cells. It is necessary to understand the cancerous pathway in order to return the cell to its normal, healthy function.


A Hierarchical Bayesian Method For Evaluating Efficacy Of Traffic Accident Remediation, Andrew Olsen, Dr. C. Shane Reese Aug 2013

A Hierarchical Bayesian Method For Evaluating Efficacy Of Traffic Accident Remediation, Andrew Olsen, Dr. C. Shane Reese

Journal of Undergraduate Research

Thousands of people die in automobile crashes each year. Departments of Transportation are working continuously to reduce the number of fatalities through highway safety projects. One of their critical tasks is to evaluate the efficacy of these projects. We have developed a Bayesian hierarchical Poisson regression model that the Utah Department of Transportation may use to evaluate its traffic-accident intervention. The use of this method is illustrated with severe crashes on a raised median on University Parkway from mile 1.20 to mile 1.96.


An Adaptive Bayesian Approach To Threat-Detection Modeling, Bradley Ferguson, Dr. C. Shane Reese Aug 2013

An Adaptive Bayesian Approach To Threat-Detection Modeling, Bradley Ferguson, Dr. C. Shane Reese

Journal of Undergraduate Research

Detection of biological and chemical threats is an important consideration in the modern national defense policy. Much of the testing and evaluation of threat detection technologies is performed without appropriate uncertainty quantification. Under the guidance of Dr. Shane Reese, my ORCA project dealt with developing a more effective and cost-efficient way of testing threat detection technologies. I utilized a Bayesian Gaussian Process model that allows for a more flexible and robust model t. I also developed an adaptive experimental design scheme that provides more information than a typical experimental design by learning the concentration levels that we are more interested …


Bayesian Semi-Parametric Modeling Of Functional Data Exploration Into Major League Baseball Analytics, Jared Fisher, Dr. Gilbert Fellingham Jun 2013

Bayesian Semi-Parametric Modeling Of Functional Data Exploration Into Major League Baseball Analytics, Jared Fisher, Dr. Gilbert Fellingham

Journal of Undergraduate Research

Most measurements follow trends over time, and those trends can be modeled. While there are many techniques for doing this, this project’s model brings a unique angle. This method can model trends with multiple peaks, from different subjects, and group them in clusters of similar curves. This permits inference on behalf of the scientist as to what is similar between subjects within a group. We have applied this algorithm to data from Major League Baseball, due to its rich nature. We hope to draw a conclusion as to who the greatest batter of all-time is, discover which players might be …


Parameter Estimation Using A Continuous, Differentiable And Asymmetric Penalized Likelihood Function, Brian Holt, Dr. Dennis Tolley Jun 2013

Parameter Estimation Using A Continuous, Differentiable And Asymmetric Penalized Likelihood Function, Brian Holt, Dr. Dennis Tolley

Journal of Undergraduate Research

Per the original ORCA proposal, work has been done to estimate relative amounts of compounds from GC-MS (gas chromatography-mass spectrometry) data using an asymmetric penalized likelihood function. The initial results of this project were presented at the CPMS Student Research Conference in March of this year1. The results consist of a simulation study where we simulate the problem of co-eluting, or overlapping, compounds and attempt to apply basic regression techniques as well as the asymmetric penalty function to see how they compare. Under certain conditions the new penalty function has less bias, but overall the function is unstable. There are …


An Integrated Screening And Optimization Strategy, Nathaniel Jackson Rohbock Jul 2012

An Integrated Screening And Optimization Strategy, Nathaniel Jackson Rohbock

Theses and Dissertations

Within statistical methods, design of experiments (DOE) is well suited to make good inference from a minimal amount of data. Two types of designs within DOE are screening designs and optimization designs. Traditionally, these approaches have been necessarily separated by a gap between the objectives of each design and the methods available. Despite being so separated, in practice these designs are frequently connected by sequential experimentation. In fact, from the genesis of a project, the experimentor often knows that both designs will be necessary to accomplish his objectives. Due to advances in the understanding of experimental designs with complex aliasing …


An Applied Investigation Of Gaussian Markov Random Fields, Jessica Lyn Olsen Jun 2012

An Applied Investigation Of Gaussian Markov Random Fields, Jessica Lyn Olsen

Theses and Dissertations

Recently, Bayesian methods have become the essence of modern statistics, specifically, the ability to incorporate hierarchical models. In particular, correlated data, such as the data found in spatial and temporal applications, have benefited greatly from the development and application of Bayesian statistics. One particular application of Bayesian modeling is Gaussian Markov Random Fields. These methods have proven to be very useful in providing a framework for correlated data. I will demonstrate the power of GMRFs by applying this method to two sets of data; a set of temporal data involving car accidents in the UK and a set of spatial …


Xprime-Em: Eliciting Expert Prior Information For Motif Exploration Using The Expectation-Maximization Algorithm, Wei Zhou Jun 2012

Xprime-Em: Eliciting Expert Prior Information For Motif Exploration Using The Expectation-Maximization Algorithm, Wei Zhou

Theses and Dissertations

Understanding the possible mechanisms of gene transcription regulation is a primary challenge for current molecular biologists. Identifying transcription factor binding sites (TFBSs), also called DNA motifs, is an important step in understanding these mechanisms. Furthermore, many human diseases are attributed to mutations in TFBSs, which makes identifying those DNA motifs significant for disease treatment. Uncertainty and variations in specific nucleotides of TFBSs present difficulties for DNA motif searching. In this project, we present an algorithm, XPRIME-EM (Eliciting EXpert PRior Information for Motif Exploration using the Expectation-Maximization Algorithm), which can discover known and de novo (unknown) DNA motifs simultaneously from a …


Estimation Of The Effects Of Parental Measures On Child Aggression Using Structural Equation Modeling, Jordan Daniel Pyper Jun 2012

Estimation Of The Effects Of Parental Measures On Child Aggression Using Structural Equation Modeling, Jordan Daniel Pyper

Theses and Dissertations

A child's parents are the primary source of knowledge and learned behaviors for developing children, and the benefits or repercussions of certain parental practices can be long lasting. Although parenting practices affect behavioral outcomes for children, families tend to be diverse in their circumstances and needs. Research attempting to ascertain cause and effect relationships between parental influences and child behavior can be difficult due to the complex nature of family dynamics and the intricacies of real life. Structural equation modeling (SEM) is an appropriate method for this research as it is able to account for the complicated nature of child-parent …


Support Vector Machines For Classification And Imputation, Spencer David Rogers May 2012

Support Vector Machines For Classification And Imputation, Spencer David Rogers

Theses and Dissertations

Support vector machines (SVMs) are a powerful tool for classification problems. SVMs have only been developed in the last 20 years with the availability of cheap and abundant computing power. SVMs are a non-statistical approach and make no assumptions about the distribution of the data. Here support vector machines are applied to a classic data set from the machine learning literature and the out-of-sample misclassification rates are compared to other classification methods. Finally, an algorithm for using support vector machines to address the difficulty in imputing missing categorical data is proposed and its performance is demonstrated under three different scenarios …


Species Identification And Strain Attribution With Unassembled Sequencing Data, Owen Eric Francis Apr 2012

Species Identification And Strain Attribution With Unassembled Sequencing Data, Owen Eric Francis

Theses and Dissertations

Emerging sequencing approaches have revolutionized the way we can collect DNA sequence data for applications in bioforensics and biosurveillance. In this research, we present an approach to construct a database of known biological agents and use this database to develop a statistical framework to analyze raw reads from next-generation sequence data for species identification and strain attribution. Our method capitalizes on a Bayesian statistical framework that accommodates information on sequence quality, mapping quality and provides posterior probabilities of matches to a known database of target genomes. Importantly, our approach also incorporates the possibility that multiple species can be present in …


Hitters Vs. Pitchers: A Comparison Of Fantasy Baseball Player Performances Using Hierarchical Bayesian Models, Scott D. Huddleston Apr 2012

Hitters Vs. Pitchers: A Comparison Of Fantasy Baseball Player Performances Using Hierarchical Bayesian Models, Scott D. Huddleston

Theses and Dissertations

In recent years, fantasy baseball has seen an explosion in popularity. Major League Baseball, with its long, storied history and the enormous quantity of data available, naturally lends itself to the modern-day recreational activity known as fantasy baseball. Fantasy baseball is a game in which participants manage an imaginary roster of real players and compete against one another using those players' real-life statistics to score points. Early forms of fantasy baseball began in the early 1960s, but beginning in the 1990s, the sport was revolutionized due to the advent of powerful computers and the Internet. The data used in this …


Using An Experimental Mixture Design To Identify Experimental Regions With High Probability Of Creating A Homogeneous Monolithic Column Capable Of Flow, Charles C. Willden Apr 2012

Using An Experimental Mixture Design To Identify Experimental Regions With High Probability Of Creating A Homogeneous Monolithic Column Capable Of Flow, Charles C. Willden

Theses and Dissertations

Graduate students in the Brigham Young University Chemistry Department are working to develop a filtering device that can be used to separate substances into their constituent parts. The device consists of a monomer and water mixture that is polymerized into a monolith inside of a capillary. The ideal monolith is completely solid with interconnected pores that are small enough to cause the constituent parts to pass through the capillary at different rates, effectively separating the substance. Although the end objective is to minimize pore sizes, it is necessary to first identify an experimental region where any combination of input variables …


Bayesian Pollution Source Apportionment Incorporating Multiple Simultaneous Measurements, Jonathan Casey Christensen Mar 2012

Bayesian Pollution Source Apportionment Incorporating Multiple Simultaneous Measurements, Jonathan Casey Christensen

Theses and Dissertations

We describe a method to estimate pollution profiles and contribution levels for distinct prominent pollution sources in a region based on daily pollutant concentration measurements from multiple measurement stations over a period of time. In an extension of existing work, we will estimate common source profiles but distinct contribution levels based on measurements from each station. In addition, we will explore the possibility of extending existing work to allow adjustments for synoptic regimes—large scale weather patterns which may effect the amount of pollution measured from individual sources as well as for particular pollutants. For both extensions we propose Bayesian methods …


Predicting Maximal Oxygen Consumption (Vo2max) Levels In Adolescents, Brent A. Shepherd Mar 2012

Predicting Maximal Oxygen Consumption (Vo2max) Levels In Adolescents, Brent A. Shepherd

Theses and Dissertations

Maximal oxygen consumption (VO2max) is considered by many to be the best overall measure of an individual's cardiovascular health. Collecting the measurement, however, requires subjecting an individual to prolonged periods of intense exercise until their maximal level, the point at which their body uses no additional oxygen from the air despite increased exercise intensity, is reached. Collecting VO2max data also requires expensive equipment and great subject discomfort to get accurate results. Because of this inherent difficulty, it is often avoided despite its usefulness. In this research, we propose a set of Bayesian hierarchical models to predict VO2max levels in adolescents, …


The Effect Of Smoking On Tuberculosis Incidence In Burdened Countries, Natalie Noel Ellison Mar 2012

The Effect Of Smoking On Tuberculosis Incidence In Burdened Countries, Natalie Noel Ellison

Theses and Dissertations

It is estimated that one third of the world's population is infected with tuberculosis. Though once thought a "dead" disease, tuberculosis is very much alive. The rise of drug resistant strains of tuberculosis, and TB-HIV coinfection have made tuberculosis an even greater worldwide threat. While HIV, poverty, and public health infrastructure are historically assumed to affect the burden of tuberculosis, recent research has been done to implicate smoking in this list. This analysis involves combining data from multiple sources in order determine if smoking is a statistically significant factor in predicting the number of incident tuberculosis cases in a country. …


Screening Designs That Minimize Model Dependence, Kenneth P. Fairchild Dec 2011

Screening Designs That Minimize Model Dependence, Kenneth P. Fairchild

Theses and Dissertations

When approaching a new research problem, we often use screening designs to determine which factors are worth exploring in more detail. Before exploring a problem, we don't know which factors are important. When examining a large number of factors, it is likely that only a handful are significant and that even fewer two-factor interactions will be significant. If there are important interactions, it is likely that they are connected with the handful of significant main effects. Since we don't know beforehand which factors are significant, we want to choose a design that gives us the highest probability a priori of …


Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson Jul 2011

Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson

Theses and Dissertations

Walmart and other “big box” stores seek to expand in rural markets, possibly due to cheap land and lack of zoning laws. In August 2000, Walmart opened a store in Ephraim, a small rural town in central Utah. It is of interest to understand how Walmart's entrance into the local market changes the sales tax revenue base for Ephraim and for the surrounding municipalities. It is thought that small “Mom and Pop” stores go out of business because they cannot compete with Walmart's prices, leading to a decrease in variety, selection, convenience, and most importantly, sales tax revenue base in …


An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey Jul 2011

An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey

Theses and Dissertations

Bayesian statistical methods have long been computationally out of reach because the analysis often requires integration of high-dimensional functions. Recent advancements in computational tools to apply Markov Chain Monte Carlo (MCMC) methods are making Bayesian data analysis accessible for all statisticians. Two such computer tools are Win-BUGS and SASR 9.2's PROC MCMC. Bayesian methodology will be introduced through discussion of fourteen statistical examples with code and computer output to demonstrate the power of these computational tools in a wide variety of settings.


Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler Jun 2011

Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler

Theses and Dissertations

University students often complete evaluations of their courses and instructors. The evaluation tool typically contains questions about the course and the instructor on an ordinal Likert scale. We assess instructor effectiveness while adjusting for known confounders. We present a probit regression model with a latent variable to measure the instructor effectiveness accounting for student specific covariates, such as student grade in the course, high school and university GPA, and ACT score.