Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Statistical Methodology

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1531 - 1560 of 1562

Full-Text Articles in Statistics and Probability

Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan Sep 2002

Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Researchers working with survival data are by now adept at handling issues associated with incomplete data, particular those associated with various forms of censoring. An extreme form of interval censoring, known as current status observation, refers to situations where the only available information on a survival random variable T is whether or not T exceeds a random independent monitoring time C. This article contains a brief review of the extensive literature on the analysis of current status data, discussing the implications of response-based sampling on these methods. The majority of the paper introduces some recent extensions of these ideas to …


Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang Aug 2002

Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …


Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick Aug 2002

Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick

U.C. Berkeley Division of Biostatistics Working Paper Series

DNA microarrays are a new and promising biotechnology which allows the monitoring of expression levels in cells for thousands of genes simultaneously. An important and common question in microarray experiments is the identification of differentially expressed genes, i.e., genes whose expression levels are associated with a response or covariate of interest. The biological question of differential expression can be restated as a problem in multiple hypothesis testing: the simultaneous test for each gene of the null hypothesis of no association between the expression levels and the responses or covariates. As a typical microarray experiment measures expression levels for thousands of …


Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins Aug 2002

Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a bivariate survival function estimator for a general right censored data structure that includes a time dependent covariate process. Firstly, an initial estimator that generalizes Dabrowska's (1988) estimator is introduced. We obtain this estimator by a general methodology of constructing estimating functions in censored data models. The initial estimator is guaranteed to improve on Dabrowska's estimator and remains consistent and asymptotically linear under informative censoring schemes if the censoring mechanism is estimated consistently. We then construct an orthogonalized estimating function which results in a more robust and efficient estimator than our initial estimator. A simulation study demonstrates the …


Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell May 2002

Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.


A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan Apr 2002

A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Clustering algorithms have been widely applied to gene expression data. For both hierarchical and partitioning clustering algorithms, selecting the number of significant clusters is an important problem and many methods have been proposed. Existing methods for selecting the number of clusters tend to find only the global patterns in the data (e.g.: the over and under expressed genes). We have noted the need for a better method in the gene expression context, where small, biologically meaningful clusters can be difficult to identify. In this paper, we define a new criteria, Mean Split Silhouette (MSS), which is a measure of cluster …


A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan Feb 2002

A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan

U.C. Berkeley Division of Biostatistics Working Paper Series

Kaufman & Rousseeuw (1990) proposed a clustering algorithm Partitioning Around Medoids (PAM) which maps a distance matrix into a specified number of clusters. A particularly nice property is that PAM allows clustering with respect to any specified distance metric. In addition, the medoids are robust representations of the cluster centers, which is particularly important in the common context that many elements do not belong well to any cluster. Based on our experience in clustering gene expression data, we have noticed that PAM does have problems recognizing relatively small clusters in situations where good partitions around medoids clearly exist. In this …


Direct Sequential Simulation Algorithms In Geostatistics, Robyn Robertson Jan 2002

Direct Sequential Simulation Algorithms In Geostatistics, Robyn Robertson

Theses : Honours

Conditional sequential simulation algorithms have been used in geostatistics for many years but we currently find new developments are being made in this field. This thesis presents two new direct sequential simulation with histogram reproduction algorithms and compares them with the efficient and widely used sequential Gaussian simulation algorithm and the original direct sequential simulation algorithm. We explore the possibility of reproducing both the semivariogram and the histogram without the need for a transformation to normal space, through optimising an objective function and placing linear constraints on the local conditional distributions. Programs from the GSLIB Fortran library are expanded to …


Discrete Predictive Analysis In Probabilistic Safety Assessment, Paul Kvam, J. Glenn Miller Jan 2002

Discrete Predictive Analysis In Probabilistic Safety Assessment, Paul Kvam, J. Glenn Miller

Department of Math & Statistics Faculty Publications

This paper presents methods for predicting future numbers of component failures for probabilistic safety assessments (PSAs). The research is motivated and illustrated by discrete failure data from the nuclear industry, including failure counts for emergency diesel generators, pumps, and motor operated valves. Failure counts are modeled with Poisson and binomial distributions. Multiple-failure environments create extra problems for predictive inference, and are a primary focus of this paper. Common cause failures (CCFs), in particular, refer to the simultaneous failure of system components due to an external event. CCF prediction is investigated, and approximate inference methods are derived for various CCF models.


Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen Nov 2001

Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen

U.C. Berkeley Division of Biostatistics Working Paper Series

Recurrent event data typically exhibit the phenomenon of intra-individual correlation, owing to not only observed covariates but also random effects. In many applications, the population can be reasonably postulated as a heterogeneous mixture of individual renewal processes, and the inference of interest is the effect of individual-level covariates. In this article, we suggest and investigate a marginal proportional hazards model for gaps between recurrent events. A connection is established between observed gap times and clustered survival data, however, with informative cluster size. We then derive a novel and general inference procedure for the latter, based on a functional formulation of …


Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch Oct 2001

Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch

U.C. Berkeley Division of Biostatistics Working Paper Series

The pool-adjacent violator-algorithm (Ayer, et al., 1955) has long been known to give the maximum likelihood estimator of a series of ordered binomial parameters, based on an independent observation from each distribution (see Barlow et al., 1972). This result has immediate application to estimation of a survival distribution based on current survival status at a set of monitoring times. This paper considers an extended problem of maximum likelihood estimation of a series of ‘ordered’ multinomial parameters. By making use of variants of the pool adjacent violator algorithm, we obtain a simple algorithm to compute the maximum likelihood estimator and demonstrate …


Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan Jul 2001

Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function of the true data generating distribution, and an estimate is obtained by applying this function to the empirical distribution. We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as …


An Extension Of Euclidean Distance Matrix Analysis Applied To Craniofacial Growth Prediction, Mark K. Batesole Jun 2000

An Extension Of Euclidean Distance Matrix Analysis Applied To Craniofacial Growth Prediction, Mark K. Batesole

Loma Linda University Electronic Theses, Dissertations & Projects

This study is [sic] introduces a new extension of Euclidean Distance Matrix Analysis (EDMA) as applied to growth prediction analysis. Using EDMA eliminates the presupposition of a set growth pattern, which is introduced by traditional superimposition techniques. When EDMA is extended to analyze prediction methodologies, using a non-age-matched growth sample, some shortcomings become evident. These involve bootstrapping techniques, relative difference in growth, and absence of clinical, real world measure. To overcome these issues, a statistical approach using the Wilcoxen signed rank test, absolute difference in growth, and a new method to evaluate the error in a prediction methodology's landmark identification …


Estimating The Probability Of Severe Convective Storms: A Local Perspective For The Central And Northern Plains, Preston W. Leftwich Jr. Dec 1999

Estimating The Probability Of Severe Convective Storms: A Local Perspective For The Central And Northern Plains, Preston W. Leftwich Jr.

National Oceanographic and Atmospheric Administration: Technical Reports and Related Materials

Summary and Conclusions

A procedure to estimate probabilities of the occurrence of severe convective storms within local areas has been described. Probabilities were based on a simulated climatology and the relative frequency of severe convective events when a selected site was contained within an operational Outlook or Watch. Combined data from five local areas were used to develop a general model for local probabilities within the central and northern Plains region. Attachment of probabilities to specific products placed values within a framework familiar to both forecasters and "end-users." Application of results in an operational scenario demonstrated representative local probabilities and …


A New Sequential Goodness Of Fit Test For The Three-Parameter Gamma Distribution With Known Shape Based On Skewness And Kurtosis, Chil Ho Park Mar 1999

A New Sequential Goodness Of Fit Test For The Three-Parameter Gamma Distribution With Known Shape Based On Skewness And Kurtosis, Chil Ho Park

Theses and Dissertations

This research presents a new sequential goodness of fit test for the three-parameter gamma distribution with a known shape. The test is accomplished by employing two new tests, sample skewness and sample kurtosis, sequentially as test statistics. Unlike the typical goodness of fit test, using parameter estimation methods such as maximum likelihood estimation and minimum distance estimation, this test using the two test statistics above does not involve a substantial degree of computational complexity. Large Monte Carlo simulation has been used to determine critical values and overall significance levels for all combinations of the two tests, and to conduct extensive …


A New Sequential Goodness-Of-Fit Test For A Family Of Two Parameter Gamma Distributions With Known Shape Based On Skewness And Q-Statistic, Jae Suk Park Mar 1999

A New Sequential Goodness-Of-Fit Test For A Family Of Two Parameter Gamma Distributions With Known Shape Based On Skewness And Q-Statistic, Jae Suk Park

Theses and Dissertations

The objective of this research is to develop a new goodness-of-fit test for the gamma distribution. The gamma distribution is widely used for reliability and failure time estimations in the real world. Several methods to measure the fit of data to a hypothesized distribution are commonly used such as the chi-squared test, and Anderson- Darling test. The most important aspect of these tests is how well the results reflect the distribution family. This research will use the sequential test with skewness and Q- statistic as test statistics for fitting a gamma distribution. The main idea of a sequential test is …


Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe Oct 1998

Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe

UW Biostatistics Working Paper Series

Often the accuracy of a new diagnostic test must be assessed when a perfect gold standard does not exist. Use of an imperfect test biases the accuracy estimates of the new test. This paper reviews existing approaches to this problem including discrepant resolution and latent class analysis. Deficiencies with these approaches are identified. A new approach is proposed that combines the results of several imperfect reference tests to define a better reference standard. We call this the composite reference standard (CRS). Using the CRS, accuracy can be assessed using multistage sampling designs. Maximum likelihood estimates of accuracy and expressions for …


A Pilot Study Testing A Proprietary Sealant For Plaque Reduction, Jennifer Rowland Jun 1998

A Pilot Study Testing A Proprietary Sealant For Plaque Reduction, Jennifer Rowland

Loma Linda University Electronic Theses, Dissertations & Projects

Enamel demineralization due to increased plaque accumulation is a well-recognized problem associated with fixed orthodontic appliances. The purpose of this study was to evaluate the efficacy of a proprietary product to be marketed by 3M Unitek for the reduction of plaque accumulation and thereby reduce enamel demineralization.

Sixty four bovine teeth were divided equally into an untreated control group and a group treated with the proprietary sealant product. Both groups were subsequently immersed in a Streptococcus mutans culture. Plaque deposits from each group were removed, suspended in trypticase soy broth, serially diluted, and plated on trypticase soy agar at 1, …


Acer Quest: The Interactive Test Analysis System. Version 2.1., Raymond J. Adams, Siek-Toon Khoo Jan 1996

Acer Quest: The Interactive Test Analysis System. Version 2.1., Raymond J. Adams, Siek-Toon Khoo

Measurement and statistics

This is a guide to using Quest. Quest offers a comprehensive test and questionnaire analysis environment by providing a data analyst with access to the most recent developments in Rasch measurement theory, as well as a range of traditional analysis procedures. It includes an easy to use control language with flexible and informative output. Quest can be used to construct and validate variables based on both dichotomous and polychotomous observations. It scores and analyses such instruments as multiple choice tests, Likert type rating scales, short answer items, and partial credit items.

Download the legacy ACER Quest software here

Download the …


Performance Indices For On-Ice Hockey Statistics, William (Bill) H. Williams Aug 1995

Performance Indices For On-Ice Hockey Statistics, William (Bill) H. Williams

Publications and Research

No abstract provided.


Several Modified Goodness-Of-Fit Tests For The Cauchy Distribution With Unknown Scale And Location Parameters, Bora H. Onen Mar 1994

Several Modified Goodness-Of-Fit Tests For The Cauchy Distribution With Unknown Scale And Location Parameters, Bora H. Onen

Theses and Dissertations

Kolmogorov-Simirnov and the Kuiper goodness-of-fit tests are studied for the Cauchy distribution with the unknown location and scale parameters. Monte Carlo simulation studies were performed using maximum likelihood estimation to calculate the critical values for standard Kolmogorov-Simirnov and the Kuiper tests. Then a reflection technique is introduced and the critical value tables are calculated for both the Reflected Kolmogorov-Simirnov and the Reflected Kuiper tests. Several sequential tests are performed by combining standard Kolmogorov-Simirnov and Kuiper in one test, standard Cramer-von Mises and the standard Kuiper in the other and finally the reflected Cramer-von Mises and the standard Kuiper in the …


A Modified Anderson Darling Goodness-Of-Fit Test For The Gamma Distribution With Unknown Scale And Location Parameters, Tamer Ozmen Mar 1993

A Modified Anderson Darling Goodness-Of-Fit Test For The Gamma Distribution With Unknown Scale And Location Parameters, Tamer Ozmen

Theses and Dissertations

A new modified Anderson-Darling goodness-of-fit test is introduced for the three-parameter Gamma distribution when the location parameter is found by minimum distance estimation and scale parameter by maximum likelihood estimation. Monte Carlo simulation studies were performed to calculate the critical values for A-D test when A-D statistic is minimized. These critical values are then used for testing whether a set of observations follows a Gamma distribution when the scale and location parameters axe unspecified and are estimated from the sample. Functional relationship between the critical values of A-D is also examined for each shape parameter by the variables, sample size …


A New Goodness-Of-Fit Test For The Weibull Distribution Based On Spacings, Mark C. Coppa Mar 1993

A New Goodness-Of-Fit Test For The Weibull Distribution Based On Spacings, Mark C. Coppa

Theses and Dissertations

The critical values for a new goodness-of-fit test based on spacings are generated for the Weibull distribution when the shape parameter is known. The critical values are used for testing whether a set of observations follow a Weibull distribution when the scale and location parameters are unknown. A Monte Carlo simulation with 10,000 iterations is used to generate the critical values for sample sizes 5(5)35 at shape parameters k equal to 0.5(0.5)1.5 and for sample sizes 5(5)20 at shape parameters k = 2.0(1.0)4.0. A Monte Carlo power study of the Z* test statistic using 5000 iterations is accomplished using nine …


Modified Anderson-Darling And Cramer-Von Mises Goodness-Of-Fit Tests For The Normal Distribution, David A. Gwinn Sr. Mar 1993

Modified Anderson-Darling And Cramer-Von Mises Goodness-Of-Fit Tests For The Normal Distribution, David A. Gwinn Sr.

Theses and Dissertations

New techniques for calculating goodness-of-fit statistics for normal distributions with parameters estimated from the sample are investigated. Samples are generated for a Normal(0,1) distribution. Critical values are calculated for five modifications to the Anderson-Darling statistic and five modifications to the Cramer-Von Mises statistic. An extensive power study is done to test the power of the new statistics versus the power of the unmodified statistics. Powers of six of the new statistics show minimal to no improvement, two of the new statistics show a marked decrease in power, and two of the new statistics show an overall increase in power over …


Transformation Of Combat Data In Support Of Battle Trace, Hyoung-Kyu Choi Mar 1992

Transformation Of Combat Data In Support Of Battle Trace, Hyoung-Kyu Choi

Theses and Dissertations

The purpose of this study was to find an appropriate way of characterizing battle data as a function of time. This study was initially designed to remove the numerical instability problem cited in Barr's battle trace methodology as suggested by TRAC-MTRY. The research focused on the instability problem and identifying/recommending a technique that improved the efficiency of computation, and enhanced an analyst's ability to meaningfully interpret the battle trace. Two Lanchester's Square Law based methodologies are introduced and analyzed. The results of this analysis indicate that the battle trace of a constant value generated by Lanchester's Square Law integration seemed …


Sets Of Typical Subsamples, Joel Atkins, G.J Sherman Sep 1990

Sets Of Typical Subsamples, Joel Atkins, G.J Sherman

Mathematical Sciences Technical Reports (MSTR)

A group theoretic condition on a set of subsamples of a random sample from a continuous random variable symmetric about 0 is shown to be sufficient to provide typical values for 0.


Design And Statistical Analysis Of Plant Protection Experiment, J F. Wallace Mar 1989

Design And Statistical Analysis Of Plant Protection Experiment, J F. Wallace

All other publications

An Australian co-operation with the national agricultural research project Thailand.

This short course is intended to cover aspects of experimental design, sampling and statistical analysis for researchers in Entomology and Plant Pathology.

The basic principles of experimental design are the same for plant protection research as they are in other areas of research. Problems in plant protection arise from the variation of the data, the complexity of the systems and interactions with environmental factors. In many cases, standard designs are quite adequate.


A Sequential Testing System For Health Services Research And Testing, Matthew B. Barkley Jan 1974

A Sequential Testing System For Health Services Research And Testing, Matthew B. Barkley

MUSC Theses and Dissertations

Use of the Sequential Probability Ratio Test (SPRT) can save an experimenter's time, material, and money. Written for the Wang 600 Programmable Calculator, the Sequential Testing System performs the SPRT for any of seven common one-sample tests, greatly facilitating the use of this long-neglected clinical tool.


Objective Measurement Of Wool : Criteria, Methods And Materials, A Ingleton Jan 1973

Objective Measurement Of Wool : Criteria, Methods And Materials, A Ingleton

Journal of the Department of Agriculture, Western Australia, Series 4

An outline of some of the technical aspects of the objective measurement of wool—processes that will mean major cost savings to the wool industry.


A Simple Method For The Construction Of Empirical Confidence Limits For Economic Forecasts, William (Bill) H. Williams, M. L. Goodman Dec 1971

A Simple Method For The Construction Of Empirical Confidence Limits For Economic Forecasts, William (Bill) H. Williams, M. L. Goodman

Publications and Research

A simple method for the construction of empirical confidence intervals for time series forecasts is described. The procedure is to go through the series making a forecast from each point in time. The comparison of these forecasts with the known actual observations will yield an empirical distribution of forecasting errors. This distribution can then be used to set confidence intervals for subsequent forecasts. The technique appears to be particularly useful when the mechanism generating the series cannot be fully identified from the available data or when limits based on more standard considerations are difficult to obtain.