Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2003

Discipline
Institution
Keyword
Publication
Publication Type

Articles 151 - 180 of 207

Full-Text Articles in Statistics and Probability

A More Efficient Way Of Obtaining A Unique Median Estimate For Circular Data, B. Sango Otieno, Christine M. Anderson-Cook May 2003

A More Efficient Way Of Obtaining A Unique Median Estimate For Circular Data, B. Sango Otieno, Christine M. Anderson-Cook

Journal of Modern Applied Statistical Methods

The procedure for computing the sample circular median occasionally leads to a non-unique estimate of the population circular median, since there can sometimes be two or more diameters that divide data equally and have the same circular mean deviation. A modification in the computation of the sample median is suggested, which not only eliminates this non-uniqueness problem, but is computationally easier and faster to work with than the existing alternative.


You Think You’Ve Got Trivials?, Shlomo S. Sawilowsky May 2003

You Think You’Ve Got Trivials?, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

Effect sizes are important for power analysis and meta-analysis. This has led to a debate on reporting effect sizes for studies that are not statistically significant. Contrary and supportive evidence has been offered on the basis of Monte Carlo methods. In this article, clarifications are given regarding what should be simulated to determine the possible effects of piecemeal publishing trivial effect sizes.


A Semiparametric Regression Model For Oligonucleotide Arrays, Jianhua Hu, Guosheng Yin May 2003

A Semiparametric Regression Model For Oligonucleotide Arrays, Jianhua Hu, Guosheng Yin

Journal of Modern Applied Statistical Methods

A semiparametric model incorporating the spline smoothing technique is proposed to study oligonucleotide gene expression data. No specific parametric functional form is assumed for mismatch probe intensities, which allows much more flexibility in the fitted model. The new approach improves the model fitting, hence the estimation of expression indexes. The method is applied to a data set of 18 HuGeneFL arrays.


Jmasm6: An Algorithm For Generating Exact Critical Values For The Kruskal-Wallis One-Way Anova, Todd C. Headrick May 2003

Jmasm6: An Algorithm For Generating Exact Critical Values For The Kruskal-Wallis One-Way Anova, Todd C. Headrick

Journal of Modern Applied Statistical Methods

A Fortran 77 subroutine is provided for computing exact critical values for the Kruskal-Wallis test on k independent groups with equal or unequal samples sizes. The subroutine requires the user to provide sorting and ranking routines and a uniform pseudo-random number generator. The program is available from the author on request.


Randomization Technique, Allocation Concealment, Masking, And Susceptibility Of Trials To Selection Bias, Vance W. Berger, Costas A. Christophi May 2003

Randomization Technique, Allocation Concealment, Masking, And Susceptibility Of Trials To Selection Bias, Vance W. Berger, Costas A. Christophi

Journal of Modern Applied Statistical Methods

It is widely believed that baseline imbalances in randomized clinical trials must necessarily be random. Yet even among masked randomized trials conducted with allocation concealment, there are mechanisms by which patients with specific covariates may be selected for inclusion into a particular treatment group. This selection bias would force imbalance in those covariates, measured or unmeasured, that are used for the patient selection. Unfortunately, few trials provide adequate information to determine even if there was allocation concealment, how the randomization was conducted, and how successful the masking may have been, let alone if selection bias was adequately controlle d. In …


Improved Multiple Comparisons With The Best In Response Surface Methodology, Laura K. Miller, Ping Sa May 2003

Improved Multiple Comparisons With The Best In Response Surface Methodology, Laura K. Miller, Ping Sa

Journal of Modern Applied Statistical Methods

A method to construct simultaneous confidence intervals about the difference in mean responses at the stationary point and at x for all x within a sphere with radius I R is proposed. Results of an efficiency study to compare the new method and the existing method by Moore and Sa (1999) are provided.


The Global Dynamics Of Isothermal Chemical Systems With Critical Nonlinearity, Yi Li, Yuanwei Qi May 2003

The Global Dynamics Of Isothermal Chemical Systems With Critical Nonlinearity, Yi Li, Yuanwei Qi

Mathematics and Statistics Faculty Publications

In this paper, we study the Cauchy problem of a cubic autocatalytic chemical reaction system u1,t = u1,xx − uα1 uβ2, u2,t = du2,xx+ uα1 uβ2 with non-negative initial data, where the exponents α,β satisfy 1<α,βd>0 is the Lewis number. Our purpose is to study the global dynamics of solutions under mild decay of initial data as |x|→∞. We show the exact large time behaviour of solutions which is universal.


Constrained Boundary Monitoring For Group Sequential Clinical Trials, Bart E. Burington, Scott S. Emerson Apr 2003

Constrained Boundary Monitoring For Group Sequential Clinical Trials, Bart E. Burington, Scott S. Emerson

UW Biostatistics Working Paper Series

Group sequential stopping rules are often used during the conduct of clinical trials in order to attain more ethical treatment of patients and to better address efficiency concerns. Because the use of such stopping rules materially affects the frequentist operating characteristics of the hypothesis test, it is necessary to choose an appropriate stopping rule during the planning of the study. It is often the case, however, that the number and timing of interim analyses are not precisely known at the time of trial design, and thus the implementation of a particular stopping rule must allow for flexible determination of the …


Comparison Of Features From Sar And Gmti Imagery Of Ground Targets, David Beckman, Samuel J. Frame Apr 2003

Comparison Of Features From Sar And Gmti Imagery Of Ground Targets, David Beckman, Samuel J. Frame

Statistics

We describe an algorithm for class-independent automated target recognition (ATR) and association using range-Doppler images of moving targets and SAR images of stationary targets. This algorithm can be used both for target identification (by comparison against a pre-existing database of measurements of all potential targets) and target association (not requiring a pre-existing database). The algorithm computes a one-dimensional signature for each received range-Doppler image; these signatures are stored in a database for comparison against other detections. The signatures used in our algorithm are range profiles, generated from the clutter-suppressed, filtered image by incoherently integrating the image energy across a number …


Cause-Effect Relationships In Analytical Surveys: An Illustration Of Statistical Issues, Gary L. Gadbury, Hans T. Schreuder Apr 2003

Cause-Effect Relationships In Analytical Surveys: An Illustration Of Statistical Issues, Gary L. Gadbury, Hans T. Schreuder

Mathematics and Statistics Faculty Research & Creative Works

Establishing cause-effect is critical in the field of natural resources where one may want to know the impact of management practices, wildfires, drought, etc. on water quality and quantity, wildlife, growth and survival of desirable trees for timber production, etc. Yet, key obstacles exist when trying to establish cause-effect in such contexts. Issues involved with identifying a causal hypothesis, and conditions needed to estimate a causal effect or to establish cause-effect are considered. Ideally one conducts an experiment and follows with a survey, or vice versa. in an experiment, the population of inference may be quite limited and in surveys, …


Machine Learning Approaches For Determining Effective Seeds For K -Means Algorithm, Kaveephong Lertwachara Apr 2003

Machine Learning Approaches For Determining Effective Seeds For K -Means Algorithm, Kaveephong Lertwachara

Doctoral Dissertations

In this study, I investigate and conduct an experiment on two-stage clustering procedures, hybrid models in simulated environments where conditions such as collinearity problems and cluster structures are controlled, and in real-life problems where conditions are not controlled. The first hybrid model (NK) is an integration between a neural network (NN) and the k-means algorithm (KM) where NN screens seeds and passes them to KM. The second hybrid (GK) uses a genetic algorithm (GA) instead of the neural network. Both NN and GA used in this study are in their simplest-possible forms.

In the simulated data sets, I investigate two …


The Government As Litigant: Further Tests Of The Case Selection Model, Theodore Eisenberg, Henry Farber Apr 2003

The Government As Litigant: Further Tests Of The Case Selection Model, Theodore Eisenberg, Henry Farber

Cornell Law Faculty Publications

We develop a model of the plaintiff's decision to file a lawsuit that has implications for how differences between the federal government and private litigants translate into differences in trial rates and plaintiff win rates at trial. Our case selection model generates a set of predictions for relative trial rates and plaintiff win rates, depending on the type of case and whether the government is defendant or plaintiff. To test the model, we use data on about 474,000 cases filed in federal district court between 1979 and 1994 in the areas of personal injury and job discrimination, in which the …


Virginia's Capital Jurors, Stephen P. Garvey, Paul Marcus Apr 2003

Virginia's Capital Jurors, Stephen P. Garvey, Paul Marcus

Cornell Law Faculty Publications

Next to Texas, no state has executed more capital defendants than Virginia. Moreover, the likelihood of a death sentence actually being carried out is greater in Virginia than it is elsewhere, while the length of time between the imposition of a death sentence and its actual execution is shorter. Virginia has thus earned a reputation among members of the defense bar as being among the worst of the death penalty states. Yet insofar as these facts about Virginia's death penalty relate primarily to the behavior of state and federal appellate courts, they suggest that what makes Virginia's death penalty unique …


A Combinatorial Technique For Face Detection Based On Color And Statistical Analysis, Harishwaran Hariharan Apr 2003

A Combinatorial Technique For Face Detection Based On Color And Statistical Analysis, Harishwaran Hariharan

Electrical & Computer Engineering Theses & Dissertations

Automatic detection of faces from video sequences is an important task in security applications. The number, location, size and orientation of human faces in a video frame are unpredictable and can vary from frame to frame. A face detection algorithm for color images in the presence of varying lighting conditions and complexity in background relying upon color and statistical analysis is presented in this thesis. The new method detects skin regions over the entire image and then classifies the skin regions as faces and non-faces. Segmentation of skin regions is performed by a novel color space merging procedure named Integrated …


Estimating The Accuracy Of Polymerase Chain Reaction-Based Tests Using Endpoint Dilution, Jim Hughes, Patricia Totten Mar 2003

Estimating The Accuracy Of Polymerase Chain Reaction-Based Tests Using Endpoint Dilution, Jim Hughes, Patricia Totten

UW Biostatistics Working Paper Series

PCR-based tests for various microorganisms or target DNA sequences are generally acknowledged to be highly "sensitive" yet the concept of sensitivity is ill-defined in the literature on these tests. We propose that sensitivity should be expressed as a function of the number of target DNA molecules in the sample (or specificity when the target number is 0). However, estimating this "sensitivity curve" is problematic since it is difficult to construct samples with a fixed number of targets. Nonetheless, using serially diluted replicate aliquots of a known concentration of the target DNA sequence, we show that it is possible to disentangle …


Design Of The Hiv Prevention Trials Network (Hptn) Protocol 054: A Cluster Randomized Crossover Trial To Evaluate Combined Access To Nevirapine In Developing Countries, Jim Hughes, Robert L. Goldenberg, Catherine M. Wilfert, Megan Valentine, Kasonde G. Mwinga, Laura A. Guay, Francis Mmiro, Jeffrey S. A. Stringer Mar 2003

Design Of The Hiv Prevention Trials Network (Hptn) Protocol 054: A Cluster Randomized Crossover Trial To Evaluate Combined Access To Nevirapine In Developing Countries, Jim Hughes, Robert L. Goldenberg, Catherine M. Wilfert, Megan Valentine, Kasonde G. Mwinga, Laura A. Guay, Francis Mmiro, Jeffrey S. A. Stringer

UW Biostatistics Working Paper Series

HPTN054 is a cluster randomized trial designed to compare two approaches to providing single dose nevirapine to HIV-seropositive mothers and their infants to prevent mother-to-child transmission of HIV in resource limited settings. A number of challenging issues arose during the design of this trial. Most importantly, the need to achieve high participation rates among pregnant, HIV-seropositive women in selected prenatal care clinics led us to develop a method of collecting anonymous and unlinked information on a key surrogate endpoint instead of pursuing linked and identified information on a clinical endpoint. In addition, since group counseling is the standard model for …


Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little Mar 2003

Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …


A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan Mar 2003

A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Estimators for the parameter of interest in semiparametric models often depend on a guessed model for the nuisance parameter. The choice of the model for the nuisance parameter can affect both the finite sample bias and efficiency of the resulting estimator of the parameter of interest. In this paper we propose a finite sample criterion based on cross validation that can be used to select a nuisance parameter model from a list of candidate models. We show that expected value of this criterion is minimized by the nuisance parameter model that yields the estimator of the parameter of interest with …


Transient Analysis And Applications Of Markov Reward Processes, Jeffrey A. Sipe Mar 2003

Transient Analysis And Applications Of Markov Reward Processes, Jeffrey A. Sipe

Theses and Dissertations

In this thesis, the problem of computing the cumulative distribution function (cdf) of the random time required for a system to first reach a specified reward threshold when the rate at which the reward accrues is controlled by a continuous time stochastic process is considered. This random time is a type of first passage time for the cumulative reward process. The major contribution of this work is a simplified, analytical expression for the Laplace-Stieltjes Transform of the cdf in one dimension rather than two. The result is obtained using two techniques: i) by converting an existing partial differential equation to …


Gaussian Mixture Reduction Of Tracking Multiple Maneuvering Targets In Clutter, Jason L. Williams Mar 2003

Gaussian Mixture Reduction Of Tracking Multiple Maneuvering Targets In Clutter, Jason L. Williams

Theses and Dissertations

The problem of tracking multiple maneuvering targets in clutter naturally leads to a Gaussian mixture representation of the Provability Density Function (PDF) of the target state vector. State-of-the-art Multiple Hypothesis Tracking (MHT) techniques maintain the mean, covariance and probability weight corresponding to each hypothesis, yet they rely on ad hoc merging and pruning rules to control the growth of hypotheses.


Statistical Process Control: An Application In Aircraft Maintenance Management, Bradley A. Beabout Mar 2003

Statistical Process Control: An Application In Aircraft Maintenance Management, Bradley A. Beabout

Theses and Dissertations

Maintenance management at the 135th Airlift Wing, Maryland Air National Guard desires a visualization tool for their maintenance performance metrics. Currently they monitor their metrics via an electronic spreadsheet. They desire a tool that presents the performance information in a graphical manner. This thesis effort focuses on the development of a visualization tool utilizing two of the seven tools offered by Statistical Process Control (SPC). This research demonstrates the application of p-charts and Pareto diagrams in the aircraft maintenance arena. P-charts are used for displaying mission capable (MC) rates and flying scheduling effectiveness (FSE) rates. Pareto diagrams are then used …


Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit Feb 2003

Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Properties of transition matrices between IBD configurations are derived for four general classes of unilineal relative pairs obtained from the grand-parent/ grand-child, half-sib, avuncular, and cousin relationships. In this setting, IBD configurations are defined as orbits of groups acting on a set of inheritance vectors. Properties of the transition matrix between IBD configurations at two linked loci are derived by relating its infinitesimal generator to the adjacency matrix of a quotient graph. The second largest eigenvalue of the infinitesimal generator and its multiplicity are key in determining the form of the transition matrix and of likelihood-based linkage tests such as …


Rank Regression In Stability Analysis, Ying Qing Chen, Annpey Pong, Biao Xing Feb 2003

Rank Regression In Stability Analysis, Ying Qing Chen, Annpey Pong, Biao Xing

U.C. Berkeley Division of Biostatistics Working Paper Series

Stability data are often collected to determine the shelf-life of certain characteristics of a pharmaceutical product, for example, a drug's potency over time. Statistical approaches such as the linear regression models are considered as appropriate to analyze the stability data. However, most of these regression models in both theory and practice rely heavily on their underlying parametric assumptions, such as normality of the continuous characteristics or their transformations. In this article, we propose and study some rank-based regression procedures for the stability data when the linear regression models are semiparametric with unspecified error structure. Numerical studies including Monte Carlo simulations …


Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles Feb 2003

Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles

U.C. Berkeley Division of Biostatistics Working Paper Series

Likelihood-based cross-validation is a statistical tool for selecting a density estimate based on n i.i.d. observations from the true density among a collection of candidate density estimators. General examples are the selection of a model indexing a maximum likelihood estimator, and the selection of a bandwidth indexing a nonparametric (e.g. kernel) density estimator. In this article, we establish asymptotic optimality of a general class of likelihood based cross-validation procedures (as indexed by the type of sample splitting used, e.g. V-fold cross-validation), in the sense that the cross-validation selector performs asymptotically as well (w.r.t. to the Kullback-Leibler distance to the true …


Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan Feb 2003

Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …


Modern Spectral Climate Patterns In Rhythmically Deposited Argillites Of The Gowganda Formation (Early Proterozoic), Southern Ontario, Canada, Gary B. Hughes, Robert Giegengack, Haralambos N. Kritikos Feb 2003

Modern Spectral Climate Patterns In Rhythmically Deposited Argillites Of The Gowganda Formation (Early Proterozoic), Southern Ontario, Canada, Gary B. Hughes, Robert Giegengack, Haralambos N. Kritikos

Statistics

Rhythmically deposited argillites of the Gowganda Formation (ca. 2.0–2.5 Ga) probably formed in a glacial setting. Drop stones and layered sedimentary couplets in the rock presumably indicate formation in a lacustrine environment with repeating freeze–thaw cycles. It is plausible that temporal variations in the thickness of sedimentary layers are related to interannual climatic variability, e.g. average seasonal temperature could have influenced melting and the amount of sediment source material carried to the lake. A sequence of layer couplet thickness measurements was made from high-resolution digitized photographs taken at an outcrop in southern Ontario, Canada. The frequency spectrum of thickness measurements …


A Method For Developing In-Silico Protein Homologs, Susan Mcclatchy Jan 2003

A Method For Developing In-Silico Protein Homologs, Susan Mcclatchy

Theses

Computational methods for identifying and screening the most promising drug receptor candidates in the human genome are of great interest to drug discovery researchers. Successful methods will accurately identify and narrow the field of potential drug receptor candidates. This study details one such method.

The method described here begins with the assumption that novel drug receptors have high sequence similarity to established drug receptors. The similarity search program FASTA3 aligns translated sequences of the human genome to known drug receptor sequences and ranks these alignments by measuring their statistical significance. Query results returned by FASTA3 are assembled into "in-silico proteins" …


Analysis Of Gene Expression Data Using Expressionist 3.1 And Genespring 4.2, Indu Shrivastava Jan 2003

Analysis Of Gene Expression Data Using Expressionist 3.1 And Genespring 4.2, Indu Shrivastava

Theses

The purpose of this study was to determine the differences in the gene expression analysis methods of two data mining tools, ExpressionisticTM 3.1 and GeneSpringTM 4.2 with focus on basic statistical analysis and clustering algorithms. The data for this analysis was derived from the hybridization of Rattus norvegicus RNA to the Affymetrix RG34A GeneChip. This analysis was derived from experiments designed to identify changes in gene expression patterns that were induced in vivo by an experimental treatment.

The tools were found to be comparable with respect to the list of statistically significant genes that were up-regulated by more …


Selecting Differentially Expressed Genes From Microarray Experiments, Margaret S. Pepe, Gary M. Longton, Garnet L. Anderson, Michel Schummer Jan 2003

Selecting Differentially Expressed Genes From Microarray Experiments, Margaret S. Pepe, Gary M. Longton, Garnet L. Anderson, Michel Schummer

UW Biostatistics Working Paper Series

High throughput technologies, such as gene expression arrays and protein mass spectrometry, allow one to simultaneously evaluate thousands of potential biomarkers that distinguish different tissue types. Of particular interest here is cancer versus normal organ tissues. We consider statistical methods to rank genes (or proteins) in regards to differential expression between tissues. Various statistical measures are considered and we argue that two measures related to the Receiver Operating Characteristic Curve are particularly suitable for this purpose. We also propose that sampling variability in the gene rankings be quantified and suggest using the “selection probability function”, the probability distribution of rankings …


Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe Jan 2003

Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve is a popular method for characterizing the accuracy of diagnostic tests when test results are not binary. Various methodologies for estimating and comparing ROC curves have been developed. One approach, due to Pepe, uses a parametric regression model with the baseline function specified up to a finite-dimensional parameter. In this article we extend the regression models by allowing arbitrary nonparametric baseline functions. We also provide asymptotic distribution theory and procedures for making statistical inference. We illustrate our approach with dataset from a prostate cancer biomarker study. Simulation studies suggest that the extra flexibility inherent …