Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (16)
- Computer Sciences (15)
- Medicine and Health Sciences (14)
- Multivariate Analysis (13)
- Statistical Models (13)
-
- Data Science (12)
- Statistical Methodology (11)
- Mathematics (10)
- Public Health (9)
- Statistical Theory (9)
- Biostatistics (8)
- Categorical Data Analysis (8)
- Clinical Epidemiology (8)
- Life Sciences (7)
- Social and Behavioral Sciences (7)
- Other Statistics and Probability (6)
- Applied Mathematics (5)
- Microarrays (5)
- Epidemiology (4)
- Genetics and Genomics (4)
- Artificial Intelligence and Robotics (3)
- Arts and Humanities (3)
- Clinical Trials (3)
- Engineering (3)
- Medical Specialties (3)
- Other Computer Sciences (3)
- Probability (3)
- Theory and Algorithms (3)
- Institution
-
- COBRA (14)
- Southern Methodist University (6)
- Utah State University (5)
- University of Nebraska - Lincoln (4)
- University of South Florida (4)
-
- Old Dominion University (3)
- University of South Carolina (3)
- Wayne State University (3)
- Brigham Young University (2)
- Loyola University Chicago (2)
- Virginia Commonwealth University (2)
- City University of New York (CUNY) (1)
- Department of Primary Industries and Regional Development, Western Australia (1)
- Embry-Riddle Aeronautical University (1)
- Florida Institute of Technology (1)
- Georgia Southern University (1)
- Illinois State University (1)
- Kennesaw State University (1)
- Louisiana Tech University (1)
- Marquette University (1)
- Minnesota State University, Mankato (1)
- Missouri University of Science and Technology (1)
- New Jersey Institute of Technology (1)
- Nova Southeastern University (1)
- Purdue University (1)
- Rose-Hulman Institute of Technology (1)
- Technological University Dublin (1)
- The Texas Medical Center Library (1)
- The University of Southern Mississippi (1)
- Universitas Negeri Malang (1)
- Publication Year
- Publication
-
- UW Biostatistics Working Paper Series (11)
- Theses and Dissertations (9)
- SMU Data Science Review (6)
- USF Tampa Graduate Theses and Dissertations (4)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (3)
-
- Journal of Modern Applied Statistical Methods (3)
- Department of Statistics: Faculty Publications (2)
- Dissertations (2)
- Electronic Theses and Dissertations (2)
- Mathematics & Statistics Faculty Publications (2)
- U.C. Berkeley Division of Biostatistics Working Paper Series (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
- Arts & Sciences Graduate Student Theses and Dissertations (1)
- Beyond: Undergraduate Research Journal (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computer Science: Faculty Publications and Other Works (1)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (1)
- Dissertations and Theses (Open Access) (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctoral Dissertations (1)
- Electrical & Computer Engineering Faculty Publications (1)
- Honors Program: Senior Projects (Public) (1)
- Journal of the Department of Agriculture, Western Australia, Series 4 (1)
- Knowledge Engineering and Data Science (1)
- Mahurin Honors College Capstone Experience/Thesis Projects (1)
- Mathematics and Statistics Faculty Research & Creative Works (1)
- Mathematics and Statistics: Faculty Publications and Other Works (1)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (1)
- Publication Type
Articles 61 - 76 of 76
Full-Text Articles in Statistics and Probability
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …
Selection Of Independent Binary Features Using Probabilities: An Example From Veterinary Medicine, Ludmila I. Kuncheva, Zoë S.J. Hoare, Peter D. Cockcroft
Selection Of Independent Binary Features Using Probabilities: An Example From Veterinary Medicine, Ludmila I. Kuncheva, Zoë S.J. Hoare, Peter D. Cockcroft
Journal of Modern Applied Statistical Methods
Supervised classification into c mutually exclusive classes based on n binary features is considered. The only information available is an n×c table with probabilities. Knowing that the best d features are not the d best, simulations were run for 4 feature selection methods and an application to diagnosing BSE in cattle and Scrapie in sheep is presented.
Special Classification Models For Lichens In The Pacific Northwest, Janeen Ardito
Special Classification Models For Lichens In The Pacific Northwest, Janeen Ardito
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
A common problem in ecological studies is that of determining where to look for rare species. This paper shows how statistical models, such as classification trees, may be used to assist in the design of probability-based surveys for rare species using information on more abundant species that are associated with the rare species. This model assisted approach to survey design involves first building models for the more abundant species. The models are then used to determine stratifications for the rare species that are associated with the more abundant species. The goal of this approach is to increase the number of …
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
UW Biostatistics Working Paper Series
Introduction: Markers that purport to distinguish subjects with a condition from those without a condition must be evaluated rigorously for their classification accuracy. A single approach to statistically evaluating and comparing markers is not yet established.
Methods: We suggest a standardization that uses the marker distribution in unaffected subjects as a reference. For an affected subject with marker value Y, the standardized placement value is the proportion of unaffected subjects with marker values that exceed Y.
Results: We apply the standardization to two illustrative datasets. In patients with pancreatic cancer placement values calculated for the CA 19-9 marker are smaller …
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang, Gary M. Longton
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang, Gary M. Longton
UW Biostatistics Working Paper Series
No single biomarker for cancer is considered adequately sensitive and specific for cancer screening. It is expected that the results of multiple markers will need to be combined in order to yield adequately accurate classification. Typically the objective function that is optimized for combining markers is the likelihood function. In this paper we consider an alternative objective function -- the area under the empirical receiver operating characteristic curve (AUC). We note that it yields consistent estimates of parameters in a generalized linear model for the risk score but does not require specifying the link function. Like logistic regression it yields …
Ip Algorithm Applied To Proteomics Data, Christopher Lee Green
Ip Algorithm Applied To Proteomics Data, Christopher Lee Green
Theses and Dissertations
Mass spectrometry has been used extensively in recent years as a valuable tool in the study of proteomics. However, the data thus produced exhibits hyper-dimensionality. Reducing the dimensionality of the data often requires the imposition of many assumptions which can be harmful to subsequent analysis. The IP algorithm is a dimension reduction algorithm, similar in purpose to latent variable analysis. It is based on the principle of maximum entropy and therefore imposes a minimum number of assumptions on the data. Partial Least Squares (PLS) is an algorithm commonly used with proteomics data from mass spectrometry in order to reduce the …
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang
UW Biostatistics Working Paper Series
We compare simple logistic regression with an alternative robust procedure for constructing linear predictors to be used for the two state classification task. Theoritical advantages of the robust procedure over logistic regression are: (i) although it assumes a generalized linear model for the dichotomous outcome variable, it does not require specification of the link function; (ii) it accommodates case-control designs even when the model is not logistic; and (iii) it yields sensible results even when the generalized linear model assumption fails to hold. Surprisingly, we find that the linear predictor derived from the logistic regression likelihood is very robust in …
Evaluating Markers For Selecting A Patient's Treatment, Xiao Song, Margaret S. Pepe
Evaluating Markers For Selecting A Patient's Treatment, Xiao Song, Margaret S. Pepe
UW Biostatistics Working Paper Series
Selecting the best treatment for a patient's disease may be facilitated by evaluating clinical characteristics or biomarker measurements at diagnosis. We consider how to evaluate the potential of such measurements to impact on treatment selection algorithms. For example, magnetic resonance neurographic imaging is potentially useful for deciding whether a patient should be treated surgically for carpal tunnel syndrome or if he/she should receive less invasive conservative therapy. We propose a graphical display, the selection impact (SI) curve, that shows the population response rate as a function of treatment selection criteria based on the marker. The curve can be useful for …
Loss-Based Estimation With Cross-Validation: Applications To Microarray Data Analysis And Motif Finding, Sandrine Dudoit, Mark J. Van Der Laan, Sunduz Keles, Annette M. Molinaro, Sandra E. Sinisi, Siew Leng Teng
Loss-Based Estimation With Cross-Validation: Applications To Microarray Data Analysis And Motif Finding, Sandrine Dudoit, Mark J. Van Der Laan, Sunduz Keles, Annette M. Molinaro, Sandra E. Sinisi, Siew Leng Teng
U.C. Berkeley Division of Biostatistics Working Paper Series
Current statistical inference problems in genomic data analysis involve parameter estimation for high-dimensional multivariate distributions, with typically unknown and intricate correlation patterns among variables. Addressing these inference questions satisfactorily requires: (i) an intensive and thorough search of the parameter space to generate good candidate estimators, (ii) an approach for selecting an optimal estimator among these candidates, and (iii) a method for reliably assessing the performance of the resulting estimator. We propose a unified loss-based methodology for estimator construction, selection, and performance assessment with cross-validation. In this approach, the parameter of interest is defined as the risk minimizer for a suitable …
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …
Selecting Differentially Expressed Genes From Microarray Experiments, Margaret S. Pepe, Gary M. Longton, Garnet L. Anderson, Michel Schummer
Selecting Differentially Expressed Genes From Microarray Experiments, Margaret S. Pepe, Gary M. Longton, Garnet L. Anderson, Michel Schummer
UW Biostatistics Working Paper Series
High throughput technologies, such as gene expression arrays and protein mass spectrometry, allow one to simultaneously evaluate thousands of potential biomarkers that distinguish different tissue types. Of particular interest here is cancer versus normal organ tissues. We consider statistical methods to rank genes (or proteins) in regards to differential expression between tissues. Various statistical measures are considered and we argue that two measures related to the Receiver Operating Characteristic Curve are particularly suitable for this purpose. We also propose that sampling variability in the gene rankings be quantified and suggest using the “selection probability function”, the probability distribution of rankings …
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Medical advances continue to provide new and potentially better means for detecting disease. Such is true in cancer, for example, where biomarkers are sought for early detection and where improvements in imaging methods may pick up the initial functional and molecular changes associated with cancer development. In other binary classification tasks, computational algorithms such as Neural Networks, Support Vector Machines and Evolutionary Algorithms have been applied to areas as diverse as credit scoring, object recognition, and peptide-binding prediction. Before a classifier becomes an accepted technology, it must undergo rigorous evaluation to determine its ability to discriminate between states. Characterization of …
The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai
The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai
UW Biostatistics Working Paper Series
The idea of using measurements such as biomarkers, clinical data, or molecular biology assays for classification and prediction is popular in modern medicine. The scientific evaluation of such measures includes assessing the accuracy with which they predict the outcome of interest. Receiver operating characteristic curves are commonly used for evaluating the accuracy of diagnostic tests. They can be applied more broadly, indeed to any problem involving classification to two states or populations (D = 0 or D = 1). We show that the ROC curve can be interpreted as a cumulative distribution function for the discriminatory measure Y in the …
Using The Id3 Symbolic Classification Algorithm To Reduce Data Density, Barry Fiachsbart, Daniel C. St. Clair, Jeff Holland
Using The Id3 Symbolic Classification Algorithm To Reduce Data Density, Barry Fiachsbart, Daniel C. St. Clair, Jeff Holland
Mathematics and Statistics Faculty Research & Creative Works
Effective data reduction is mandatory for modeling complex domains. The work described here demonstrates how to use a symbolic classifier algorithm from machine learning to effectively reduce large amounts of data. The algorithm, Quirdan's ID3, uses input data records and corresponding classifications to produce a decision tree. The resulting tree can be used to classify previously unseen inputs. Alternatively, the attributes found in the tree can be used as the basis to develop other system modeling techniques such as neural networks or mathematical programming algorithms. This approach has been used to effectively reduce data from a large complex domain. The …
Discriminant Function Analysis, Kuo Hsiung Su
Discriminant Function Analysis, Kuo Hsiung Su
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
The technique of discriminant function analysis was originated by R.A. Fisher and first applied by Barnard (1935). Two very useful summaries of the recent work in this technique can be found in Hodges (1950) and in Tosuoka and Tiedeman (1954). The techniques have been used primarily in the fields of anthropology, psychology, biology, medicine, and education, and have only begun to be applied to other fields in recent years.
Classification and discriminant function analyses are two phases in the attempt to predict which of several populations an observation might be a member of, on the basis of multivariate measurements. Both …
Objective Measurement Of Wool : Criteria, Methods And Materials, A Ingleton
Objective Measurement Of Wool : Criteria, Methods And Materials, A Ingleton
Journal of the Department of Agriculture, Western Australia, Series 4
An outline of some of the technical aspects of the objective measurement of wool—processes that will mean major cost savings to the wool industry.