Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Theory (134)
- Applied Statistics (110)
- Social and Behavioral Sciences (98)
- Statistical Methodology (45)
- Biostatistics (34)
-
- Statistical Models (29)
- Medicine and Health Sciences (24)
- Mathematics (22)
- Multivariate Analysis (22)
- Life Sciences (19)
- Survival Analysis (18)
- Applied Mathematics (16)
- Microarrays (14)
- Longitudinal Data Analysis and Time Series (13)
- Genetics and Genomics (12)
- Public Health (11)
- Epidemiology (10)
- Computational Biology (8)
- Computer Sciences (8)
- Numerical Analysis and Computation (8)
- Other Statistics and Probability (8)
- Bioinformatics (7)
- Genetics (7)
- Categorical Data Analysis (6)
- Clinical Trials (6)
- Law (6)
- Medical Specialties (6)
- Disease Modeling (5)
- Institution
-
- COBRA (101)
- Wayne State University (95)
- Brigham Young University (8)
- University of Nebraska - Lincoln (8)
- Missouri University of Science and Technology (7)
-
- Cornell University Law School (6)
- Wright State University (5)
- Loma Linda University (4)
- Air Force Institute of Technology (3)
- Department of Primary Industries and Regional Development, Western Australia (3)
- Marquette University (3)
- Old Dominion University (3)
- California Polytechnic State University, San Luis Obispo (2)
- Claremont Colleges (2)
- Cleveland State University (2)
- Dartmouth College (2)
- Montclair State University (2)
- New Jersey Institute of Technology (2)
- Southern Illinois University Carbondale (2)
- University of Dayton (2)
- University of Kentucky (2)
- East Tennessee State University (1)
- Indiana State University (1)
- Institute of Business Administration (1)
- Kennesaw State University (1)
- St. John Fisher University (1)
- Syracuse University (1)
- The University of San Francisco (1)
- University of Central Florida (1)
- University of Massachusetts Boston (1)
- Keyword
-
- Bootstrap (9)
- Classification (5)
- Empirical legal studies (5)
- Longitudinal data (5)
- Power (5)
-
- Sample size (5)
- Factor analysis (4)
- Prediction (4)
- Sensitivity (4)
- Type I error rate (4)
- Bayesian (3)
- Biased sampling (3)
- Confidence intervals (3)
- Discriminant analysis (3)
- Estimating equations (3)
- Gamma distribution (3)
- Informative follow-up (3)
- Interim analyses (3)
- Linear regression (3)
- Logistic regression (3)
- Missing data (3)
- Models (3)
- Nonnormality (3)
- Null distribution (3)
- Operating characteristics (3)
- P-value (3)
- Permutation test (3)
- Robustness (3)
- Sampling times process (3)
- Semiparametric regression (3)
- Publication
-
- Journal of Modern Applied Statistical Methods (93)
- UW Biostatistics Working Paper Series (32)
- U.C. Berkeley Division of Biostatistics Working Paper Series (29)
- Harvard University Biostatistics Working Paper Series (15)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (15)
-
- Theses and Dissertations (11)
- Department of Statistics: Faculty Publications (8)
- Mathematics and Statistics Faculty Research & Creative Works (7)
- Cornell Law Faculty Publications (6)
- Loma Linda University Electronic Theses, Dissertations & Projects (4)
- Mathematics and Statistics Faculty Publications (4)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (3)
- The University of Michigan Department of Biostatistics Working Paper Series (3)
- All Maxine Goodman Levin School of Urban Affairs Publications (2)
- Articles and Preprints (2)
- COBRA Preprint Series (2)
- Dartmouth Scholarship (2)
- Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works (2)
- Electronic Theses and Dissertations (2)
- Fisheries Occasional Publications (2)
- Mathematics & Statistics Faculty Publications (2)
- Mathematics Faculty Publications (2)
- Mathematics Faculty Research Publications (2)
- Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series (2)
- Pomona Faculty Publications and Research (2)
- Statistics (2)
- Theses (2)
- UPenn Biostatistics Working Papers (2)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (1)
- All-Inclusive List of Electronic Theses and Dissertations (1)
- Publication Type
Articles 211 - 240 of 279
Full-Text Articles in Statistics and Probability
Causal Inference In Longitudinal Studies With History-Restricted Marginal Structural Models, Romain Neugebauer, Mark J. Van Der Laan, Ira B. Tager
Causal Inference In Longitudinal Studies With History-Restricted Marginal Structural Models, Romain Neugebauer, Mark J. Van Der Laan, Ira B. Tager
U.C. Berkeley Division of Biostatistics Working Paper Series
Causal Inference based on Marginal Structural Models (MSMs) is particularly attractive to subject-matter investigators because MSM parameters provide explicit representations of causal effects. We introduce History-Restricted Marginal Structural Models (HRMSMs) for longitudinal data for the purpose of defining causal parameters which may often be better suited for Public Health research. This new class of MSMs allows investigators to analyze the causal effect of a treatment on an outcome based on a fixed, shorter and user-specified history of exposure compared to MSMs. By default, the latter represents the treatment causal effect of interest based on a treatment history defined by the …
Application Of The Time-Dependent Roc Curves For Prognostic Accuracy With Multiple Biomarkers, Yingye Zheng, Tianxi Cai, Ziding Feng
Application Of The Time-Dependent Roc Curves For Prognostic Accuracy With Multiple Biomarkers, Yingye Zheng, Tianxi Cai, Ziding Feng
UW Biostatistics Working Paper Series
The rapid advancement in molecule technology has lead to the discovery of many markers that have potential applications in disease diagnosis and prognosis. In a prospective cohort study, information on a panel of biomarkers as well as the disease status for a patient are routinely collected over time. Such information is useful to predict patients' prognosis and select patients for targeted therapy. In this paper, we develop procedures for constructing a composite test with optimal discrimination power when there are multiple markers available to assist in prediction and characterize the accuracy of the resulting test by extending the time-dependent receiver …
The Sensitivity And Specificity Of Markers For Event Times, Tianxi Cai, Margaret S. Pepe, Thomas Lumley, Yingye Zheng, Nancy Swords Jenny
The Sensitivity And Specificity Of Markers For Event Times, Tianxi Cai, Margaret S. Pepe, Thomas Lumley, Yingye Zheng, Nancy Swords Jenny
Harvard University Biostatistics Working Paper Series
No abstract provided.
Survival Ensembles, Torsten Hothorn, Peter Buhlmann, Sandrine Dudoit, Annette M. Molinaro, Mark J. Van Der Laan
Survival Ensembles, Torsten Hothorn, Peter Buhlmann, Sandrine Dudoit, Annette M. Molinaro, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a unified and flexible framework for ensemble learning in the presence of censoring. For right-censored data, we introduce a random forest algorithm and a generic gradient boosting algorithm for the construction of prognostic models. The methodology is utilized for predicting the survival time of patients suffering from acute myeloid leukemia based on clinical and genetic covariates. Furthermore, we compare the diagnostic capabilities of the proposed censored data random forest and boosting methods applied to the recurrence free survival time of node positive breast cancer patients with previously published findings.
The Bayesian Two-Sample T-Test, Mithat Gonen, Wesley O. Johnson, Yonggang Lu, Peter H. Westfall
The Bayesian Two-Sample T-Test, Mithat Gonen, Wesley O. Johnson, Yonggang Lu, Peter H. Westfall
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
In this article we show how the pooled-variance two-sample t-statistic arises from a Bayesian formulation of the two-sided point null testing problem, with emphasis on teaching. We identify a reasonable and useful prior giving a closed-form Bayes factor that can be written in terms of the distribution of the two-sample t-statistic under the null and alternative hypotheses respectively. This provides a Bayesian motivation for the two-sample t-statistic, which has heretofore been buried as a special case of more complex linear models, or given only roughly via analytic or Monte Carlo approximations. The resulting formulation of the Bayesian test is easy …
Nonparametric Estimation Of The Case Fatality Ratio With Competing Risks Data: An Application To Severe Acute Respiratory Syndome (Sars) , Nicholas P. Jewell, Xiudong Lei, A. C. Ghani, C. A. Donnelly, G. M. Leung, L. M. Ho, B. Cowling, A. J. Hedley
Nonparametric Estimation Of The Case Fatality Ratio With Competing Risks Data: An Application To Severe Acute Respiratory Syndome (Sars) , Nicholas P. Jewell, Xiudong Lei, A. C. Ghani, C. A. Donnelly, G. M. Leung, L. M. Ho, B. Cowling, A. J. Hedley
U.C. Berkeley Division of Biostatistics Working Paper Series
For diseases with some level of associated mortality, the case fatality ratio measures the proportion of diseased individuals who die from the disease. In principle, it is straightforward to estimate this quantity from individual follow-up data that provides times from onset to death or recovery. In particular, in a competing risks context, the case fatality ratio is defined by the limiting value of the sub-distribution function, associated with death, at infinity. When censoring is present, however, estimation of this quantity is complicated by the possibility of little information in the right tail of of the sub-distribution function, requiring use of …
Statistical Analysis Of Longitudinal And Multivariate Discrete Data, Deepak Mav
Statistical Analysis Of Longitudinal And Multivariate Discrete Data, Deepak Mav
Mathematics & Statistics Theses & Dissertations
Correlated multivariate Poisson and binary variables occur naturally in medical, biological and epidemiological longitudinal studies. Modeling and simulating such variables is difficult because the correlations are restricted by the marginal means via Fréchet bounds in a complicated way. In this dissertation we will first discuss partially specified models and methods for estimating the regression and correlation parameters. We derive the asymptotic distributions of these parameter estimates. Using simulations based on extensions of the algorithm due to Sim (1993, Journal of Statistical Computation and Simulation, 47, pp. 1–10), we study the performance of these estimates using infeasibility, coverage probabilities of the …
Additivity Of Information Value In Two-Act Linear Loss Decisions With Normal Priors, Jeffrey Keisler
Additivity Of Information Value In Two-Act Linear Loss Decisions With Normal Priors, Jeffrey Keisler
Management Science and Information Systems Faculty Publication Series
For the two-act linear loss decision problem with normal priors, conditions are derived for which the expected value of perfect information about two independent risks is super-additive in value. Several applications show how a variety of decision problems can reduce to the canonical problem, and how the general results obtained here can be translated simply to prescriptions for specific situations.
Resampling Based Multiple Testing Procedure Controlling Tail Probability Of The Proportion Of False Positives, Mark J. Van Der Laan, Merrill D. Birkner, Alan E. Hubbard
Resampling Based Multiple Testing Procedure Controlling Tail Probability Of The Proportion Of False Positives, Mark J. Van Der Laan, Merrill D. Birkner, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. In this article we propose a new resampling based multiple testing procedure asymptotically controlling the probability that the proportion of false positives among the set of rejections exceeds q at level alpha, where q and alpha are user supplied numbers. The procedure involves 1) specifying a conditional distribution for a guessed set of true null hypotheses, given the data, which asymptotically is degenerate at the true set of …
Instrument Cross-Comparisons And Automated Quality Control Of Atmospheric Radiation Measurement Data, S. Moore, Gary B. Hughes
Instrument Cross-Comparisons And Automated Quality Control Of Atmospheric Radiation Measurement Data, S. Moore, Gary B. Hughes
Statistics
Within the Atmospheric Radiation Measurement (ARM) instrument network, several different systems often measure the same quantity at the same site. For example, several ARM instruments measure time-series profiles of the atmosphere that were previously available only from balloon-borne radiosonde systems. These instruments include the Radar Wind Profilers (RWP) with Radio-Acoustic Sounding Systems (RASS), the Atmospheric Emitted Radiance Interferometer (AERI), the Microwave Radiometer Profiler (MWRP), and the Raman Lidar (RL). ARM researchers have described methods for direct cross-comparison of time-series profiles (Coulter and Lesht 1996; Turner et al. 1996) and we have extended this concept to the development of methods for …
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For two continuous-scale diagnostic tests, it is of interest to compare their sensitivities at a predetermined level of specificity. In this paper we propose three new intervals for the difference between two sensitivities at a fixed level of specificity. These intervals are easy to compute. We also conduct simulation studies to compare the relative performance of the new intervals with the existing normal approximation based interval proposed by Wieand et al (1989). Our simulation results show that the newly proposed intervals perform better than the existing normal approximation based interval in terms of coverage accuracy and interval length.
Improvements To And Status Of The Data Quality Health And Status System, K. Kehoe, K. Sonntag, R. Peppler, B. Burkholder, C. Shafer, M. Zaman, T. Thompson, S. Moore, Gary B. Hughes, K. Doty
Improvements To And Status Of The Data Quality Health And Status System, K. Kehoe, K. Sonntag, R. Peppler, B. Burkholder, C. Shafer, M. Zaman, T. Thompson, S. Moore, Gary B. Hughes, K. Doty
Statistics
The Atmospheric Radiation Measurement (ARM) Data Quality Office (DQO) has made a number of improvements and additions over the past year to its main tool for inspecting and assessing ARM data quality—the Data Quality Health and Status (DQ HandS) system (http://dq.arm.gov/). Among the improvements and additions, some of which are shown below, are the inclusion of ARM Mobile Facility (AMF) data; a new plot browser to facilitate the viewing of DQ HandS diagnostic plots; an improved method for writing and databasing weekly data quality assessment reports; a new automated daily alert, an improved method for searching ARM report databases (see …
Frequentist Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
Frequentist Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Group sequential stopping rules are often used as guidelines in the monitoring of clinical trials in order to address the ethical and efficiency issues inherent in human testing of a new treatment or preventive agent for disease. Such stopping rules have been proposed based on a variety of different criteria, both scientific (e.g., estimates of treatment effect) and statistical (e.g., frequentist type I error, Bayesian posterior probabilities, stochastic curtailment). It is easily shown, however, that a stopping rule based on one of those criteria induces a stopping rule on all other criteria. Thus the basis used to initially define a …
Bayesian Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
Bayesian Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Clincal trial designs often incorporate a sequential stopping rule to serve as a guide in the early termination of a study. When choosing a particular stopping rule, it is most common to examine frequentist operating characteristics such as type I error, statistical power, and precision of confi- dence intervals (Emerson, et al. [1]). Increasingly, however, clinical trials are designed and analyzed in the Bayesian paradigm. In this paper we describe how the Bayesian operating characteristics of a particular stopping rule might be evaluated and communicated to the scientific community. In particular, we consider a choice of probability models and a …
On The Use Of Stochastic Curtailment In Group Sequential Clinical Trials, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
On The Use Of Stochastic Curtailment In Group Sequential Clinical Trials, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Many different criteria have been proposed for the selection of a stopping rule for group sequen- tial trials. These include both scientific (e.g., estimates of treatment effect) and statistical (e.g., frequentist type I error, Bayesian posterior probabilities, stochastic curtailment) measures of the evidence for or against beneficial treatment effects. Because a stopping rule based on one of those criteria induces a stopping rule on all other criteria, the utility of any particular scale relates to the ease with which it allows a clinical trialist to search for sequential sampling plans having de- sirable operating characteristics. In this paper we examine …
A Causal Inference Approach For Constructing Transcriptional Regulatory Networks, Biao Xing, Mark J. Van Der Laan
A Causal Inference Approach For Constructing Transcriptional Regulatory Networks, Biao Xing, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Transcriptional regulatory networks specify the interactions among regulatory genes and between regulatory genes and their target genes. Discovering transcriptional regulatory networks helps us to understand the underlying mechanism of complex cellular processes and responses. In this paper, we describe a causal inference approach for constructing transcriptional regulatory networks using gene expression data, promoter sequences and information on transcription factor binding sites. The method rst identies active transcription factors under each individual experiment using a feature selection approach similar to Bussemaker et al. (2001), Keles et al. (2002) and Conlon et al. (2003). Transcription factors are viewed as `treatments' and gene …
A Statistical Framework For The Analysis Of Microarray Probe-Level Data, Zhijin Wu, Rafael A. Irizarry
A Statistical Framework For The Analysis Of Microarray Probe-Level Data, Zhijin Wu, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Microarrays are an example of the powerful high through-put genomics tools that are revolutionizing the measurement of biological systems. In this and other technologies, a number of critical steps are required to convert the raw measures into the data relied upon by biologists and clinicians. These data manipulations, referred to as preprocessing, have enormous influence on the quality of the ultimate measurements and studies that rely upon them. Many researchers have previously demonstrated that the use of modern statistical methodology can substantially improve accuracy and precision of gene expression measurements, relative to ad-hoc procedures introduced by designers and manufacturers of …
Acute Toxicity Testing Without Animals: More Scientific And Less Of A Gamble, Gillian R. Langley
Acute Toxicity Testing Without Animals: More Scientific And Less Of A Gamble, Gillian R. Langley
Application of Alternative Methods Collection
In this report, we argue specifically that acute toxicity data should not be sought from animal tests. The underlying principle of such tests on rats and mice is that the results can be effectively extrapolated to humans. In fact, after nearly 80 years of use of these tests, the predictivity of rodent data for human acute toxic effects has been disputed but never proven.
Customization Of Discriminant Function Analysis For Prediction Of Solar Flares, Evelyn A. Schumer
Customization Of Discriminant Function Analysis For Prediction Of Solar Flares, Evelyn A. Schumer
Theses and Dissertations
This research is an extension to the research conducted by K. Leka and G. Barnes of the Colorado Research Associates Division, Northwest Research Associates, Inc. in Boulder, Colorado (CORA) in which they found no single photospheric solar parameter they considered could sufficiently identify a flare-producing active region (AR). Their research then explored the possibility a linear combination of parameters used in a multivariable discriminant function (DF) could adequately predict solar activity. The purpose of this research is to extend the DF research conducted by Leka and Barnes by refining the method of statistical discriminant analysis (DA) with the goal of …
Sectorial Convergence Of U-Statistics, Anna Gadidov
Sectorial Convergence Of U-Statistics, Anna Gadidov
Faculty Articles
In this note we show that almost sure convergence to zero of symmetrized U-statistics indexed by a linear sector in ℤd+ is equivalent to convergence along the diagonal of ℤd+, as it is considered in Latała and Zinn [Ann. Probab. 28 (2000) 1908–1924]. Comparisons with similar results for sums of multi-indexed i.i.d. random variables are also made.
Judge-Jury Agreement In Criminal Cases: A Partial Replication Of Kalven And Zeisel's The American Jury, Theodore Eisenberg, Paula L. Hannaford-Agor, Valerie P. Hans, Nicole L. Waters, G. Thomas Munsterman, Stewart J. Schwab, Martin T. Wells
Judge-Jury Agreement In Criminal Cases: A Partial Replication Of Kalven And Zeisel's The American Jury, Theodore Eisenberg, Paula L. Hannaford-Agor, Valerie P. Hans, Nicole L. Waters, G. Thomas Munsterman, Stewart J. Schwab, Martin T. Wells
Cornell Law Faculty Publications
This study uses a new criminal case data set to partially replicate Kalven and Zeisel's classic study of judge-jury agreement. The data show essentially the same rate of judge-jury agreement as did Kalven and Zeisel for cases tried almost 50 years ago. This study also explores judge-jury agreement as a function of evidentiary strength (as reported by both judges and juries), evidentiary complexity (as reported by both judges and juries), legal complexity (as reported by judges), and locale. Regardless of which adjudicator's view of evidentiary strength is used, judges tend to convict more than juries in cases of "middle" evidentiary …
The Fate Of Firms: Explaining Mergers And Bankruptcies, Clas Bergström, Theodore Eisenberg, Stefan Sundgren, Martin T. Wells
The Fate Of Firms: Explaining Mergers And Bankruptcies, Clas Bergström, Theodore Eisenberg, Stefan Sundgren, Martin T. Wells
Cornell Law Faculty Publications
Using a uniquely complete data set of more than 50,000 observations of approximately 16,000 corporations, we test theories that seek to explain which firms become merger targets and which firms go bankrupt. We find that merger activity is much greater during prosperous periods than during recessions. In bad economic times, firms in industries with high bankruptcy rates are less likely to file for bankruptcy than they are in better years, supporting the market illiquidity arguments made by Shleifer and Vishny (1992). At the firm level, we find that, among poorly performing firms, the likelihood of merger increases with poorer performance, …
Implementation Of Estimating-Function Based Inference Procedures With Mcmc Sampler, Lu Tian, Jun S. Liu, L. J. Wei
Implementation Of Estimating-Function Based Inference Procedures With Mcmc Sampler, Lu Tian, Jun S. Liu, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Fixed-Width Output Analysis For Markov Chain Monte Carlo, Galin L. Jones, Murali Haran, Brian S. Caffo, Ronald Neath
Fixed-Width Output Analysis For Markov Chain Monte Carlo, Galin L. Jones, Murali Haran, Brian S. Caffo, Ronald Neath
Johns Hopkins University, Dept. of Biostatistics Working Papers
Markov chain Monte Carlo is a method of producing a correlated sample in order to estimate features of a complicated target distribution via simple ergodic averages. A fundamental question in MCMC applications is when should the sampling stop? That is, when are the ergodic averages good estimates of the desired quantities? We consider a method that stops the MCMC sampling the first time the width of a confidence interval based on the ergodic averages is less than a user-specified value. Hence calculating Monte Carlo standard errors is a critical step in assessing the output of the simulation. In particular, we …
Designs In Partially Controlled Studies: Messages From A Review, Fan Li, Constantine E. Frangakis
Designs In Partially Controlled Studies: Messages From A Review, Fan Li, Constantine E. Frangakis
Johns Hopkins University, Dept. of Biostatistics Working Papers
The ability to evaluate effects of factors on outcomes is increasingly important for a class of studies that control some but not all of the factors. Although important advances have been made in methods of analysis for such partially controlled studies,work on designs for such studies has been relatively limited. To help understand why, we review main designs that have been used for such partially controlled studies. Based on the review, we give two complementary reasons that explain the limited work on such designs, and suggest a new direction in this area.
Variable Selection For 1d Regression Models, David J. Olive, Douglas M. Hawkins
Variable Selection For 1d Regression Models, David J. Olive, Douglas M. Hawkins
Articles and Preprints
Variable selection, the search for j relevant predictor variables from a group of p candidates, is a standard problem in regression analysis. The class of 1D regression models is a broad class that includes generalized linear models. We show that existing variable selection algorithms, originally meant for multiple linear regression and based on ordinary least squares and Mallows’ Cp, can also be used for 1D models. Graphical aids for variable selection are also provided.
2d Quantitative Structure Activity Relationship Modeling Of Methylphenidate Analogues Using Algorithm And Partial Least Square Regression, Noureen Wadhwaniya
2d Quantitative Structure Activity Relationship Modeling Of Methylphenidate Analogues Using Algorithm And Partial Least Square Regression, Noureen Wadhwaniya
Theses
Quantitative Structure-Activity Relationship (QSAR) analysis attempts to develop a predictive model of biological activity based on molecular descriptors. 2D QSAR uses descriptors, such as topological indices, that are independent of molecular conformation. A genetic algorithm - partial least squares (GA-PLS) approach was used to identify the molecular descriptors that correlate to the biological activity (binding affinity) of a set of 80 methylphenidate analogues and to construct a predictive model. The GA code was implemented using the fitness function (1-(n-1)(1-q2)/ (n - c)), where n is the number of compounds, c is the optimal number of components, and q …
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
UW Biostatistics Working Paper Series
Identification of differentially expressed genes and clustering of genes are two important and complementary objectives addressed with gene expression data. For the differential expression question, many "per-gene" analytic methods have been proposed. These methods can generally be characterized as using a regression function to independently model the observations for each gene; various adjustments for multiplicity are then used to interpret the statistical significance of these per-gene regression models over the collection of genes analyzed. Motivated by this common structure of per-gene models, we propose a new model-based clustering method -- the clustering of regression models method, which groups genes that …
Insights Into Latent Class Analysis, Margaret S. Pepe, Holly Janes
Insights Into Latent Class Analysis, Margaret S. Pepe, Holly Janes
UW Biostatistics Working Paper Series
Latent class analysis is a popular statistical technique for estimating disease prevalence and test sensitivity and specificity. It is used when a gold standard assessment of disease is not available but results of multiple imperfect tests are. We derive analytic expressions for the parameter estimates in terms of the raw data, under the conditional independence assumption. These expressions indicate explicitly how observed two- and three-way associations between test results are used to infer disease prevalence and test operating characteristics. Although reasonable if the conditional independence model holds, the estimators have no basis when it fails. We therefore caution against using …
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
UW Biostatistics Working Paper Series
Introduction: Markers that purport to distinguish subjects with a condition from those without a condition must be evaluated rigorously for their classification accuracy. A single approach to statistically evaluating and comparing markers is not yet established.
Methods: We suggest a standardization that uses the marker distribution in unaffected subjects as a reference. For an affected subject with marker value Y, the standardized placement value is the proportion of unaffected subjects with marker values that exceed Y.
Results: We apply the standardization to two illustrative datasets. In patients with pancreatic cancer placement values calculated for the CA 19-9 marker are smaller …