Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2010

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 241 - 270 of 371

Full-Text Articles in Statistics and Probability

Publication Bias In Reports Of Animal Stroke Studies Leads To Major Overstatement Of Efficacy, Emily Sena, H. Bart Van Der Worp, Philip M.W. Bath, David W. Howells, Malcolm Macleod Mar 2010

Publication Bias In Reports Of Animal Stroke Studies Leads To Major Overstatement Of Efficacy, Emily Sena, H. Bart Van Der Worp, Philip M.W. Bath, David W. Howells, Malcolm Macleod

Validation of Animal Experimentation Collection

The consolidation of scientific knowledge proceeds through the interpretation and then distillation of data presented in research reports, first in review articles and then in textbooks and undergraduate courses, until truths become accepted as such both amongst “experts” and in the public understanding. Where data are collected but remain unpublished, they cannot contribute to this distillation of knowledge. If these unpublished data differ substantially from published work, conclusions may not reflect adequately the underlying biological effects being described. The existence and any impact of such “publication bias” in the laboratory sciences have not been described. Using the CAMARADES (Collaborative Approach …


An Analysis Of Nonignorable Nonresponse In A Survey With A Rotating Panel Design, Caterina Giusti, Roderick J. Little Mar 2010

An Analysis Of Nonignorable Nonresponse In A Survey With A Rotating Panel Design, Caterina Giusti, Roderick J. Little

The University of Michigan Department of Biostatistics Working Paper Series

Missing values to income questions are common in survey data. When the probabilities of nonresponse are assumed to depend on the observed information and not on the underlining unobserved amounts, the missing income values are missing at random (MAR), and methods such as sequential multiple imputation can be applied. However, the MAR assumption is often considered questionable in this context, since missingness of income is thought to be related to the value of income itself, after conditioning on available covariates. In this article we describe a sensitivity analysis based on a pattern-mixture model for deviations from MAR, in the context …


Statistical Learning And Behrens-Fisher Distribution Methods For Heteroscedastic Data In Microarray Analysis, Nabin K. Manandhr-Shrestha Mar 2010

Statistical Learning And Behrens-Fisher Distribution Methods For Heteroscedastic Data In Microarray Analysis, Nabin K. Manandhr-Shrestha

USF Tampa Graduate Theses and Dissertations

The aim of the present study is to identify the di®erentially expressed genes be- tween two di®erent conditions and apply it in predicting the class of new samples using the microarray data. Microarray data analysis poses many challenges to the statis- ticians because of its high dimensionality and small sample size, dubbed as "small n large p problem". Microarray data has been extensively studied by many statisticians and geneticists. Generally, it is said to follow a normal distribution with equal vari- ances in two conditions, but it is not true in general. Since the number of replications is very small, …


Panel Count Data Regression With Informative Observation Times, Petra Buzkova Mar 2010

Panel Count Data Regression With Informative Observation Times, Petra Buzkova

UW Biostatistics Working Paper Series

When patients are monitored for potentially recurrent events such as infections or tumor metastases, it is common for clinicians to ask patients to come back sooner for follow-up based on the results of the most recent exam. This means that subjects’ observation times will be irregular and related to subject-specific factors. Previously proposed methods for handling such panel count data assume that the dependence between the events process and the observation time process is time-invariant. This article considers situations where the observation times are predicted by time-varying factors, such as the outcome observed at the last visit or cumulative exposure. …


Extensions Of Nearest Shrunken Centroid Method For Classification, Tomohiko Funai Mar 2010

Extensions Of Nearest Shrunken Centroid Method For Classification, Tomohiko Funai

Theses and Dissertations

Stylometry assumes that the essence of the individual style of an author can be captured using a number of quantitative criteria, such as the relative frequencies of noncontextual words (e.g., or, the, and, etc.). Several statistical methodologies have been developed for authorship analysis. Jockers et al. (2009) utilize Nearest Shrunken Centroid (NSC) classification, a promising classification methodology in DNA microarray analysis for authorship analysis of the Book of Mormon. Schaalje et al. (2010) develop an extended NSC classification to remedy the problem of a missing author. Dabney (2005) and Koppel et al. (2009) suggest other modifications of NSC. This paper …


Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang Mar 2010

Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider likelihood ratio tests (LRT) and their modifications for homogeneity in admixture models. The admixture model is a special case of two component mixture model, where one component is indexed by an unknown parameter while the parameter value for the other component is known. It has been widely used in genetic linkage analysis under heterogeneity, in which the kernel distribution is binomial. For such models, it is long recognized that testing for homogeneity is nonstandard and the LRT statistic does not converge to a conventional 2 distribution. In this paper, we investigate the asymptotic behavior of the LRT for …


Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly Mar 2010

Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly

Theses and Dissertations

In an ever-increasing data rich environment, actionable information must be extracted, filtered, and correlated from massive amounts of disparate often free text sources. The usefulness of the retrieved information depends on how we accomplish these steps and present the most relevant information to the analyst. One method for extracting information from free text is Latent Dirichlet Allocation (LDA), a document categorization technique to classify documents into cohesive topics. Although LDA accounts for some implicit relationships such as synonymy (same meaning) it often ignores other semantic relationships such as polysemy (different meanings), hyponym (subordinate), meronym (part of), and troponomys (manner). To …


Poverty, Vulnerability, And Provision Of Healthcare In Afghanistan, Jean-Francois Trani, Parul Bakhshi, Ayan A. Noor, Dominque Lopez, Ashraf Mashkoor Mar 2010

Poverty, Vulnerability, And Provision Of Healthcare In Afghanistan, Jean-Francois Trani, Parul Bakhshi, Ayan A. Noor, Dominque Lopez, Ashraf Mashkoor

Brown School Faculty Publications

This paper presents findings on conditions of healthcare delivery in Afghanistan. There is an ongoing debate about barriers to healthcare in low-income as well as fragile states. In 2002, the Government of Afghanistan established a Basic Package of Health Services (BPHS), contracting primary healthcare delivery to non-state providers. The priority was to give access to the most vulnerable groups: women, children, disabled persons, and the poorest households. In 2005, we conducted a nationwide survey, and using a logistic regression model, investigated provider choice. We also measured associations between perceived availability and usefulness of healthcare providers. Our results indicate that the …


Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway Mar 2010

Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway

Theses and Dissertations

The normal distribution is a commonly seen distribution in nature, education, and business. Data that are mounded or bell shaped are easily found across various fields of study. Although there is high utility with the normal distribution; often the full range can not be observed. The truncated normal distribution accounts for the inability to observe the full range and allows for inferring back to the original population. Depending on the amount of truncation, the truncated normal has several distinct shapes. A simulation study evaluating the performance of the maximum likelihood estimators and method of moment estimators is conducted and a …


Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi Mar 2010

Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi

Theses and Dissertations

In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …


Finite Element Approximations For Stokes-Darcy Flow With Beavers-Joseph Interface Conditions, Yanzhao Cao, Max Gunzburger, Xiaolong Hu, Fei Hua, Xiaoming Wang, Weidong Zhao Mar 2010

Finite Element Approximations For Stokes-Darcy Flow With Beavers-Joseph Interface Conditions, Yanzhao Cao, Max Gunzburger, Xiaolong Hu, Fei Hua, Xiaoming Wang, Weidong Zhao

Mathematics and Statistics Faculty Research & Creative Works

Numerical solutions using finite element methods are considered for transient flow in a porous medium coupled to free flow in embedded conduits. Such situations arise, for example, for groundwater flows in karst aquifers. the coupled flow is modeled by the Darcy equation in a porous medium and the Stokes equations in the conduit domain. on the interface between the matrix and conduit, Beavers-Joseph interface conditions, instead of the simplified Beavers-Joseph-Saffman conditions, are imposed. Convergence and error estimates for finite element approximations are obtained. Numerical experiments illustrate the validity of the theoretical results. © 2010 Society for Industrial and Applied Mathematics.


Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan Mar 2010

Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal and repeated measures data analysis, often the goal is to determine the effect of a treatment or aspect on a particular outcome (e.g. disease progression). We consider semiparametric repeated measures regression model, where the parametric component models effect of the variable of interest and any modification by other covariates. The expectation of this parametric component over the other covariates is a measure of variable importance. Here we present a targeted maximum likelihood estimator of the finite dimensional regression parameter, which is easily estimated using standard software for generalized estimating equations. The targeted maximum likelihood method provides double robust …


Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei Mar 2010

Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. Van Der Laan Mar 2010

Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Current methods used to analyze time to event data either, rely on highly parametric assumptions which result in biased estimates of parameters which are purely chosen out of convenience, or are highly unstable because they ignore the global constraints of the true model. By using Targeted Maximum Likelihood Estimation one may consistently estimate parameters which directly answer the statistical question of interest. Targeted Maximum Likelihood Estimators are substitution estimators, which rely on estimating the underlying distribution. However, unlike other substitution estimators, the underlying distribution is estimated specifically to reduce bias in the estimate of the parameter of interest. We will …


Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu Mar 2010

Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu

The University of Michigan Department of Biostatistics Working Paper Series

The linear mixed effects model (LMM) is widely used in the analysis of clustered or longitudinal data. In the practice of LMM, the inference on the structure of the random effects component is of great importance, not only to yield proper interpretation of subject-specific effects but also to draw valid statistical conclusions. This task of inference becomes significantly challenging when a large number of fixed effects and random effects are involved in the analysis. The difficulty of variable selection arises from the need of simultaneously regularizing both mean model and covariance structures, with possible parameter constraints between the two. In …


Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu Mar 2010

Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The basic observational unit in this paper is a function. Data are assumed to have a natural hierarchy of basic units. A simple example is when functions are recorded at multiple visits for the same subject. Di et al. (2009) proposed Multilevel Functional Principal Component Analysis (MFPCA) for this type of data structure when functions are densely sampled. Here we consider the case when functions are sparsely sampled and may contain as few as 2 or 3 observations per function. As with MFPCA, we exploit the multilevel structure of covariance operators and data reduction induced by the use of principal …


A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao Mar 2010

A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Power Analysis Using Sas And R, Peter Osmena Mar 2010

Statistical Power Analysis Using Sas And R, Peter Osmena

Statistics

Statistical power is something that has to be considered when designing an experiment. The power is used to determine the usefulness of the test. In this paper the concept of power and what it is will be discussed. A general ANOVA test and a Chi Squared test will be discussed in greater depth. Computers make these power calculations relatively easy to compute, considering they are used right. SAS and R both have the capabilities to make these calculations for a different variety of tests. Calculating power for a general ANOVA test and a Chi Squared test using these programs are …


Monopoly, Regulation, And Innovation, Matt Bogard Mar 2010

Monopoly, Regulation, And Innovation, Matt Bogard

Economics Faculty Publications

Recently the Justice department has started investigations into alleged anti-trust violations by Monsanto. This has helped fuel a lot of already hyped discontent with one of the world’s leaders in innovative solutions for sustainable agriculture. This article discusses how the regulatory environment could possibly have contributed to more concentration and power in the biotech industry. Increasing regulation would likely have the opposite effect of creating a level playing field in the agriculture industry. From AgWeb, March 27,2010 http://www.agweb.com/blog/Economic_Sense_190/Monopoly_Regulation__and_Innovation_10771/


Culture, Acculturation, And Social Capital: Latinos And Use Of Mental Health Services, Edward Mcfield Jr. Mar 2010

Culture, Acculturation, And Social Capital: Latinos And Use Of Mental Health Services, Edward Mcfield Jr.

Loma Linda University Electronic Theses, Dissertations & Projects

Studies suggest that the prevalence of mental illness in Latinos is similar to that of other groups; however, Latinos are less likely than non-Latino whites to access mental health services, and when they do, the quality of care is poor. To better understand the factors that influence use of mental health services among Latinos, a descriptive, cross-sectional, correlational study was conducted (N= 340; 219 non-consumers and 121 consumers), which examined the association between social capital, acculturation, cultural beliefs or explanatory models of illness, stigma, need, and mental health service use. An innovative integrative model drawing from Andersen's Behavioral …


Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ Mar 2010

Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ

Master's Theses

View John Huân Vũ's thesis presentation at http://youtu.be/y3bzNmkTr-c.

In 2001, the ACM and IEEE Computing Curriculum stated that it was necessary to address "the need to develop implementation models that are international in scope and could be practiced in universities around the world." With increasing connectivity through the internet, the move towards a global economy and growing use of technology places software internationalization as a more important concern for developers. However, there has been a "clear shortage in terms of numbers of trained persons applying for entry-level positions" in this area. Eric Brechner, Director of Microsoft Development Training, suggested …


Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan Feb 2010

Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the asymptotic study of adaptive group sequential designs in the case of randomized clinical trials with binary treatment, binary outcome and no covariate. By adaptive design, we mean in this setting a clinical trial design that allows the investigator to dynamically modify its course through data-driven adjustment of the randomization probability based on data accrued so far, without negatively impacting on the statistical integrity of the trial. By adaptive group sequential design, we refer to the fact that group sequential testing methods can be equally well applied on top of adaptive designs. Prior to collection …


Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan Feb 2010

Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Given causal graph assumptions, intervention-specific counterfactual distributions of the data can be defined by the so called G-computation formula, which is obtained by carrying out these interventions on the likelihood of the data factorized according to the causal graph. The obtained G-computation formula represents the counterfactual distribution the data would have had if this intervention would have been enforced on the system generating the data. A causal effect of interest can now be defined as some difference between these counterfactual distributions indexed by different interventions. For example, the interventions can represent static treatment regimens or individualized treatment rules that assign …


Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson Feb 2010

Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson

UW Biostatistics Working Paper Series

After a non-inferiority clinical trial, a new therapy may be accepted as effective, even if its treatment effect is slightly smaller than the current standard. It is therefore possible that, after a series of trials where the new therapy is slightly worse than the preceding drugs, an ineffective or harmful therapy might be incorrectly declared efficacious; this is known as “bio-creep.” Several factors may influence the rate at which bio-creep occurs, including the distribution of the effects of the new agents being tested and how that changes over time, the choice of active comparator, the method used to model the …


Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson Feb 2010

Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson

UW Biostatistics Working Paper Series

In group sequential clinical trials, it is necessary to estimate the amount of information present at interim analysis times relative to the amount of information that would be present at the final analysis. If only one measurement is made per individual, this is often the ratio of sample sizes available at the interim and final analyses. However, as discussed by Wu and Lan (1992), when the statistic of interest is a change over time, as with longitudinal data, such an approach overstates the information. In this paper, we discuss other problems that can result in overestimating the information, such as …


Effects Of Socioeconomic Status On Brain Development, And How Cognitive Neuroscience May Contribute To Levelling The Playing Field, Rajeev Raizada, Mark M. Kishiyama Feb 2010

Effects Of Socioeconomic Status On Brain Development, And How Cognitive Neuroscience May Contribute To Levelling The Playing Field, Rajeev Raizada, Mark M. Kishiyama

Dartmouth Scholarship

The study of socioeconomic status (SES) and the brain finds itself in a circumstance unusual for Cognitive Neuroscience: large numbers of questions with both practical and scientific importance exist, but they are currently under-researched and ripe for investigation. This review aims to highlight these questions, to outline their potential significance, and to suggest routes by which they might be approached. Although remarkably few neural studies have been carried out so far, there exists a large literature of previous behavioural work. This behavioural research provides an invaluable guide for future neuroimaging work, but also poses an important challenge for it: how …


Research Poster: Climate Prediction Downscaling Of Temperature And Precipitation In The Great Basin Region, Ramesh Vellore, Benjamin J. Hatchett, Darko Koracin Feb 2010

Research Poster: Climate Prediction Downscaling Of Temperature And Precipitation In The Great Basin Region, Ramesh Vellore, Benjamin J. Hatchett, Darko Koracin

2010 Annual Nevada NSF EPSCoR Climate Change Conference

Research poster


Research Poster: Hydrological Impacts Of Climate Change On Colorado Basin, Peng Jiang, Zhongbo Yu Feb 2010

Research Poster: Hydrological Impacts Of Climate Change On Colorado Basin, Peng Jiang, Zhongbo Yu

2010 Annual Nevada NSF EPSCoR Climate Change Conference

Research poster


Research Poster: An Overview Of Progress In Nsf Epscor Project Entitled, “Reducing Cloud Uncertainties In Climate Models”, Subhashree Mishra, David L. Mitchell, W. Patrick Arnott Feb 2010

Research Poster: An Overview Of Progress In Nsf Epscor Project Entitled, “Reducing Cloud Uncertainties In Climate Models”, Subhashree Mishra, David L. Mitchell, W. Patrick Arnott

2010 Annual Nevada NSF EPSCoR Climate Change Conference

Research poster


Applied Statistics: Experience & Cerification In Quality Assurance, Huey D. Dodson Feb 2010

Applied Statistics: Experience & Cerification In Quality Assurance, Huey D. Dodson

Statistics

The composition of my senior project can be broken down into two parts. The first part of my project, without which the second could not be pursued, involved a 13 week internship at a produce processing facility where I took part in several projects varying in scope and type. The second part was to acquire certification as a Quality Process Analyst from the American Society for Quality.

This document is structured to represent the dichotomous nature of my project; the first section is dedicated to my internship experience, and the second dedicated to the certification examination preparation and completion.