Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,804 Full-Text Articles 23,873 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,804 full-text articles. Page 480 of 486.

Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang 2010 Fred Hutchinson Cancer Research Center

Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider likelihood ratio tests (LRT) and their modifications for homogeneity in admixture models. The admixture model is a special case of two component mixture model, where one component is indexed by an unknown parameter while the parameter value for the other component is known. It has been widely used in genetic linkage analysis under heterogeneity, in which the kernel distribution is binomial. For such models, it is long recognized that testing for homogeneity is nonstandard and the LRT statistic does not converge to a conventional 2 distribution. In this paper, we investigate the asymptotic behavior of the LRT for …


Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly 2010 Air Force Institute of Technology

Augmenting Latent Dirichlet Allocation And Rank Threshold Detection With Ontologies, Laura A. Isaly

Theses and Dissertations

In an ever-increasing data rich environment, actionable information must be extracted, filtered, and correlated from massive amounts of disparate often free text sources. The usefulness of the retrieved information depends on how we accomplish these steps and present the most relevant information to the analyst. One method for extracting information from free text is Latent Dirichlet Allocation (LDA), a document categorization technique to classify documents into cohesive topics. Although LDA accounts for some implicit relationships such as synonymy (same meaning) it often ignores other semantic relationships such as polysemy (different meanings), hyponym (subordinate), meronym (part of), and troponomys (manner). To …


Poverty, Vulnerability, And Provision Of Healthcare In Afghanistan, Jean-Francois Trani, Parul Bakhshi, Ayan A. Noor, Dominque Lopez, Ashraf Mashkoor 2010 Washington University in St. Louis, George Warren Brown School

Poverty, Vulnerability, And Provision Of Healthcare In Afghanistan, Jean-Francois Trani, Parul Bakhshi, Ayan A. Noor, Dominque Lopez, Ashraf Mashkoor

Brown School Faculty Publications

This paper presents findings on conditions of healthcare delivery in Afghanistan. There is an ongoing debate about barriers to healthcare in low-income as well as fragile states. In 2002, the Government of Afghanistan established a Basic Package of Health Services (BPHS), contracting primary healthcare delivery to non-state providers. The priority was to give access to the most vulnerable groups: women, children, disabled persons, and the poorest households. In 2005, we conducted a nationwide survey, and using a logistic regression model, investigated provider choice. We also measured associations between perceived availability and usefulness of healthcare providers. Our results indicate that the …


Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway 2010 Brigham Young University - Provo

Parameter Estimation And Hypothesis Testing For The Truncated Normal Distribution With Applications To Introductory Statistics Grades, James T. Hattaway

Theses and Dissertations

The normal distribution is a commonly seen distribution in nature, education, and business. Data that are mounded or bell shaped are easily found across various fields of study. Although there is high utility with the normal distribution; often the full range can not be observed. The truncated normal distribution accounts for the inability to observe the full range and allows for inferring back to the original population. Depending on the amount of truncation, the truncated normal has several distinct shapes. A simulation study evaluating the performance of the maximum likelihood estimators and method of moment estimators is conducted and a …


Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi 2010 Virginia Commonwealth University

Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi

Theses and Dissertations

In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …


Finite Element Approximations For Stokes-Darcy Flow With Beavers-Joseph Interface Conditions, Yanzhao Cao, Max Gunzburger, Xiaolong Hu, Fei Hua, Xiaoming Wang, Weidong Zhao 2010 Missouri University of Science and Technology

Finite Element Approximations For Stokes-Darcy Flow With Beavers-Joseph Interface Conditions, Yanzhao Cao, Max Gunzburger, Xiaolong Hu, Fei Hua, Xiaoming Wang, Weidong Zhao

Mathematics and Statistics Faculty Research & Creative Works

Numerical solutions using finite element methods are considered for transient flow in a porous medium coupled to free flow in embedded conduits. Such situations arise, for example, for groundwater flows in karst aquifers. the coupled flow is modeled by the Darcy equation in a porous medium and the Stokes equations in the conduit domain. on the interface between the matrix and conduit, Beavers-Joseph interface conditions, instead of the simplified Beavers-Joseph-Saffman conditions, are imposed. Convergence and error estimates for finite element approximations are obtained. Numerical experiments illustrate the validity of the theoretical results. © 2010 Society for Industrial and Applied Mathematics.


Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. van der Laan 2010 Division of Biostatistics, School of Public Health, University of California, Berkeley

Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal and repeated measures data analysis, often the goal is to determine the effect of a treatment or aspect on a particular outcome (e.g. disease progression). We consider semiparametric repeated measures regression model, where the parametric component models effect of the variable of interest and any modification by other covariates. The expectation of this parametric component over the other covariates is a measure of variable importance. Here we present a targeted maximum likelihood estimator of the finite dimensional regression parameter, which is easily estimated using standard software for generalized estimating equations. The targeted maximum likelihood method provides double robust …


Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei 2010 Dana Farber Cancer Institute

Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. van der Laan 2010 University of California - Berkeley

Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Current methods used to analyze time to event data either, rely on highly parametric assumptions which result in biased estimates of parameters which are purely chosen out of convenience, or are highly unstable because they ignore the global constraints of the true model. By using Targeted Maximum Likelihood Estimation one may consistently estimate parameters which directly answer the statistical question of interest. Targeted Maximum Likelihood Estimators are substitution estimators, which rely on estimating the underlying distribution. However, unlike other substitution estimators, the underlying distribution is estimated specifically to reduce bias in the estimate of the parameter of interest. We will …


Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu 2010 University of Michigan

Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu

The University of Michigan Department of Biostatistics Working Paper Series

The linear mixed effects model (LMM) is widely used in the analysis of clustered or longitudinal data. In the practice of LMM, the inference on the structure of the random effects component is of great importance, not only to yield proper interpretation of subject-specific effects but also to draw valid statistical conclusions. This task of inference becomes significantly challenging when a large number of fixed effects and random effects are involved in the analysis. The difficulty of variable selection arises from the need of simultaneously regularizing both mean model and covariance structures, with possible parameter constraints between the two. In …


Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu 2010 Division of Public Health Sciences, Fred Hutchinson Cancer Research Center

Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The basic observational unit in this paper is a function. Data are assumed to have a natural hierarchy of basic units. A simple example is when functions are recorded at multiple visits for the same subject. Di et al. (2009) proposed Multilevel Functional Principal Component Analysis (MFPCA) for this type of data structure when functions are densely sampled. Here we consider the case when functions are sparsely sampled and may contain as few as 2 or 3 observations per function. As with MFPCA, we exploit the multilevel structure of covariance operators and data reduction induced by the use of principal …


A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao 2010 Harvard University and Dana Farber Cancer Institute

A New Class Of Dantzig Selectors For Censored Linear Regression Models, Yi Li, Lee Dicker, Sihai Dave Zhao

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Power Analysis Using Sas And R, Peter Osmena 2010 California Polytechnic State University, San Luis Obispo

Statistical Power Analysis Using Sas And R, Peter Osmena

Statistics

Statistical power is something that has to be considered when designing an experiment. The power is used to determine the usefulness of the test. In this paper the concept of power and what it is will be discussed. A general ANOVA test and a Chi Squared test will be discussed in greater depth. Computers make these power calculations relatively easy to compute, considering they are used right. SAS and R both have the capabilities to make these calculations for a different variety of tests. Calculating power for a general ANOVA test and a Chi Squared test using these programs are …


Monopoly, Regulation, And Innovation, Matt Bogard 2010 Western Kentucky University

Monopoly, Regulation, And Innovation, Matt Bogard

Economics Faculty Publications

Recently the Justice department has started investigations into alleged anti-trust violations by Monsanto. This has helped fuel a lot of already hyped discontent with one of the world’s leaders in innovative solutions for sustainable agriculture. This article discusses how the regulatory environment could possibly have contributed to more concentration and power in the biotech industry. Increasing regulation would likely have the opposite effect of creating a level playing field in the agriculture industry. From AgWeb, March 27,2010 http://www.agweb.com/blog/Economic_Sense_190/Monopoly_Regulation__and_Innovation_10771/


Culture, Acculturation, And Social Capital: Latinos And Use Of Mental Health Services, Edward McField Jr. 2010 Loma Linda University

Culture, Acculturation, And Social Capital: Latinos And Use Of Mental Health Services, Edward Mcfield Jr.

Loma Linda University Electronic Theses, Dissertations & Projects

Studies suggest that the prevalence of mental illness in Latinos is similar to that of other groups; however, Latinos are less likely than non-Latino whites to access mental health services, and when they do, the quality of care is poor. To better understand the factors that influence use of mental health services among Latinos, a descriptive, cross-sectional, correlational study was conducted (N= 340; 219 non-consumers and 121 consumers), which examined the association between social capital, acculturation, cultural beliefs or explanatory models of illness, stigma, need, and mental health service use. An innovative integrative model drawing from Andersen's Behavioral …


Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ 2010 California Polytechnic State University, San Luis Obispo

Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ

Master's Theses

View John Huân Vũ's thesis presentation at http://youtu.be/y3bzNmkTr-c.

In 2001, the ACM and IEEE Computing Curriculum stated that it was necessary to address "the need to develop implementation models that are international in scope and could be practiced in universities around the world." With increasing connectivity through the internet, the move towards a global economy and growing use of technology places software internationalization as a more important concern for developers. However, there has been a "clear shortage in terms of numbers of trained persons applying for entry-level positions" in this area. Eric Brechner, Director of Microsoft Development Training, suggested …


Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. van der Laan 2010 Laboratoire MAP5, Université Paris Descartes and CNRS

Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the asymptotic study of adaptive group sequential designs in the case of randomized clinical trials with binary treatment, binary outcome and no covariate. By adaptive design, we mean in this setting a clinical trial design that allows the investigator to dynamically modify its course through data-driven adjustment of the randomization probability based on data accrued so far, without negatively impacting on the statistical integrity of the trial. By adaptive group sequential design, we refer to the fact that group sequential testing methods can be equally well applied on top of adaptive designs. Prior to collection …


Targeted Maximum Likelihood Based Causal Inference, Mark J. van der Laan 2010 University of California - Berkeley

Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Given causal graph assumptions, intervention-specific counterfactual distributions of the data can be defined by the so called G-computation formula, which is obtained by carrying out these interventions on the likelihood of the data factorized according to the causal graph. The obtained G-computation formula represents the counterfactual distribution the data would have had if this intervention would have been enforced on the system generating the data. A causal effect of interest can now be defined as some difference between these counterfactual distributions indexed by different interventions. For example, the interventions can represent static treatment regimens or individualized treatment rules that assign …


Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson 2010 University of Washington - Seattle Campus

Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson

UW Biostatistics Working Paper Series

After a non-inferiority clinical trial, a new therapy may be accepted as effective, even if its treatment effect is slightly smaller than the current standard. It is therefore possible that, after a series of trials where the new therapy is slightly worse than the preceding drugs, an ineffective or harmful therapy might be incorrectly declared efficacious; this is known as “bio-creep.” Several factors may influence the rate at which bio-creep occurs, including the distribution of the effects of the new agents being tested and how that changes over time, the choice of active comparator, the method used to model the …


Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson 2010 University of Washington

Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson

UW Biostatistics Working Paper Series

In group sequential clinical trials, it is necessary to estimate the amount of information present at interim analysis times relative to the amount of information that would be present at the final analysis. If only one measurement is made per individual, this is often the ratio of sample sizes available at the interim and final analyses. However, as discussed by Wu and Lan (1992), when the statistic of interest is a change over time, as with longitudinal data, such an approach overstates the information. In this paper, we discuss other problems that can result in overestimating the information, such as …


Digital Commons powered by bepress