Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,804 Full-Text Articles 23,873 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,804 full-text articles. Page 406 of 486.

Statistical Models For Gene And Transcripts Quantification And Identification Using Rna-Seq Technology, Han Wu 2013 Purdue University

Statistical Models For Gene And Transcripts Quantification And Identification Using Rna-Seq Technology, Han Wu

Open Access Dissertations

RNA-Seq has emerged as a powerful technique for transcriptome study. As much as the improved sensitivity and coverage, RNA-Seq also brings challenges for data analysis. The massive amount of sequence reads data, excessive variability, uncertainties, and bias and noises stemming from multiple sources all make the analysis of RAN-Seq data difficult. Despite much progress, RNA-Seq data analysis still has much room for improvement, especially on the quantification of gene and transcript expression levels. The quantification of gene expression level is a direct inference problem, whereas the quantification of the transcript expression level is an indirect problem, because the label of …


Non-Parametric Spatial Models, Cheng Liu 2013 Purdue University

Non-Parametric Spatial Models, Cheng Liu

Open Access Dissertations

Covariance functions play a central role in spatial statistics. Parametric covariance functions have been used in most of the existing works on the analysis of spatial data. The primary reason for this is that the classes of parametric covariance functions guarantee that the fitted covariance function is positive definite. In this dissertation, I undertake two non-parametric approaches to modelling the covariance functions.

Our approach is motivated by problems that arise in spatial data analysis in recent years. First, it is nontrivial to choose a parametric family among many parametric families of covariance function. A non-parametric covariance function circumvents this problem. …


A Jackknife Empirical Likelihood Approach To Goodness Of Fit U-Statistic Testing With Side Information, Qun Lin 2013 Purdue University

A Jackknife Empirical Likelihood Approach To Goodness Of Fit U-Statistic Testing With Side Information, Qun Lin

Open Access Dissertations

Motivated by applications to goodness of fit U-statistics testing, the jackknife empirical likelihood of Jing, et al. (2009) is justified with an alternative approach, and the Wilks theorem for vector U-statistics is proved. This generalizes Owen's empirical likelihood theorem for a vector mean to a vector U-statistics-based mean and includes the jackknife empirical likelihood of U-statistics with side information as a special case. The results are generalized to allow for the constraints to use estimated criteria functions and for the number of constraints to grow with the sample size. The latter is needed to handle naturally occurring nuisance parameters in …


Generation And Statistical Modeling Of Active Protein Chimeras: A Sequence Based Approach, Nicholas Fico 2013 Purdue University

Generation And Statistical Modeling Of Active Protein Chimeras: A Sequence Based Approach, Nicholas Fico

Open Access Dissertations

Generation of active protein chimeras is a valuable tool to probe the functional space of proteins. Statistical modeling is the next logical step, allowing us to build a model of gene fragment replaceability between species. In this thesis I begin to develop the statistical tools that are needed to systematically describe combinatorial protein libraries. I present three sets of diverse chimeric protein libraries developed using sequence information. The statistical model of the human N-Ras and human K-Ras-4B genes reveal a set previously unidetifed surface residues on the N-Ras G-Domain that may be involved in cellular localization. Statistical modeling of a …


Disk Diffusion Breakpoint Determination Using A Bayesian Nonparametric Variation Of The Errors-In-Variables Model, Glen Richard DePalma 2013 Purdue University

Disk Diffusion Breakpoint Determination Using A Bayesian Nonparametric Variation Of The Errors-In-Variables Model, Glen Richard Depalma

Open Access Dissertations

Drug dilution (MIC) and disk diffusion (DIA) are the two most common antimicrobial susceptibility tests used by hospitals and clinics to determine an unknown pathogen's susceptibility to various antibiotics. Both tests use breakpoints to classify the pathogen as either susceptible, indeterminant, or resistant to each drug under consideration. While the determination of these drug-specific MIC classification breakpoints is straightforward, determination of comparable DIA breakpoints is not. It is this issue that motivates this research.

Traditionally, the error-rate bounded (ERB) method has been used to calibrate the two tests. This procedure involves determining DIA breakpoints which minimize the observed discrepancies between …


Asymptotically Unbiased Estimator Of The Informational Energy With Knn, Angel Caţaron, Răzvan Andonie, Chinmei Y. Chueh 2013 Transylvania University

Asymptotically Unbiased Estimator Of The Informational Energy With Knn, Angel Caţaron, Răzvan Andonie, Chinmei Y. Chueh

All Faculty Scholarship for the College of the Sciences

Motivated by machine learning applications (e.g., classification, function approximation, feature extraction), in previous work, we have introduced a non- parametric estimator of Onicescu’s informational energy. Our method was based on the k-th nearest neighbor distances between the n sample points, where k is a fixed positive integer. In the present contribution, we discuss mathematical properties of this estimator. We show that our estimator is asymptotically unbiased and consistent. We provide further experimental results which illustrate the convergence of the estimator for standard distributions.


Characterizations Of Distribution Of Ratio Of Rayleigh Random Variables, Gholamhossein Hamedani 2013 Marquette University

Characterizations Of Distribution Of Ratio Of Rayleigh Random Variables, Gholamhossein Hamedani

Mathematics, Statistics and Computer Science Faculty Research and Publications

Various characterizations of the distribution of the ratio of two independent Rayleigh random variables are presented. These characterizations are based, on a truncated moment; on hazard function; and on certain functions of order statistics.


Bootstrapped Deattenuated Correlation With Missing Data, Anna Veprinsky 2013 Old Dominion University

Bootstrapped Deattenuated Correlation With Missing Data, Anna Veprinsky

Psychology Theses & Dissertations

Issues with correlation attenuation due to measurement error are well documented. A corresponding correction, the deattenuated correlation, has been known for over a century. For over a decade, researchers have been investigated the deattenuated correlation identifying factors impacting its performance. Nonetheless, the deattenuated correlation is underutilized. In addition, there is limited research concerning confidence intervals for the deattenuated correlation. Here, the bootstrapped deattenuated correlation with corresponding confidence intervals is investigated for simulation conditions not previously considered simultaneously: missing data and non-normal distributions. The bootstrap deattenuated correlation was assessed for relative bias, standard error, and 95% coverage probability for the percentile …


Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak 2013 Indiana University Fairbanks School of Public Health, Department of Biostatistics

Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak

COBRA Preprint Series

Often when a longitudinal change is studied in a population of interest we find that changes over time are heterogeneous (in terms of time and/or covariates' effect) and a traditional linear mixed effect model [Laird and Ware, 1982] on the entire population assuming common parametric form for covariates and time may not be applicable to the entire population. This is usually the case in studies when there are many possible predictors influencing the response trajectory. For example, Raudenbush [2001] used depression as an example to argue that it is incorrect to assume that all the people in a given population …


Operation Export A Study Of Small- To Mid-Size Export Manufacturing Businesses In The Philippines, Anthony J. Familia, Dr. Paul J. Fields 2013 Brigham Young University

Operation Export A Study Of Small- To Mid-Size Export Manufacturing Businesses In The Philippines, Anthony J. Familia, Dr. Paul J. Fields

Journal of Undergraduate Research

Market globalization gives developing countries a better opportunity to start successful export manufacturing operations than ever before. However, because virtually no research exists on export manufacturing companies in developing countries, many micro-entrepreneurs are unable to tap into these vast international markets. Consequently, potential export manufacturers typically start small family businesses that struggle to stay afloat in their own local economies.


The Identity Of Socially Responsible Business, Ryan Quinn, Dr. David A. Whetten 2013 Brigham Young University

The Identity Of Socially Responsible Business, Ryan Quinn, Dr. David A. Whetten

Journal of Undergraduate Research

In the popular press of today’s business world, socially responsible business is a hot topic. Periodicals, books, mutual funds, associations, and countless other media, groups, and people tout socially responsible business as the “right” thing for forward-looking companies to do. Underlying all of the claims of benefits that come from social responsibility is the assumption that becoming a socially responsible company is a simple task that any firm can accomplish. The purpose of this study was to question that assumption. It may (or may not) be a simple task to acquire a socially responsible image, but to become a truly …


Data Analysis Using Regression Modeling: Visual Display And Setup Of Simple And Complex Statistical Models, Emil N. Coman, Maria A. Coman, Eugen Iordache, Russell Barbour, Lisa Dierker 2013 Ethel Donaghue TRIPP Center, UConn Health Center

Data Analysis Using Regression Modeling: Visual Display And Setup Of Simple And Complex Statistical Models, Emil N. Coman, Maria A. Coman, Eugen Iordache, Russell Barbour, Lisa Dierker

Yale Day of Data

We present visual modeling solutions for testing simple and more advanced statistical hypotheses in any research field. All models can be directly specified in analytical software like Mplus or R.

Data analysis in any substantive field can be easily accomplished by translating statistical tests in the intuitive language of regression-based path diagrams with observed and unobserved variables. All models we presented can be directly specified and estimated in analytical software.

Students can particularly benefit from being taught the simple regression modeling setup of the path analytical method, as it empowers them to apply the techniques to any data to test …


Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty 2013 Fred Hutchinson Cancer Rsrch Center

Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty

UW Biostatistics Working Paper Series

The evaluation of biomarkers to improve risk prediction is a common theme in modern research. Since its introduction in 2008, the net reclassification index (NRI) (Pencina et al. 2008, Pencina et al. 2011) has gained widespread use as a measure of prediction performance with over 1,200 citations as of June 30, 2013. The NRI is considered by some to be more sensitive to clinically important changes in risk than the traditional change in the AUC (Delta AUC) statistic (Hlatky et al. 2009). Recent statistical research has raised questions, however, about the validity of conclusions based on the NRI. (Hilden and …


Three Classification Techniques In Explaining Health Insurance Coverage, David Boyack Dahl, Dr. Scott D. Grimshaw 2013 Brigham Young University

Three Classification Techniques In Explaining Health Insurance Coverage, David Boyack Dahl, Dr. Scott D. Grimshaw

Journal of Undergraduate Research

Classification is the assignment of objects to categories using a decision rule based on observed characteristics. Several classification techniques are available to form the decision rule. Each builds the decision rule by modeling the relationship between the actual categories and observed characteristics. The decision rule can then be used to classify other objects for which the true category is unknown. Three popular classification techniques are logistic regression, discriminant analysis, and classification trees. While all three of these classification techniques seem to work with some success, the literature has not come to uniformly accept one technique.


Development Of A Reliable Confidence Interval For The Kappa Statistic, With Application To The Semiconductor Industry, Rachelle Curtis, Dr. G. Bruce Schallje 2013 Brigham Young University

Development Of A Reliable Confidence Interval For The Kappa Statistic, With Application To The Semiconductor Industry, Rachelle Curtis, Dr. G. Bruce Schallje

Journal of Undergraduate Research

In the Semiconductor industry, it is necessary to compare the performance of different test-tapes to assure that each produces similar quality work. Test tapes sort each die, or computer chip, into one of a fixed number of bins, depending on how well it functions. When validating a new test tape, several die are tested using both the old and the new test tape and the bin assignments are compared. Kappa, a statistical measure of agreement, compares the number of matching bin assignments to the number expected if the test tapes are independent. Unfortunately, the closer Kappa is to one (perfect …


A Statistical Report Of Byu Premedical Students From 2002-2007, Daniel Chan, David Kaiser 2013 Brigham Young University

A Statistical Report Of Byu Premedical Students From 2002-2007, Daniel Chan, David Kaiser

Journal of Undergraduate Research

In 2006 the average number of BYU medical school applications for any one applicant was 17. With the chance of matriculation per application only 3.8%, yet the cost of each application running as high as $700 each, it is important for BYU premedical students to make informed decisions when choosing to which of 142 medical schools to apply.


Transcription Factor Binding Site Identification Using Mathematical Algorithms, Colin Rogerson, Dr. W. Evan Johnson 2013 Brigham Young University

Transcription Factor Binding Site Identification Using Mathematical Algorithms, Colin Rogerson, Dr. W. Evan Johnson

Journal of Undergraduate Research

The purpose of our research project was, originally, first to identify the likely binding sites of the Estrogen Receptor transcription factor, and second to identify likely co-factors that interact with Estrogen Receptor in the binding process. We planned to do this using a computational algorithm which scanned sequences of DNA, and by utilizing the position weight matrices of many transcription factors, statistically identify likely binding sites and co-factors.


An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger, Dr. C. Shane Reese 2013 Brigham Young University

An Adaptive Bayesian Approach To Dose-Response Modeling, Thomas J. Leininger, Dr. C. Shane Reese

Journal of Undergraduate Research

The Food and Drug Administration is responsible for testing and screening clinical drugs before market entry. The screening process involves different phases which determine a drug’s potency, recommended dosage, and potential side effects. Current statistical designs used to model the efficacy of a drug at varying dosage levels are robust yet inefficient.


Making Sense Of Trends And Data, Singapore Management University 2013 Singapore Management University

Making Sense Of Trends And Data, Singapore Management University

Perspectives@SMU

How do you make sense of data when it is all unpredictable?


Bayesian Decision Theoretic Approach To Directional Multiple Hypotheses Problems, Naveen K. Bansal, Klaus J. Miescke 2013 Marquette University

Bayesian Decision Theoretic Approach To Directional Multiple Hypotheses Problems, Naveen K. Bansal, Klaus J. Miescke

Mathematics, Statistics and Computer Science Faculty Research and Publications

A multiple hypothesis problem with directional alternatives is considered in a decision theoretic framework. Skewness in the alternatives is considered, and it is shown that this skewness permits the Bayes rules to possess certain advantages when one direction of the alternatives is more important or more probable than the other direction. Bayes rules subject to constraints on certain directional false discovery rates are obtained, and their performances are compared with a traditional FDR rule through simulation. We also analyzed a gene expression data using our methodology, and compare the results to that of a FDR method.


Digital Commons powered by bepress