Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 6091 - 6120 of 12830

Full-Text Articles in Statistics and Probability

Generalized Clusterwise Regression For Simultaneous Estimation Of Optimal Pavement Clusters And Performance Models, Mukesh Khadka May 2017

Generalized Clusterwise Regression For Simultaneous Estimation Of Optimal Pavement Clusters And Performance Models, Mukesh Khadka

UNLV Theses, Dissertations, Professional Papers, and Capstones

The existing state-of-the-art approach of Clusterwise Regression (CR) to estimate pavement performance models (PPMs) pre-specifies explanatory variables without testing their significance; as an input, this approach requires the number of clusters for a given data set. Time-consuming ‘trial and error’ methods are required to determine the optimal number of clusters. A common objective function is the minimization of the total sum of squared errors (SSE). Given that SSE decreases monotonically as a function of the number of clusters, the optimal number of clusters with minimum SSE always is the total number of data points. Hence, the minimization of SSE is …


Demographics, Patterns Of Care, And Survival In Pediatric Medulloblastoma, Emily V. Dressler, Therese A. Dolecek, Meng Liu, John L. Villano May 2017

Demographics, Patterns Of Care, And Survival In Pediatric Medulloblastoma, Emily V. Dressler, Therese A. Dolecek, Meng Liu, John L. Villano

Internal Medicine Faculty Publications

We evaluated the American College of Surgeon’s National Cancer Data Base (NCDB) to describe current hospital-based epidemiologic frequency, survival, and patterns of care of pediatric medulloblastoma. We analyzed NCDB 1998–2011 data on medulloblastoma for children ages 0–19 years using logistic and poisson regression, Kaplan–Meier survival estimates, and Cox proportional hazards models. 3647 cases of medulloblastoma in those aged 0–19 years were identified. Chemotherapy was received by 79 and 74% received radiation, with 65% receiving both therapies. Those who received radiation were more likely to be older than four, while those who received chemotherapy were more likely to be age four …


A Bayesian Variable Selection Method With Applications To Spatial Data, Xiahan Tang May 2017

A Bayesian Variable Selection Method With Applications To Spatial Data, Xiahan Tang

Graduate Theses and Dissertations

This thesis first describes the general idea behind Bayes Inference, various sampling methods based on Bayes theorem and many examples. Then a Bayes approach to model selection, called Stochastic Search Variable Selection (SSVS) is discussed. It was originally proposed by George and McCulloch (1993). In a normal regression model where the number of covariates is large, only a small subset tend to be significant most of the times. This Bayes procedure specifies a mixture prior for each of the unknown regression coefficient, the mixture prior was originally proposed by Geweke (1996). This mixture prior will be updated as data becomes …


Spatiotemporal Analyses Of Recycled Water Production, Jana E. Archer May 2017

Spatiotemporal Analyses Of Recycled Water Production, Jana E. Archer

Electronic Theses and Dissertations

Increased demands on water supplies caused by population expansion, saltwater intrusion, and drought have led to water shortages which may be addressed by use of recycled water as recycled water products. Study I investigated recycled water production in Florida and California during 2009 to detect gaps in distribution and identify areas for expansion. Gaps were detected along the panhandle and Miami, Florida, as well as the northern and southwestern regions in California. Study II examined gaps in distribution, identified temporal change, and located areas for expansion for Florida in 2009 and 2015. Production increased in the northern and southern regions …


Assessing The Relationship Between Change Blindness And The Anchoring Effect, Melissa Schoenlein May 2017

Assessing The Relationship Between Change Blindness And The Anchoring Effect, Melissa Schoenlein

Honors Projects

A lack of consideration for all aspects of a question prompts fragmented decision making. These decisions, as they leave out fundamental information, repeatedly then lead to a potentially problematic reaction to the target question or stimuli. The anchoring heuristic propels one to make a decision, usually an estimate, based on a presented “fact”, often ignoring additional background and environmental clues. Reducing the rate of occurrence of the anchoring bias is thought to lead to an increase in holistic decision making. To promote this reduction, the purpose of this research was to affirmexamine the relationship between the susceptibility to change blindness …


Peptide Identification: Refining A Bayesian Stochastic Model, Theophilus Barnabas Kobina Acquah May 2017

Peptide Identification: Refining A Bayesian Stochastic Model, Theophilus Barnabas Kobina Acquah

Electronic Theses and Dissertations

Notwithstanding the challenges associated with different methods of peptide identification, other methods have been explored over the years. The complexity, size and computational challenges of peptide-based data sets calls for more intrusion into this sphere. By relying on the prior information about the average relative abundances of bond cleavages and the prior probability of any specific amino acid sequence, we refine an already developed Bayesian approach in identifying peptides. The likelihood function is improved by adding additional ions to the model and its size is driven by two overall goodness of fit measures. In the face of the complexities associated …


Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch May 2017

Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch

Electronic Theses and Dissertations

Missing data is one of the challenges we are facing today in modeling valid statistical models. It reduces the representativeness of the data samples. Hence, population estimates, and model parameters estimated from such data are likely to be biased.

However, the missing data problem is an area under study, and alternative better statistical procedures have been presented to mitigate its shortcomings. In this paper, we review causes of missing data, and various methods of handling missing data. Our main focus is evaluating various multiple imputation (MI) methods from the multiple imputation of chained equation (MICE) package in the statistical software …


Denoising Tandem Mass Spectrometry Data, Felix Offei May 2017

Denoising Tandem Mass Spectrometry Data, Felix Offei

Electronic Theses and Dissertations

Protein identification using tandem mass spectrometry (MS/MS) has proven to be an effective way to identify proteins in a biological sample. An observed spectrum is constructed from the data produced by the tandem mass spectrometer. A protein can be identified if the observed spectrum aligns with the theoretical spectrum. However, data generated by the tandem mass spectrometer are affected by errors thus making protein identification challenging in the field of proteomics. Some of these errors include wrong calibration of the instrument, instrument distortion and noise. In this thesis, we present a pre-processing method, which focuses on the removal of noisy …


A Comparison Of Statistical Methods Relating Pairwise Distance To A Binary Subject-Level Covariate, Rachael Stone May 2017

A Comparison Of Statistical Methods Relating Pairwise Distance To A Binary Subject-Level Covariate, Rachael Stone

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

A community ecologist provided a motivating data set involving a certain animal species with two behavior groups, along with a pairwise genetic distance matrix among individuals. Many community ecologists have analyzed similar data sets with a method known as the Hopkins method, testing for an association between the subject-level covariate (behavior group) and the pairwise distance. This community ecologist wanted to know if they used the Hopkins method, would their results be meaningful? Their question inspired this thesis work, where a different data set was used for confidentiality reasons. Multiple methods (Hopkins method, ADONIS, ANOSIM, and Distance Regression) were used …


Statistical Methods For Two Problems In Cancer Research: Analysis Of Rna-Seq Data From Archival Samples And Characterization Of Onset Of Multiple Primary Cancers, Jialu Li May 2017

Statistical Methods For Two Problems In Cancer Research: Analysis Of Rna-Seq Data From Archival Samples And Characterization Of Onset Of Multiple Primary Cancers, Jialu Li

Dissertations and Theses (Open Access)

My dissertation is focused on quantitative methodology development and application for two important topics in translational and clinical cancer research.

The first topic was motivated by the challenge of applying transcriptome sequencing (RNA-seq) to formalin-fixation and paraffin-embedding (FFPE) tumor samples for reliable diagnostic development. We designed a biospecimen study to directly compare gene expression results from different protocols to prepare libraries for RNA-seq from human breast cancer tissues, with randomization to fresh-frozen (FF) or FFPE conditions. To comprehensively evaluate the FFPE RNA-seq data quality for expression profiling, we developed multiple computational methods for assessment, such as the uniformity and continuity …


Detecting And Evaluating Therapy Induced Changes In Radiomics Features Measured From Non-Small Cell Lung Cancer To Predict Patient Outcomes, Xenia J. Fave May 2017

Detecting And Evaluating Therapy Induced Changes In Radiomics Features Measured From Non-Small Cell Lung Cancer To Predict Patient Outcomes, Xenia J. Fave

Dissertations and Theses (Open Access)

The purpose of this study was to investigate whether radiomics features measured from weekly 4-dimensional computed tomography (4DCT) images of non-small cell lung cancers (NSCLC) change during treatment and if those changes are prognostic for patient outcomes or dependent on treatment modality. Radiomics features are quantitative metrics designed to evaluate tumor heterogeneity from routine medical imaging. Features that are prognostic for patient outcome could be used to monitor tumor response and identify high-risk patients for adaptive treatment. This would be especially valuable for NSCLC due to the high prevalence and mortality of this disease.

A novel process was designed to …


Telephone Polls And Pps Sampling: A Potential Boon To The Polling Industry, Jade Mckay Burt May 2017

Telephone Polls And Pps Sampling: A Potential Boon To The Polling Industry, Jade Mckay Burt

Undergraduate Honors Capstone Projects

In the wake of the 2016 election, the polling industry has no shortage of critics. While these are difficult times for the industry as a whole, there are exciting innovations happening that will serve to benefit and revitalize the industry for years. One of these exciting innovations is Probability Proportional to Size (PPS) sampling. I will elaborate on what PPS sampling is and provide a mathematical foundation for its use in polling. I also discuss what some of the myriad of issues plaguing the polling industry are and then show how PPS sampling can be used to remedy many of …


A Distribution Of The First Order Statistic When The Sample Size Is Random, Vincent Z. Forgo Mr May 2017

A Distribution Of The First Order Statistic When The Sample Size Is Random, Vincent Z. Forgo Mr

Electronic Theses and Dissertations

Statistical distributions also known as probability distributions are used to model a random experiment. Probability distributions consist of probability density functions (pdf) and cumulative density functions (cdf). Probability distributions are widely used in the area of engineering, actuarial science, computer science, biological science, physics, and other applicable areas of study. Statistics are used to draw conclusions about the population through probability models. Sample statistics such as the minimum, first quartile, median, third quartile, and maximum, referred to as the five-number summary, are examples of order statistics. The minimum and maximum observations are important in extreme value theory. This paper will …


Irrational Eigenvalues Of The Discrete Laplacian: A Study Of Simplical Complexes, Brian Bollen May 2017

Irrational Eigenvalues Of The Discrete Laplacian: A Study Of Simplical Complexes, Brian Bollen

Psychology

We study the behavior of eigenvalues of the discrete Laplacian of an abstract simplicial complex K when subdividing a single face of K . We show that if K is a simplex, performing this kind of restricted subdivision twice on a single face produces irrational eigenvalues for the discrete Laplacian.


Inference On The Stress-Strength Model From Weibull Gamma Distribution, Mahmoud Mansour, Rashad El-Sagheer, M. A. W. Mahmoud Prof. May 2017

Inference On The Stress-Strength Model From Weibull Gamma Distribution, Mahmoud Mansour, Rashad El-Sagheer, M. A. W. Mahmoud Prof.

Basic Science Engineering

No abstract provided.


Combinatorial Games On Graphs, Trevor K. Williams May 2017

Combinatorial Games On Graphs, Trevor K. Williams

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Combinatorial Games are intriguing and have a tendency to engross students and lead them into a serious study of mathematics. The engaging nature of games is the basis for this thesis. Two combinatorial games and some educational tools are presented which were developed by the author in the pursuit of the solution of these games.


Does Educational Level Influence Postmenopausal Breast Cancer Mortality Among Asians In U.S.?, Sfurti Maheshwari May 2017

Does Educational Level Influence Postmenopausal Breast Cancer Mortality Among Asians In U.S.?, Sfurti Maheshwari

UNLV Theses, Dissertations, Professional Papers, and Capstones

Studies on mortality from postmenopausal breast cancer (PMBC) by education level have not shown consistent results among US women. For US Asians, often seen as a “model” minority in terms of affluence and education this relationship has never been studied despite PMBC being the most common cancer in the country.

We analyzed 2008-2012 California Vital Statistics data and population data from the American Community Survey 2012 to compute age and education adjusted mortality ratios using negative binomial regression model for White (as a reference category) and Asian women. In total 3,277,106 (80%) White women and 852,376 (20%) Asian women died …


Functional Human Grin2b Promoter Polymorphism And Variation Of Mental Processing Speed In Older Adults, Yang Jiang, Ming Kuan Lin, Gregory A. Jicha, Xiuhua Ding, Sabrina L. Mcilwrath, David W. Fardo, Lucas S. Broster, Frederick A. Schmitt, Richard J. Kryscio, Robert H. Lipsky Apr 2017

Functional Human Grin2b Promoter Polymorphism And Variation Of Mental Processing Speed In Older Adults, Yang Jiang, Ming Kuan Lin, Gregory A. Jicha, Xiuhua Ding, Sabrina L. Mcilwrath, David W. Fardo, Lucas S. Broster, Frederick A. Schmitt, Richard J. Kryscio, Robert H. Lipsky

Behavioral Science Faculty Publications

We investigated the role of a single nucleotide polymorphism rs3764030 (G > A) within the human GRIN2B promoter in mental processing speed in healthy, cognitively intact, older adults. In vitro DNA-binding and reporter gene assays of different allele combinations in transfected cells showed that the A allele was a gain-of-function variant associated with increasing GRIN2B mRNA levels. We tested the hypothesis that individuals with A allele will have better memory performance (i.e. faster reaction times) in older age. Twenty-eight older adults (ages 65-86) from a well-characterized longitudinal cohort were recruited and performed a modified delayed match-to-sample task. The rs3764030 polymorphism was …


High-Dimensional Repeated Measures, Martin Happ, Solomon W. Harrar, Arne C. Bathke Apr 2017

High-Dimensional Repeated Measures, Martin Happ, Solomon W. Harrar, Arne C. Bathke

Statistics Faculty Publications

Recently, new tests for main and simple treatment effects, time effects, and treatment by time interactions in possibly high-dimensional multigroup repeated-measures designs with unequal covariance matrices have been proposed. Technical details for using more than one between-subject and more than one within-subject factor are presented in this article. Furthermore, application to electroencephalography (EEG) data of a neurological study with two whole-plot factors (diagnosis and sex) and two subplot factors (variable and region) is shown with the R package HRM (high-dimensional repeated measures).


The Technological Revolution And Data Science, Leslie Walcott Apr 2017

The Technological Revolution And Data Science, Leslie Walcott

Honors Theses

What was once only depicted in science fiction is now a reality: computers are taking jobs from humans. As technology improves, automation is transforming the workplace. They say a “fourth industrial revolution” is inevitable within the next ten years. In the industrial revolution, the jobs lost were unskilled laborers, such as coal miners, textiles manufacturers, or cotton workers. There was no argument for whether or not a machine could do the jobs more efficiently--it was fact. The term technological unemployment means the loss of jobs caused by technological change. The headline, “Factory workers replaced by automation,” is not particularly startling …


An Empirical Look At The Controversy Surrounding The Nobel Prize For Magnetic Resonance Imaging, Anthony Breitzman Apr 2017

An Empirical Look At The Controversy Surrounding The Nobel Prize For Magnetic Resonance Imaging, Anthony Breitzman

College of Science & Mathematics Departmental Research

Disputes between researchers over who deserves credit for technological breakthroughs are not unusual. Few such disputes, however, have attracted as much attention as the arguments surrounding the award of the 2003 Nobel Prize for Medicine. This prize was awarded to Paul Lauterbur and Peter Mansfield to honor “discoveries concerning the development of magnetic resonance imaging” – i.e. MRI. Soon after the award, another scientist, Raymond Damadian, took out full-page advertisements in national newspapers, decrying the award and stating that he should have been included alongside Lauterbur and Mansfield. This technical report examines Damadian’s claim from a strictly empirical perspective, by …


Estimating Autoantibody Signatures To Detect Autoimmune Disease Patient Subsets, Zhenke Wu, Livia Casciola-Rosen, Ami A. Shah, Antony Rosen, Scott L. Zeger Apr 2017

Estimating Autoantibody Signatures To Detect Autoimmune Disease Patient Subsets, Zhenke Wu, Livia Casciola-Rosen, Ami A. Shah, Antony Rosen, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Autoimmune diseases are characterized by highly specific immune responses against molecules in self-tissues. Different autoimmune diseases are characterized by distinct immune responses, making autoantibodies useful for diagnosis and prediction. In many diseases, the targets of autoantibodies are incompletely defined. Although the technologies for autoantibody discovery have advanced dramatically over the past decade, each of these techniques generates hundreds of possibilities, which are onerous and expensive to validate. We set out to establish a method to greatly simplify autoantibody discovery, using a pre-filtering step to define subgroups with similar specificities based on migration of labeled, immunoprecipitated proteins on sodium dodecyl sulfate …


Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane Apr 2017

Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane

Theses

Alzheimer Disease (AD) is difficult to diagnose by using genetic testing or other traditional methods. Unlike diseases with simple genetic risk components, there exists no single marker determining as to whether someone will develop AD. Furthermore, AD is highly heterogeneous and different subgroups of individuals develop the disease due to differing factors. Traditional diagnostic methods using perceivable cognitive deficiencies are often too little too late due to the brain having suffered damage from decades of disease progression. In order to observe AD at early stages prior to the observation of cognitive deficiencies, biomarkers with greater accuracy are required. By using …


Optimal Experimental Design To Characterize A Wave Source Using Dosimeter Measurements, Renee L. Gooding Apr 2017

Optimal Experimental Design To Characterize A Wave Source Using Dosimeter Measurements, Renee L. Gooding

Mathematics & Statistics ETDs

When modeling physical phenomena we want to solve the inverse problem by estimating the parameters that characterize the source model that we are interested in. In this thesis, we focus on the optimal placement of a finite number of individual sensors, called dosimeters, in two and three dimensions with a time dependent Gaussian wave source. Using a computational model along with experimental data, we design an iterative process to determine the optimal placement of an additional sensor such that the noise in the measurements has a minimal effect on the parameter estimation. First, we estimate the parameters that characterize the …


Deterministic And Probabilistic Methods For Seismic Source Inversion, Juan Pablo Madrigal Cianci Apr 2017

Deterministic And Probabilistic Methods For Seismic Source Inversion, Juan Pablo Madrigal Cianci

Mathematics & Statistics ETDs

The national Earthquake Information Center (NEIC) reports an occurrence of about 13,000 earthquakes every year, spanning different values on the Richter scale from very mild (2) to "giant earthquakes'' (8 and above). Being able to study these earthquakes provides useful information for a wide range of applications in geophysics. In the present work we study the characteristics of an earthquake by performing seismic source inversion; a mathematical problem that, given some recorded data, produces a set of parameters that when used as input in a mathematical model for the earthquake generates synthetic data that closely resembles the measured data. There …


Cancer Modeling: From Optimal Cell Renewal To Immunotherapy, Cesar L. Alvarado Apr 2017

Cancer Modeling: From Optimal Cell Renewal To Immunotherapy, Cesar L. Alvarado

Mathematics & Statistics ETDs

Cancer is a disease caused by mutations in normal cells. According to the National Cancer Institute, in 2016, an estimated 1.6 million people were diagnosed and approximately 0.5 million people died from the disease in the United States. There are many factors that shape cancer at the cellular and organismal level, including genetic, immunological, and environmental components. In this thesis, we show how mathematical modeling can be used to provide insight into some of the key mechanisms underlying cancer dynamics. First, we use mathematical modeling to investigate optimal homeostatic cell renewal in tissues such as the small intestine with an …


Modeling Trait Evolutionary Processes With More Than One Gene, Huan Jiang Apr 2017

Modeling Trait Evolutionary Processes With More Than One Gene, Huan Jiang

Mathematics & Statistics ETDs

Phylogenetic comparative methods have been used to test evolutionary signals through trait evolutionary processes. Traditionally, biologists use one phylogenetic tree as a tool to handle dependent data for the traits of interest and hence utilize one gene only. However, it is more informative if the evolutionary processes of a trait are presented by phylogenetic trees reconstructed by the DNA alignments from more than one gene. In this work, we explain and develop two methods involving modeling the trait evolutionary processes: (a) two gene trees via the Brownian motion (BM) model; and (b) two gene trees via the Ornstein-Uhlenbeck (OU) model. …


A General Approach For Predicting The Behavior Of The Supreme Court Of The United States, Daniel Katz Apr 2017

A General Approach For Predicting The Behavior Of The Supreme Court Of The United States, Daniel Katz

All Faculty Scholarship

Building on developments in machine learning and prior work in the science of judicial prediction, we construct a model designed to predict the behavior of the Supreme Court of the United States in a generalized, out-of-sample context. To do so, we develop a time-evolving random forest classifier that leverages unique feature engineering to predict more than 240,000 justice votes and 28,000 cases outcomes over nearly two centuries (1816-2015). Using only data available prior to decision, our model outperforms null (baseline) models at both the justice and case level under both parametric and non-parametric tests. Over nearly two centuries, we achieve …


Positive Affect Predicts Cerebral Glucose Metabolism In Late Middle-Aged Adults., Christopher Nicholas, Siobhan M Hoscheidt, Lindsay R Clark, Annie M Racine, Sara E Berman, Rebecca L Koscik, N Maritza Dowling, Sanjay Asthana, Bradley T Christian, Mark A Sager, Sterling C Johnson Apr 2017

Positive Affect Predicts Cerebral Glucose Metabolism In Late Middle-Aged Adults., Christopher Nicholas, Siobhan M Hoscheidt, Lindsay R Clark, Annie M Racine, Sara E Berman, Rebecca L Koscik, N Maritza Dowling, Sanjay Asthana, Bradley T Christian, Mark A Sager, Sterling C Johnson

GW Biostatistics Center

Positive affect is associated with a number of health benefits; however, few studies have examined the relationship between positive affect and cerebral glucose metabolism, a key energy source for neuronal function and a possible index of brain health. We sought to determine if positive affect was associated with cerebral glucose metabolism in late middle-aged adults (n = 133). Participants completed the positive affect subscale of the Center for Epidemiological Studies Depression Scale at two time points over a two-year period and underwent 18F-fluorodeoxyglucose-positron emission tomography scanning. After controlling for age, sex, perceived health status, depressive symptoms, anti-depressant use, family …


Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun Apr 2017

Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun

Biostatistics Faculty Publications

In contrast to feature selection and gene set analysis, bi-level selection is a process of selecting not only important gene sets but also important genes within those gene sets. Depending on the order of selections, a bi-level selection method can be classified into three categories – forward selection, which first selects relevant gene sets followed by the selection of relevant individual genes; backward selection which takes the reversed order; and simultaneous selection, which performs the two tasks simultaneously usually with the aids of a penalized regression model. To test the existence of subtype-specific prognostic genes for non-small cell lung cancer …