Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 511 - 540 of 618

Full-Text Articles in Statistics and Probability

Wage-Productivity Analysis Of U.S. Domestic Airlines, Jesse Lucas, Khairul Azuar, Justin Tan, Syed Ilyas Jan 2020

Wage-Productivity Analysis Of U.S. Domestic Airlines, Jesse Lucas, Khairul Azuar, Justin Tan, Syed Ilyas

Introduction to Research Methods RSCH 202

This study examines the impact of wages on productivity by examining US domestic airlines.

Current literature places emphasis on jobs conducted in-flight, specifically pilots and cabin crew. This paper considers all job titles involved in the operations of the airline, including executives and management. Existing research focuses on factors such as governance, domestic economic level, and personal attributes such as intrinsic motivation, gender, and age. There is insufficient research regarding the relationship between wage and productivity. Thus, it is uncertain if high wage leads to high productivity. Preliminary findings suggest higher wage equates to higher productivity.


Joint Simulation Of Continuous And Categorical Variables For Mineral Resource Modeling And Recoverable Reserves Calculation, Sentle Augustinus Hlajoane Jan 2020

Joint Simulation Of Continuous And Categorical Variables For Mineral Resource Modeling And Recoverable Reserves Calculation, Sentle Augustinus Hlajoane

Dissertations, Master's Theses and Master's Reports

Spatial variability and uncertainty of continuous variables (grade) and categorical variables (rock-types) in mineral evaluation significantly impact the economics of mining projects. The conventional approach of simulating grades using deterministic rock- types is problematic since spatial variability, and uncertainty of grades at rock-type contacts are not well captured in deposits where the grade changes gradually between rock-types. Therefore, jointly simulating these variables can improve confidence (reduce uncertainty) in a resource model. Also, resource classification and recoverable reserve calculation can significantly improve the understanding of the deposit and its economic viability. This research utilized the Plural-Gaussian geostatistical simulation to jointly simulate …


Supp & Mapp: Adaptable Structure-Based Representations For Mir Tasks, Claire Savard, Erin H. Bugbee, Melissa R, Mcguirl, Katherine M. Kinnaird Jan 2020

Supp & Mapp: Adaptable Structure-Based Representations For Mir Tasks, Claire Savard, Erin H. Bugbee, Melissa R, Mcguirl, Katherine M. Kinnaird

Statistical and Data Sciences: Faculty Publications

Accurate and flexible representations of music data are paramount to addressing MIR tasks, yet many of the existing approaches are difficult to interpret or rigid in nature. This work introduces two new song representations for structure-based retrieval methods: Surface Pattern Preservation (SuPP), a continuous song representation, and Matrix Pattern Preservation (MaPP), SuPP’s discrete counterpart. These representations come equipped with several user-defined parameters so that they are adaptable for a range of MIR tasks. Experimental results show MaPP as successful in addressing the cover song task on a set of Mazurka scores, with a mean precision of 0.965 and recall of …


Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu Jan 2020

Nonparametric Misclassification Simulation And Extrapolation Method And Its Application, Congjian Liu

College of Graduate Studies: Theses & Dissertations

The misclassification simulation extrapolation (MC-SIMEX) method proposed by Küchenho et al. is a general method of handling categorical data with measurement error. It consists of two steps, the simulation and extrapolation steps. In the simulation step, it simulates observations with varying degrees of measurement error. Then parameter estimators for varying degrees of measurement error are obtained based on these observations. In the extrapolation step, it uses a parametric extrapolation function to obtain the parameter estimators for data with no measurement error. However, as shown in many studies, the parameter estimators are still biased as a result of the parametric extrapolation …


Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil Jan 2020

Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil

College of Graduate Studies: Theses & Dissertations

Modern research strategies rely predominantly on three steps, data collection, data analysis, and inference. In research, if the data is not collected as designed, researchers may face challenges of having incomplete data, especially when it is non-ignorable. These situations affect the subsequent steps of evaluation and make them difficult to perform. Inference with incomplete data is a challenging task in data analysis and clinical trials when missing data related to the condition under the study. Moreover, results obtained from incomplete data are prone to biases. Parameter estimation with non-ignorable missing data is even more challenging to handle and extract useful …


Projecting Regions Of North Atlantic Right Whale, Eubalaena Glacialis, Habitat Suitability In The Gulf Of Maine In 2050, Camille Ross Jan 2020

Projecting Regions Of North Atlantic Right Whale, Eubalaena Glacialis, Habitat Suitability In The Gulf Of Maine In 2050, Camille Ross

Honors Theses

North Atlantic right whales (Eubalaena glacialis) are endangered. Understanding the role environmental conditions play in habitat suitability is key to determining the regions in need of protection for conservation of the species, particularly as climate change shifts suitable habitat. This thesis uses three species distribution modeling algorithms, together with historical data on whale abundance(1993 to 2009) and environmental covariates to build monthly ensemble models of past E. glacialis habitat suitability in the Gulf of Maine. Then, the models are projected onto the year 2050 for a range of climate scenarios. Specifically, the distribution of the species was modeled …


Webexpo : Vers Une Meilleure Interprétation Des Mesures D'Exposition Professionnelle Aux Substances Chimiques Sur Les Lieux De Travail, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc Jan 2020

Webexpo : Vers Une Meilleure Interprétation Des Mesures D'Exposition Professionnelle Aux Substances Chimiques Sur Les Lieux De Travail, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc

Rapports de recherche scientifique

Une grande partie de l'activité des hygiénistes du travail consiste à mesurer les niveaux d'exposition professionnelle des travailleurs. La plupart des rapports d'évaluation de l'exposition font état d'une importante variabilité spatiale et temporelle quant à l'intensité de l'exposition, laquelle fluctue souvent du simple au décuple en dépit de conditions apparemment similaires. Il en résulte depuis toujours un défi de taille en ce qui concerne l'interprétation des niveaux mesurés par rapport aux valeurs limites d’exposition professionnelle (VLEP). Il existe désormais un cadre consensuel – issu d'une élaboration progressive au cours des deux dernières décennies – concernant l'analyse des niveaux d’exposition par …


Webexpo: Towards A Better Interpretation Of Measurements Of Occupational Exposure To Chemicals In The Workplace, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc Jan 2020

Webexpo: Towards A Better Interpretation Of Measurements Of Occupational Exposure To Chemicals In The Workplace, Jérôme Lavoué, Lawrence Joseph, Tracy L. Kirkham, France Labrèche, Gautier Mater, Frédéric Clerc

Rapports de recherche scientifique

Une grande partie de l'activité des hygiénistes du travail consiste à mesurer les niveaux d'exposition professionnelle des travailleurs. La plupart des rapports d'évaluation de l'exposition font état d'une importante variabilité spatiale et temporelle quant à l'intensité de l'exposition, laquelle fluctue souvent du simple au décuple en dépit de conditions apparemment similaires. Il en résulte depuis toujours un défi de taille en ce qui concerne l'interprétation des niveaux mesurés par rapport aux valeurs limites d’exposition professionnelle (VLEP). Il existe désormais un cadre consensuel – issu d'une élaboration progressive au cours des deux dernières décennies – concernant l'analyse des niveaux d’exposition par …


Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg Jan 2020

Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg

Theses and Dissertations

Dropout is a common issue in randomized controlled clinical trials and can negatively impact the internal validity of a study and potentially bias the treatment effect. When subjects discontinue study participation, they are not being given the opportunity to gain from the investigational therapy as if they had remained in the study, defeating one of the main purposes of clinical trials, providing treatment. Specifically in unblinded studies, such as the wait-list control (WLC) design, dropout is often due to group membership. Subjects allocated to the control group often dropout at higher rates than in the treatment group. Adaptive designs have …


Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero Jan 2020

Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero

Theses and Dissertations

Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …


Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii Jan 2020

Applications Of Dynamic Linear Models To Random Allocation Models, Albert H. Lee Iii

Theses and Dissertations

Although advances in modern computational algorithms have provided researchers the ability to work problems which were once too computationally complex to solve, problems with high computation or large parameter spaces still remain. Problems such as those involving Time Series can be such problems. Chapter 1 looks at the the use of Exponentially Weighted Moving Averages developed by \citep{holt2004forecasting, winters1960forecasting} which were thought to provide sufficient solutions to these Time Series. A discussion is provided which illustrates the shortcomings of the EWMA and how its infinite number of possible starting values provides the modeler with an endless number of possible solutions …


Utilizing Design Structure For Improving Design Selection And Analysis, Ahlam Ali Alzharani Jan 2020

Utilizing Design Structure For Improving Design Selection And Analysis, Ahlam Ali Alzharani

Theses and Dissertations

Recent work has shown that the structure for design plays a role in the simplicity or complexity of data analysis. To increase the knowledge of research in these areas, this dissertation aims to utilize design structure for improving design selection and analysis. In this regard, minimal dependent sets and block diagonal structure are both important concepts that are relevant to the orthogonality of the columns of a design. We are interested in finding ways to improve the data analysis especially for active effect detection by utilizing minimal dependent sets and block diagonal structure for design.

We introduce a new classification …


Stochastic Technique For Solutions Of Non-Linear Fin Equation Arising In Thermal Equilibrium Model, Iftikhar Ahmad, Hina Qureshi, Muhammad Bilal, Muhammad Usman Jan 2020

Stochastic Technique For Solutions Of Non-Linear Fin Equation Arising In Thermal Equilibrium Model, Iftikhar Ahmad, Hina Qureshi, Muhammad Bilal, Muhammad Usman

Mathematics Faculty Publications

In this study, a stochastic numerical technique is used to investigate the numerical solution of heat transfer temperature distribution system using feed forward artificial neural networks. Mathematical model of fin equation is formulated with the help of artificial neural networks. The effect of the heat on a rectangular fin with thermal conductivity and temperature de-pendent internal heat generation is calculated through neural networks optimization with optimizers like active set technique, interior point technique, pattern search, genetic algorithm and a hybrid approach of pattern search - interior point technique, genetic algorithm - active set technique, genetic algorithm - interior point technique, …


Simulation Based Inference In Epidemic Models, Tejitha Dharmagadda Jan 2020

Simulation Based Inference In Epidemic Models, Tejitha Dharmagadda

Theses, Dissertations and Culminating Projects

From ancient times to the modern day, public health has been an area of great interest. Studies on the nature of disease epidemics began around 400 BC and has been a continuous area of study for the well-being of individuals around the world. For over 100 years, epidemiologists and mathematicians have developed numerous mathematical models to improve our understanding of infectious disease dynamics with an eye on controlling and preventing disease outbreak and spread. In this thesis, we discuss several types of mathematical compartmental models such as the SIR, and SIS models. To capture the noise inherent in the real-world, …


A Comparative Spatial And Climate Analysis Of Human Granulocytic Anaplasmosis And Human Babesiosis In New York State (2013-2018), Collin J. O'Connor Jan 2020

A Comparative Spatial And Climate Analysis Of Human Granulocytic Anaplasmosis And Human Babesiosis In New York State (2013-2018), Collin J. O'Connor

Legacy Theses & Dissertations (2009 - 2024)

Human granulocytic anaplasmosis (HGA) and human babesiosis are tick-borne diseases spread by Ixodes scapularis (the blacklegged or deer tick) and are the result of infection with Anaplasma phagocytophilum and Babesia microti, respectively. In New York State (NYS), incidence rates of these diseases increased concordantly until around 2013, when rates of HGA began to increase more rapidly than human babesiosis, and the spatial extent of the diseases diverged. Surveillance data of tick-borne pathogens (2007 to 2018) and reported human cases of HGA (n=4,297) and human babesiosis (n=2,986) (2013 to 2018) from the New York State Department of Health (NYSDOH) showed a …


Three Essays On Model Selection, Fangning Li Jan 2020

Three Essays On Model Selection, Fangning Li

Legacy Theses & Dissertations (2009 - 2024)

In empirical research, we often need to address the issue of what model to use given a collection of candidate models. Conventionally, we use model selection to choose one best model from the collection of candidate models based on some model selection criteria. Model averaging is a generalization of model selection in the sense that it assigns weights to candidate models and uses a weighted average to construct an aggregated model. Usually model averaging provides better performance than model selection which chooses a single candidate model based on AIC or BIC.


Parsimonious Covariate Selection For Interval Censored Data, Yi Cui Jan 2020

Parsimonious Covariate Selection For Interval Censored Data, Yi Cui

Legacy Theses & Dissertations (2009 - 2024)

Interval censored outcomes widely arise in many clinical trials and observational studies. In many cases, subjects are only followed-up periodically. As a result, the event of interest is known only to occur within a certain interval. We provided a method to select the parsimonious set of covariates associated with the interval censored outcome. First, the iterative sure independence screening (ISIS) method was applied to all interval censored time points across subjects to simultaneously select a set of potentially important covariates; then multiple testing approaches were used to improve the selection accuracy through refining the selection criteria, i.e. determining a refined …


An Analysis Of Income And Other Associative Demographics To Charity Donation And Volunteerism, Christine Lynn Klotz Jan 2020

An Analysis Of Income And Other Associative Demographics To Charity Donation And Volunteerism, Christine Lynn Klotz

Legacy Theses & Dissertations (2009 - 2024)

Nongovernmental organizations have a vast institutional presence across the United States. Each year, there is an ever-increasing body of charitable organizations which span, enhance and characterize the civil sphere. Overall, charities occupy an important institutional role in society, and the individuals who help to sustain charities encompass a vital social role. This paper is particularly concerned with analyzing charitable donation and volunteering dynamics on the individual level. Using the 2014 General Social Survey data on charitableness, this paper estimates the probability of engaging in volunteerism and charitable donation within the nonprofit sector based on income level. These results suggest that …


The Marshall-Olkin Exponentiated Generalized G Family Of Distributions: Properties, Applications, And Characterizations, Haitham M. Yousof, Mahdi Rasekhi, Morad Alizadeh, Gholamhossein Hamedani Jan 2020

The Marshall-Olkin Exponentiated Generalized G Family Of Distributions: Properties, Applications, And Characterizations, Haitham M. Yousof, Mahdi Rasekhi, Morad Alizadeh, Gholamhossein Hamedani

Mathematical and Statistical Science Faculty Research and Publications

In this paper, we propose and study a new class of continuous distributions called the Marshall-Olkin exponentiated generalized G (MOEG-G) family which extends the Marshall-Olkin-G family introduced by Marshall and Olkin [A. W. Marshall, I. Olkin, Biometrika 84 (1997), 641-652]. Some of its mathematical properties including explicit expressions for the ordinary and incomplete moments, generating function, order statistics and probability weighted moments are derived. Some characteristics of the new family are presented. Maximum likelihood estimation for the model parameters under uncensored and censored data is addressed in Section 5 as well as a simulation study to assess the performance of …


Algebraic And Geometric Properties Of Hierarchical Models, Aida Maraj Jan 2020

Algebraic And Geometric Properties Of Hierarchical Models, Aida Maraj

Theses and Dissertations--Mathematics

In this dissertation filtrations of ideals arising from hierarchical models in statistics related by a group action are are studied. These filtrations lead to ideals in polynomial rings in infinitely many variables, which require innovative tools. Regular languages and finite automata are used to prove and explicitly compute the rationality of some multivariate power series that record important quantitative information about the ideals. Some work regarding Markov bases for non-reducible models is shown, together with advances in the polyhedral geometry of binary hierarchical models.


Unitary And Symmetric Structure In Deep Neural Networks, Kehelwala Dewage Gayan Maduranga Jan 2020

Unitary And Symmetric Structure In Deep Neural Networks, Kehelwala Dewage Gayan Maduranga

Theses and Dissertations--Mathematics

Recurrent neural networks (RNNs) have been successfully used on a wide range of sequential data problems. A well-known difficulty in using RNNs is the vanishing or exploding gradient problem. Recently, there have been several different RNN architectures that try to mitigate this issue by maintaining an orthogonal or unitary recurrent weight matrix. One such architecture is the scaled Cayley orthogonal recurrent neural network (scoRNN), which parameterizes the orthogonal recurrent weight matrix through a scaled Cayley transform. This parametrization contains a diagonal scaling matrix consisting of positive or negative one entries that can not be optimized by gradient descent. Thus the …


Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich Jan 2020

Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich

Theses and Dissertations--Mathematics

Despite the recent success of various machine learning techniques, there are still numerous obstacles that must be overcome. One obstacle is known as the vanishing/exploding gradient problem. This problem refers to gradients that either become zero or unbounded. This is a well known problem that commonly occurs in Recurrent Neural Networks (RNNs). In this work we describe how this problem can be mitigated, establish three different architectures that are designed to avoid this issue, and derive update schemes for each architecture. Another portion of this work focuses on the often used technique of batch normalization. Although found to be successful …


Cancer Phylogenetic Analysis Based On Rna-Seq Data, Tingting Zhai Jan 2020

Cancer Phylogenetic Analysis Based On Rna-Seq Data, Tingting Zhai

Theses and Dissertations--Statistics

Studying tumor evolution is a major task to understand the biological mechanism of carcinogenesis, develop new cancer therapies, and prevent drug resistance. We focus on two important questions in tumor evolution. The first question is to quantify intra-tumor heterogeneity, where multiple subclones of tumor cells with distinct transcriptomic profiles. Another question is to estimate the temporal order of alteration of key cancer pathways during tumor evolution. We present a new statistical method to 1) reconstruct the evolutionary history and population frequency of the subclonal lineages of tumor cells and 2) infer temporal order of pathway alterations in tumor evolution for …


Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li Jan 2020

Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li

Theses and Dissertations--Statistics

Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …


Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu Jan 2020

Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu

Theses and Dissertations--Statistics

The Bayesian adjustment for confounding (BAC) is a Bayesian model averaging method to select and adjust for confounding factors when evaluating the average causal effect of an exposure on a certain outcome. We extend the BAC method to time-to-event outcomes. Specifically, the posterior distribution of the exposure effect on a time-to-event outcome is calculated as a weighted average of posterior distributions from a number of candidate proportional hazards models, weighing each model by its ability to adjust for confounding factors. The Bayesian Information Criterion based on the partial likelihood is used to compare different models and approximate the Bayes factor. …


Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui Jan 2020

Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui

Theses and Dissertations--Statistics

In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.

In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …


Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu Jan 2020

Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu

Theses and Dissertations--Statistics

A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …


Measuring Variability In Model Performance Measures, Matthew Rutledge Jan 2020

Measuring Variability In Model Performance Measures, Matthew Rutledge

Theses and Dissertations--Statistics

As data become increasingly available, statisticians are confronted with both larger sample sizes and larger numbers of predictors. While both of these factors are beneficial in building better predictive models and allowing for better inference, models can become difficult to interpret and often include variables of little practical significance. This dissertation provides methods that assist model builders to better understand and select from a collection of candidate models. We study the asymptotic distribution of AIC and propose a graphical tool to assist practitioners in comparing and contrasting candidate models. Real-world examples show how this graphic might be used and a …


The Economic Determinants Of American Professional Sports Franchise Valuations, Ryan Flora Jan 2020

The Economic Determinants Of American Professional Sports Franchise Valuations, Ryan Flora

Mahurin Honors College Capstone Experience/Thesis Projects

This thesis seeks to analyze the impact of regional identities on American professional sports team valuations. Regional identities are classified as any name of a team that is not tied directly to the city that they reside in. For example, the Carolina Panthers have a regional identity because they are not based out of “Carolina”, they are based out of Charlotte, North Carolina. Another example would be the Arizona Cardinals, whose name encompasses the whole state of Arizona rather than Phoenix, the city they are based out of. The leagues that will be involved in this study are the National …


Artificial Neural Network Models For Pattern Discovery From Ecg Time Series, Mehakpreet Kaur Jan 2020

Artificial Neural Network Models For Pattern Discovery From Ecg Time Series, Mehakpreet Kaur

College of Graduate Studies: Theses & Dissertations

Artificial Neural Network (ANN) models have recently become de facto models for deep learning with a wide range of applications spanning from scientific fields such as computer vision, physics, biology, medicine to social life (suggesting preferred movies, shopping lists, etc.). Due to advancements in computer technology and the increased practice of Artificial Intelligence (AI) in medicine and biological research, ANNs have been extensively applied not only to provide quick information about diseases, but also to make diagnostics accurate and cost-effective. We propose an ANN-based model to analyze a patient's electrocardiogram (ECG) data and produce accurate diagnostics regarding possible heart diseases …