Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (374)
- Statistical Methodology (350)
- Social and Behavioral Sciences (254)
- Statistical Theory (216)
- Medicine and Health Sciences (197)
-
- Biostatistics (192)
- Data Science (181)
- Multivariate Analysis (162)
- Life Sciences (145)
- Computer Sciences (144)
- Longitudinal Data Analysis and Time Series (141)
- Applied Mathematics (136)
- Probability (129)
- Survival Analysis (128)
- Engineering (123)
- Categorical Data Analysis (112)
- Public Health (111)
- Mathematics (110)
- Business (108)
- Economics (93)
- Other Statistics and Probability (83)
- Artificial Intelligence and Robotics (78)
- Environmental Sciences (72)
- Design of Experiments and Sample Surveys (66)
- Epidemiology (66)
- Numerical Analysis and Computation (66)
- Genetics and Genomics (54)
- Institution
-
- COBRA (242)
- University of Kentucky (58)
- Southern Methodist University (52)
- Central Bank of Nigeria (29)
- City University of New York (CUNY) (29)
-
- Virginia Commonwealth University (29)
- University of Nebraska - Lincoln (28)
- Old Dominion University (27)
- World Maritime University (26)
- Illinois State University (25)
- University of Arkansas, Fayetteville (24)
- Georgia Southern University (22)
- Claremont Colleges (20)
- East Tennessee State University (20)
- Air Force Institute of Technology (19)
- Purdue University (19)
- University of Nevada, Las Vegas (17)
- Kennesaw State University (16)
- Technological University Dublin (16)
- The Texas Medical Center Library (16)
- The University of Akron (16)
- Embry-Riddle Aeronautical University (15)
- California Polytechnic State University, San Luis Obispo (14)
- University of Louisville (14)
- Utah State University (14)
- Clemson University (13)
- Michigan Technological University (12)
- Western Michigan University (12)
- Wright State University (12)
- Loma Linda University (11)
- Keyword
-
- Statistics (73)
- Machine Learning (32)
- Machine learning (28)
- Regression (28)
- Simulation (18)
-
- Modeling (17)
- Prediction (15)
- Bayesian (14)
- Classification (13)
- Logistic regression (13)
- Mathematics (13)
- Deep Learning (12)
- Forecasting (12)
- Survival analysis (12)
- COVID-19 (11)
- Data Science (11)
- Gene expression (11)
- Models (11)
- Causal inference (10)
- Mathematical models (10)
- R (10)
- Epidemiology (9)
- Genetics (9)
- Model selection (9)
- Psychology (9)
- Time series (9)
- Artificial Intelligence (8)
- Cross-validation (8)
- Deep learning (8)
- Linear regression (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (60)
- Theses and Dissertations (59)
- Electronic Theses and Dissertations (48)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (47)
- Harvard University Biostatistics Working Paper Series (43)
-
- Theses and Dissertations--Statistics (42)
- The University of Michigan Department of Biostatistics Working Paper Series (40)
- SMU Data Science Review (37)
- UW Biostatistics Working Paper Series (31)
- CBN Journal of Applied Statistics (JAS) (29)
- World Maritime University Dissertations (26)
- Annual Symposium on Biomathematics and Ecology Education and Research (21)
- College of Graduate Studies: Theses & Dissertations (21)
- Articles (17)
- Dissertations (17)
- Graduate Theses and Dissertations (17)
- CMC Senior Theses (16)
- Williams Honors College, Honors Research Projects (16)
- COBRA Preprint Series (15)
- Dissertations and Theses (Open Access) (15)
- Psychology Faculty Publications (14)
- Statistical Science Theses and Dissertations (14)
- Dissertations, Master's Theses and Master's Reports (12)
- All Dissertations (11)
- Loma Linda University Electronic Theses, Dissertations & Projects (11)
- Master's Theses (11)
- Dissertations, Theses, and Capstone Projects (10)
- LSU New Orleans Theses and Dissertations (10)
- Publications (10)
- Publications and Research (9)
- Publication Type
- File Type
Articles 511 - 540 of 1308
Full-Text Articles in Statistical Models
Shrinkage Priors For Isotonic Probability Vectors And Binary Data Modeling, Philip S. Boonstra, Daniel R. Owen, Jian Kang
Shrinkage Priors For Isotonic Probability Vectors And Binary Data Modeling, Philip S. Boonstra, Daniel R. Owen, Jian Kang
The University of Michigan Department of Biostatistics Working Paper Series
This paper outlines a new class of shrinkage priors for Bayesian isotonic regression modeling a binary outcome against a predictor, where the probability of the outcome is assumed to be monotonically non-decreasing with the predictor. The predictor is categorized into a large number of groups, and the set of differences between outcome probabilities in consecutive categories is equipped with a multivariate prior having support over the set of simplexes. The Dirichlet distribution, which can be derived from a normalized cumulative sum of gamma-distributed random variables, is a natural choice of prior, but using mathematical and simulation-based arguments, we show that …
Rejoinder On ‘A Selective Overview Of Sparse Sufficient Dimension Reduction’, Lu Li, Xuewong Meggie Wen, Zhou Yu
Rejoinder On ‘A Selective Overview Of Sparse Sufficient Dimension Reduction’, Lu Li, Xuewong Meggie Wen, Zhou Yu
Mathematics and Statistics Faculty Research & Creative Works
No abstract provided.
Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown
Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown
Murray State Theses and Dissertations
Data and algorithmic modeling are two different approaches used in predictive analytics. The models discussed from these two approaches include the proportional odds logit model (POLR), the vector generalized linear model (VGLM), the classification and regression tree model (CART), and the random forests model (RF). Patterns in the data were analyzed using trigonometric polynomial approximations and Fast Fourier Transforms. Predictive modeling is used frequently in statistics and data science to find the relationship between the explanatory (input) variables and a response (output) variable. Both approaches prove advantageous in different cases depending on the data set. In our case, the data …
Projecting Regions Of North Atlantic Right Whale, Eubalaena Glacialis, Habitat Suitability In The Gulf Of Maine In 2050, Camille Ross
Projecting Regions Of North Atlantic Right Whale, Eubalaena Glacialis, Habitat Suitability In The Gulf Of Maine In 2050, Camille Ross
Honors Theses
North Atlantic right whales (Eubalaena glacialis) are endangered. Understanding the role environmental conditions play in habitat suitability is key to determining the regions in need of protection for conservation of the species, particularly as climate change shifts suitable habitat. This thesis uses three species distribution modeling algorithms, together with historical data on whale abundance(1993 to 2009) and environmental covariates to build monthly ensemble models of past E. glacialis habitat suitability in the Gulf of Maine. Then, the models are projected onto the year 2050 for a range of climate scenarios. Specifically, the distribution of the species was modeled …
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Theses and Dissertations
Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …
Modelling Interactions Among Offenders: A Latent Space Approach For Interdependent Ego-Networks, Isabella Gollini, Alberto Caimo, Paolo Campana
Modelling Interactions Among Offenders: A Latent Space Approach For Interdependent Ego-Networks, Isabella Gollini, Alberto Caimo, Paolo Campana
Articles
Illegal markets are notoriously difficult to study. Police data offer an increasingly exploited source of evidence. However, their secondary nature poses challenges for researchers. A key issue is that researchers often have to deal with two sets of actors: targeted and non-targeted. This work develops a latent space model for interdependent ego-networks purposely created to deal with the targeted nature of police evidence. By treating targeted offenders as egos and their contacts as alters, the model (a) leverages on the full information available and (b) mirrors the specificity of the data collection strategy. The paper then applies this approach to …
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Theses and Dissertations--Statistics
Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Theses and Dissertations--Statistics
The Bayesian adjustment for confounding (BAC) is a Bayesian model averaging method to select and adjust for confounding factors when evaluating the average causal effect of an exposure on a certain outcome. We extend the BAC method to time-to-event outcomes. Specifically, the posterior distribution of the exposure effect on a time-to-event outcome is calculated as a weighted average of posterior distributions from a number of candidate proportional hazards models, weighing each model by its ability to adjust for confounding factors. The Bayesian Information Criterion based on the partial likelihood is used to compare different models and approximate the Bayes factor. …
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Theses and Dissertations--Statistics
A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Theses and Dissertations--Statistics
Kinetic modeling of the time dependence of metabolite concentrations including the unstable isotope labeled species is an important approach to simulate metabolic pathway dynamics. It is also essential for quantitative metabolic flux analysis using tracer data. However, as the metabolic networks are complex including extensive compartmentation and interconnections, the parameter estimation for enzymes that catalyze individual reactions needed for kinetic modeling is challenging. As the pa- rameter space is large and multi-dimensional while kinetic data are comparatively sparse, the estimation procedure (especially the point estimation methods) often en- counters multiple local maximum such that standard maximum likelihood methods may yield …
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Theses and Dissertations--Statistics
Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …
Enhancing Models And Measurements Of Traffic-Related Air Pollutants For Health Studies Using Dispersion Modeling And Bayesian Data Fusion, Stuart A. Batterman, Veronica J. Berrocal, Chad Milando, Owais Gilani, Saravanan Arunachalam, K. Max Zhang
Enhancing Models And Measurements Of Traffic-Related Air Pollutants For Health Studies Using Dispersion Modeling And Bayesian Data Fusion, Stuart A. Batterman, Veronica J. Berrocal, Chad Milando, Owais Gilani, Saravanan Arunachalam, K. Max Zhang
Faculty Journal Articles
Research Report 202 describes a study led by Dr. Stuart Batterman at the University of Michigan, Ann Arbor and colleagues. The investigators evaluated the ability to predict traffic-related air pollution using a variety of methods and models, including a line source air pollution dispersion model and sophisticated spatiotemporal Bayesian data fusion methods. Exposure assessment for traffic-related air pollution is challenging because the pollutants are a complex mixture and vary greatly over space and time. Because extensive direct monitoring is difficult and expensive, a number of modeling approaches have been developed, but each model has its own limitations and errors.
Dr. …
Parameter Estimation Of A Seasonal Poisson Inar(1) Model With Different Monthly Means, Turaj Vazifedan, Homa Jalaeian Taghadomi, Xixi Wang, Mujde Erten-Unal
Parameter Estimation Of A Seasonal Poisson Inar(1) Model With Different Monthly Means, Turaj Vazifedan, Homa Jalaeian Taghadomi, Xixi Wang, Mujde Erten-Unal
Civil & Environmental Engineering Faculty Publications
Analysing seasonality in count time series is an essential application of statistics to predict phenomena in different fields like economics, agriculture, healthcare, environment, and climatic change. However, the information in the existing literature is scarce regarding the performances of relevant statistical models. This study provides the Yule-Walker (Y-W), Conditional Least Squares (CLS), and Maximum Likelihood Estimation (MLE) for First-order Non-negative Integer-valued Autoregressive, INAR(1), process with Poisson innovations with different monthly means. The performance of Y-W, CLS, and MLE are assessed by the Monte Carlo simulation method. The performance of this model is compared with another seasonal INAR(1) model by reproducing …
An Examination Of Covid-19 Statistical Modeling, Shane Vaughan
An Examination Of Covid-19 Statistical Modeling, Shane Vaughan
Williams Honors College, Honors Research Projects
The 2019 novel coronavirus, also known as COVID-19, is an infectious disease which was first reported in late 2019 and soon spread to become a global pandemic, prompting major action from world governments. Soon after, many institutions began attempts to analyze and predict the spread and severity of the disease via statistical modeling. Some information is not available for public consumption; however, a number of institutions have published the results of their analyses and some have made public repositories of the code used to build the models. This research paper attempts use these and other resources to examine the modeling …
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Theses and Dissertations
Humans are exposed to multiple chemicals every day. Epidemiological studies have shown that chemical mixtures are associated with cancers, allergies, neurodevelopmental disorders, and other adverse health effects. To assess these associations, investigators are increasingly using chemical mixture approaches like weighted quantile sum (WQS) regression. In these studies, the research objectives are to determine whether a mixture of correlated chemicals is associated with an adverse health outcome and to identify the important chemicals. However, as experimental equipment measures each exposure to a chemical-specific detection limit, the exposures are unknown between zero and the detection limit. Indeed, the number of exposures below …
K-Means Stock Clustering Analysis Based On Historical Price Movements And Financial Ratios, Shu Bin
K-Means Stock Clustering Analysis Based On Historical Price Movements And Financial Ratios, Shu Bin
CMC Senior Theses
The 2015 article Creating Diversified Portfolios Using Cluster Analysis proposes an algorithm that uses the Sharpe ratio and results from K-means clustering conducted on companies' historical financial ratios to generate stock market portfolios. This project seeks to evaluate the performance of the portfolio-building algorithm during the beginning period of the COVID-19 recession. S&P 500 companies' historical stock price movement and their historical return on assets and asset turnover ratios are used as dissimilarity metrics for K-means clustering. After clustering, stock with the highest Sharpe ratio from each cluster is picked to become a part of the portfolio. The economic and …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
Modeling The Galactic Compact Binary Neutron Star Population And Studying The Double Pulsar System, Nihan Pol
Modeling The Galactic Compact Binary Neutron Star Population And Studying The Double Pulsar System, Nihan Pol
Graduate Theses, Dissertations, and Problem Reports (ETD)
Binary neutron star (BNS) systems consisting of at least one neutron star provide an avenue for testing a broad range of physical phenomena ranging from tests of General Relativity to probing magnetospheric physics to understanding the behavior of matter in the densest environments in the Universe. Ultra-compact BNS systems with orbital periods less than few tens of minutes emit gravitational waves with frequencies ~mHz and are detectable by the planned space-based Laser Interferometer Space Antenna (LISA), while merging BNS systems produce a chirping gravitational wave signal that can be detected by the ground-based Laser Interferometer Gravitational-Wave Observatory (LIGO). Thus, BNS …
A Mathematical Model For Malaria With Age-Heterogeneous Biting Rate, Sho Kawakami
A Mathematical Model For Malaria With Age-Heterogeneous Biting Rate, Sho Kawakami
All Graduate Theses, Dissertations, and Other Capstone Projects
We propose a mathematical model for malaria with age-heterogeneous biting rate from mosquitos. The existence of the model, the local behavior of the disease free equilibrium are explored. Furthermore the model is extended to an optimal control problem and the corresponding adjoint equations and optimality conditions are derived. Age dependent parameter values are estimated and numerical simulations are carried out for the model. The new model better accounts for difference in biting rates of mosquitos to different age groups, and improvements in stability to the explicit algorithm. The optimal control is also shown to depend on the age distribution of …
Statistical Analysis Of Demographic Effects On Insurance Coverage Of Perinatal And Neonatal Morbidity, Madeline Durbin
Statistical Analysis Of Demographic Effects On Insurance Coverage Of Perinatal And Neonatal Morbidity, Madeline Durbin
Undergraduate Honors Thesis Projects
In the United States of America, Ohio has one of the worst neonatal and perinatal death rates. Within Ohio, Montgomery County has an above average neonatal and perinatal death rate. This statistic can be lowered if more women in Montgomery County have health insurance. They would be more likely to seek out prenatal health care, since they would no longer have to pay as much money out-of-pocket. This would allow medical professionals to be able to diagnose and treat any potential issues in the mother or child earlier. Having health insurance would also prevent mothers-to-be from seeking out other potentially …
Phenotype Extraction: Estimation And Biometrical Genetic Analysis Of Individual Dynamics, Kevin L. Mckee
Phenotype Extraction: Estimation And Biometrical Genetic Analysis Of Individual Dynamics, Kevin L. Mckee
Theses and Dissertations
Within-person data can exhibit a virtually limitless variety of statistical patterns, but it can be difficult to distinguish meaningful features from statistical artifacts. Studies of complex traits have previously used genetic signals like twin-based heritability to distinguish between the two. This dissertation is a collection of studies applying state-space modeling to conceptualize and estimate novel phenotypic constructs for use in psychiatric research and further biometrical genetic analysis. The aims are to: (1) relate control theoretic concepts to health-related phenotypes; (2) design statistical models that formally define those phenotypes; (3) estimate individual phenotypic values from time series data; (4) consider hierarchical …
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
Theses and Dissertations
Diversity of intrinsic neural attributes and network connections is known to exist in many areas of the brain and is thought to significantly affect neural coding. Recent theoretical and experimental work has argued that in uncoupled networks, coding is most accurate at intermediate levels of heterogeneity. I explore this phenomenon through two distinct approaches: a theoretical mathematical modeling approach and a data-driven statistical modeling approach.
Through the mathematical approach, I examine firing rate heterogeneity in a feedforward network of stochastic neural oscillators utilizing a high-dimensional model. The firing rate heterogeneity stems from two sources: intrinsic (different individual cells) and network …
Sex And Age Differences In Prevalence And Risk Factors For Prediabetes In Mexican-Americans, Kristina Vatcheva, Belinda M. Reininger, Susan P. Fisher-Hoch, Joseph B. Mccormick
Sex And Age Differences In Prevalence And Risk Factors For Prediabetes In Mexican-Americans, Kristina Vatcheva, Belinda M. Reininger, Susan P. Fisher-Hoch, Joseph B. Mccormick
School of Mathematical & Statistical Sciences Faculty Publications
AIMS:
Over 1/3 of Americans have prediabetes, while 9.4% have type 2 diabetes. The aim of our study was to estimate the prevalence of prediabetes in Mexican Americans, with known 28.2% prevalence of type 2 diabetes, by age and sex and to identify critical socio-demographic and clinical factors associated with prediabetes.
METHODS:
Data were collected between 2004 and 2017 from the Cameron County Hispanic Cohort in Texas. Weighted crude and sex- and age- stratified prevalences were calculated. Survey weighted logistic regression analyses were conducted to identify risk factors for prediabetes.
RESULTS:
The prevalence of prediabetes (32%) was slightly higher than …
Aggregate Loss Model With Poisson-Tweedie Loss Frequency, Si Chen
Aggregate Loss Model With Poisson-Tweedie Loss Frequency, Si Chen
Theses and Dissertations (Comprehensive)
The aggregate loss model has applications in various areas such as financial risk management and actuarial science. The aggregate loss is the summation of all random losses occurred in a period, and it is governed by both the loss severity and the loss frequency. While the impact of the loss severity on aggregate loss is well studied, less focus is paid on the influence of loss frequency on aggregate loss, which motivates our study. In this thesis, we enrich the aggregate loss framework by introducing the Poisson-Tweedie distribution as a candidate for modelling loss frequency, prove the closedness of Poisson-Tweedie …
Identifying Customer Churn In After-Market Operations Using Machine Learning Algorithms, Vitaly Briker, Richard Farrow, William Trevino, Brent Allen
Identifying Customer Churn In After-Market Operations Using Machine Learning Algorithms, Vitaly Briker, Richard Farrow, William Trevino, Brent Allen
SMU Data Science Review
This paper presents a comparative study on machine learning methods as they are applied to product associations, future purchase predictions, and predictions of customer churn in aftermarket operations. Association rules are used help to identify patterns across products and find correlations in customer purchase behaviour. Studying customer behaviour as it pertains to Recency, Frequency, and Monetary Value (RFM) helps inform customer segmentation and identifies customers with propensity to churn. Lastly, Flowserve’s customer purchase history enables the establishment of churn thresholds for each customer group and assists in constructing a model to predict future churners. The aim of this model is …
Personalized Detection Of Anxiety Provoking News Events Using Semantic Network Analysis, Jacquelyn Cheun Phd, Luay Dajani, Quentin B. Thomas
Personalized Detection Of Anxiety Provoking News Events Using Semantic Network Analysis, Jacquelyn Cheun Phd, Luay Dajani, Quentin B. Thomas
SMU Data Science Review
In the age of hyper-connectivity, 24/7 news cycles, and instant news alerts via social media, mental health researchers don't have a way to automatically detect news content which is associated with triggering anxiety or depression in mental health patients. Using the Associated Press news wire, a semantic network was built with 1,056 news articles containing over 500,000 connections across multiple topics to provide a personalized algorithm which detects problematic news content for a given reader. We make use of Semantic Network Analysis to surface the relationship between news article text and anxiety in readers who struggle with mental health disorders. …
Ordinal Hyperplane Loss, Bob Vanderheyden
Ordinal Hyperplane Loss, Bob Vanderheyden
Doctor of Data Science and Analytics Dissertations
This research presents the development of a new framework for analyzing ordered class data, commonly called “ordinal class” data. The focus of the work is the development of classifiers (predictive models) that predict classes from available data. Ratings scales, medical classification scales, socio-economic scales, meaningful groupings of continuous data, facial emotional intensity and facial age estimation are examples of ordinal data for which data scientists may be asked to develop predictive classifiers. It is possible to treat ordinal classification like any other classification problem that has more than two classes. Specifying a model with this strategy does not fully utilize …
The Epsilon-Skew Rayleigh Distribution, By John Greene, John M. Greene
The Epsilon-Skew Rayleigh Distribution, By John Greene, John M. Greene
Theses and Dissertations
In this dissertation, a new family of skew distributions is introduced and developed, the Epsilon Skew Rayleigh. The members of this family are bimodal skewed distributions with location, scale and skewness parameters. There exist two unimodal parameter cases. The distribution can be skewed or symmetric. This distribution family has many applications including population demographics, signal dynamics, ocean wave heights and hardware failure rates. The effects of the parameters are described and developed. We derive the moment generating and maximum likelihood functions, as well as the expected value, median, modes, variance, skewness and kurtosis. The properties of a random variable with …
Seasonal Time Series Models With Application To Weather And Lake Level Data, Mengqing Qin
Seasonal Time Series Models With Application To Weather And Lake Level Data, Mengqing Qin
Graduate Theses/Dissertations
This work studies seasonal time series models with application to lake level and weather data. The thesis includes related time series concepts, integrated autoregressive moving average models (abbreviated as ARIMA), parameter estimation, model diagnostics, and forecasting. The studied time series models are applied to the data of daily lake level in Beaver Lake (1988-2017) and the data of daily maximum temperature in New York Central Park (1870-2017). Due to seasonality of the data, three different approaches are proposed to the modeling: regression method, functional ARIMA method and multiplicative seasonal ARIMA method. The forecasted values of the year 2018 are compared …
Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova
Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova
College of Education and Human Sciences: Dissertations, Theses, and Student Research
When assessing a certain characteristic or trait using a multiple item measure, quality of that measure can be assessed by examining the reliability. To avoid multiple time points, reliability can be represented by internal consistency, which is most commonly calculated using Cronbach’s coefficient alpha. Almost every time human participants are involved in research, there is missing data involved. Missing data means that even though complete data were expected to be collected, some data are missing. Missing data can follow different patterns as well as be the result of different mechanisms. One traditional way to deal with missing data is listwise …