Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Open Access Theses & Dissertations

Discipline
Keyword
Publication Year

Articles 1 - 30 of 121

Full-Text Articles in Statistics and Probability

Advancing Effective Connectivity Analysis: Robust And Sparse Group Dynamic Causal Modeling Via Extended Parametric Empirical Bayes, Godfred Arhin Dec 2025

Advancing Effective Connectivity Analysis: Robust And Sparse Group Dynamic Causal Modeling Via Extended Parametric Empirical Bayes, Godfred Arhin

Open Access Theses & Dissertations

Dynamic Causal Modeling (DCM) provides a principled framework for estimating effective connectivity in neuroimaging data, with Parametric Empirical Bayes (PEB) enabling hierarchical inference across sessions and subjects. However, standard PEB assumes Gaussian distributions at all hierarchical levels and employs conventional shrinkage priors that do not enforce true sparsity, limiting robustness to outlying subjects and reducing sensitivity to parsimonious connectivity structures. This thesis introduces a robust, sparsity-inducing hierarchical extension to DCM-PEB that addresses these limitations through two key innovations. First, a Student-t group-level likelihood replaces the conventional Gaussian likelihood, automatically downweighting outlying subject-level parameters using adaptive precision weights while retaining computational …


Copula-Based Tests For Assessing The Association Between Genetic Variants And Mixed Phenotypes, Martin Amoah Dec 2025

Copula-Based Tests For Assessing The Association Between Genetic Variants And Mixed Phenotypes, Martin Amoah

Open Access Theses & Dissertations

High-dimensional omics studies increasingly involve heterogeneous data types and phenotypes, which traditional association methods struggle to model jointly due to incompatible marginal distributions and complex dependence structures. This thesis develops a unified copula-based framework for assessing associations between genetic variants and mixed phenotypes by decoupling flexible marginal models from their joint dependence structure. While previous copula-based approaches in this setting have focused largely on continuous and binary traits, we extend these methods to a broader class of phenotype pairs. Specifically, we introduce new association tests for bivariate outcomes involving ordinal–continuous, nominal–continuous, and survival–continuous combinations. The proposed methodology derives joint density …


Instance-Adaptive Gated Fusion Of Multi-Transform Image Representations, Prince Appiah Dec 2025

Instance-Adaptive Gated Fusion Of Multi-Transform Image Representations, Prince Appiah

Open Access Theses & Dissertations

This dissertation proposes the Instance-Adaptive Gated Fusion (IAGF) framework, a novel deep learning architecture for adaptive and interpretable fusion of multiple time–series image transformations. While existing methods rely on static concatenation or dataset-level optimization, IAGF introduces a learnable gating mechanism that dynamically assigns per-instance weights to Recurrence Plots (RP), Gramian Angular Summation Fields (GASF), and Gramian Angular Difference Fields (GADF). The gating layer performs a convex fusion of transformation-specific embeddings under a softmax constraint, ensuring mathematical stability and interpretability. An entropy-regularized objective prevents dominance collapse and promotes balanced exploration of transformations during training. Comprehensive experiments across eighteen benchmark datasets, spanning …


Multi-Hop Hybrid Graph Neural Network, James Arthur Dec 2025

Multi-Hop Hybrid Graph Neural Network, James Arthur

Open Access Theses & Dissertations

Graph-structured data appear across diverse domains, such as social networks, citation graphs, biological systems, and knowledge bases. Graph Neural Networks (GNNs) have emerged as a powerful framework for learning on such data, yet existing architectures face significant challenges. Graph Convolutional Networks (GCNs) suffer from over-smoothing as depth increases, Graph Attention Networks (GATs) introduce computational and statistical instabilities, and naïve multi-hop propagation inflates memory and computation while failing to adapt to topology. These limitations motivate the development of a new framework that is both expressive and scalable. This dissertation proposes the Multi-Hop Hybrid Graph Neural Network (MHHGNN), a novel architecture that …


A Unified Framework For Embedding-Based Synthetic Data Generation With High Cardinality Categorical Features, Cesar Iram Vazquez Dec 2025

A Unified Framework For Embedding-Based Synthetic Data Generation With High Cardinality Categorical Features, Cesar Iram Vazquez

Open Access Theses & Dissertations

High-cardinality categorical variables remain difficult to model in tabular data, where classical encoders encounter sparsity, susceptibility to leakage, and the loss of meaningful relational structure. This dissertation develops a unified framework for learning, evaluating, and synthesizing representations of such variables using both traditional encoders and modern embedding methods, including Word2Vec, FastText, Node2Vec, TF–IDF/SVD, and supervised entity embeddings. The framework is applied across three benchmark datasets (Adult, PetFinder, Breast Cancer) and a hierarchical educational case study (IPEDS/CIP). Embedding quality is examined through both downstream predictive performance and structure-focused diagnostics that quantify neighborhood behavior and geometric coherence. To assess whether synthetic data …


Rethinking Iterative Proportional Fitting: Scalable And Hybrid Approaches To Joint Distribution Fitting, William Ofosu Agyapong Aug 2025

Rethinking Iterative Proportional Fitting: Scalable And Hybrid Approaches To Joint Distribution Fitting, William Ofosu Agyapong

Open Access Theses & Dissertations

The Iterative Proportional Fitting (IPF) algorithm is widely used in contingency table estimation, survey weighting, and synthetic population generation due to its simplicity and strong theoretical foundation for matching observed marginal distributions. However, in high-dimensional settings, IPF faces substantial computational and memory demands, as well as statistical instability caused by sparse contingency tables. Moreover, IPF is less useful in modern population synthesis tasks that require both scalability and realism because, despite its superiority in matching known marginal distributions, it cannot produce realistic out-of-sample data points. To address these limitations, we first propose a blockwise IPF framework, in which the feature …


A Multi-Modal Method For Synthetic Data Generation In Social Network Analysis, Hortencia Josefina Hernandez Aug 2025

A Multi-Modal Method For Synthetic Data Generation In Social Network Analysis, Hortencia Josefina Hernandez

Open Access Theses & Dissertations

Social network analysis (SNA) research is often rife with data collection pitfalls, frequently leading to incomplete and missing data. With the growing use of SNA-based research, researchers must address the challenge of missing data and synthetic data generation in these settings. Missing data occurs due to longitudinal non-response or lack of response to sensitive or difficult-to-answer questions. Synthetic data generation in SNA settings addresses the lack of representation that is often present in large-scale SNA studies. This dissertation investigates synthetic data generation methods to address these challenges and develops a novel algorithm that leverages information from multi-modal data, e.g., databases …


Simultaneous Selection Of Inflations And Variables In Multiple Inflations Poisson Model (Mip), John Koomson Aug 2025

Simultaneous Selection Of Inflations And Variables In Multiple Inflations Poisson Model (Mip), John Koomson

Open Access Theses & Dissertations

Count data frequently arise in biomedical, economic, and social science research and are often characterized by structural excesses at specific count levels. To accommodate such patterns, Su et al. (2013), among others, introduced the Multiple-Inflation Poisson (MIP) model, which allows for multiple inflated counts within the distribution. However, two critical challenges remain in modeling such data: (i) identifying the true inflation points where excess counts occur, and (ii) selecting the relevant covariates that explain variation in the inflation and count process. This dissertation addresses these issues by advancing the MIP model through a novel methodology that enables the simultaneous selection …


Kernel Density Estimation And Convolution, Nicholas Tenkorang May 2025

Kernel Density Estimation And Convolution, Nicholas Tenkorang

Open Access Theses & Dissertations

Kernel Density Estimation (KDE) is a widely used technique for estimating the probability density function of a random variable. In this study, we revisit KDE through the lens of convolution and extend this perspective to special cases such as positive, bounded and heavy tailed random variables. Building on this foundation, we propose a novel simulation-based density estimation method that generates new data by adding noise to observed values and then smoothing the resulting histogram using splines. A minor adjustment to natural cubic splines is required to ensure nonnegative estimates. The noise is drawn from a class of bounded polynomial kernel …


Yield Prediction Of Pv Solar Energy Systems And Its Application In The Energy Grid For Operational Efficiency, Pablo Bustamante May 2025

Yield Prediction Of Pv Solar Energy Systems And Its Application In The Energy Grid For Operational Efficiency, Pablo Bustamante

Open Access Theses & Dissertations

The generation of power from photovoltaic (PV) solar panels is influenced by a multitude of factors. These include the tilt and orientation of the solar panels, the latitude of their location, and the prevailing climate and weather conditions. Additionally, shading at specific locations, particularly if the panels are not part of a solar facility, can significantly impact their efficiency. The quality and efficiency of the panels themselves, along with the preventive maintenance of both the solar panels and associated components such as inverters and trackers, are also critical. Finally, the overall system design and installation play a vital role in …


A Profile Wald Test In M-Estimation, Reagan Kesseku May 2025

A Profile Wald Test In M-Estimation, Reagan Kesseku

Open Access Theses & Dissertations

Despite the growing popularity of machine learning-based inference, classical statistical inference remains highly relevant in modern data science due to its interpretability and theoretical rigor. Among its core tools, the likelihood ratio test, Wald test, and score test are foundational methods for hypothesis testing within the maximum likelihood framework. Although these tests are asymptotically equivalent under regularity conditions, each offers distinct advantages depending on the context, computational demands, and the availability of parameter estimates. In this dissertation, we introduce a fourth method, the Profile Wald Test (PWT), within the broader M-estimation framework. The PWT is based on profile estimators of …


Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta May 2025

Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta

Open Access Theses & Dissertations

Prostate cancer (PrCa) remains a critical challenge in precision oncology due to several reasons including its apparent heterogenous condition, recurrence following treatment and rapid progressive forms. Therefore, identifying patients at risk of progression is essential to fast-track therapeutic decisions and improve outcomes. Despite recent advances in genomic and molecular profiling, conventional PrCa risk assessment tools heavily rely on a few clinical parameters, neglecting the prognostic potential of genomic biomarkers in the presence of clinical biomarkers. This study presents a computational pipeline to harmonize and evaluate the prognostic value of clinicogenomic profiles of patients in modelling progression free survival (PFS). PFS, …


Optimal Experimental Plan For Multi-Level Stress Testing Under Progressively Hybrid Censoring, David Kojo Amakye May 2025

Optimal Experimental Plan For Multi-Level Stress Testing Under Progressively Hybrid Censoring, David Kojo Amakye

Open Access Theses & Dissertations

Reliability analysis is essential for understanding how products perform over time, particularly in environments where failure data is limited or costly to obtain. One effective approach is multi-level stress testing, where test units are subjected to varying levels of stress to accelerate failures and extract more information within constrained timeframes. This study presents a novel optimization-based framework for designing life-testing experiments under progressively Type-II hybrid censoring, assuming Weibull lifetime distributions. Leveraging a Variable Neighborhood Search (VNS) algorithm, we determine efficient allocations of test units and censoring parameters across multiple stress levels to enhance the precision of estimates of the model …


Optimal Accelerated Life-Testing Plans Under Progressive Type Ii First Failure Censoring Scheme, Emmanuella Duah May 2025

Optimal Accelerated Life-Testing Plans Under Progressive Type Ii First Failure Censoring Scheme, Emmanuella Duah

Open Access Theses & Dissertations

Designing optimal accelerated life testing (ALT) plans under progressive Type-II first failure censoring involves complex computational challenges, especially when trying to balance efficiency, precision, and practicality. This research introduces a novel optimization framework aimed at determining the best test configuration that minimizes estimation uncertainty within constrained experimental conditions. We developed a tailored meta-heuristic strategy based on an enhanced Variable Neighborhood Search (VNS) algorithm, which efficiently navigates the complex landscape of censoring schemes and stress allocations. Unlike conventional methods that typically focus on simpler censoring or stress-level structures, this approach simultaneously optimizes both the censoring points and sample allocations across various …


Computing Optimal Progressive Hybrid Censoring Schmes Using An Mcmc Type Probabilistic Approach, Irene Yemotiorkor Odoi May 2025

Computing Optimal Progressive Hybrid Censoring Schmes Using An Mcmc Type Probabilistic Approach, Irene Yemotiorkor Odoi

Open Access Theses & Dissertations

Progressive hybrid censoring schemes play a crucial role in optimizing life-testing experiments by balancing test duration and statistical efficiency. This study presents a Markov Chain Monte Carlo (MCMC)-based probabilistic approach for determining optimal progressive hybrid censoring schemes, incorporating a time-dependent component that enhances traditional progressive censoring methods.We implement our approach using three distinct probability distributions the multinomial, hypergeometric, and uniform distributions to simulate censoring schemes. The optimal censoring schemes are then identified based on three optimality criteria: A-optimality, D-optimality, and T-optimality, ensuring robust selection by minimizing estimator variance, maximizing Fisher information, and optimizing test duration, respectively..To evaluate the effectiveness of …


Computing Optimal Multi-Level Stress Testing Plans Using A Combined Variable Neighborhood Search Algorithm Under Progressive Type-Ii Censoring Scheme, Michael Obuobi May 2025

Computing Optimal Multi-Level Stress Testing Plans Using A Combined Variable Neighborhood Search Algorithm Under Progressive Type-Ii Censoring Scheme, Michael Obuobi

Open Access Theses & Dissertations

In multi-level stress life tests under Type-II progressive censoring, determining optimal allocation poses significant computational challenges due to the vast solution space. Efficient methods are essential for exploring the admissible censoring schemes effectively. This thesis introduces a novel meta-heuristic algorithm, the Combined Variable Neighborhood Search (CVNS), which computes optimal schemes at different stress levels simultaneously. Unlike methods focusing on marginal stress levels or one-step progressive censoring, this approach leverages a unified framework to ensure enhanced computational efficiency and solution quality. By integrating the components of the design parameters into a cohesive optimization process, the algorithm effectively reduces computational time while …


A Study Of End-Cut Preference In Tree-Based Modeling, Xiangya Wang May 2025

A Study Of End-Cut Preference In Tree-Based Modeling, Xiangya Wang

Open Access Theses & Dissertations

Decision trees, particularly those built using the Classification and Regression Trees (CART) algorithm, are widely used for their interpretability and flexibility. However, the greedy nature of the CART splitting procedure gives rise to the end-cut preference (ECP) phenomenon, wherein split points near the extremes of predictor ranges are favored. This study offers a comprehensive investigation of ECP, exploring its theoretical underpinnings, practical manifestations, and implications for both single decision trees and ensemble methods such as Random Forests. Through theoretical analysis and simulation studies, we examine how ECP affects tree structure, variable selection, and predictive accuracy across tree-structured, linear, and nonlinear …


D-Optimal Joint Best Linear Unbiased Predictors In Progressively Type-Ii Ordered Statistics, Tamim Alam Dec 2024

D-Optimal Joint Best Linear Unbiased Predictors In Progressively Type-Ii Ordered Statistics, Tamim Alam

Open Access Theses & Dissertations

Reliability and life-testing experiments play a crucial role in understanding the longevity and performance of systems and components, particularly in high-stakes applications such as engineering, manufacturing, and quality control. In this thesis, we focus on the prediction of future unobserved failure times by employing joint predictors based on progressively Type-II censored data obtained from such life-testing experiments. Specifically, we derive explicit analytical expressions for the joint best linear unbiased predictors (BLUPs) of two future order statistics under the D-optimality criterion. The derivation involves minimizing the determinant of the variance-covariance matrix of the predictors within the context of progressively Type-II censored …


Enhancing Predictive Accuracy In Noisy Data: A Robust Gradient Boosting Approach, Gabriela Alexa Acuna Dec 2024

Enhancing Predictive Accuracy In Noisy Data: A Robust Gradient Boosting Approach, Gabriela Alexa Acuna

Open Access Theses & Dissertations

This thesis aims to enhance the gradient boosting technique, a well-known machine learning method, in a regression setup. Gradient boosting techniques implement sequential weak learners, in this case, shallow trees that contribute to the model with a small percentage. The traditional gradient boosting approach uses regression trees that choose thresholds based on minimizing variance and makes predictions based on the mean of the target variable for the observations contained within each node. The residual sum of squares (RSS) and mean metrics are sensitive to outliers, which makes the model's predictions less robust. Outliers influence the predictions, resulting in a model …


Robust Multivariate Estimation And Inference With The Minimum Density Power Divergence Estimator, Ebenezer Nkum Aug 2024

Robust Multivariate Estimation And Inference With The Minimum Density Power Divergence Estimator, Ebenezer Nkum

Open Access Theses & Dissertations

The estimation of the location vector and scatter matrix plays a crucial role in many multivariate statistical methods. However, the classical likelihood-based estimation is greatly influenced by outliers, potentially leading to unreliable decisions. Hence, a fundamental challenge in multivariate statistics is to develop robust alternatives that can maintain performancein the presence of outliers and deviations from the assumed data distribution. Unfortunately, methods with good global robustness often substantially sacrifice efficiency. To address this, we propose the adoption of Minimum Density Power Divergence (MDPD) estimation, a well-established robust technique known for its efficiency and statistical robustness to outliers and model violations. …


Random Forest For High-Dimensional Data, George Ekow Quaye Aug 2024

Random Forest For High-Dimensional Data, George Ekow Quaye

Open Access Theses & Dissertations

The exponential growth of data has led to a rapid increase in high-dimensional datasets across various domains, presenting significant challenges in data analysis, particularly in predictive modeling tasks. Traditional Random Forest (RF), while robust, often struggles with datasets filled with numerous noisy or non-informative features, compromising both performance and accuracy. This study introduces an advanced algorithm, High-Dimensional Random Forests (HDRF), designed to address these challenges by integrating robust multivariate feature selection techniques directly into the decision tree construction process. Unlike standard RF, HDRF incorporates ridge regression-based variable screening at each decision split, enhancing its ability to identify and utilize the …


Metrics For Comparison Of Complex Networks, Clarissa Reyes Dec 2023

Metrics For Comparison Of Complex Networks, Clarissa Reyes

Open Access Theses & Dissertations

Heuristic network statistics are used as a preliminary approach to identify change across networks. In networks where there is known node correspondence (KNC), conventional network comparison methods include taking a norm of the difference matrix, or calculating dissimilarity measures like DeltaCon and cut distance. Since different KNC measures provide varying insight to the network comparison problem, we propose employing Rank Score Characteristic Functions (RSCFs) and the rank-score process as a method for reaching a consensus when ranking quantified change across multiple pairs of networks â?? which is particularly useful for ranking change across subpopulations or subgraphs. Additionally, we propose a …


Integrating Machine Learning Methods For Medical Diagnosis, Jazmin Quezada Dec 2023

Integrating Machine Learning Methods For Medical Diagnosis, Jazmin Quezada

Open Access Theses & Dissertations

Abstract:The rapid advancement of machine learning techniques has revolutionized the field of medical diagnosis by offering powerful tools to analyze complex data sets and make accurate predictions. In this proposed method, we present a novel approach that integrates machine learning and optimization models to enhance the accuracy of medical diagnoses. Our method focuses on fine-tuning and optimizing the parameters of machine learning algorithms commonly used in medical diagnosis, such as logistic regression, support vector machines, and neural networks. By employing optimization techniques, we systematically explore the parameter space of these algorithms to discover the most optimal configurations. Moreover, by representing …


Single-Index Multinomial Model For Analyzing Crime Data, Kwabena Gyamfi Duodu Aug 2023

Single-Index Multinomial Model For Analyzing Crime Data, Kwabena Gyamfi Duodu

Open Access Theses & Dissertations

We develop a flexible single-index multinomial model for analyzing crime data. In additionto the number of crimes reported, the data also includes covariates such as location, time of day, weather, and other demographic factors. We provide an estimation algorithm and develop R code for the single-index multinomial model. Using simulations, we evaluate the performance of the proposed estimation algorithm. When applied to crime data, the single-index multinomial model provides important insights into crime trends and risk variables, assisting in the development of tailored crime prevention programs. Policymakers and law enforcement organizations can use the model's projections to more efficiently allocate …


Robust Penalized Density Power Divergence Regression With Scad Penalty For High Dimensional Data Analysis, Maxwell Kwesi Mac-Ocloo Aug 2023

Robust Penalized Density Power Divergence Regression With Scad Penalty For High Dimensional Data Analysis, Maxwell Kwesi Mac-Ocloo

Open Access Theses & Dissertations

Amidst the exponential surge in big data, managing high-dimensional datasets across diverse fields and industries has emerged as a significant challenge. Conventional statistical methods struggle to handle their complexity, making analysis intricate. In response, we've formulated a robust estimator tailored to counter outliers and heavy-tailed errors. Our approach integrates the SCAD penalty into the Density Power Divergence method, effectively reducing insignificant coefficients to zero. This enhances analysis precision and result reliability.We benchmark our robust and penalized model against existing techniques like Huber, Tukey, LASSO, LAD, and LAD-LASSO. Employing both simulated and UCI machine learning repository datasets, we assess method performance …


Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu Aug 2023

Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu

Open Access Theses & Dissertations

The goal of classification is to develop a model that can be used to accurately assign new observations to labeled classes based on the patterns learned from the training data. K-nearest Neighbors algorithm (KNN) is a popular and widely used algorithm for classification, however, its performance can be adversely affected by the presence of outliers in a dataset. In this study we have modified this existing KNN algorithm that can alleviate the effect of outliers in a dataset, thereby improving the performance of the KNN algorithm. We compared the performances of the Modified KNN method and the Existing KNN algorithm …


Robust Mahalanobis K-Means Algorithm In Comparison With Other Existing Clustering Methods., Eleazer Tabi Serebour Aug 2023

Robust Mahalanobis K-Means Algorithm In Comparison With Other Existing Clustering Methods., Eleazer Tabi Serebour

Open Access Theses & Dissertations

This study enhances K-means Mahalanobis clustering using Density Power Divergence (DPD) for outlier handling and detection. Through the utilization of simulations and the analysis of real-world data, our approach consistently outperforms standard K-means, Mahalanobis K-means, Fuzzy C-means, and others in clustering datasets with outliers. While our method performs similarly to others on spherical datasets, it ranks second to DBSCAN for arbitrary shapes. We showcase its superiority on real-life datasets (Iris flower and wheat seed), demonstrating resilient outlier identification. By navigating various structures and cluster characteristics, our Modified Mahalanobis K-means method proves adaptable and robust, offering insights into diverse clustering scenarios. …


Theoretical And Computational Aspects Of Robust Cluster Analysis For Multivariate And High-Dimensional Datasets, Andrews Tawiah Anum May 2023

Theoretical And Computational Aspects Of Robust Cluster Analysis For Multivariate And High-Dimensional Datasets, Andrews Tawiah Anum

Open Access Theses & Dissertations

Multivariate and high-dimensional datasets typically contain subgroups that may not be immediately apparent. To reveal these groups, cluster analysis is performed. Cluster analysis is an unsupervised machine learning technique commonly employed to partition a dataset into distinct categories referred to as clusters. The k-means algorithm is a prominent distance-based clustering method. Despite overwhelming popularity, the algorithm is not invariant under non-singular affine transformations and is not robust, i.e., can be unduly influenced by outliers. To address these deficiencies, we propose an alternative model-based clustering procedure by minimizing a “trimmed” variant of the negative log-likelihood function. We develop a “concentration step”, …


Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data, Jagannath Das May 2023

Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data, Jagannath Das

Open Access Theses & Dissertations

Due to advanced technology and wide source of data collection, high-dimensional data is available in several fields, including healthcare, bioinformatics, medicine, epidemiology, economics, finance, sociology, and climatology. In those datasets, outliers are generally encountered due to technical errors, heterogeneous sources, or the effect of some confounding variables. As outliers are often difficult to detect in high-dimensional data, the standard approaches may fail to model such data and produce misleading information. In this thesis, we studied Huber and Tukey's M-estimators for linear regression that automatically down-weight outliers and provide a good fit. We also investigated two variable selection methods -- LASSO …


Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions, Tahiru Mahama May 2023

Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions, Tahiru Mahama

Open Access Theses & Dissertations

Marginal Integration (MI) is a statistical method that is extensively employed to estimatecomponent functions of the nonparametric additive models. The shortcoming of the purely additive model is that interaction between predictor variables is often ignored, and it may produce poor performance in some real applications. As a result, this research considers the second-order interactions in the regression models. The primary objective is to use marginal integration techniques to estimate the nonparametric additive functions. We compare this model with other models/estimators such as the Generalized Additive Model (GAM), Generalized Additive Model with Selection (GAMSEL), Robust Marginal Integration (RMI), Ordinary Least Squares …