Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection,
2023
University of Nebraska-Lincoln
Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild
Department of Statistics: Dissertations, Theses, and Student Research
The words people choose to use hold a lot of power, whether that be in spreading truth or deception. As listeners and readers, we do our best to understand how words are being used. There are many current methods in computer science literature attempting to embed words into numerical information for statistical analyses. Some of these embedding methods, such as Bag of Words, treat words as independent, while others, such as Word2Vec, attempt to gain information about the context of words. It is of interest to compare how well these various methods of translating text into numerical data work specifically …
Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data,
2023
University of Texas at El Paso
Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data, Jagannath Das
Open Access Theses & Dissertations
Due to advanced technology and wide source of data collection, high-dimensional data is available in several fields, including healthcare, bioinformatics, medicine, epidemiology, economics, finance, sociology, and climatology. In those datasets, outliers are generally encountered due to technical errors, heterogeneous sources, or the effect of some confounding variables. As outliers are often difficult to detect in high-dimensional data, the standard approaches may fail to model such data and produce misleading information. In this thesis, we studied Huber and Tukey's M-estimators for linear regression that automatically down-weight outliers and provide a good fit. We also investigated two variable selection methods -- LASSO …
Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions,
2023
University of Texas at El Paso
Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions, Tahiru Mahama
Open Access Theses & Dissertations
Marginal Integration (MI) is a statistical method that is extensively employed to estimatecomponent functions of the nonparametric additive models. The shortcoming of the purely additive model is that interaction between predictor variables is often ignored, and it may produce poor performance in some real applications. As a result, this research considers the second-order interactions in the regression models. The primary objective is to use marginal integration techniques to estimate the nonparametric additive functions. We compare this model with other models/estimators such as the Generalized Additive Model (GAM), Generalized Additive Model with Selection (GAMSEL), Robust Marginal Integration (RMI), Ordinary Least Squares …
Outlier Detection In Multivariate And High-Dimensional Datasets,
2023
University of Texas at El Paso
Outlier Detection In Multivariate And High-Dimensional Datasets, Yuanhong Wu
Open Access Theses & Dissertations
Accurate detection of outliers is crucial in the field of statistical analysis. Using classical statisticalmodels without considering the presence of outliers in the data can lead to misleading outcomes. There exist a myriad of procedures to detect outliers in statistics. We concentrate on the statistical techniques that can robustly identify outliers in data sets. To this end, we pursue two aims. First, we give an extensive overview of robust statistical methods which are still popular in recent years for outlier detection. We provide the definitions, algorithms and also discuss some important properties of these methods. Second, two real examples are …
Spatially Adaptive Estimation Of Spectrum,
2023
University of Texas at El Paso
Spatially Adaptive Estimation Of Spectrum, Yi Xie
Open Access Theses & Dissertations
A time series may be analyzed either in the time or in the frequency domain. When working in the frequency domain, the main objective is to estimate the underlying spectrum. Various approaches have been proposed to this end, but most are based on smoothing the periodogram using a single smoothing parameter across all Fourier frequencies. Such a global smoothing parameter may result in a biased estimate. To improve the estimation, in this paper, we smooth the log periodogram by placing a dynamic shrinkage prior, such that varying degrees of smoothing may be applied to different regions of the Fourier frequencies, …
Flexible Models For The Estimation Of Treatment Effect,
2023
University of Texas at El Paso
Flexible Models For The Estimation Of Treatment Effect, Habeeb Abolaji Bashir
Open Access Theses & Dissertations
Estimation of treatment effect is an important problem which is well studied in the literature. While the regression models are one of the most commonly used techniques for the estimation of treatment effect, they are prone to model misspecification. To minimize the model misspecification bias, flexible nonparametric models are introduced for the estimation. Continuing this line of research, we propose two flexible nonparametric models that allow the treatment effect to vary across different levels of covariates. We provide estimation algorithms for both these models. Using simulations and data analysis, we illustrate the usefulness of the proposed methods.
Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data.,
2023
University of Texas at El Paso
Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data., Peter Kwadwo Asante
Open Access Theses & Dissertations
This dissertation aims to assess the performance of Ornstein-Uhlenbeck-type models by examining the fractal characteristics of time series data from various sources, including finance, volcanic and earthquake events, US COVID-19 reported cases and deaths, and two simulated time series with differing properties. The time series data is categorized as either a Gaussian or a Lévy process (Lévy walk or Lévy flight) by using three scaling methods: Rescaled range analysis, Detrended fluctuation analysis, and Diffusion entropy analysis. The outcomes of this analysis indicate that the financial indices are classified as Lévy walks, while the volcanic, earthquake, and COVID-19 data are classified …
Nonparametric Estimation Of Elliptical Copulas,
2023
University of Texas at El Paso
Nonparametric Estimation Of Elliptical Copulas, Panfeng Liang
Open Access Theses & Dissertations
Elliptical copulas provide flexibility in modeling the dependence structure of a random vector. They are often parameterized with a correlation matrix and a scalar function, called generator. The estimation of the generator can be challenging, because it is a functional parameter. In this dissertation, we provide a rigorous approach to estimating the generator in a Bayesian framework, which is simpler, more robust, and outperforms existing estimation methods in the literature. Based on the proposed framework in this dissertation, other researchers may modify the model for other types of generators in their own research.
Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision,
2023
University of Texas at El Paso
Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision, Mayra Eydie Pacheco
Open Access Theses & Dissertations
No abstract provided.
Hispanic Human Capital And Financial Aid Application In The West Census Region,
2023
California State University, Monterey Bay
Hispanic Human Capital And Financial Aid Application In The West Census Region, Benjamin Lundy-Paine
Capstone Projects and Master's Theses
As of 2021, very few Hispanic residents in the United States held a college degree in comparison to non-Hispanic residents. Research has shown that, particularly for Hispanic students, financial aid increases college persistence. Hispanic Free Application for Federal Student Aid (FAFSA) submission rates rank among the lowest, preventing many Hispanic students from receiving financial assistance. This issue is most prevalent West Census Region (WCR), where there is the highest concentration of Hispanic residents. To understand what barriers may be preventing Hispanic submission in the WCR this Capstone used logistic regression models to analyze student-level data from the National Center for …
A Brascamp-Lieb–Rary Of Examples,
2023
Macalester College
A Brascamp-Lieb–Rary Of Examples, Anina Peersen
Mathematics, Statistics, and Computer Science Honors Projects
This paper focuses on the Brascamp-Lieb inequality and its applications in analysis, fractal geometry, computer science, and more. It provides a beginner-level introduction to the Brascamp-Lieb inequality alongside re- lated inequalities in analysis and explores specific cases of extremizable, simple, and equivalent Brascamp-Lieb data. Connections to computer sci- ence and geometric measure theory are introduced and explained. Finally, the Brascamp-Lieb constant is calculated for a chosen family of linear maps.
Mixing Measures For Trees Of Fixed Diameter,
2023
Macalester College
Mixing Measures For Trees Of Fixed Diameter, Ari Holcombe Pomerance
Mathematics, Statistics, and Computer Science Honors Projects
A mixing measure is the expected length of a random walk in a graph given a set of starting and stopping conditions. We determine the tree structures of order n with diameter d that minimize and maximize for a few mixing measures. We show that the maximizing tree is usually a broom graph or a double broom graph and that the minimizing tree is usually a seesaw graph or a double seesaw graph.
Gentrification And Crime In The Twin Cities: Insights And Challenges Through A Statistical Lens,
2023
Macalester College
Gentrification And Crime In The Twin Cities: Insights And Challenges Through A Statistical Lens, Erin G. Franke
Mathematics, Statistics, and Computer Science Honors Projects
Gentrification is a complex process of urban redevelopment that typically involves an in-migration of educated people to neighborhoods experiencing a period of disinvestment. While gentrification is widely regarded for its potential to displace long-time businesses and residents of the neighborhood, its impact on crime is highly controversial. There is not a consensus on the relationship between gentrification and crime across criminological theory and past statistical studies have also shown contradictory results. Measuring gentrification on the tract level with census data, we seek to understand gentrification’s relationship with violent crime and theft in the Twin Cities. Using a Poisson model with …
Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers,
2023
East Tennessee State University
Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham
Electronic Theses and Dissertations
The abundance, accessibility, and scale of data have engendered an era where machine learning can quickly and accurately solve complex problems, identify complicated patterns, and uncover intricate trends. One research area where many have applied these techniques is the stock market. Yet, financial domains are influenced by many factors and are notoriously difficult to predict due to their volatile and multivariate behavior. However, the literature indicates that public sentiment data may exhibit significant predictive qualities and improve a model’s ability to predict intricate trends. In this study, momentum SVM classification accuracy was compared between datasets that did and did not …
A Machine Learning Approach To Obese-Inflammatory Phenotyping,
2023
The University of Texas Rio Grande Valley
A Machine Learning Approach To Obese-Inflammatory Phenotyping, Tania Mayleth Vargas
Theses and Dissertations
Obesity is the accumulation of an abnormal, or excessive, amount of fat in the body, which can have negative effects on overall health. This excess accumulation of macronutrients in adipose tissue can cause the release of inflammatory mediators, leading to a proinflammatory state. Inflammation is a known risk factor for various health conditions, including cardiovascular diseases, metabolic syndrome, and diabetes. This study sought to examine the use of data mining methods, particularly clustering algorithms, to identify inflammatory biomarker phenotypes and their association with obesity in a local adolescent population. The algorithms evaluated in this study included: k-means, Ward's hierarchical …
Small But Mighty: Examing The Utility Of Microstatistics In Modeling Ice Hockey,
2023
Liberty University
Small But Mighty: Examing The Utility Of Microstatistics In Modeling Ice Hockey, Matt Palmer
Senior Honors Theses
As research into hockey analytics continues, an increasing number of metrics are being introduced into the knowledge base of the field, creating a need to determine whether various stats are useful or simply add noise to the discussion. This paper examines microstatistics – manually tracked metrics which go beyond the NHL’s publicly released stats – both through the lens of meta-analytics (which attempt to objectively assess how useful a metric is) and modeling game probabilities. Results show that while there is certainly room for improvement in understanding and use of microstats in modeling, the metrics overall represent an area of …
Bayesian Semi-Mechanistic Dose-Finding Designs For Phase I Oncology Trials,
2023
The University of Texas MD Anderson Cancer Center
Bayesian Semi-Mechanistic Dose-Finding Designs For Phase I Oncology Trials, Chao Yang
Dissertations and Theses (Open Access)
Bayesian adaptive designs are getting more popular in research and in practice because they are flexible and efficient in evaluating an experimental drug. In oncology, despite the great advances in novel dose-finding designs, the high failure rates of clinical cancer drug development from phase I to III trials call for further improvements on novel designs, in addition to the need to promote and adopt novel designs in practice. Because anticancer agents often have a narrow therapeutic index, an accurate identification of the maximum tolerated dose (MTD) in a phase I trial is crucial for identifying a tolerable and efficacious dose …
Inference For Multiple Utility In Time-Dependent Choice Pairs Under Copula-Based Models,
2023
Old Dominion University
Inference For Multiple Utility In Time-Dependent Choice Pairs Under Copula-Based Models, Sasanka Adikari
Mathematics & Statistics Theses & Dissertations
Models for discrete choice experiments (DCE) are frequently used to analyze consumer choices about products and services. A family of DCE, best-worst scaling experiments, offers more in-depth insights into consumer preferences by eliciting a best and worst choice from a set of options, rather than just a single preference. Traditional approaches often assume that choices are mutually exclusive over time, which may not always be the case. This dissertation proposes a novel model for DCE that takes into account the changing nature of consumer choices over time and the priority constraint of transition probabilities. The model introduces a copula combination …
Jackknife Empirical Likelihood Tests For Equality Of Generalized Lorenz Curves,
2023
California State University, San Bernardino
Jackknife Empirical Likelihood Tests For Equality Of Generalized Lorenz Curves, Anton Butenko
Electronic Theses, Projects, and Dissertations
A Lorenz curve is a graphical representation of the distribution of income or wealth within a population. The generalized Lorenz curve can be created by scaling the values on the vertical axis of a Lorenz curve by the average output of the distribution. In this thesis, we propose two nonparametric methods for testing the equality of two generalized Lorenz curves. Both methods are based on empirical likelihood and utilize a U -statistic. We derive the limiting distribution of the likelihood ratio, which is shown to follow a chi-squared distribution with one degree of freedom. We conduct simulations to compare the …
Distance Correlation Based Feature Selection In Random Forest,
2023
California State University - San Bernardino
Distance Correlation Based Feature Selection In Random Forest, Jose Munoz-Lopez
Electronic Theses, Projects, and Dissertations
The Pearson correlation coefficient is a commonly used measure of correlation, but it has limitations as it only measures the linear relationship between two numerical variables. In 2007, Szekely et al. introduced the distance correlation, which measures all types of dependencies between random vectors X and Y in arbitrary dimensions, not just the linear ones. In this thesis, we propose a filter method that utilizes distance correlation as a criterion for feature selection in Random Forest regression. We conduct extensive simulation studies to evaluate its performance compared to existing methods under various data settings, in terms of the prediction mean …
