Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2221 - 2250 of 12804

Full-Text Articles in Statistics and Probability

Identifying And Analyzing Multi-Star Systems Among Tess Planetary Candidates Using Gaia, Katie E. Bailey May 2023

Identifying And Analyzing Multi-Star Systems Among Tess Planetary Candidates Using Gaia, Katie E. Bailey

Electronic Theses and Dissertations

Exoplanets represent a young, rapidly advancing subfield of astrophysics where much is still unknown. It is therefore important to analyze trends among their parameters to learn more about these systems. More complexity is added to these systems with the presence of additional stellar companions. To study these complex systems, one can employ programming languages such as Python to parse databases such as those constructed by TESS and Gaia to bridge the gap between exoplanets and stellar companions. Data can then be analyzed for trends in these multi-star exoplanet systems and in juxtaposition to their single-star counterparts. This research was able …


Examining Political Discourse On Online 8kun And Reddit Forums, Braden Mindrum May 2023

Examining Political Discourse On Online 8kun And Reddit Forums, Braden Mindrum

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A recent example of political violence in the United States was that of the January 6, 2021, Capitol attack in connection with the certification of Joseph R. Biden’s victory over Donald J. Trump in the 2020 US presidential election. This thesis analyzes the events of January 6, 2021, through the lens of social media discourse. This thesis presents a workflow that acquired over 5 million 8kun and Reddit posts from various apolitical and political forums in the three months preceding and following the Capitol attack on January 6, 2021. Techniques from text analysis are then used to group forums according …


Investigating The Effect Of Greediness On The Coordinate Exchange Algorithm For Generating Optimal Experimental Designs, William Thomas Gullion May 2023

Investigating The Effect Of Greediness On The Coordinate Exchange Algorithm For Generating Optimal Experimental Designs, William Thomas Gullion

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Design of Experiments (DoE) is the field of statistics concerned with helping researchers maximize the amount of information they gain from their experiments. Recently, researchers have been turning to optimal experimental designs instead of classical/catalog experimental designs. One of the most popular algorithms used today to generate optimal designs is the Coordinate Exchange (CEXCH) Algorithm. CEXCH is known to be a greedy algorithm, which means it tends to favor immediate, locally best designs instead of globally optimal designs. Previous research demonstrated that this tradeoff was efficacious in that it reduced the cost of a single run of CEXCH and allowed …


Subgroup Identification Via Interaction Tree And Mixed Model For Repeated Measures With Application To Alzheimer’S Disease, Zhichen Xu May 2023

Subgroup Identification Via Interaction Tree And Mixed Model For Repeated Measures With Application To Alzheimer’S Disease, Zhichen Xu

Arts & Sciences Graduate Student Theses and Dissertations

No abstract provided.


Theoretical And Computational Aspects Of Robust Cluster Analysis For Multivariate And High-Dimensional Datasets, Andrews Tawiah Anum May 2023

Theoretical And Computational Aspects Of Robust Cluster Analysis For Multivariate And High-Dimensional Datasets, Andrews Tawiah Anum

Open Access Theses & Dissertations

Multivariate and high-dimensional datasets typically contain subgroups that may not be immediately apparent. To reveal these groups, cluster analysis is performed. Cluster analysis is an unsupervised machine learning technique commonly employed to partition a dataset into distinct categories referred to as clusters. The k-means algorithm is a prominent distance-based clustering method. Despite overwhelming popularity, the algorithm is not invariant under non-singular affine transformations and is not robust, i.e., can be unduly influenced by outliers. To address these deficiencies, we propose an alternative model-based clustering procedure by minimizing a “trimmed” variant of the negative log-likelihood function. We develop a “concentration step”, …


Dynamics Of Inertial And Non-Inertial Particles In Geophysical Flows, Nishanta Baral May 2023

Dynamics Of Inertial And Non-Inertial Particles In Geophysical Flows, Nishanta Baral

Theses, Dissertations and Culminating Projects

We consider the dynamics of inertial and non-inertial particles in various flows. We investigate the underlying structures of the flow field by examining their Lagrangian coherent structures (LCS), which are found by computing finitetime Lyapunov exponents (FTLE). We compare the behavior of massless noninertial particles using the velocity fields from four models, the Duffing oscillator, the Bickley jet, the double-gyre flow, and a quasi-geostrophic geophysical flow model, with that of inertial particles. For inertial particles with finite size and mass, we use the Maxey-Riley equation to describe the particle’s motion. We explore the preferential aggregation of inertial particles and demonstrate …


Parameter Optimization For Excitable Cell Models, Amrit Parmar May 2023

Parameter Optimization For Excitable Cell Models, Amrit Parmar

Theses, Dissertations and Culminating Projects

The electrophysiology of nodose ganglia neurons is of great interest in the analysis of cell membrane currents and action potential behavior. This behavior was initially outlined in the Hodgkin-Huxley conductance model [1] using a system of nonlinear differential equations. Later, Schild et al. [2] developed an extension of the Hodgkin-Huxley model to provide a more exhaustive description of ion channels involved in nodose neuronal action potential activity. We consider a variety of methods to fit the parameters of both the Hodgkin-Huxley and Schild et al. models to an empirical stimulus response dataset. Our methods were validated using synthetic datasets, as …


Effects Of Functional Network Model Definition On Biomarker Outcome Prediction, Xinyang Feng May 2023

Effects Of Functional Network Model Definition On Biomarker Outcome Prediction, Xinyang Feng

Arts & Sciences Graduate Student Theses and Dissertations

Machine learning (ML) models are widely used to investigate the human connectome and to predict and understand behavior, emotion, and cognition. Prior research has organized pediatric connectome data using adult functional network models. However, this assumes that adult functional network models are appropriate and useful for prediction developmental outcomes from pediatric connectome data. We hypothesize that the application of adult brain network models could result in poor model fit, limiting the generalizability of results. Here, we test whether prediction of biological age is improved by concordant brain network models matching underlying functional connectome data. To quantify the difference in age …


Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski May 2023

Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski

Honors Scholar Theses

Challenging conventional wisdom is at the very core of baseball analytics. Using data and statistical analysis, the sets of rules by which coaches make decisions can be justified, or possibly refuted. One of those sets of rules relates to the construction of a batting order. Through data collection, data adjustment, the construction of a baseball simulator, and the use of a Monte Carlo Simulation, I have assessed thousands of possible batting orders to determine the roster-specific strategies that lead to optimal run production for the 2023 UConn baseball team. This paper details a repeatable process in which basic player statistics …


Characterization Of Public Opinion On Severity Of Mental Illness And Hiv Based On Individual Traits Using Hierarchical Multi-Category Probit Models, Md Moinul Ahsan May 2023

Characterization Of Public Opinion On Severity Of Mental Illness And Hiv Based On Individual Traits Using Hierarchical Multi-Category Probit Models, Md Moinul Ahsan

Graduate Theses and Dissertations

In this thesis, we focus on modeling categorical response variables from public opinion datasets. A hierarchical probit model was used to analyze these different variables. Particularly for multinomial data, we tried different covariate settings to see the model’s performance. For that purpose, we tried two different estimation techniques. The first algorithm uses identified parameters by fixing the first diagonal element of the covariance matrix at 1. The second algorithm uses one unidentifiable parameter and subsequently identifies the parameters by fixing the trace of the covariance matrix. The results from the simulation study confirm that the trace-restricted algorithm performs better with …


Baseball’S Evolution In The 21st Century, And How It Exemplifies Human Response To Change, Jonathan Sharpe May 2023

Baseball’S Evolution In The 21st Century, And How It Exemplifies Human Response To Change, Jonathan Sharpe

Honors Projects

The game of baseball has changed a lot in the past twenty years. It can be primarily attributed to the explosion in data analytics and how they are used to evaluate baseball players. This led to different player profiles being preferred and eventually led to the development of players changing. As a result, the strategies employed have also evolved and turned into a different game than seen only a couple of decades ago. This paper will explore the changes that the game has seen. On the other hand, Major League Baseball has also implemented its own changes to try and …


Effects Of Land Use On Soil Microbial Communities In Tropical Montane Forests Of Malaysian Borneo, Yang Kai Tang May 2023

Effects Of Land Use On Soil Microbial Communities In Tropical Montane Forests Of Malaysian Borneo, Yang Kai Tang

Graduate Theses and Dissertations

Land use, such as logging and forest conversion to agriculture, can modify soil physicochemical and biological properties, and affect soil health. To understand how land use change can impact soil properties and canopy structure, we used a land use gradient in Malaysian Borneo consisting of six sites, including old growth forests, mixed forests, and agriculture fields. Specifically, we aimed to answer the following questions: (1) How do soil physicochemical properties vary across land use types? (2) Does bacterial diversity and composition vary across different land use types? (3) Does fungal diversity and composition vary across different land use types? We …


Machine Learning-Based Data And Model Driven Bayesian Uncertanity Quantification Of Inverse Problems For Suspended Non-Structural System, Zhiyuan Qin May 2023

Machine Learning-Based Data And Model Driven Bayesian Uncertanity Quantification Of Inverse Problems For Suspended Non-Structural System, Zhiyuan Qin

All Dissertations

Inverse problems involve extracting the internal structure of a physical system from noisy measurement data. In many fields, the Bayesian inference is used to address the ill-conditioned nature of the inverse problem by incorporating prior information through an initial distribution. In the nonparametric Bayesian framework, surrogate models such as Gaussian Processes or Deep Neural Networks are used as flexible and effective probabilistic modeling tools to overcome the high-dimensional curse and reduce computational costs. In practical systems and computer models, uncertainties can be addressed through parameter calibration, sensitivity analysis, and uncertainty quantification, leading to improved reliability and robustness of decision and …


Do Firms Respond To Peer Disclosures? Evidence From Disclosures Of Clinical Trial Results, Vedran Capkun, Yun Lou, Clemens A. Otto, Yin Wang May 2023

Do Firms Respond To Peer Disclosures? Evidence From Disclosures Of Clinical Trial Results, Vedran Capkun, Yun Lou, Clemens A. Otto, Yin Wang

Research Collection School Of Accountancy

Using data on the registration of clinical trials and the disclosure of trial results, we examine how firms respond to peer disclosures. We find that firms are less likely to disclose their own trial results if the results of a larger number of closely related trials are disclosed by their peers. This relation is stronger if the firms face higher competition (as measured by the number of competing trials). It is weaker if the firms are further along in their research than the peers (as measured by the trials’ phase) and if the peers’ disclosures convey more negative news (as …


The Last Drought Frontier: Building A Drought Index For The State Of Alaska, Olivia Campbell May 2023

The Last Drought Frontier: Building A Drought Index For The State Of Alaska, Olivia Campbell

School of Natural Resources: Dissertations, Theses, and Student Research

Drought is characterized by periods of below average precipitation. There are five major types of drought recognized in the literature: meteorological, hydrological, agricultural, socioeconomic, and ecological. A relatively new concept in the drought literature is “snow drought.” A key part of the definition of drought is that it is not always accompanied by extreme heat. This means drought can occur even in cold climates, cold seasons, and higher latitudes and altitudes, like Alaska. Drought is a natural part of climate variability, but Alaska’s climate is changing faster than any other state in the United States. Alaska is no stranger to …


Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild May 2023

Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild

Department of Statistics: Dissertations, Theses, and Student Research

The words people choose to use hold a lot of power, whether that be in spreading truth or deception. As listeners and readers, we do our best to understand how words are being used. There are many current methods in computer science literature attempting to embed words into numerical information for statistical analyses. Some of these embedding methods, such as Bag of Words, treat words as independent, while others, such as Word2Vec, attempt to gain information about the context of words. It is of interest to compare how well these various methods of translating text into numerical data work specifically …


Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data, Jagannath Das May 2023

Comparison Of Different Robust Methods In Linear Regression And Applications In Cardiovascular Data, Jagannath Das

Open Access Theses & Dissertations

Due to advanced technology and wide source of data collection, high-dimensional data is available in several fields, including healthcare, bioinformatics, medicine, epidemiology, economics, finance, sociology, and climatology. In those datasets, outliers are generally encountered due to technical errors, heterogeneous sources, or the effect of some confounding variables. As outliers are often difficult to detect in high-dimensional data, the standard approaches may fail to model such data and produce misleading information. In this thesis, we studied Huber and Tukey's M-estimators for linear regression that automatically down-weight outliers and provide a good fit. We also investigated two variable selection methods -- LASSO …


Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions, Tahiru Mahama May 2023

Generalized Additive Model Using Marginal Integration Estimation Techniques With Interactions, Tahiru Mahama

Open Access Theses & Dissertations

Marginal Integration (MI) is a statistical method that is extensively employed to estimatecomponent functions of the nonparametric additive models. The shortcoming of the purely additive model is that interaction between predictor variables is often ignored, and it may produce poor performance in some real applications. As a result, this research considers the second-order interactions in the regression models. The primary objective is to use marginal integration techniques to estimate the nonparametric additive functions. We compare this model with other models/estimators such as the Generalized Additive Model (GAM), Generalized Additive Model with Selection (GAMSEL), Robust Marginal Integration (RMI), Ordinary Least Squares …


Outlier Detection In Multivariate And High-Dimensional Datasets, Yuanhong Wu May 2023

Outlier Detection In Multivariate And High-Dimensional Datasets, Yuanhong Wu

Open Access Theses & Dissertations

Accurate detection of outliers is crucial in the field of statistical analysis. Using classical statisticalmodels without considering the presence of outliers in the data can lead to misleading outcomes. There exist a myriad of procedures to detect outliers in statistics. We concentrate on the statistical techniques that can robustly identify outliers in data sets. To this end, we pursue two aims. First, we give an extensive overview of robust statistical methods which are still popular in recent years for outlier detection. We provide the definitions, algorithms and also discuss some important properties of these methods. Second, two real examples are …


Spatially Adaptive Estimation Of Spectrum, Yi Xie May 2023

Spatially Adaptive Estimation Of Spectrum, Yi Xie

Open Access Theses & Dissertations

A time series may be analyzed either in the time or in the frequency domain. When working in the frequency domain, the main objective is to estimate the underlying spectrum. Various approaches have been proposed to this end, but most are based on smoothing the periodogram using a single smoothing parameter across all Fourier frequencies. Such a global smoothing parameter may result in a biased estimate. To improve the estimation, in this paper, we smooth the log periodogram by placing a dynamic shrinkage prior, such that varying degrees of smoothing may be applied to different regions of the Fourier frequencies, …


Flexible Models For The Estimation Of Treatment Effect, Habeeb Abolaji Bashir May 2023

Flexible Models For The Estimation Of Treatment Effect, Habeeb Abolaji Bashir

Open Access Theses & Dissertations

Estimation of treatment effect is an important problem which is well studied in the literature. While the regression models are one of the most commonly used techniques for the estimation of treatment effect, they are prone to model misspecification. To minimize the model misspecification bias, flexible nonparametric models are introduced for the estimation. Continuing this line of research, we propose two flexible nonparametric models that allow the treatment effect to vary across different levels of covariates. We provide estimation algorithms for both these models. Using simulations and data analysis, we illustrate the usefulness of the proposed methods.


Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data., Peter Kwadwo Asante May 2023

Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data., Peter Kwadwo Asante

Open Access Theses & Dissertations

This dissertation aims to assess the performance of Ornstein-Uhlenbeck-type models by examining the fractal characteristics of time series data from various sources, including finance, volcanic and earthquake events, US COVID-19 reported cases and deaths, and two simulated time series with differing properties. The time series data is categorized as either a Gaussian or a Lévy process (Lévy walk or Lévy flight) by using three scaling methods: Rescaled range analysis, Detrended fluctuation analysis, and Diffusion entropy analysis. The outcomes of this analysis indicate that the financial indices are classified as Lévy walks, while the volcanic, earthquake, and COVID-19 data are classified …


Nonparametric Estimation Of Elliptical Copulas, Panfeng Liang May 2023

Nonparametric Estimation Of Elliptical Copulas, Panfeng Liang

Open Access Theses & Dissertations

Elliptical copulas provide flexibility in modeling the dependence structure of a random vector. They are often parameterized with a correlation matrix and a scalar function, called generator. The estimation of the generator can be challenging, because it is a functional parameter. In this dissertation, we provide a rigorous approach to estimating the generator in a Bayesian framework, which is simpler, more robust, and outperforms existing estimation methods in the literature. Based on the proposed framework in this dissertation, other researchers may modify the model for other types of generators in their own research.


Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision, Mayra Eydie Pacheco May 2023

Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision, Mayra Eydie Pacheco

Open Access Theses & Dissertations

No abstract provided.


Hispanic Human Capital And Financial Aid Application In The West Census Region, Benjamin Lundy-Paine May 2023

Hispanic Human Capital And Financial Aid Application In The West Census Region, Benjamin Lundy-Paine

Capstone Projects and Master's Theses

As of 2021, very few Hispanic residents in the United States held a college degree in comparison to non-Hispanic residents. Research has shown that, particularly for Hispanic students, financial aid increases college persistence. Hispanic Free Application for Federal Student Aid (FAFSA) submission rates rank among the lowest, preventing many Hispanic students from receiving financial assistance. This issue is most prevalent West Census Region (WCR), where there is the highest concentration of Hispanic residents. To understand what barriers may be preventing Hispanic submission in the WCR this Capstone used logistic regression models to analyze student-level data from the National Center for …


A Brascamp-Lieb–Rary Of Examples, Anina Peersen May 2023

A Brascamp-Lieb–Rary Of Examples, Anina Peersen

Mathematics, Statistics, and Computer Science Honors Projects

This paper focuses on the Brascamp-Lieb inequality and its applications in analysis, fractal geometry, computer science, and more. It provides a beginner-level introduction to the Brascamp-Lieb inequality alongside re- lated inequalities in analysis and explores specific cases of extremizable, simple, and equivalent Brascamp-Lieb data. Connections to computer sci- ence and geometric measure theory are introduced and explained. Finally, the Brascamp-Lieb constant is calculated for a chosen family of linear maps.


Mixing Measures For Trees Of Fixed Diameter, Ari Holcombe Pomerance May 2023

Mixing Measures For Trees Of Fixed Diameter, Ari Holcombe Pomerance

Mathematics, Statistics, and Computer Science Honors Projects

A mixing measure is the expected length of a random walk in a graph given a set of starting and stopping conditions. We determine the tree structures of order n with diameter d that minimize and maximize for a few mixing measures. We show that the maximizing tree is usually a broom graph or a double broom graph and that the minimizing tree is usually a seesaw graph or a double seesaw graph.


Gentrification And Crime In The Twin Cities: Insights And Challenges Through A Statistical Lens, Erin G. Franke May 2023

Gentrification And Crime In The Twin Cities: Insights And Challenges Through A Statistical Lens, Erin G. Franke

Mathematics, Statistics, and Computer Science Honors Projects

Gentrification is a complex process of urban redevelopment that typically involves an in-migration of educated people to neighborhoods experiencing a period of disinvestment. While gentrification is widely regarded for its potential to displace long-time businesses and residents of the neighborhood, its impact on crime is highly controversial. There is not a consensus on the relationship between gentrification and crime across criminological theory and past statistical studies have also shown contradictory results. Measuring gentrification on the tract level with census data, we seek to understand gentrification’s relationship with violent crime and theft in the Twin Cities. Using a Poisson model with …


Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham May 2023

Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham

Electronic Theses and Dissertations

The abundance, accessibility, and scale of data have engendered an era where machine learning can quickly and accurately solve complex problems, identify complicated patterns, and uncover intricate trends. One research area where many have applied these techniques is the stock market. Yet, financial domains are influenced by many factors and are notoriously difficult to predict due to their volatile and multivariate behavior. However, the literature indicates that public sentiment data may exhibit significant predictive qualities and improve a model’s ability to predict intricate trends. In this study, momentum SVM classification accuracy was compared between datasets that did and did not …


A Machine Learning Approach To Obese-Inflammatory Phenotyping, Tania Mayleth Vargas May 2023

A Machine Learning Approach To Obese-Inflammatory Phenotyping, Tania Mayleth Vargas

Theses and Dissertations

Obesity is the accumulation of an abnormal, or excessive, amount of fat in the body, which can have negative effects on overall health. This excess accumulation of macronutrients in adipose tissue can cause the release of inflammatory mediators, leading to a proinflammatory state. Inflammation is a known risk factor for various health conditions, including cardiovascular diseases, metabolic syndrome, and diabetes. This study sought to examine the use of data mining methods, particularly clustering algorithms, to identify inflammatory biomarker phenotypes and their association with obesity in a local adolescent population. The algorithms evaluated in this study included: k-means, Ward's hierarchical …