Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (1395)
- Statistical Theory (1191)
- Statistical Models (374)
- Applied Mathematics (362)
- Mathematics (341)
-
- Statistical Methodology (274)
- Data Science (194)
- Computer Sciences (171)
- Biostatistics (155)
- Engineering (154)
- Medicine and Health Sciences (151)
- Probability (150)
- Multivariate Analysis (143)
- Life Sciences (139)
- Business (127)
- Longitudinal Data Analysis and Time Series (122)
- Categorical Data Analysis (115)
- Economics (113)
- Law (102)
- Other Statistics and Probability (93)
- Education (91)
- Environmental Sciences (86)
- Artificial Intelligence and Robotics (81)
- Design of Experiments and Sample Surveys (77)
- Econometrics (76)
- Psychology (65)
- Numerical Analysis and Scientific Computing (58)
- Institution
-
- Wayne State University (1097)
- Wright State University (138)
- Utah State University (103)
- Cornell University Law School (75)
- California Polytechnic State University, San Luis Obispo (59)
-
- Air Force Institute of Technology (55)
- Old Dominion University (55)
- University of Kentucky (54)
- Montclair State University (53)
- University of Arkansas, Fayetteville (51)
- Southern Methodist University (46)
- Western Kentucky University (46)
- University of Nebraska - Lincoln (43)
- Central Bank of Nigeria (41)
- City University of New York (CUNY) (39)
- Virginia Commonwealth University (39)
- Claremont Colleges (36)
- Illinois State University (35)
- Kennesaw State University (35)
- University of Richmond (34)
- Georgia Southern University (31)
- Louisiana Tech University (29)
- Stephen F. Austin State University (25)
- University of New Mexico (25)
- East Tennessee State University (24)
- University of Nevada, Las Vegas (24)
- Prairie View A&M University (22)
- The University of Akron (19)
- Technological University Dublin (17)
- Michigan Technological University (16)
- Keyword
-
- Statistics (122)
- Empirical legal studies (55)
- Simulation (49)
- Machine learning (42)
- Regression (40)
-
- Bias (35)
- Logistic regression (33)
- Monte Carlo simulation (30)
- Bootstrap (29)
- Power (29)
- Bayesian (28)
- Machine Learning (27)
- Reliability (27)
- Confidence interval (26)
- Western Kentucky University (26)
- Mean squared error (24)
- Estimation (22)
- Maximum likelihood estimation (22)
- Missing data (22)
- Monte Carlo (22)
- Nonparametric (21)
- Robustness (21)
- Sample size (21)
- Type I error (21)
- Effect size (20)
- Multicollinearity (20)
- Statistical analysis (20)
- Confidence intervals (19)
- Enrollment (19)
- Pure sciences (19)
- Publication Year
- Publication
-
- Journal of Modern Applied Statistical Methods (1093)
- Mathematics and Statistics Faculty Publications (137)
- Theses and Dissertations (91)
- Cornell Law Faculty Publications (75)
- Electronic Theses and Dissertations (58)
-
- Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works (49)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (46)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (42)
- CBN Journal of Applied Statistics (JAS) (40)
- Graduate Theses and Dissertations (38)
- Department of Math & Statistics Faculty Publications (34)
- Theses and Dissertations--Statistics (32)
- College of Graduate Studies: Theses & Dissertations (29)
- SMU Data Science Review (29)
- Master's Theses (27)
- Mathematics & Statistics Theses & Dissertations (26)
- WKU Administration Documents (26)
- Annual Symposium on Biomathematics and Ecology Education and Research (25)
- Symposium of Student Scholars (24)
- Applications and Applied Mathematics: An International Journal (AAM) (22)
- Articles (22)
- Statistics (21)
- Publications and Research (20)
- Williams Honors College, Honors Research Projects (19)
- Department of Statistics: Dissertations, Theses, and Student Research (17)
- CMC Senior Theses (16)
- Dissertations, Master's Theses and Master's Reports (16)
- Mathematics Senior Capstone Papers (15)
- Statistical Science Theses and Dissertations (15)
- Doctoral Dissertations (14)
- Publication Type
- File Type
Articles 391 - 420 of 2918
Full-Text Articles in Applied Statistics
Formula 101 Using 2022 Formula One Season Data To Understand The Race Results, Christopher Garcia, Oliver Lopez
Formula 101 Using 2022 Formula One Season Data To Understand The Race Results, Christopher Garcia, Oliver Lopez
Student Scholar Symposium Abstracts and Posters
The reason why I am interested in Formula One is that my friend showed me what Formula One was all about. It became interesting to see the action of the sport, including the battles the drivers have during the race and how fast they go through a corner. Also, when qualifying comes around, they push their car to the absolute limit to gain a few seconds off their opponents. The drivers only in the top 10 receive points from the winner getting 25 points, the last driver in the top 10 getting 1 point, and those below the top ten …
Quantifying The Effect Of Socio-Economic Predictors And The Built Environment On Mental Health Events In Little Rock, Ar, Alfieri Ek, Grant Drawve, Samantha Robinson, Jyotishka Datta
Quantifying The Effect Of Socio-Economic Predictors And The Built Environment On Mental Health Events In Little Rock, Ar, Alfieri Ek, Grant Drawve, Samantha Robinson, Jyotishka Datta
Sociology and Criminology Faculty Publications and Presentations
Law enforcement agencies continue to grow in the use of spatial analysis to assist in identifying patterns of outcomes. Despite the critical nature of proper resource allocation for mental health incidents, there has been little progress in statistical modeling of the geo-spatial nature of mental health events in Little Rock, Arkansas. In this article, we provide insights into the spatial nature of mental health data from Little Rock, Arkansas between 2015 and 2018, under a supervised spatial modeling framework. We provide evidence of spatial clustering and identify the important features influencing such heterogeneity via a spatially informed hierarchy of generalized …
A Monte Carlo Analysis Of Nonprobability Sampling & Post Hoc Corrections, Julia Hong
A Monte Carlo Analysis Of Nonprobability Sampling & Post Hoc Corrections, Julia Hong
Masters Theses & Specialist Projects
Nonprobability samples are often used in place of probability samples because the former are less trouble and less expensive. Unfortunately, it is difficult to determine how well a sample represents population parameters when using nonprobability samples. Researchers attempt to mitigate the disadvantages of nonprobability sampling by performing post hoc corrections, but this adjustment may not successfully undo the effects of nonprobability sampling. To examine these effects, a Monte Carlo simulation was conducted to create a pseudo-population from which samples were drawn. Forty-one conditions were replicated 10,000 times each, with each sample consisting of 100 observations. A post-stratification adjustment was made …
Dynamics Of Inertial And Non-Inertial Particles In Geophysical Flows, Nishanta Baral
Dynamics Of Inertial And Non-Inertial Particles In Geophysical Flows, Nishanta Baral
Theses, Dissertations and Culminating Projects
We consider the dynamics of inertial and non-inertial particles in various flows. We investigate the underlying structures of the flow field by examining their Lagrangian coherent structures (LCS), which are found by computing finitetime Lyapunov exponents (FTLE). We compare the behavior of massless noninertial particles using the velocity fields from four models, the Duffing oscillator, the Bickley jet, the double-gyre flow, and a quasi-geostrophic geophysical flow model, with that of inertial particles. For inertial particles with finite size and mass, we use the Maxey-Riley equation to describe the particle’s motion. We explore the preferential aggregation of inertial particles and demonstrate …
Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski
Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski
Honors Scholar Theses
Challenging conventional wisdom is at the very core of baseball analytics. Using data and statistical analysis, the sets of rules by which coaches make decisions can be justified, or possibly refuted. One of those sets of rules relates to the construction of a batting order. Through data collection, data adjustment, the construction of a baseball simulator, and the use of a Monte Carlo Simulation, I have assessed thousands of possible batting orders to determine the roster-specific strategies that lead to optimal run production for the 2023 UConn baseball team. This paper details a repeatable process in which basic player statistics …
Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild
Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild
Department of Statistics: Dissertations, Theses, and Student Research
The words people choose to use hold a lot of power, whether that be in spreading truth or deception. As listeners and readers, we do our best to understand how words are being used. There are many current methods in computer science literature attempting to embed words into numerical information for statistical analyses. Some of these embedding methods, such as Bag of Words, treat words as independent, while others, such as Word2Vec, attempt to gain information about the context of words. It is of interest to compare how well these various methods of translating text into numerical data work specifically …
A Machine Learning Approach To Obese-Inflammatory Phenotyping, Tania Mayleth Vargas
A Machine Learning Approach To Obese-Inflammatory Phenotyping, Tania Mayleth Vargas
Theses and Dissertations
Obesity is the accumulation of an abnormal, or excessive, amount of fat in the body, which can have negative effects on overall health. This excess accumulation of macronutrients in adipose tissue can cause the release of inflammatory mediators, leading to a proinflammatory state. Inflammation is a known risk factor for various health conditions, including cardiovascular diseases, metabolic syndrome, and diabetes. This study sought to examine the use of data mining methods, particularly clustering algorithms, to identify inflammatory biomarker phenotypes and their association with obesity in a local adolescent population. The algorithms evaluated in this study included: k-means, Ward's hierarchical …
Small But Mighty: Examing The Utility Of Microstatistics In Modeling Ice Hockey, Matt Palmer
Small But Mighty: Examing The Utility Of Microstatistics In Modeling Ice Hockey, Matt Palmer
Senior Honors Theses
As research into hockey analytics continues, an increasing number of metrics are being introduced into the knowledge base of the field, creating a need to determine whether various stats are useful or simply add noise to the discussion. This paper examines microstatistics – manually tracked metrics which go beyond the NHL’s publicly released stats – both through the lens of meta-analytics (which attempt to objectively assess how useful a metric is) and modeling game probabilities. Results show that while there is certainly room for improvement in understanding and use of microstats in modeling, the metrics overall represent an area of …
Inference For Multiple Utility In Time-Dependent Choice Pairs Under Copula-Based Models, Sasanka Adikari
Inference For Multiple Utility In Time-Dependent Choice Pairs Under Copula-Based Models, Sasanka Adikari
Mathematics & Statistics Theses & Dissertations
Models for discrete choice experiments (DCE) are frequently used to analyze consumer choices about products and services. A family of DCE, best-worst scaling experiments, offers more in-depth insights into consumer preferences by eliciting a best and worst choice from a set of options, rather than just a single preference. Traditional approaches often assume that choices are mutually exclusive over time, which may not always be the case. This dissertation proposes a novel model for DCE that takes into account the changing nature of consumer choices over time and the priority constraint of transition probabilities. The model introduces a copula combination …
Jackknife Empirical Likelihood Tests For Equality Of Generalized Lorenz Curves, Anton Butenko
Jackknife Empirical Likelihood Tests For Equality Of Generalized Lorenz Curves, Anton Butenko
Electronic Theses, Projects, and Dissertations
A Lorenz curve is a graphical representation of the distribution of income or wealth within a population. The generalized Lorenz curve can be created by scaling the values on the vertical axis of a Lorenz curve by the average output of the distribution. In this thesis, we propose two nonparametric methods for testing the equality of two generalized Lorenz curves. Both methods are based on empirical likelihood and utilize a U -statistic. We derive the limiting distribution of the likelihood ratio, which is shown to follow a chi-squared distribution with one degree of freedom. We conduct simulations to compare the …
Time Series Analysis Of Longitudinally Collected Standard Autoperimetry Data In Glaucoma Patients, Carlyn Childress
Time Series Analysis Of Longitudinally Collected Standard Autoperimetry Data In Glaucoma Patients, Carlyn Childress
Honors College Theses
Glaucoma is a group of eye diseases in which damage gradually occurs to the optic nerve, which often leads to partial or complete loss of vision. As the second leading cause of blindness, there is no cure for glaucoma. Early detection and the tracking of its progression is key to managing the effects of glaucoma. Ordinary Least Squares Regression (OLSR), the most commonly used methodology for tracking glaucoma progression, is inappropriate as the longitudinally collected perimetry data from the glaucoma patients appears to be temporally correlated. Time series models, that account for temporal correlation, are better methods to analyze Mean …
Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash
Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash
Symposium of Student Scholars
Employee attrition is a relevant issue that every business employer must consider when gauging the effectiveness of their employees. Whether or not an employee chooses to leave their job can come from a multitude of factors. As a result, employers need to develop methods in which they can measure attrition by calculating the several qualities of their employees. Factors like their age, years with the company, which department they work in, their level of education, their job role, and even their marital status are all considered by employers to assist in predicting employee attrition. This project will be analyzing a …
Reducing Restaurant Inventory Costs Through Sales Forecasting, Tyler Mason, Chris Schoen, Trevor Gilbert, Jonathan Enriquez
Reducing Restaurant Inventory Costs Through Sales Forecasting, Tyler Mason, Chris Schoen, Trevor Gilbert, Jonathan Enriquez
Senior Design Project For Engineers
Family Restaurant is a local restaurant in the greater Atlanta area that serves a variety of dishes that include an assortment of 19 different proteins. Currently, Family Restaurant places protein orders based on business intuition, and tends to over-stock and sometimes under-stock. To minimize inventory costs by reducing over-stocking and preventing under-stocking of proteins, we applied Facebook Prophet (FB Prophet), ARIMA, and XG Boost machine learning models to predict protein demand and then fed these results into a Fixed Time Period inventory model to make an overall order suggestion based on the specified time period. We trained our models on …
Two Sample Statistical Test For Location Parameters, Narinder Kumar, Arun Kumar
Two Sample Statistical Test For Location Parameters, Narinder Kumar, Arun Kumar
Journal of Modern Applied Statistical Methods
A class of distribution-free tests for the homogeneity of location parameters is proposed and compared with different competitors in terms of Pitman asymptotic relative efficiency. A numerical example is provided and a simulation study is made to check the performance of the tests.
Interpretable Learning In Multivariate Big Data Analysis For Network Monitoring, José Camacho, Rasmus Bro, David Kotz
Interpretable Learning In Multivariate Big Data Analysis For Network Monitoring, José Camacho, Rasmus Bro, David Kotz
Dartmouth Scholarship
There is an increasing interest in the development of new data-driven models useful to assess the performance of communication networks. For many applications, like network monitoring and troubleshooting, a data model is of little use if it cannot be interpreted by a human operator. In this paper, we present an extension of the Multivariate Big Data Analysis (MBDA) methodology, a recently proposed interpretable data analysis tool. In this extension, we propose a solution to the automatic derivation of features, a cornerstone step for the application of MBDA when the amount of data is massive. The resulting network monitoring approach allows …
Here Come The Floods: Classification Of Rain-On-Snow Induced Flooding In Nevada, Emma Watts
Here Come The Floods: Classification Of Rain-On-Snow Induced Flooding In Nevada, Emma Watts
Student Research Symposium
Given Nevada’s history of destructive flooding resulting from rain falling on mountainous snowpack, often called rain-on-snow (ROS) events, there is a great need to incorporate these events and their residual effects in infrastructure design methods. Examining relationships between USGS streamflow measurements and climate variables (specifically precipitation, temperature, and snowpack) obtained from neighboring SNOTEL stations provides means by which to classify ROS-induced floods from ROS events. Using both temperature and snowpack-based criterion to classify ROS events, this project differentiates between non-ROS and ROS-induced floods in a subset of USGS stations across the Sierra Nevada and reveals that ROS-induced floods produce, on …
A Graphical User Interface Using Spatiotemporal Interpolation To Determine Fine Particulate Matter Values In The United States, Kelly M. Entrekin
A Graphical User Interface Using Spatiotemporal Interpolation To Determine Fine Particulate Matter Values In The United States, Kelly M. Entrekin
Honors College Theses
Fine particulate matter or PM2.5 can be described as a pollution particle that has a diameter of 2.5 micrometers or smaller. These pollution particle values are measured by monitoring sites installed across the United States throughout the year. While these values are helpful, a lot of areas are not accounted for as scientists are not able to measure all of the United States. Some of these unmeasured regions could be reaching high PM2.5 values over time without being aware of it. These high values can be dangerous by causing or worsening health conditions, such as cardiovascular and lung diseases. Within …
State Gross Domestic Product Predictions Using Hierarchical Clustering And Multivariate Time Series, Austin Dae Nietfeld
State Gross Domestic Product Predictions Using Hierarchical Clustering And Multivariate Time Series, Austin Dae Nietfeld
Mathematics Senior Capstone Papers
This research was conducted to determine the weight certain taxes and expenditures have over state Gross Domestic Product(GDP) as well as how accurately these predictors can predict future GDP. The motivation behind this project comes from a desire to find the most efficient way to increase the GDP of states with poorer economies. This will improve the quality of life of citizens of these states. To come to a consensus as to what predictors are most influential, Hierarchical Clustering will be used to split the states into four groups. The average of each tax, expenditure and GDP from 2015-2020 will …
Regression Analysis Of Injuries On Nfl Quarterbacks, Julie Weems
Regression Analysis Of Injuries On Nfl Quarterbacks, Julie Weems
Mathematics Senior Capstone Papers
Risk assessment is an important aspect of many careers such as first responders and the military. This is no different for people who play sports, especially people who are in contact sports such as football. These players’ lives can be changed forever with one bad hit. The goal of this research is to analyze the probability of an injury for the National Football League’s (NFL) quarterbacks. It is hard to predict when, what, and where an injury will occur, because of this very little work has been done on the subject matter in a general form. The goal of this …
Firefighter Safety, Haynes Mandino
Firefighter Safety, Haynes Mandino
Mathematics Senior Capstone Papers
There are close to 1.2 million career and volunteer firefighters across the United States. In the year 2020 alone 62 of these firefighters died and 64,875 were injured. The following research was performed to determine if the firefighter profession has become safer due to new standards and regulations. Each year the National Fire Protection Agency(NFPA) and the Federal Emergency Management Agency(FEMA) collect data on the number of firefighter deaths and injuries, in order to determine if the standards and regulations are keeping firefighters safe. Statistical hypothesis testing and linear regression were performed on the data to show if in fact …
Does The Three Point Shot Affect Winning Percentage, Marcamus Winn
Does The Three Point Shot Affect Winning Percentage, Marcamus Winn
Mathematics Senior Capstone Papers
The three-point shot, introduced in the late 1970s, is a shot that occurs typically 24 feet away from the basket at the professional level. Strategically the game of basketball was originally based on two-point field goals. Recently, there has been a noticeable trend in the popularity of the three-point shot amongst professional teams. Nowadays, three point shot attempts account for more than a third of average NBA shot selection. Statistical analysis is becoming integral to athletics. Statistics has become a critical component to the development of not only on court basketball strategies, but also team structure as well. There are …
Mktg 666: Mktg 666 Research Methods 2 Seminar, Saim Kashmiri
Mktg 666: Mktg 666 Research Methods 2 Seminar, Saim Kashmiri
GMAS Course Syllabi
No abstract provided.
Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley
Modeling The Probability Of A Successful Stolen Base Attempt In Major League Baseball, Cade Stanley
Senior Theses
In Major League Baseball (MLB), the outcome of a stolen base attempt has important implications. Success moves the runner closer to scoring, while failure records an out and removes the runner from the basepaths altogether. Therefore, it is important that the decision by a coach or player to steal a base is well-informed. In this thesis, I explore a statistical approach to making this decision. I train logistic regression and random forest models, using data about the game situation and about the runner, pitcher, and catcher involved in the stolen base attempt, to estimate the probability that a stolen base …
Moral Injury To Inform Analysis Of Post-Traumatic Stress Disorder, Amanda Julia Manea
Moral Injury To Inform Analysis Of Post-Traumatic Stress Disorder, Amanda Julia Manea
Senior Theses
Post-traumatic stress disorder (PTSD) is a mental health condition that almost one out of ten veterans struggle with. Although the National Center for PTSD has made extensive progress in characterizing and developing new treatments for PTSD, most veterans still experience symptoms of PTSD following treatment. Novel avenues of investigation, such as developing algorithms to review electronic health record (EHR) data and better understanding moral injury, are being pursued to address the gap that still exists when it comes to treating veterans. Moral injury is the individual evaluation of exposure to a potentially morally injurious event (PMIE) and can lead to …
Self-Learning Algorithms For Intrusion Detection And Prevention Systems (Idps), Juan E. Nunez, Roger W. Tchegui Donfack, Rohit Rohit, Hayley Horn
Self-Learning Algorithms For Intrusion Detection And Prevention Systems (Idps), Juan E. Nunez, Roger W. Tchegui Donfack, Rohit Rohit, Hayley Horn
SMU Data Science Review
Today, there is an increased risk to data privacy and information security due to cyberattacks that compromise data reliability and accessibility. New machine learning models are needed to detect and prevent these cyberattacks. One application of these models is cybersecurity threat detection and prevention systems that can create a baseline of a network's traffic patterns to detect anomalies without needing pre-labeled data; thus, enabling the identification of abnormal network events as threats. This research explored algorithms that can help automate anomaly detection on an enterprise network using Canadian Institute for Cybersecurity data. This study demonstrates that Neural Networks with Bayesian …
Finite Mixture Modeling For Hierarchically Structured Data With Application To Keystroke Dynamics, Andrew Simpson, Semhar Michael
Finite Mixture Modeling For Hierarchically Structured Data With Application To Keystroke Dynamics, Andrew Simpson, Semhar Michael
SDSU Data Science Symposium
Keystroke dynamics has been used to both authenticate users of computer systems and detect unauthorized users who attempt to access the system. Monitoring keystroke dynamics adds another level to computer security as passwords are often compromised. Keystrokes can also be continuously monitored long after a password has been entered and the user is accessing the system for added security. Many of the current methods that have been proposed are supervised methods in that they assume that the true user of each keystroke is known apriori. This is not always true for example with businesses and government agencies which have internal …
Two-Stage Approach For Forensic Handwriting Analysis, Ashlan J. Simpson, Danica M. Ommen
Two-Stage Approach For Forensic Handwriting Analysis, Ashlan J. Simpson, Danica M. Ommen
SDSU Data Science Symposium
Trained experts currently perform the handwriting analysis required in the criminal justice field, but this can create biases, delays, and expenses, leaving room for improvement. Prior research has sought to address this by analyzing handwriting through feature-based and score-based likelihood ratios for assessing evidence within a probabilistic framework. However, error rates are not well defined within this framework, making it difficult to evaluate the method and can lead to making a greater-than-expected number of errors when applying the approach. This research explores a method for assessing handwriting within the Two-Stage framework, which allows for quantifying error rates as recommended by …
Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad
Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad
SDSU Data Science Symposium
Additive manufacturing (AM) is the process of building components through an iterative process of adding material in specific designs. AM has a wide range of process parameters that influence the quality of the component. This work applies Gaussian mixture models to detect clusters of similar stress values within and across components manufactured with varying process parameters. Further, a mixture of regression models is considered to simultaneously find groups and also fit regression within each group. The results are compared with a previous naive approach.
A Characterization Of Bias Introduced Into Forensic Source Identification When There Is A Subpopulation Structure In The Relevant Source Population., Dylan Borchert, Semhar Michael, Christopher Saunders
A Characterization Of Bias Introduced Into Forensic Source Identification When There Is A Subpopulation Structure In The Relevant Source Population., Dylan Borchert, Semhar Michael, Christopher Saunders
SDSU Data Science Symposium
In forensic source identification the forensic expert is responsible for providing a summary of the evidence that allows for a decision maker to make a logical and coherent decision concerning the source of some trace evidence of interest. The academic consensus is usually that this summary should take the form of a likelihood ratio (LR) that summarizes the likelihood of the trace evidence arising under two competing propositions. These competing propositions are usually referred to as the prosecution’s proposition, that the specified source is the actual source of the trace evidence, and the defense’s proposition, that another source in a …
Session 8: Ensemble Of Score Likelihood Ratios For The Common Source Problem, Federico Veneri, Danica M. Ommen
Session 8: Ensemble Of Score Likelihood Ratios For The Common Source Problem, Federico Veneri, Danica M. Ommen
SDSU Data Science Symposium
Machine learning-based Score Likelihood Ratios have been proposed as an alternative to traditional Likelihood Ratios and Bayes Factor to quantify the value of evidence when contrasting two opposing propositions.
Under the common source problem, the opposing proposition relates to the inferential problem of assessing whether two items come from the same source. Machine learning techniques can be used to construct a (dis)similarity score for complex data when developing a traditional model is infeasible, and density estimation is used to estimate the likelihood of the scores under both propositions.
In practice, the metric and its distribution are developed using pairwise comparisons …