Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2521 - 2550 of 12804

Full-Text Articles in Statistics and Probability

A Bootstrap Test For Informative Intra-Cluster Group Sizes In Clustered Data, Hasika K. Wickrama Senevirathne, Sandipan Dutta Jan 2023

A Bootstrap Test For Informative Intra-Cluster Group Sizes In Clustered Data, Hasika K. Wickrama Senevirathne, Sandipan Dutta

College of Sciences Posters

Clustered data are frequently observed in various domains of scientific and social studies. In a typical clustered data, units within a cluster are correlated while units between different clusters are independent. An example of such clustered data can be found in dental studies where individuals are treated as clusters and the teeth in an individual are the units within a cluster. While analyzing such clustered data, it has been observed that the number of units present in a cluster can be informative in terms of being associated with the outcome from that cluster. Specifically, when the aim is to compare …


Copula-Based Models For Bivariate And Multivariate Zero-Inflated Count Time Series Data, Dimuthu Fernando, Norou Diawara Jan 2023

Copula-Based Models For Bivariate And Multivariate Zero-Inflated Count Time Series Data, Dimuthu Fernando, Norou Diawara

College of Sciences Posters

Count time series data have multiple applications. The applications can be found in areas of finance, climate, public health and crime data analyses. In some scenarios, count time series come as multivariate vectors that exhibit not only serial dependence within each time series but also with cross correlation among the series. When considering these observed counts, analysis presents crucial challenges when a value, say zero, occurs more often than usual. There is presence of zero-inflation in the data.

In this presentation, we mainly focus on modeling bivariate zero-inflated count time series model based on a joint distribution of the two …


Early Detection Of Covid-19 In Female Athletes Using Wearable Technology, Liliana I. Rentería, Casey E. Greenwalt, Sarah Johnson, Shiloah Shiloah Kviatkovsky, Marine Dupuit, Elisa Angeles, Sachin Narayanan, Tucker Zeleny, Michael J. Ormsbee Jan 2023

Early Detection Of Covid-19 In Female Athletes Using Wearable Technology, Liliana I. Rentería, Casey E. Greenwalt, Sarah Johnson, Shiloah Shiloah Kviatkovsky, Marine Dupuit, Elisa Angeles, Sachin Narayanan, Tucker Zeleny, Michael J. Ormsbee

Department of Statistics: Faculty Publications

Background: Heart rate variability (HRV), respiratory rate (RR), and resting heart rate (RHR) are common variables measured by wrist-worn activity trackers to monitor health, fitness, and recovery in athletes. Variations in RR are observed in lower-respiratory infections, and preliminary data suggest changes in HRV and RR are linked to early detection of COVID-19 infection in nonathletes.

Hypothesis: Wearable technology measuring HRV, RR, RHR, and recovery will be successful for early detection of COVID-19 in NCAA Division I female athletes.

Study Design: Cohort study.

Level of Evidence: Level 2.

Methods: Female athletes wore WHOOP, Inc. bands …


Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg Jan 2023

Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg

Electronic Theses and Dissertations

Using a multi-level ensemble design, we forecast international stock market returns with a novel high-dimensional data set of aggregated cross sectional firm-level predictors. The method includes considerations of model uncertainty, parameter instability, model density and non-linearities with machine learning, shrinkage and model averaging. We provide evidence that it is important to systematically focus on all four sources of forecast failure, shed light on the sparsity/density debate in the stock return forecasting dialogue and contribute interesting findings on the efficacy dimensionality reduction with principal components analysis and partial least squares. The robustness of the approach is demonstrated through applications in four …


Second Order, Unconditionally Stable, Linear Ensemble Algorithms For The Magnetohydrodynamics Equations, John Carter, Daozhi Han, Nan Jiang Jan 2023

Second Order, Unconditionally Stable, Linear Ensemble Algorithms For The Magnetohydrodynamics Equations, John Carter, Daozhi Han, Nan Jiang

Mathematics and Statistics Faculty Research & Creative Works

We Propose Two Unconditionally Stable, Linear Ensemble Algorithms with Pre-Computable Shared Coefficient Matrices Across Different Realizations for the Magnetohydrodynamics Equations. the Viscous Terms Are Treated by a Standard Perturbative Discretization. the Nonlinear Terms Are Discretized Fully Explicitly within the Framework of the Generalized Positive Auxiliary Variable Approach (GPAV). Artificial Viscosity Stabilization that Modifies the Kinetic Energy is Introduced to Improve Accuracy of the GPAV Ensemble Methods. Numerical Results Are Presented to Demonstrate the Accuracy and Robustness of the Ensemble Algorithms.


Lie Rings And Hall Basis Elements With Two Generators, Emma Schmidt, James Turner Jan 2023

Lie Rings And Hall Basis Elements With Two Generators, Emma Schmidt, James Turner

Summer Research

A Lie ring is a set L with three operations: addition, subtraction, and the Lie bracket. The first two follow the usual rules of addition and subtraction.

Lie rings are said to be free if there are no further relations beyond these. A subset S in L is said to generate Lie ring L if every element of L can be expressed as a sum or difference of iterated brackets on elements of S.


A Conceptual Framework For Knowledge Exchange In A Wildland Fire Research And Practice Context, Colin B. Mcfayden, Lynn M. Johnston, Douglas G. Woolford, Colleen George, Den Boychuk, Daniel Johnston, B. Mike Wotton, Joshua M. Johnston Jan 2023

A Conceptual Framework For Knowledge Exchange In A Wildland Fire Research And Practice Context, Colin B. Mcfayden, Lynn M. Johnston, Douglas G. Woolford, Colleen George, Den Boychuk, Daniel Johnston, B. Mike Wotton, Joshua M. Johnston

Statistical and Actuarial Sciences Publications

No abstract provided.


Invasion Dynamics Of The European Collared-Dove In North America Are Explained By Combined Effects Of Habitat And Climate, Yiran Shao, Danielle Ethier, Simon Bonner Jan 2023

Invasion Dynamics Of The European Collared-Dove In North America Are Explained By Combined Effects Of Habitat And Climate, Yiran Shao, Danielle Ethier, Simon Bonner

Statistical and Actuarial Sciences Publications

Global biodiversity is increasingly threatened by the spread of invasive species. Understanding the mechanisms influencing the initial colonization and persistence of invaders is therefore needed if conservation actions are to prevent new invasions or strive to slow their spread. The Eurasian Collared-Dove (Streptopelia decaocto, EUCO) is one of the most successful avian invasive species in North America; however, to our knowledge, no study has simultaneously examined the role that climate-matching, human activity, directional propagation, and local density have in this invasion process. Our research expands upon a cellular-automata-based hierarchical model developed to assess directional invasion dynamics to further quantify the …


Monty Hall, Admin Stem For Success Jan 2023

Monty Hall, Admin Stem For Success

STEM for Success Showcase

No abstract provided.


Tennessee Brewconomy: Navigating The Wholesale Beer Tax Landscape, Lauren E. Dansbury Jan 2023

Tennessee Brewconomy: Navigating The Wholesale Beer Tax Landscape, Lauren E. Dansbury

Science University Research Symposium (SURS)

In 2013, Tennessee transitioned from a price-based wholesale tax model to a per barrelage assessment. This research delves into the repercussions of this tax reform, assessing its impact on the brewing and wholesale distribution sectors. Despite the shift, Tennessee maintains the nation's highest wholesale beer tax for 16 consecutive years. The study examines the opportunity costs associated with this elevated tax, exploring alternative uses for the funds. Utilizing data on annual revenue collected by wholesalers from 2019 to 2022, segmented by city and county, the research provides actionable insights advocating for a reduction in the wholesale tax. The argument posits …


Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo Jan 2023

Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo

Theses and Dissertations--Statistics

Tolerance intervals in a regression setting allow the user to quantify, with a specified degree of confidence, bounds for a specified proportion of the sampled population when conditioned on a set of covariate values. While methods are available for tolerance intervals in fully-parametric regression settings, the construction of tolerance intervals for semiparametric regression models has been treated in a limited capacity. The first project fills this gap and develops likelihood-based approaches for the construction of pointwise one-sided and two-sided tolerance intervals for semiparametric regression models. A numerical approach is also presented for constructing simultaneous tolerance intervals. An appealing facet of …


Enforcement Penalties At The Itc, Andrea R. Hugill, John C. Jarosz, Katherine D. Cappaert Jan 2023

Enforcement Penalties At The Itc, Andrea R. Hugill, John C. Jarosz, Katherine D. Cappaert

Northwestern Journal of International Law & Business

The U.S. International Trade Commission (“ITC” or “Commission”) has grown in importance as a venue for U.S. companies to pursue intellectual property (“IP”) violators and to block the sale or importation of goods from overseas that infringe U.S. IP rights. Once a violation of the Section 337 of the Tariff Act of 1930 is found, an order halting further infringement, including importation, is almost always entered. In theory, potentially sizeable penalties may be imposed on entities that do not comply with the terms of an import restriction. In practice, the terms of an import restriction are almost always honored, but …


Forecasting Remission Time Of A Treatment Method For Leukemia As An Application To Statistical Inference Approach, Ahmed Galal Atia, Mahmoud Mansour, Rashad Mohamed El-Sagheer, B. S. El-Desouky Jan 2023

Forecasting Remission Time Of A Treatment Method For Leukemia As An Application To Statistical Inference Approach, Ahmed Galal Atia, Mahmoud Mansour, Rashad Mohamed El-Sagheer, B. S. El-Desouky

Basic Science Engineering

In this paper, Weibull-Linear Exponential distribution (WLED) has been investigated whether being it is a well-fit distribution to a clinical real data. These data represent the duration of remission achieved by a certain drug used in the treatment of leukemia for a group of patients. The statistical inference approach is used to estimate the parameters of the WLED through the set of the fitted data. The estimated parameters are utilized to evaluate the survival and hazard functions and hence assessing the treatment method through forecasting the duration of remission times of patients. A two-sample prediction approach has been applied to …


Classification Of Adult Income Using Decision Tree, Roland Fiagbe Jan 2023

Classification Of Adult Income Using Decision Tree, Roland Fiagbe

Data Science and Data Mining

Decision tree is a commonly used data mining methodology for performing classification tasks. It is a tree-based supervised machine learning algorithm that is used to classify or make predictions in a path of how previous questions are answered. Generally, the decision tree algorithm categorizes data into branch-like segments that develop into a tree that contains a root, nodes, and leaves. This project seeks to explore the decision tree methodology and apply it to the Adult Income dataset from the UCI Machine Learning Repository, to determine whether a person makes over 50K per year and determine the necessary factors that improve …


Massachusetts Prevalence Of Opioid Use Disorder Estimation Revisited: Comparing A Bayesian Approach To Standard Capture-Recapture Methods, Jianing Wang, Nathan Doogan, Katherine L. Thompson, Dana Bernson, Daniel Feaster, Jennifer Villani, Redonna Chandler, Laura F. White, David Kline, Joshua A. Barocas Jan 2023

Massachusetts Prevalence Of Opioid Use Disorder Estimation Revisited: Comparing A Bayesian Approach To Standard Capture-Recapture Methods, Jianing Wang, Nathan Doogan, Katherine L. Thompson, Dana Bernson, Daniel Feaster, Jennifer Villani, Redonna Chandler, Laura F. White, David Kline, Joshua A. Barocas

Statistics Faculty Publications

Background: The National Survey on Drug Use and Health (NSDUH) estimated the prevalence of opioid use disorder (OUD) among the civilian, noninstitutionalized people aged 12 years or older in Massachusetts as 1.2% between 2015 and 2017. Accurate estimation of the prevalence of OUD is critical to the success of treatment and resource planning. Various indirect estimation approaches have been used but are subject to data availability and infrastructure-related issues.

Methods: We used 2015 data from the Massachusetts Public Health Data Warehouse (PHD) to compare the results of two approaches to estimating OUD prevalence in the Massachusetts population. First, we used …


An Adaptive Algorithm For `The Secretary Problem': Alternate Proof Of The Divergence Of A Maximizer Sequence, Andrew Benfante, Xiang Xu Jan 2023

An Adaptive Algorithm For `The Secretary Problem': Alternate Proof Of The Divergence Of A Maximizer Sequence, Andrew Benfante, Xiang Xu

OUR Journal: ODU Undergraduate Research Journal

This paper presents an alternate proof of the divergence of the unique maximizer sequence {𝑥∗ 𝑛} of a function sequence {𝐹𝑛(𝑥)} that is derived from an adaptive algorithm based on the now classic optimal stopping problem, known by many names but here ‘the secretary problem’. The alternate proof uses a result established by Nguyen, Xu, and Zhao (n.d.) regarding the uniqueness of maximizer points of a generalized function sequence {𝑆𝜇,𝜎 𝑛 } and relies on the strict monotonicity of 𝐹𝑛(𝑥) as 𝑛 increases in order to show divergence of {𝑥∗ 𝑛}. Towards this, limits of the exponentiated Gaussian CDF are …


Turnover, Covid-19, And Reasons For Leaving And Staying Within Governmental Public Health, Jonathan P. Leider, Gulzar H. Shah, Valerie A. Yeager, Jingjing Yin, Kusuma Madamala Jan 2023

Turnover, Covid-19, And Reasons For Leaving And Staying Within Governmental Public Health, Jonathan P. Leider, Gulzar H. Shah, Valerie A. Yeager, Jingjing Yin, Kusuma Madamala

Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications

Background and Objectives:

Public health workforce recruitment and retention continue to challenge public health agencies. This study aims to describe the trends in intention to leave and retire and analyze factors associated with intentions to leave and intentions to stay.

Design:

Using national-level data from the 2017 and 2021 Public Health Workforce Interests and Needs Surveys, bivariate analyses of intent to leave were conducted using a Rao-Scott adjusted chi-square and multivariate analysis using logistic regression models.

Results:

In 2021, 20% of employees planned to retire and 30% were considering leaving. In contrast, 23% of employees planned to retire and 28% …


Aircraft Damage Classification By Using Machine Learning Methods, Tüzün Tolga İnan Jan 2023

Aircraft Damage Classification By Using Machine Learning Methods, Tüzün Tolga İnan

International Journal of Aviation, Aeronautics, and Aerospace

Safety is the most significant factor that affected incidents (non-fatal) and accidents (fatal) in civil aviation history related to scheduled flights. In the history of scheduled flights, the total incident and accident number until 2022 is 1988. In this study, 677 of them are taken into consideration since 11 September 2001. The purpose of this study is to reveal the factors that can classify type of aircraft damages such as none, minor and substantial in all-time incidents and accidents. ML algorithms with different configurations are applied for the classification process. The RFE and PCA are used to find the most …


Stochastic Optimization To Reduce Aircraft Taxi-In Time At Igia, New Delhi, Rajib Das, Saileswar Ghosh, Rajendra Desai, Pijus Kanti Bhuin, Stuti Agarwal Jan 2023

Stochastic Optimization To Reduce Aircraft Taxi-In Time At Igia, New Delhi, Rajib Das, Saileswar Ghosh, Rajendra Desai, Pijus Kanti Bhuin, Stuti Agarwal

International Journal of Aviation, Aeronautics, and Aerospace

Since there is an uncertainty in the arrival times of flights, pre-scheduled allocation of runways and stands and the subsequent first-come-first-served treatment results in a sub-optimal allocation of runways and stands, this is the prime reason for the unusual delays in taxi-in times at IGIA, New Delhi.

We simulated the arrival pattern of aircraft and utilized stochastic optimization to arrive at the best runway-stands allocation for a day. Optimization is done using a GRG Non-Linear algorithm in the Frontline Systems Analytic Solver platform. We applied this model to eight representative scenarios of two different days. Our results show that without …


A Deep Bilstm Machine Learning Method For Flight Delay Prediction Classification, Desmond B. Bisandu, Irene Moulitsas Jan 2023

A Deep Bilstm Machine Learning Method For Flight Delay Prediction Classification, Desmond B. Bisandu, Irene Moulitsas

Journal of Aviation/Aerospace Education & Research

This paper proposes a classification approach for flight delays using Bidirectional Long Short-Term Memory (BiLSTM) and Long Short-Term Memory (LSTM) models. Flight delays are a major issue in the airline industry, causing inconvenience to passengers and financial losses to airlines. The BiLSTM and LSTM models, powerful deep learning techniques, have shown promising results in a classification task. In this study, we collected a dataset from the United States (US) Bureau of Transportation Statistics (BTS) of flight on-time performance information and used it to train and test the BiLSTM and LSTM models. We set three criteria for selecting highly important features …


Meta-Analysis Of Mesenchymal Stem Cell Gene Expression Data From Obese And Non-Obese Patients, Dakota William Shields Jan 2023

Meta-Analysis Of Mesenchymal Stem Cell Gene Expression Data From Obese And Non-Obese Patients, Dakota William Shields

Masters Theses

"The prevalence of gene expression microarray datasets in public repositories gives opportunity to analyze biologically interesting datasets without running the laboratory aspect in house. Such experimentation is expensive in terms of finances, time, and expertise, which often results in low numbers of replicates. Meta-analysis techniques attempt to overcome issues due to few biological or technical replicates by combining separate experiments together to increase statistical power. Proper statistical considerations help to offset issues like simultaneous testing of thousands of genes, unintended hybridization, and other noises.

Microarrays contain light intensities from tens of thousands of hybridized probes giving a measure of gene …


The Vertical Evolution And Horizontal Correlation Of Information Science Education Research In China, Siyu Chen, Xiaofeng Zhu, Xumu Jiang Jan 2023

The Vertical Evolution And Horizontal Correlation Of Information Science Education Research In China, Siyu Chen, Xiaofeng Zhu, Xumu Jiang

Journal of Scientific Information Research

[Purpose/significance]The comprehensive and accurate visualization of the information science education research in our country will help clarify the research context and inspire future research.[Method/process] Based on the relevant literature on "information education" in the CNKI database from 1980 to 2022, from the high-frequency keywords of the literature, the evolutionary context and development direction of our country's information science education research hotspots were longitudinally studied; from the three levels of system including our county's information science reseanch subject and research authar, our country's information science research research subject and research auther, horizontally explored and analyzed the mutual influence and mutual connection …


Graphs Without A 2c3-Minor And Bicircular Matroids Without A U3,6-Minor, Daniel Slilaty Jan 2023

Graphs Without A 2c3-Minor And Bicircular Matroids Without A U3,6-Minor, Daniel Slilaty

Mathematics and Statistics Faculty Publications

In this note we characterize all graphs without a 2C3-minor. A consequence of this result is a characterization of the bicircular matroids with no U3,6-minor.


Making Data-Driven Decisions For Investing In Restaurant Business: A Case Study Based On Zomato Dataset, Rachna Shah Jan 2023

Making Data-Driven Decisions For Investing In Restaurant Business: A Case Study Based On Zomato Dataset, Rachna Shah

All Graduate Theses, Dissertations, and Other Capstone Projects

In today’s fast-paced world, where time is a precious commodity, the ability to order a wide array of cuisines from the comfort of your home or office impacts your quality of life. With an increasing number of food delivery services, with just a few taps on the smartphone or clicks on the computer, we can enjoy the food we want. The importance of this convenience cannot be overstated, as it allows people to save time and effort that would otherwise be spent on cooking, grocery shopping, or dining out. As the food delivery system grows and develops, its economic framework …


The Influence Of Urban Forms And Street Infrastructure On Pedestrian-Motorist Collisions, Taylor J. Foreman Jan 2023

The Influence Of Urban Forms And Street Infrastructure On Pedestrian-Motorist Collisions, Taylor J. Foreman

College of Graduate Studies: Theses & Dissertations

Unwalkable cities are afflicted by serious issues such as increasing rates of pedestrian traffic accidents, public health concerns, and the denied right to have an accessible city. This study examines how different types of urban forms and street infrastructure contribute to the prevalence of traffic accidents in two major metropolitan cities in the United States: Atlanta, Georgia, and Boston, Massachusetts. This study utilizes geospatial analysis through the Average Nearest Neighbor and Optimized Hot Spot Analysis tools to determine the spatial distribution of traffic accidents throughout both cities. Additionally, statistical tests were conducted to explore the relationships between the number of …


The Effect Of Age, Syntax Complexity, And Cognitive Ability On The Rate Of Semantic Illusions, Sara Anne Goring Jan 2023

The Effect Of Age, Syntax Complexity, And Cognitive Ability On The Rate Of Semantic Illusions, Sara Anne Goring

CGU Theses & Dissertations

Semantic illusions are recognition errors that occur when an individual fails to notice that information contradicts their prior knowledge (Barton & Sanford, 1993; Erickson & Mattson, 1981). For example, after hearing the question, “If a plane crashes while flying over state lines, where should the survivors be buried?” many start to consider the legality or appropriateness of the scenario despite knowing “survivors” should not be buried. Having more knowledge does not necessarily prevent individuals from overlooking illusory information/misinformation. Older adults tend to have greater crystallized intelligence than young adults, yet these age groups appear to detect illusory information at equivalent …


Fitting Time Series Models To Fisheries Data To Ascertain Age, Kathleen S. Kirch, Norou Diawara, Cynthia M. Jones Jan 2023

Fitting Time Series Models To Fisheries Data To Ascertain Age, Kathleen S. Kirch, Norou Diawara, Cynthia M. Jones

OES Faculty Publications

The ability of government agencies to assign accurate ages of fish is important to fisheries management. Accurate ageing allows for most reliable age-based models to be used to support sustainability and maximize economic benefit. Assigning age relies on validating putative annual marks by evaluating accretional material laid down in patterns in fish ear bones, typically by marginal increment analysis. These patterns often take the shape of a sawtooth wave with an abrupt drop in accretion yearly to form an annual band and are typically validated qualitatively. Researchers have shown key interest in modeling marginal increments to verify the marks do, …


การเปรียบเทียบอัลกอริทึมวิเคราะห์หาสาเหตุการตายโดยการสัมภาษณ์เมื่อมีข้อมูลไม่สมบูรณ์, ณัชชา สุวันทารัตน์ Jan 2023

การเปรียบเทียบอัลกอริทึมวิเคราะห์หาสาเหตุการตายโดยการสัมภาษณ์เมื่อมีข้อมูลไม่สมบูรณ์, ณัชชา สุวันทารัตน์

Chulalongkorn University Theses and Dissertations (Chula ETD)

การศึกษาการเปรียบเทียบประสิทธิภาพของอัลกอริทึมในการวิเคราะห์หาสาเหตุการตายจากการสัมภาษณ์ (Verbal Autopsy) เมื่อมีระดับความไม่สมบูรณ์ของข้อมูลที่ต่างกัน ทำโดยเปรียบเทียบ 5 อัลกอริทึมในการแปลผลหาสาเหตุการตาย คือ InSilicoVA InterVA-5 Tariff Naïve Bayes Classifiers (NBC) และ Random Forest (RF) ในการจำแนกหาสาเหตุการตายเมื่อไม่มีการหายของข้อมูล มีการหายของข้อมูล 5% มีการหายของข้อมูล 10% และ มีการหายของข้อมูล 20% โดยเกณฑ์วัดในการประเมินประสิทธิภาพของอัลกอริทึมคือ ระยะเวลาในการทำงาน ค่าความถูกต้อง ค่า Chance-Concordance Corrected (CCC) ค่า Cause Specific Mortality Fraction (CMSF) accuracy ค่า Sensitivity และ ค่า Specificity ซึ่งข้อมูลที่นำการวิเคราะห์ในงานวิจัยนี้คือข้อมูล PHMRC Gold Standard ที่เป็นข้อมูลของผู้ใหญ่จำนวน 7,841 ตัวอย่าง จากการศึกษาพบว่าข้อมูลสมบูรณ์ทำให้มีประสิทธิภาพดีกว่าข้อมูลที่มีการทำให้หาย มีค่าตัวชี้วัดดีกว่าในทุกระดับการหายของข้อมูลและตัวชี้วัดข้างต้นที่กล่าวมาจะค่อยๆ ลดลงตามระดับการหายที่เพิ่มขึ้น VA อัลกอริทึมที่มีประสิทธิภาพในการแปลผลดีที่สุดคือ RF ในทุกระดับการหายของข้อมูลจากทุกดัชนีชี้วัดยกเว้นระยะเวลาในการทำงานและกลุ่มที่ใช้วิธีคิดแบบ data-driven อัลกอริทึมเช่น RF, Tariff และ NBC จะมีประสิทธิภาพในการแปลผลดีกว่า InSilicoVA และ InterVA 5 อาจเป็นเพราะเมื่อข้อมูลที่ใช้ในชุดเรียนรู้ในมีความครบถ้วนและสะท้อนกับข้อมูลชุดทดสอบแตกต่างกับวิธีที่ใช้หลักทางสถิติที่จะใช้ข้อมูลในการแปลผลข้อมูลมาจากแพทย์หรือผู้เชี่ยวชาญมากกว่าให้อัลกอริทึมการแปลผลเรียนรู้จากชุดเรียนรู้


Model Criteria Selection For Predictive Purpose In Logistic Regression, Pattharapon Yuttharsaknukul Jan 2023

Model Criteria Selection For Predictive Purpose In Logistic Regression, Pattharapon Yuttharsaknukul

Chulalongkorn University Theses and Dissertations (Chula ETD)

This study investigates the performance of various model selection criteria for binary logistic regression models in diverse data settings. The research compares traditional criteria (AIC, AICc, BIC, FIC) and proposes criteria (pAIC, pAICc, pBIC) designed to improve predictive ability and prevent overfitting. A simulation study systematically manipulates factors like imbalanced ratios, collinearity, number of observations, and number of variables to evaluate the effectiveness of these criteria across various scenarios. The performance is assessed using four key metrics: F1-score, false positive rate, false negative rate, and Area Under the Curve (AUC). The findings reveal a complex interplay between data characteristics and …


ผลของความเชื่อก่อนหน้าและจำนวนค่าผิดปกติที่มีต่อการประมาณค่าสหสัมพันธ์และการเปลี่ยนแปลงความเชื่อของบุคคล, สุทัตตา เหรียญทอง Jan 2023

ผลของความเชื่อก่อนหน้าและจำนวนค่าผิดปกติที่มีต่อการประมาณค่าสหสัมพันธ์และการเปลี่ยนแปลงความเชื่อของบุคคล, สุทัตตา เหรียญทอง

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยในอดีตพบว่าบุคคลตีความสหสัมพันธ์จากแผนภาพการกระจายอย่างมีอคติตามความเชื่อของตน หากแผนภาพการกระจายมีลักษณะทางภาพที่สนับสนุนความเชื่อเดิมของบุคคลอยู่ด้วยจะส่งผลให้บุคคลนั้นเปลี่ยนแปลงความเชื่อน้อยลงแม้ว่าแผนภาพนั้นจะแสดงความสัมพันธ์ที่ไม่สอดคล้องกับความเชื่อของบุคคล อย่างไรก็ตาม งานวิจัยที่ผ่านมายังไม่มีการนำค่าผิดปกติมาศึกษาร่วมกับความเชื่อ งานวิจัยนี้จึงมีวัตถุประสงค์เพื่อศึกษาอิทธิพลของความเชื่อก่อนหน้าและจำนวนค่าผิดปกติต่อการประมาณค่าสหสัมพันธ์และการเปลี่ยนแปลงความเชื่อ โดยผู้วิจัยคาดว่าการมีค่าผิดปกติจะยิ่งทำให้แผนภาพดูมีความสัมพันธ์ต่ำลงและจะทำให้บุคคลเปลี่ยนแปลงความเชื่อน้อยลงในกรณีที่บุคคลนั้นเชื่อว่าคู่ตัวแปรไม่มีความสัมพันธ์กัน การศึกษานี้ทำการทดลองผ่านแบบสอบถามออนไลน์โดยให้ผู้เข้าร่วมวิจัยแต่ละคนตอบแบบสอบถามความเชื่อเกี่ยวกับความสัมพันธ์ระหว่างตัวแปรและประมาณค่าสหสัมพันธ์จากแผนภาพการกระจายทั้ง 12 แผนภาพ ซึ่งมาจากการผสมกันระหว่างเงื่อนไขการทดลอง ได้แก่ ความเชื่อก่อนหน้า (ไม่สอดคล้อง/สอดคล้อง), จำนวนค่าผิดปกติ (0/2/4 จุด) และระดับสหสัมพันธ์ (0.4/0.6) ผลพบว่าความเชื่อก่อนหน้าและจำนวนค่าผิดปกติมีอิทธิพลทำนายต่อการประมาณค่าสหสัมพันธ์และการเปลี่ยนแปลงความเชื่อ โดยกรณีความเชื่อก่อนหน้าไม่สอดคล้องกับข้อมูลจะประมาณค่าสหสัมพันธ์ต่ำกว่าความเป็นจริงและจะประมาณต่ำกว่ากรณีความเชื่อก่อนหน้าสอดคล้องกับข้อมูล และเมื่อจำนวนค่าผิดปกติเพิ่มขึ้นบุคคลกลับประมาณค่าสหสัมพันธ์สูงขึ้นซึ่งไม่ตรงกับสมมติฐานที่ตั้งไว้ และไม่ว่าสหสัมพันธ์จะเป็นระดับกลาง (0.4) หรือสูง (0.6) บุคคลต่างก็ประมาณค่าสหสัมพันธ์ผิดพลาดไปตามความเชื่อของตน ส่วนประเด็นการเปลี่ยนแปลงความเชื่อผลพบว่า กรณีความเชื่อก่อนหน้าไม่สอดคล้องกับข้อมูลจะเปลี่ยนแปลงความเชื่อเพิ่มขึ้นจากเดิมมากกว่ากรณีความเชื่อก่อนหน้าสอดคล้องกับข้อมูล และจะเปลี่ยนแปลงความเชื่อมากขึ้นเมื่อสหสัมพันธ์สูงขึ้น (0.6) โดยจำนวนค่าผิดปกติมีอิทธิพลเล็กน้อยต่อการเปลี่ยนแปลงความเชื่อที่เพิ่มขึ้นของกรณีความเชื่อก่อนหน้าไม่สอดคล้องกับข้อมูลซึ่งไม่ตรงกับสมมติฐานที่ตั้งไว้ และจากการตรวจสอบเพิ่มเติมพบว่าจำนวนค่าผิดปกติมีอิทธิพลทางอ้อมต่อการเปลี่ยนแปลงความเชื่อโดยมีการประมาณค่าสหสัมพันธ์เป็นตัวแปรส่งผ่านซึ่งอิทธิพลทางอ้อมนี้ถูกกำกับโดยความเชื่อก่อนหน้าและระดับสหสัมพันธ์ โดยที่กรณีความเชื่อก่อนหน้าไม่สอดคล้องกับข้อมูลจะมีขนาดอิทธิพลทางอ้อมสูงกว่ากรณีความเชื่อก่อนหน้าสอดคล้องกับข้อมูล