Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (172)
- Biostatistics (160)
- Public Health (143)
- Applied Statistics (113)
- Epidemiology (112)
-
- Social and Behavioral Sciences (108)
- Mathematics (79)
- Life Sciences (76)
- Statistical Models (67)
- Data Science (65)
- Applied Mathematics (54)
- Computer Sciences (51)
- Statistical Methodology (46)
- Health Services Research (45)
- Other Statistics and Probability (41)
- Public Affairs, Public Policy and Public Administration (38)
- Engineering (32)
- Public Health Education and Promotion (32)
- Environmental Public Health (31)
- Health Policy (31)
- Nutrition (31)
- Probability (31)
- Women's Health (31)
- Business (30)
- Occupational Health and Industrial Hygiene (30)
- Categorical Data Analysis (29)
- Education (29)
- Clinical Trials (25)
- Institution
-
- University of South Carolina (62)
- Universitas Indonesia (31)
- Missouri University of Science and Technology (26)
- Chulalongkorn University (19)
- University of Nebraska - Lincoln (19)
-
- University of South Florida (18)
- University of New Mexico (17)
- Roseman University of Health Sciences (16)
- Utah State University (16)
- University of Arkansas, Fayetteville (15)
- University of Kentucky (15)
- Air Force Institute of Technology (14)
- Georgia Southern University (14)
- Clemson University (11)
- Prairie View A&M University (11)
- Virginia Commonwealth University (11)
- Central Bank of Nigeria (10)
- City University of New York (CUNY) (9)
- Louisiana State University (9)
- Southern Methodist University (9)
- University of Nevada, Las Vegas (9)
- Bethel University (8)
- Smith College (8)
- University of Denver (8)
- Old Dominion University (7)
- Northern Illinois University (6)
- University of Mississippi (6)
- Washington University in St. Louis (6)
- Wayne State University (6)
- DePauw University (5)
- Keyword
-
- COVID-19 (28)
- Statistics (20)
- Machine learning (17)
- Machine Learning (9)
- Dietary inflammatory index (8)
-
- Mortality (8)
- Psychology (8)
- Risk (8)
- Inflammation (7)
- Data science (6)
- Morgridge College of Education (6)
- Regression (6)
- Research Methods and Information Science (6)
- Research Methods and Statistics (6)
- Women (6)
- Classification (5)
- Deep Learning (5)
- Epidemiology (5)
- Exercise (5)
- Nutrition (5)
- Obesity (5)
- Pregnancy (5)
- Survival analysis (5)
- Biomarkers (4)
- Deep learning (4)
- Forecasting (4)
- HIV (4)
- Health (4)
- Humans (4)
- Mathematics (4)
- Publication
-
- Faculty Publications (57)
- Theses and Dissertations (32)
- Kesmas (31)
- Mathematics and Statistics Faculty Research & Creative Works (21)
- Chulalongkorn University Theses and Dissertations (Chula ETD) (19)
-
- Department of Statistics: Faculty Publications (17)
- Annual Research Symposium (16)
- Electronic Theses and Dissertations (15)
- Mathematics & Statistics ETDs (15)
- USF Tampa Graduate Theses and Dissertations (14)
- Applications and Applied Mathematics: An International Journal (AAM) (11)
- All Dissertations (10)
- Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications (10)
- CBN Journal of Applied Statistics (JAS) (10)
- Graduate Theses and Dissertations (10)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (8)
- Psychology Student Works (8)
- LSU Doctoral Dissertations (7)
- SMU Data Science Review (7)
- Conference on Applied Statistics in Agriculture and Natural Resources (6)
- Statistical and Data Sciences: Faculty Publications (6)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (6)
- Arts & Sciences Graduate Student Theses and Dissertations (5)
- Dissertations and Theses (Open Access) (5)
- Harrisburg University Research Symposium: Highlighting Research, Innovation, & Creativity (5)
- Journal of Modern Applied Statistical Methods (5)
- Legacy Theses & Dissertations (2009 - 2024) (5)
- Open Access Theses & Dissertations (5)
- Publications (5)
- Research outputs 2022 to 2026 (5)
- Publication Type
- File Type
Articles 541 - 570 of 595
Full-Text Articles in Statistics and Probability
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Statistical and Data Sciences: Faculty Publications
Objective. Caregivers frequently report poor quality of life(QOL) in children with sleep-disordered breathing (SDB).Our objective is to assess the correlation between care-giver- and child-reported QOL in children with mild SDBand identify factors associated with differences between caregiver and child report.
Study Design. Analysis of baseline data from a multi-institutional randomized trialSetting. Pediatric Adenotonsillectomy Trial for Snoring, where children with mild SDB (obstructive apnea-hypopnea index\3) were randomized to observation or adenotonsillectomy.
Methods. The Pediatric Quality of Life Inventory (Peds QL)assessed baseline global QOL in participating children 5 to12 years old and their caregivers. Caregiver and child scores were compared. Multivariable regression …
Graph Neural Networks For Improved Interpretability And Efficiency, Patrick Pho
Graph Neural Networks For Improved Interpretability And Efficiency, Patrick Pho
Electronic Theses and Dissertations, 2020-2023
Attributed graph is a powerful tool to model real-life systems which exist in many domains such as social science, biology, e-commerce, etc. The behaviors of those systems are mostly defined by or dependent on their corresponding network structures. Graph analysis has become an important line of research due to the rapid integration of such systems into every aspect of human life and the profound impact they have on human behaviors. Graph structured data contains a rich amount of information from the network connectivity and the supplementary input features of nodes. Machine learning algorithms or traditional network science tools have limitation …
Change Point Detection For Streaming Data Using Support Vector Methods, Charles Harrison
Change Point Detection For Streaming Data Using Support Vector Methods, Charles Harrison
Electronic Theses and Dissertations, 2020-2023
Sequential multiple change point detection concerns the identification of multiple points in time where the systematic behavior of a statistical process changes. A special case of this problem, called online anomaly detection, occurs when the goal is to detect the first change and then signal an alert to an analyst for further investigation. This dissertation concerns the use of methods based on kernel functions and support vectors to detect changes. A variety of support vector-based methods are considered, but the primary focus concerns Least Squares Support Vector Data Description (LS-SVDD). LS-SVDD constructs a hypersphere in a kernel space to bound …
Nidus Idearum. Scilogs, Ix: Neutrosophia Perennis, Florentin Smarandache
Nidus Idearum. Scilogs, Ix: Neutrosophia Perennis, Florentin Smarandache
Branch Mathematics and Statistics Faculty and Staff Publications
In this ninth book of scilogs collected from my nest of ideas, one may find new and old questions and solutions, – in email messages to research colleagues, or replies, and personal notes, some handwritten on the planes to, and from international conferences, about topics on Neutrosophy and its applications, such as: Neutrosophic Bipolar Set, Linguistic Neutrosophic Set, Neutrosophic Resonance Frequency, n-ary HyperAlgebra, n-ary NeutroHyperAlgebra, n-ary AntiHyperAlgebra, Plithogenic Crisp Graph, Plithogenic Fuzzy Graph, Plithogenic Intuitionistic Fuzzy Graph, Plithogenic Neutrosophic Graph, Plithogenic Real Number Graph, Plithogenic Complex Number Graph, Plithogenic Neutrosophic Number Graph, and many more.
Exchanging ideas with: Tareq Al-Shami, …
Reducing Print Time While Minimizing Loss In Mechanical Properties In Consumer Fdm Parts, Long Le, Mitchel A. Rabsatt, Hamid Eisazadeh, Mona Torabizadeh
Reducing Print Time While Minimizing Loss In Mechanical Properties In Consumer Fdm Parts, Long Le, Mitchel A. Rabsatt, Hamid Eisazadeh, Mona Torabizadeh
Mechanical & Aerospace Engineering Faculty Publications
Fused deposition modeling (FDM), one of various additive manufacturing (AM) technologies, offers a useful and accessible tool for prototyping and manufacturing small volume functional parts. Polylactic acid (PLA) is among the commonly used materials for this process. This study explores the mechanical properties and print time of additively manufactured PLA with consideration to various process parameters. The objective of this study is to optimize the process parameters for the fastest print time possible while minimizing the loss in ultimate strength. Design of experiments (DOE) was employed using a split-plot design with five factors. Analysis of variance (ANOVA) was employed to …
Modeling Joint Survival Probabilities Of Runs Scored And Balls Faced In Limited Overs Cricket Using Copulas, Lochana K. Palayangoda, Hasika W. Senevirathne, Ananda B. Manage
Modeling Joint Survival Probabilities Of Runs Scored And Balls Faced In Limited Overs Cricket Using Copulas, Lochana K. Palayangoda, Hasika W. Senevirathne, Ananda B. Manage
Mathematics & Statistics Faculty Publications
In limited overs cricket, the goal of a batsman is to score a maximum number of runs within a limited number of balls. Therefore, the number of runs scored and the number of balls faced are the two key statistics used to evaluate the performance of a batsman. In cricket, as the batsmen play as pairs, having longer partnerships is also key to building strong innings. Moreover, having a steady opening partnership is extremely important as a team aims to build such a stronger innings. In this study, we have shown a way to evaluate the performance of opening partnerships …
The Online Ordering Behaviors Among Participants In The Oklahoma Women, Infants, And Children Program: A Cross-Sectional Analysis, Qi Zhang, Kayoung Park, Junzhou Zhang, Chuanyi Tang
The Online Ordering Behaviors Among Participants In The Oklahoma Women, Infants, And Children Program: A Cross-Sectional Analysis, Qi Zhang, Kayoung Park, Junzhou Zhang, Chuanyi Tang
Community & Environmental Health Faculty Publications
The Special Supplemental Nutrition Program for Women, Infants, and Children (WIC) is a nutrition assistance program in the United States (U.S.). Participants in the program redeem their prescribed food benefits in WIC-authorized grocery stores. Online ordering is an innovative method being pilot-tested in some stores to facilitate WIC participants' food benefit redemption, which has become especially important in the COVID-19 pandemic. The present research aimed to examine the online ordering (OO) behaviors among 726 WIC households who adopted WIC OO in a grocery chain, XYZ (anonymous) store, in Oklahoma (OK). These households represented approximately 5% of WIC households who redeemed …
Maintenance Optimization In A Digital Twin For Industry 4.0, Abhijit Gosavi, Vy Khoi Le
Maintenance Optimization In A Digital Twin For Industry 4.0, Abhijit Gosavi, Vy Khoi Le
Engineering Management and Systems Engineering Faculty Research & Creative Works
The advent of Internet of Things and artificial intelligence in the era of Industry 4.0 has transformed decision-making within production systems. In particular, many decisions that previously required significant human activity are now made automatically with minimal human intervention via so-called digital twins (DTs). In the context of maintenance and reliability modeling, this naturally calls for new paradigms that can be seamlessly integrated within DTs for decision-making. The input data for time to failure needed in reliability computations are directly collected from the work center in a digital setting and often do not satisfy a known distribution. A neural network …
Applying Machine Learning Algorithms For Face Mask Detections, Mackenzie Frato
Applying Machine Learning Algorithms For Face Mask Detections, Mackenzie Frato
Williams Honors College, Honors Research Projects
Goal: Apply multiple machine learning techniques to Face Mask images to detect if a student is wear a Face Mask and/or wearing it incorrectly or not at all. Methodology: Use 2-3 different machine learning techniques to develop this program. Will choose these techniques as I research over the semester. The best technique will be the final one used, but many will be explored. Validation techniques will be used to see which is the best technique. Timeline: Choose Dataset - October 1st, Choose techniques - October 31st, Research techniques/validation - November 31st, Begin writing code - December 13th, Finish code - …
Behavioral Predictive Analytics Towards Personalization For Self-Management – A Use Case On Linking Health-Related Social Needs, Bon Sy, Michael Wassil, Helene Connelly, Alisha Hassan
Behavioral Predictive Analytics Towards Personalization For Self-Management – A Use Case On Linking Health-Related Social Needs, Bon Sy, Michael Wassil, Helene Connelly, Alisha Hassan
Publications and Research
The objective of this research is to investigate the feasibility of applying behavioral predictive analytics to optimize patient engagement in diabetes self-management, and to gain insights on the potential of infusing a chatbot with NLP technology for discovering health-related social needs. In the U.S., less than 25% of patients actively engage in self-health management even though self-health management has been reported to associate with improved health outcomes and reduced healthcare costs. The proposed behavioral predictive analytics relies on manifold clustering to identify subpopulations segmented by behavior readiness characteristics that exhibit non-linear properties. For each subpopulation, an individualized auto-regression model and …
Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling
Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling
Williams Honors College, Honors Research Projects
This study uses various statistical analyses to evaluate the justification of rule changes for Major League Baseball that were implemented within the Minor Leagues during the 2021 minor league season. The primary focus of the study is predicting how some of these Minor League rule changes could affect the stolen base success rate and the number of attempts per game within the Major Leagues. A survey was conducted to evaluate how fans feel about stolen bases within the current game and if rules should be altered to increase the number of stolen bases that occur. Additionally, recorded Major and Minor …
A Comparison Of Stacking Methods To Estimate Survival Using Residual Lifetime Data From Prevalent Cohort Studies, Zhaoheng Li
A Comparison Of Stacking Methods To Estimate Survival Using Residual Lifetime Data From Prevalent Cohort Studies, Zhaoheng Li
Mathematics, Statistics, and Computer Science Honors Projects
Prevalent cohort studies are widely used for their cost-efficiency and convenience. However, in such studies, only the residual lifetime can be observed. Traditionally, researchers rely on self-reported onset times to infer the underlying survival distribution, which may introduce additional bias that confounds downstream analysis. This study compares two stacking procedures and one mixture model approach that uses only residual lifetime data while leveraging the strengths of different estimators. Our simulation results show that the two stacked estimators outperform the nonparametric maximum likelihood estimator (NPMLE) and the mixture model, allowing robust and accurate estimations for underlying survival distributions.
การเปรียบเทียบวิธีการคัดเลือกตัวแปรแบบรวมกลุ่ม สำหรับข้อมูลที่มีลักษณะการจำแนกแบบไบนารี, กรชนก ชมเชย
การเปรียบเทียบวิธีการคัดเลือกตัวแปรแบบรวมกลุ่ม สำหรับข้อมูลที่มีลักษณะการจำแนกแบบไบนารี, กรชนก ชมเชย
Chulalongkorn University Theses and Dissertations (Chula ETD)
งานศึกษานี้เปรียบเทียบวิธีการคัดเลือกตัวแปรแบบเดียว (Single-Feature Selection) และแบบรวมกลุ่ม (Ensemble Feature Selection) ซึ่งแบ่งเป็น 2 รูปแบบคือ รูปแบบการรวมลำดับความสำคัญของตัวแปรแล้วตามด้วยการเลือกจำนวนตัวแปรที่มีความสำคัญตามเกณฑ์ที่ระบุ (Design CT: Combination followed by Thresholding) และรูปแบบการการเลือกจำนวนตัวแปรที่มีความสำคัญตามเกณฑ์ที่ระบุแล้วตามด้วยการรวมเซตของตัวแปรที่มีความสำคัญดังกล่าว (Design TC: Thresholding followed by Combination) ผู้ศึกษาได้ใช้การคัดเลือกตัวแปรจากประเภท Filter Wrapper และ Embedded โดยใช้ 10-fold cross validation ในการเปรียบเทียบค่าเฉลี่ยของ F1-score แทนประสิทธิภาพการทำนายและค่าเบี่ยงเบนของ F1-score แทนค่าความเสถียรของการทำนาย ผ่านข้อมูล 3 ชุดได้แก่ Parkinson's Disease dataset (จำนวนตัวแปรต้น(P)=ขนาดข้อมูล(N)), LSVT Voice Rehabilitation dataset (P>N) และ Colon Cancer dataset (P>>N) ใช้ XGBoost เป็นตัวแบบทำนาย จากการศึกษาภายใต้ขอบเขตดังกล่าวพบว่า การคัดเลือกตัวแปรแบบวิธีเดียวด้วย RFE จะให้ผลดีในชุดข้อมูลที่มีมิติมาก P>>N ในเกณฑ์ 2.5% 5% และ 10% แต่การคัดเลือกแบบรวมกลุ่มจะให้ผลการทำนายที่ต่างกันภายใต้ลักษณะมิติของชุดข้อมูลและเกณฑ์ที่เลือกใช้ สำหรับการรวมลำดับความสำคัญของตัวแปรในรูปแบบ Design CT ด้วยค่ากลางและค่าเฉลี่ยเลขคณิตที่เกณฑ์ log2(P) จะให้ผลการทำนายดีกว่าวิธีอื่นใน Design CT ในชุดข้อมูล P>>N แต่สำหรับชุดข้อมูล P=N และ P>N ผลการทำนายจากแต่ละวิธีใน Design CT เพิ่มประสิทธิภาพการทำนายเล็กน้อย และสำหรับ Design TC การรวมเซตของตัวแปรต้นที่มีความสำคัญด้วยวิธีอินเตอร์เซกและมัลติอินเตอร์เซกจะให้ผลดีกว่าวิธียูเนียน สำหรับชุดข้อมูล P>>N ในทุกเกณฑ์ …
การปรับปรุงความสามารถในการพยากรณ์แบบไบนารี่โดยใช้การเรียนรู้เมตาแบบถ่วงน้ำหนักแบบปรับสำหรับการจำแนกความยากจนระดับครัวเรือนในประเทศไทย, ธารินทร์ สุขเนาว์
การปรับปรุงความสามารถในการพยากรณ์แบบไบนารี่โดยใช้การเรียนรู้เมตาแบบถ่วงน้ำหนักแบบปรับสำหรับการจำแนกความยากจนระดับครัวเรือนในประเทศไทย, ธารินทร์ สุขเนาว์
Chulalongkorn University Theses and Dissertations (Chula ETD)
งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาปัจจัยที่มีความสัมพันธ์กับความยากจนในระดับครัวเรือนและเสนอวิธีการเปรียบเทียบและปรับปรุงความสามารถในการพยากรณ์แบบไบนารี่โดยใช้การเรียนรู้เมตาแบบถ่วงน้ำหนักแบบปรับจากการคำนวนค่าถ่วงน้ำหนักวิธีที่ดีที่สุดสำหรับการจำแนกความยากจนระดับครัวเรือนในประเทศไทย โดยนำเสนอวิธีการสองขั้นตอน คือนำตัววัดประสิทธิภาพการทำนายมาใช้ในการคำนวณค่าถ่วงน้ำหนักแบบปรับ ซึ่งนำมาใช้เสมือนเป็นค่าถ่วงน้ำหนักเริ่มต้นที่ให้กับแต่ละตัวแบบ จากนั้นจึงทำนายผลด้วยวิธีการวิเคราะห์การถดถอยลอจิสติกอีกขั้นตอนหนึ่ง งานวิจัยนี้ศึกษาการคำนวณค่าถ่วงน้ำหนักแบบปรับจากตัววัดประสิทธิภาพการทำนายใน 3 กรณี ได้แก่ 1. การใช้ค่า AUC 2. การใช้ค่า F1-Score โดยพิจารณาจุดตัด 0.5 และ 3. การใช้ค่า F1-Score โดยพิจารณาค่าจุดตัดที่เหมาะสมที่สุดจากดัชนีโยเดนที่สูงสุด นอกจากนี้ เนื่องจากชุดข้อมูลสำรวจประชากรรายครัวเรือนในระดับพื้นที่มีความไม่สมดุลของระดับความยากจน จึงใช้เทคนิค SMOTE ในการจัดการกับข้อมูลที่ไม่สมดุล ทั้งนี้ ผู้วิจัยได้ทำการเปรียบเทียบผลลัพธ์จากชุดข้อมูลก่อนและหลังใช้เทคนิค SMOTE ผลการศึกษาพบว่า ปัจจัยที่มีความสัมพันธ์กับความยากจนในระดับครัวเรือนสูงมีหลายปัจจัย อาทิ อายุของหัวหน้าครัวเรือน จำนวนผู้ที่ได้รับบัตรสวัสดิการแห่งรัฐในครัวเรือน,ค่าใช้จ่ายเพื่อการบริโภคในครัวเรือน เป็นต้น และวิธีการคำนวณค่าถ่วงน้ำหนักแบบปรับจากตัววัดประสิทธิภาพ F1-Score ที่จุดตัด 0.5 มีประสิทธิภาพสูงสุดจากการพิจารณาด้วยค่าความแม่นยำในชุดข้อมูลตั้งต้นก่อนใช้เทคนิค SMOTE อย่างไรก็ตาม จากการทดสอบในชุดข้อมูลที่มีการจัดการกับข้อมูลที่ไม่สมดุลด้วยวิธี SMOTE พบว่า ประสิทธิภาพในการทำนายไม่ปรากฏว่าวิธีการคำนวณค่าถ่วงน้ำหนักแบบปรับจากตัววัดประสิทธิภาพแบบใดแบบหนึ่งที่มีประสิทธิภาพสูงสุดอย่างชัดเจน
การจำลองข้อมูลเพื่อประเมินประสิทธิภาพของการเลือกตัวอย่างแบบมีระบบชนิดผสม, นภสร รัตนวุฒิขจร
การจำลองข้อมูลเพื่อประเมินประสิทธิภาพของการเลือกตัวอย่างแบบมีระบบชนิดผสม, นภสร รัตนวุฒิขจร
Chulalongkorn University Theses and Dissertations (Chula ETD)
งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพของตัวประมาณค่าเฉลี่ยที่ได้จากการเลือกตัวอย่างแบบมีระบบชนิดผสม (Mixed Systematic Random Sampling : MRSS) กับการเลือกตัวอย่างแบบมีระบบชนิดวงกลม (Circular Systematic Sampling : CSS) และการเลือกตัวอย่างแบบมีระบบโดยใช้ช่วงเศษส่วน (Fractional Interval) สำหรับกรณีช่วงของการเลือกตัวอย่างไม่เป็นจำนวนเต็ม เมื่อประชากรมีแนวโน้มเชิงเส้น ด้วยค่าความคลาดเคลื่อนกำลังสองเฉลี่ย (Mean Square Error : MSE) และเปรียบเทียบประสิทธิภาพของการเลือกตัวอย่างแบบมีระบบทั้ง 3 วิธีด้วยค่าประสิทธิภาพสัมพัทธ์ (Relative Efficiency : RE) โดยการจำลองข้อมูลของประชากรเป็น 3 ขนาด แบ่งเป็น ขนาดเล็กหลักร้อย ได้แก่ 300, 500 และ 700 ขนาดกลางหลักพัน ได้แก่ 3,000, 5,000 และ 7,000 ขนาดใหญ่หลักหมื่น ได้แก่ 30,000, 50,000 และ 70,000 ด้วยโปรแกรม R กำหนดขนาดตัวอย่างที่ทำให้สัดส่วนระหว่างขนาดประชากรและขนาดตัวอย่างไม่เป็นจำนวนเต็ม ทำซ้ำทั้งหมด 1,000 ครั้ง พบว่าการเลือกตัวอย่างแบบมีระบบด้วยวิธี MRSS มีค่า MSE สูงกว่าการเลือกตัวอย่างอีกทั้ง 2 วิธี แต่เมื่อค่า g = 2 จะทำให้ค่าของ MSE ของการเลือกตัวอย่างทั้ง 3 วิธีมีค่ามากขึ้น โดยที่ค่า MSE ของการเลือกตัวอย่างแบบมีระบบชนิดผสมมีค่าต่ำกว่าการเลือกตัวอย่างแบบมีระบบชนิดวงกลมและวิธีใช้ช่วงเศษส่วน ทั้งนี้เป็นผลเนื่องมาจากค่า g เป็นค่าที่กำหนดความเป็นเชิงเส้น เมื่อค่า g เพิ่มมากขึ้น ความเป็นเชิงเส้นของประชากรจะลดลง ทำให้ตัวประมาณค่าเฉลี่ยตัวอย่างที่คำนวณได้มีค่าต่างจากค่าเฉลี่ยประชากรมากตามไปด้วย จึงสามารถสรุปได้ว่าตัวประมาณค่าเฉลี่ยที่ได้จากวิธีการเลือกตัวอย่างแบบมีระบบชนิดผสม มีแนวโน้มที่จะให้ค่า MSE สูงที่สุด เมื่อเทียบกับการเลือกตัวอย่างแบบมีระบบชนิดวงกลมและการเลือกตัวอย่างแบบมีระบบโดยใช้ช่วงเศษส่วน
Containing Compounding Container Congestion, Curtis Salinger
Containing Compounding Container Congestion, Curtis Salinger
CMC Senior Theses
The Covid-19 pandemic caused major disruptions throughout the container shipping supply chain. Professor Dongping Song of Liverpool University wrote a paper discussing the logistical vulnerabilities in the supply chain, including the issue of congestion in ports. This paper examines the Port of Los Angeles from 2018-2021 as it relates to Song’s paper to see how its operations were impacted during the Covid-19 timeframe. It is found that labor shortages, chassis shortages, and change in trade behavior each contributed to the congestion. Unfortunately, the implemented policies were insufficient to bolster the port against sustained challenges and congestion continues to worsen.
Informative G-Prior For Linear Models, Yu-Fang Chien
Informative G-Prior For Linear Models, Yu-Fang Chien
Graduate Research Theses & Dissertations
Zellner's objective g-prior has been widely used in linear regression models due to its simple interpretation and computational tractability in evaluating marginal likelihoods. However, the g-prior further allows portioning the prior variability explained by the linear predictor versus that of pure noise. Here, a novel, yet remarkably simple g-prior speci_cation is proposed when a subject-matter expert has information on the marginal distribution of the response yi. The approach is extended for use in mixed models with some surprising, but intuitive results. Also, this formulation of g-prior is compared with other approaches via simulation studies.
Forecasting Bitcoin, Ethereum And Litecoin Prices Using Machine Learning, Sai Prabhu Jaligama
Forecasting Bitcoin, Ethereum And Litecoin Prices Using Machine Learning, Sai Prabhu Jaligama
Graduate Research Theses & Dissertations
This research aims to predict the cryptocurrencies Bitcoin, Litecoin and Ethereum using Time Series Modelling with daily data of closing price from 16th of October 2018 to 9th of September 2021for a total of 1073 days. Augmented Dickey Fuller test was first used to check stationarity of the time series, then two forecasting algorithms called ARIMA, and PROPHET were used to make predictions. The findings show similar results for both the models for each of Bitcoin, Ethereum and Litecoin. The results achieved show modelling cryptocurrencies which are volatile using a single variable produces satisfying results.
Impact Of Public And Private Investments On Economic Growth Of Developing Countries, Faruque Ahamed
Impact Of Public And Private Investments On Economic Growth Of Developing Countries, Faruque Ahamed
Graduate Research Theses & Dissertations
This paper aims to study the impact of public and private investments on the economic growth of developing countries. The study uses panel data from 39 developing countries covering the periods 1990-2019. The study is based on the neoclassical growth models or exogenous growth models in which land, labor, capital accumulation, etc., and technology proved substantial for economic growth. The paper uses the impact on overall GDP growth and GDP per capita growth. The study used a mixed-effect regression model and a Bayesian logistic regression model to derive the findings. For private investments, domestic credit has a positive association, but …
Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair
Development Of Regional Landslide Susceptibility Models: A First Step Towards Model Transferability, Gina M. Belair
Graduate Student Theses, Dissertations, & Professional Papers
Landslides are a globally pervasive problem with the potential to cause significant fatalities and economic losses. Although landslides are widespread, many at-risk regions may not have the high-quality data or resources used in most landslide susceptibility analyses. This study aims to develop regional susceptibility relationships that are versatile and use publicly available data and open-sourced software. Logistic Regression and Frequency Ratio susceptibility relationships were developed in 23 regions in Washington, Utah, North Carolina, and Kentucky, with a region referring to a unique area and data combination. Regions were diverse in their geology, morphology, climate, and nature and quality of their …
A Non-Deterministic Deep Learning Based Surrogate For Ice Sheet Modeling, Hannah Jordan
A Non-Deterministic Deep Learning Based Surrogate For Ice Sheet Modeling, Hannah Jordan
Graduate Student Theses, Dissertations, & Professional Papers
Surrogate modeling is a new and expanding field in the world of deep learning, providing a computationally inexpensive way to approximate results from computationally demanding high-fidelity simulations. Ice sheet modeling is one of these computationally expensive models, the model used in this study currently requires between 10 and 20 minutes to complete one simulation. While this process is adequate for certain applications, the ability to use sampling approaches to perform statistical inference becomes infeasible. This issue can be overcome by using a surrogate model to approximate the ice sheet model, bringing the time to produce output down to a tenth …
Stigma And Discrimination’S Effect On Hiv Testing Of Pregnant Women In Nigeria, Charles Echezona Nzelu
Stigma And Discrimination’S Effect On Hiv Testing Of Pregnant Women In Nigeria, Charles Echezona Nzelu
Walden Dissertations and Doctoral Studies
The utilization of HIV testing services among pregnant women in Nigeria has not been optimal. Although much is known about the determinants of HIV testing among pregnant women, there is a gap in knowledge on determinants for pregnant women infected with the virus, specifically whether stigma and discrimination are barriers. The purpose of this study was to examine the effect of stigmatizing attitudes and personal knowledge of discriminatory practices towards persons living with HIV/AIDS on the decision by pregnant Nigerian women aged 15-49 years to test for HIV during antenatal visits or childbirth. The health belief model served as the …
Length Of Stay In A Homeless Shelter And Mitigating Homelessness, Uwemedimo S. Etteyit
Length Of Stay In A Homeless Shelter And Mitigating Homelessness, Uwemedimo S. Etteyit
Walden Dissertations and Doctoral Studies
AbstractHomelessness is a major public health issue in the United States. Every night, thousands of people have no residence to call their own. Most homeless persons turn to homeless shelters for help. Despite the homeless shelters, the problem of homelessness persists. This study examined the concept that the length of time spent at a homeless shelter is related to the homeless persons mitigating their homelessness through home placement, jobs, and healthcare access. Homelessness was examined using the socioecological model with its attendant levels of influence. On the intrapersonal level, socioeconomic status, education, old age, veteran status, and disability were factors. …
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
College of Graduate Studies: Theses & Dissertations
In reinforcement learning the process of selecting an action during the exploration or exploitation stage is difficult to optimize. The purpose of this thesis is to create an action selection process for an agent by employing a low discrepancy action selection (LDAS) method. This should allow the agent to quickly determine the utility of its actions by prioritizing actions that are dissimilar to ones that it has already picked. In this way the learning process should be faster for the agent and result in more optimal policies.
การเปรียบเทียบอัลกอริทึมระหว่างการสุ่มตัวอย่างแบบทอมสันและอัลกอริทึมความเชื่อมั่นขอบเขตบน สำหรับการเรียนรู้แบบเสริมแรงในเกมเป่ายิ้งฉุบ, ธันยวุฒิ อักขระสมชีพ
การเปรียบเทียบอัลกอริทึมระหว่างการสุ่มตัวอย่างแบบทอมสันและอัลกอริทึมความเชื่อมั่นขอบเขตบน สำหรับการเรียนรู้แบบเสริมแรงในเกมเป่ายิ้งฉุบ, ธันยวุฒิ อักขระสมชีพ
Chulalongkorn University Theses and Dissertations (Chula ETD)
งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพระหว่างอัลกอริทึมการสุ่มตัวอย่างแบบทอมสันและอัลกอริทึมความเชื่อมั่นขอบเขตบน ในตัวแบบการเรียนรู้แบบเสริมแรงกับการตัดสินใจเชิงพฤติกรรมของมนุษย์ ทั้งสองอัลกอริทึมเป็นอัลกอริทึมที่มีประสิทธิภาพในการแก้ไขปัญหาแบนดิทหลายแขน แต่ไม่ชัดเจนว่าทั้งสองอัลกอริทึมจะมีประสิทธิภาพอย่างไรกับปัญหาการตัดสินใจเชิงพฤติกรรมของมนุษย์ที่ความซับซ้อนทางด้านพฤติกรรม งานวิจัยนี้จำลองเกมเป่ายิ้งฉุบแทนปัญหาการตัดสินใจของมนุษย์ โดยมีองค์ประกอบเชิงพฤติกรรม 2 องค์ประกอบ คือ พฤติกรรมการใช้กลยุทธตามเข็มนาฬิกาแบบผสม และพฤติกรรมการใช้กลยุทธยุติการสูญเสีย โดยตัวแบบเกมเป่ายิ้งฉุบถูกจำลองขึ้นตามกระบวนการตัดสินใจแบบมาร์คอฟ ตัวแทนตัวแบบจากทั้งสองอัลกอริทึมจะแก้ไขปัญหาดังกล่าวและวัดประสิทธิภาพด้วยผลรางวัลสะสมภายใต้เงื่อนไขการจำลองในรูปแบบต่าง ๆ ผลการเปรียบเทียบประสิทธิภาพพบว่า ตัวแทนตัวแบบจากอัลกอริทึมความเชื่อมั่นขอบเขตบนมีประสิทธิภาพดีกว่าตัวแทนตัวแบบจากอัลกอริทึมการสุ่มตัวอย่างแบบทอมสันในการจำลองส่วนใหญ่ ยกเว้นกรณีการจำลองที่รูปแบบพฤติกรรมของมนุษย์มีความชัดเจนเป็นระยะเวลายาว ตัวแทนตัวแบบจากอัลกอริทึมการสุ่มตัวอย่างแบบทอมสันมีประสิทธิภาพดีกว่าตัวแทนตัวแบบจากอัลกอริทึมความเชื่อมั่นขอบเขตบน
การพัฒนาเวิร์กโฟลว์สําหรับตัวแบบต้นไม้จําแนกประเภทที่ดีที่สุด, พงศ์ทวัส ฮั่นวัฒนวงศ์
การพัฒนาเวิร์กโฟลว์สําหรับตัวแบบต้นไม้จําแนกประเภทที่ดีที่สุด, พงศ์ทวัส ฮั่นวัฒนวงศ์
Chulalongkorn University Theses and Dissertations (Chula ETD)
งานวิจัยนี้มีวัตถุประสงค์เพื่อพัฒนาเวิร์กโฟลว์สำหรับสร้างต้นไม้จำแนกประเภทที่ดีที่สุด ด้วยตัวแบบเชิงเส้นจำนวนเต็มแบบผสม ทำการประเมินประสิทธิภาพของตัวแบบต้นไม้จำแนกประเภทที่ดีที่สุดบนชุดข้อมูลเยอรมันเครดิต และขยายตัวแบบให้รองรับชุดข้อมูลที่ตัวแปรต้นมีค่าสูญหายจำนวนมาก จากการพัฒนาเวิร์กโฟลว์พบว่าการสร้างต้นไม้จำแนกประเภทที่ดีที่สุดโดยใช้ตัวแบบเชิงเส้นจำนวนเต็มแบบผสมในงานวิจัยของ Lin และ Tang (2021) และกำหนดค่าพารามิเตอร์ความซับซ้อนตั้งต้นเป็นค่าบวกใกล้เคียงศูนย์ให้ผลลัพธ์เป็นที่น่าพอใจ จากการเปรียบเทียบประสิทธิภาพระหว่างตัวแบบต้นไม้จําแนกประเภทที่ดีที่สุดกับต้นไม้ตัดสินใจบนชุดข้อมูลเยอรมันเครดิต พบว่าต้นไม้จำแนกประเภทที่ดีที่สุดให้อัตราความถูกต้องสูงกว่าต้นไม้ตัดสินใจทั้งบนชุดข้อมูลสร้างตัวแบบและบนชุดข้อมูลทวนสอบ 0.4% ถึง 3.2% ข้อดีของการพัฒนาเวิร์กโฟลว์โดยใช้โปรแกรมหาคำตอบสำหรับปัญหาเชิงเส้นจำนวนเต็มแบบผสม คือความสามารถในการขยายตัวแบบให้รองรับเงื่อนไขเพิ่มเติมได้ ในงานวิจัยนี้จึงเสนอตัวแบบต้นไม้จำแนกประเภทที่ดีที่สุดที่ถูกขยายให้รองรับชุดข้อมูลที่มีตัวแปรต้นสูญหายจำนวนมาก และแสดงให้เห็นว่าตัวแบบที่ถูกขยายสามารถทำงานอย่างมีประสิทธิผลบนเวิร์กโฟลว์ที่พัฒนาขึ้น
การศึกษาเปรียบเทียบตัวแบบจำลองการถดถอยโดยความไม่แน่นอนเพื่อลดเวลาในกระบวนการทดสอบวัดค่ากระแสไฟฟ้าเขียนที่เหมาะสมที่สุดของฮาร์ดไดรฟ์, ภัทรดิศ ดำรงค์ศักดิ์
การศึกษาเปรียบเทียบตัวแบบจำลองการถดถอยโดยความไม่แน่นอนเพื่อลดเวลาในกระบวนการทดสอบวัดค่ากระแสไฟฟ้าเขียนที่เหมาะสมที่สุดของฮาร์ดไดรฟ์, ภัทรดิศ ดำรงค์ศักดิ์
Chulalongkorn University Theses and Dissertations (Chula ETD)
ฮาร์ดไดรฟ์ (HDD) เป็นอุปกรณ์บันทึกข้อมูลแม่เหล็กที่มีความแม่นยำสูง ดังนั้นจึงมีค่าใช้จ่ายสูง และเสียเวลาในการวัดค่ากระแสไฟฟ้าเขียนที่เหมาะสมที่สุดฮาร์ดไดรฟ์ หากจ่ายกระแสไฟฟ้าเขียนไม่เหมาะสมจะส่งผลกระทบต่อประสิทธิภาพการทำงานของฮาร์ดไดรฟ์ ซึ่งเราใช้วิธีการเงื่อนไขการทดสอบแบบปรับตัว (Adaptive Test Condition) เป็นเทคนิคที่ปรับวิธีการทดสอบแบบดั้งเดิม ตามรูปแบบข้อมูลพารามิเตอร์ เพื่อปรับปรุงวิธีการทดสอบปัจจุบัน และลดเวลาการทดสอบ งานวิทยานิพนธ์นี้มีวัตถุประสงค์เพื่อศึกษาและเปรียบเทียบวิธีการใช้ตัวแบบจำลองการถดถอยโดยความไม่แน่นอนสำหรับการลดช่วงการวัดค่ากระไฟฟ้าเขียนที่เหมาะสมที่สุด สำหรับการลดเวลาการทดสอบวัคค่ากระแสไฟฟ้าเขียน (write current test) โดยการคำนวณช่วงความเชื่อมั่นของผลทำนายที่ระดับความเชื่อมั่นที่ยอมรับได้ โดยใช้ค่าความไม่แน่นอนของข้อมูล (Data uncertainty) ที่ผ่านวิธีปรับการเทียบมาตรฐาน (Recalibration) แล้วนำมาลดช่วงวัดที่ได้จากการทดสอบฮาร์ดไดรฟ์ จากนั้นนำช่วงเชื่อมั่นของผลทำนายนั้นมาลดช่วงการวัดค่ากระแสไฟฟ้าเขียน โดยการศึกษา และเปรียบเทียบใช้ตัวแบบจำลองการถดถอยโดยความไม่แน่นอน ได้แก่ NGBoost, XGB-Distribution และ CatBoost ซึ่งผลลัพธ์ของงานวิทยานิพนธ์คือ CatBoost สามารถลดเวลาในการทดสอบวัคค่ากระแสไฟฟ้าเขียนสูงสุดที่ช่วงความเชื่อมั่นของผลทำนาย ณ ระดับความเชื่อมั่นที่ยอมรับได้ ซึ่งครอบคลุมสัดส่วน 0.9 ของทุกชุดการทดสอบ
การวิเคราะห์ความคงทนของตัวแบบการเรียนรู้เชิงลึกต่อการโจมตีแบบพอยซันนิ่งแบบแกนส์ในงานภาพทางการแพทย์, ภาคภูมิ สิงขรภูมิ
การวิเคราะห์ความคงทนของตัวแบบการเรียนรู้เชิงลึกต่อการโจมตีแบบพอยซันนิ่งแบบแกนส์ในงานภาพทางการแพทย์, ภาคภูมิ สิงขรภูมิ
Chulalongkorn University Theses and Dissertations (Chula ETD)
ปัจจุบันเทคโนโลยี deep learning ได้เข้ามีส่วนช่วยในการพัฒนางานทางด้านการแพทย์และสาธารณสุขเป็นอย่างมาก ด้วยการใช้สถาปัตยกรรมที่ล้ำสมัยและพารามิเตอร์ที่ถูกสอนด้วยข้อมูลขนาดใหญ่ แต่ทว่า model เหล่านี้สามารถถูกโจมตีได้ด้วย adversarial attack เพราะว่า model เหล่านี้ยังต้องพึ่งพารามิเตอร์ในการสร้างเอาต์พุตและลักษณะที่ไม่สามารถอธิบายได้ของ model นั้นก็ทำให้ยากที่จะหาทางแก้หากถูกโจมตีแล้ว ในทุกๆวันมีการใช้ model เหล่านี้เยอะมากขึ้นเพื่อช่วยบุคลากรทางการแพทย์ แต่ด้วยงานที่ต้องคำนึงถึงชีวิตของผู้คนเป็นหลักการทดสอบความปลอดภัยและความคงทนของตัว model จึงจำเป็น การโจมตีสามารถแบ่งได้ออกเป็นสองประเภทคือ evasion atttack และ poisoning attack ที่มีความยืดหยุ่นกว่า evasion attack ทั้งในเรื่องของการสร้างข้อมูลแปลกปลอมใหม่ขึ้นมาและวิธีการโจมตีทำให้การทดสอบความคงทนต่อ poisoning attack ในงานทางการแพทย์นั้นสำคัญเป็นอย่างยิ่ง วิทยานิพนธ์ฉบับนี้ศึกษาความคงทนของ deep learning model ที่มีสถาปัตยกรรมล้ำสมัยที่ถูกพัฒนามาเพื่องานจำแนกภาพเอกซเรย์ปอดแบบไบนารีภายใต้การโจมตีแบบ poisoninng attack การโจมตีนั้นจะใช้ GANs ในการสร้างข้อมูลสังเคราะห์ปลอมขึ้นมาและทำการติดป้ายกำกับที่ผิดให้ในรูปแบบของ black box และใช้ปริมาณของตัววัดที่ลดลงเมื่อนำข้อมูลนี้ไปอัพเดท model เป็นตัวบ่งชี้ถึงคความคงทนของแต่ละสถาปัตยกรรมที่่แตกต่างกันออกไป จากการทดลองเราพบว่าสถาปัตยกรรม ConvNext นั้นมีความคงทนมากที่สุดและอาจจะสื่อได้ว่าเทคโนโลยีที่มาจาก Transformer นั้นมีส่วนช่วยสนับสนุนความคงทนของ model
ตัวแบบการเรียนรู้ของเครื่องอิทธิพลผสมสำหรับการวิเคราะห์การรอดชีพเวลาไม่ต่อเนื่อง, มนัสพร ตรีรุ่งโรจน์
ตัวแบบการเรียนรู้ของเครื่องอิทธิพลผสมสำหรับการวิเคราะห์การรอดชีพเวลาไม่ต่อเนื่อง, มนัสพร ตรีรุ่งโรจน์
Chulalongkorn University Theses and Dissertations (Chula ETD)
การวิเคราะห์การรอดชีพไม่ต่อเนื่องจะศึกษาบนข้อมูลตามยาวซึ่งชุดข้อมูลตามยาวมักถูกจัดเก็บเป็นตารางโดยข้อมูลแต่ละแถวแสดงถึงการจัดเก็บข้อมูลของบุคคลหนึ่ง ณ เวลาหนึ่งๆ ดังนั้น ข้อมูลจากบุคคลเดียวกันจึงประกอบไปด้วยข้อมูลหลายแถวซึ่งมีความสัมพันธ์กัน การใช้อัลกอริทึมการเรียนรู้ของเครื่องสำหรับการวิเคราะห์ชุดข้อมูลดังกล่าวมักมองข้ามความสัมพันธ์ของข้อมูลที่เกิดจากคนเดียวกัน แต่จะสมมติว่าข้อมูลแต่ละแถวเป็นอิสระต่อกัน งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาการวิเคราะห์การรอดชีพไม่ต่อเนื่องโดยเปรียบเทียบผลลัพธ์จากการพิจารณาความสัมพันธ์ของข้อมูลระหว่างบุคคลคนเดียวกัน โดยใช้ตัวแบบการสุ่มป่าไม้, CatBoost และโครงข่ายประสาทเทียม ที่พิจารณาเฉพาะอิทธิพลคงที่ และตัวแบบการเรียนรู้ของเครื่องอิทธิพลผสมที่พิจารณาทั้งอิทธิพลคงที่และอิทธิพลสุ่ม เพื่อพยากรณ์การเกิดเหตุการณ์บนข้อมูลการรอดชีพ 2 ชุด คือ ข้อมูลท่อน้ำดีอักเสบปฐมภูมิ และข้อมูลการคัดกรองและผลการคัดกรองโรคเบาหวานของประชากรไทย ซึ่งเป็นข้อมูลที่ขาดความสมดุลสูง ผลการศึกษาพบว่าสำหรับตัวแบบอิทธิพลคงที่ การพิจารณาความสัมพันธ์ของข้อมูลระหว่างบุคคลคนเดียวกันให้ประสิทธิภาพการพยากรณ์ที่ดีขึ้นเฉพาะเมื่อใช้ตัวแบบ CatBoost ในขณะที่ตัวแบบอิทธิพลผสมไม่ได้ให้ประสิทธิภาพการพยากรณ์ที่ดีขึ้นเสมอไปเมื่อเทียบกับตัวแบบที่พิจารณาเฉพาะอิทธิพลคงที่ โดยสรุป งานวิจัยนี้ได้แสดงให้เห็นว่าการพิจารณาความสัมพันธ์ของข้อมูลไม่ได้ส่งผลให้ประสิทธิภาพการพยากรณ์ดีขึ้นเสมอไป ทั้งบนตัวแบบอิทธิพลคงที่และตัวแบบอิทธิพลผสม ขึ้นอยู่ข้อจำกัดและปัจจัยต่างๆ เช่น ลักษณะข้อมูล ตัวแบบ การกำหนดตัวแปรอิทธิพลสุ่ม และวิธีการสกัดอิทธิพลคงที่จากตัวแบบ อย่างไรก็ตาม การใช้ตัวแบบอิทธิพลผสมร่วมกับการเรียนรู้ของเครื่องเป็นอีกหนึ่งวิธีการที่น่าลอง และสามารถทำให้ประสิทธิภาพการทำงานดีขึ้นจากการใช้เทคนิคการเรียนรู้ของเครื่องเพียงอย่างเดียว
การเปรียบเทียบวิธีการคัดเลือกตัวแปรสำหรับการถดถอยโลจิสติกในข้อมูลที่มีมิติสูง, รัชพงศ์ ปรัชญาเศรษฐ
การเปรียบเทียบวิธีการคัดเลือกตัวแปรสำหรับการถดถอยโลจิสติกในข้อมูลที่มีมิติสูง, รัชพงศ์ ปรัชญาเศรษฐ
Chulalongkorn University Theses and Dissertations (Chula ETD)
Regularization เป็นวิธีการป้องกันปัญหา overfitting ด้วยการเพิ่มฟังก์ชันการลงโทษไปในตัวแบบเพื่อให้เกิดการคัดกรองตัวแปรเข้าสู่ตัวแบบ งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาและเปรียบเทียบประสิทธิภาพของวิธีการคัดกรองตัวแปรสำหรับการวิเคราะห์การถดถอยโลจิสติกในข้อมูลที่มีมิติสูง ด้วยการใช้ฟังก์ชันการลงโทษในรูปแบบ (1) L0-regularization (2) L1-regularization (3) L0L2-regularization การวิจัยนี้ใช้การจำลองข้อมูลเพื่อทำการทดสอบ 18 กรณี โดยกำหนดค่าที่ต่างกันประกอบด้วย จำนวนตัวแปรอิสระมีจำนวน 200, 500 และ 1000 ตัวแปร ความสัมพันธ์ของตัวแปรอิสระมีค่าเท่ากับ 0, 0.5 และ 0.9 อัตราส่วนสัญญาณต่อสัญญาณรบกวนมีค่าเท่ากับ 1 และ 6 โดยจำลองข้อมูลแต่ละกรณีจำนวน 100 ชุด ในการศึกษานี้มุ่งเน้นที่การเปรียบเทียบประสิทธิภาพในการคัดกรองตัวแปรของตัวแบบ และประสิทธิภาพในการทำนายของตัวแบบ ซึ่งเปรียบเทียบประสิทธิภาพในแต่ละวิธีด้วย ความผิดพลาดในการตรวจจับเชิงบวก ค่าเฉลี่ยแบบฮาร์โมนิคของค่าความแม่นยำและค่าความไว และ พื้นที่ใต้เส้นโค้ง จากการศึกษาพบว่าวิธี L0 มีความแม่นยำในการคัดกรองตัวแปรมากที่สุดเมื่อพิจารณาด้วยความผิดพลาดในการตรวจจับเชิงบวก เมื่อพิจารณาด้วยค่าเฉลี่ยแบบฮาร์โมนิคของค่าความแม่นยำและค่าความไว พบว่าวิธี L1 และ L0L2 มีประสิทธิภาพในการคัดกรองตัวแปรใกล้เคียงกัน แต่วิธี L0L2 จะมีประสิทธิภาพสูงกว่าเมื่อความสัมพันธ์ระหว่างตัวแปรอิสระมีค่าสูง และเมื่อพิจารณาประสิทธิภาพในการทำนายของตัวแบบด้วยพื้นที่ใต้เส้นโค้ง พบว่าวิธี L1 จะมีประสิทธิภาพสูงที่สุดในทุกกรณี