Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,804 Full-Text Articles 23,873 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,804 full-text articles. Page 86 of 486.

Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles 2024 University of New Hampshire, Durham

Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles

Honors Theses and Capstones

With the analytics revolution in sports in the past 20 years, it seems that everything that can be quantified is. In basketball though, trying to break the game down into a set of numbers comes with a unique problem. While we've come up with a good set of advanced numbers to measure offensive efficiency, defense is fundamentally harder to quantify. The game is played five on five, but it has often been popular or convenient to model defense as a set of five one on one games. As defenses became more complex into the 2010s, this methodology became more insignificant. …


Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb 2024 University of New Hampshire, Durham

Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb

Honors Theses and Capstones

In this paper, I explore some of the ways in which artificial intelligence might enhance the sentencing process through recidivism prediction technology. Notably, this technology can increase the accuracy of risk predictions and the speed with which sentencing decisions are reached. I then show, however, that the recidivism prediction technology is likely to turn into what data scientist Cathy O’Neil calls a Weapon of Math Destruction. The potential harmfulness of this technology is due not to the inherent nature of the technology, but the symbiotic relationship it will have with our already harmful criminal justice system. I argue that the …


Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe 2024 University of Central Florida

Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe

Data Science and Data Mining

Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …


การศึกษาเปรียบเทียบแบบจำลองโครงข่ายปฏิปักษ์เชิงกำเนิดในการสร้างภาพความละเอียดสูงเพื่อการตรวจจับวัตถุขนาดเล็ก, ภัทรชนน สกุลคู 2024 คณะพาณิชยศาสตร์และการบัญชี

การศึกษาเปรียบเทียบแบบจำลองโครงข่ายปฏิปักษ์เชิงกำเนิดในการสร้างภาพความละเอียดสูงเพื่อการตรวจจับวัตถุขนาดเล็ก, ภัทรชนน สกุลคู

Chulalongkorn University Theses and Dissertations (Chula ETD)

การตรวจจับวัตถุขนาดเล็ก (Small Object Detection) เป็นหนึ่งในความท้าทายด้านคอมพิวเตอร์วิทัศน์ เนื่องจากภาพที่มีความละเอียดต่ำมักมีข้อจำกัดด้านการระบุขอบเขตและรายละเอียดของวัตถุ งานวิจัยนี้มุ่งเน้นการเปรียบเทียบประสิทธิภาพของแบบจำลอง Super-Resolution GAN คือ SRGAN ESRGAN Real-ESRGAN และ BSRGAN กับวิธีการสร้างภาพความละเอียดสูงแบบดั้งเดิม คือ Bilinear และ Bicubic เพื่อศึกษาว่าการเพิ่มความละเอียดของภาพสามารถช่วยให้การตรวจจับวัตถุขนาดเล็กมีความแม่นยำสูงขึ้นหรือไม่ โดยการทดลองดำเนินการกับภาพต้นฉบับขนาด 160 × 160 พิกเซล และสร้างภาพความละเอียดสูงขนาด 640 × 640 พิกเซลก่อนนำไปประเมินผลการตรวจจับวัตถุผ่านตัวชี้วัด mAP@50 และ mAP@50-95 ในสามชุดข้อมูล ได้แก่ ภาพสายเคเบิล Microglia และวัชพืช Ridderzuring ผลการทดลองพบว่า SRGAN ให้ค่า mAP@50 สูงสุดในทุกชุดข้อมูล ในขณะที่ Real-ESRGAN และ BSRGAN มีค่าต่ำกว่าวิธีอื่นในบางกรณี อย่างไรก็ตาม แม้ว่า SRGAN จะมีความแม่นยำสูงสุด แต่ใช้เวลาในการสร้างภาพมากกว่าวิธี Bicubic ประมาณ 4-7 เท่า ทำให้ต้องพิจารณาความสมดุลระหว่างความเร็วและความแม่นยำ ทั้งนี้ ภาพต้นฉบับความละเอียดสูง (HR) ยังคงให้ค่าคะแนนสูงสุดในทุกชุดข้อมูล ซึ่งสะท้อนว่าการใช้ Super-Resolution สามารถลดช่องว่างระหว่างภาพความละเอียดต่ำและภาพต้นฉบับได้ และสามารถนำไปประยุกต์ใช้กับงานเกี่ยวกับการตรวจจับวัตถุจากภาพความละเอียดต่ำได้จริง


Parameter Tuning Of Information Directed Sampling In Credit Scoring Problems Under Ungeneralizable Contextual Logistic Bandit Model, Sorawit Panjapiyakul 2024 Faculty of Commerce and Accountancy

Parameter Tuning Of Information Directed Sampling In Credit Scoring Problems Under Ungeneralizable Contextual Logistic Bandit Model, Sorawit Panjapiyakul

Chulalongkorn University Theses and Dissertations (Chula ETD)

This study investigates the tuning parameter of Information Directed Sampling (IDS) in credit scoring problems under an ungeneralizable contextual logistic bandit framework. Decision-making scenarios, such as credit scoring and underwriting, involve balancing the tradeoff between exploration and exploitation, which is essential for optimizing learning efficiency in the environment while minimizing costs. The IDS algorithm offers a principled framework that leverages mutual information to enhance the decision-making process. However, its learning performance is dependent on the tuning parameter, denoted as gamma, which plays a key role in balancing information gain and expected regret. Extensive simulation experiments were conducted to identify the …


การศึกษาเปรียบเทียบการแทนค่าน้ำหนักสูญหายด้วยวิธีการเรียนรู้เชิงลึกกับวิธีดั้งเดิมด้วยวิธีการจำลองข้อมูลในบริบทของผู้ป่วยในและหอผู้ป่วยหนักภายในโรงพยาบาล, เมธัส ม่วงนาค 2024 คณะพาณิชยศาสตร์และการบัญชี

การศึกษาเปรียบเทียบการแทนค่าน้ำหนักสูญหายด้วยวิธีการเรียนรู้เชิงลึกกับวิธีดั้งเดิมด้วยวิธีการจำลองข้อมูลในบริบทของผู้ป่วยในและหอผู้ป่วยหนักภายในโรงพยาบาล, เมธัส ม่วงนาค

Chulalongkorn University Theses and Dissertations (Chula ETD)

การสูญหายของข้อมูลในเวชระเบียนโรงพยาบาล โดยเฉพาะในหอผู้ป่วยหนัก (ICU) และหอผู้ป่วยใน (IPD) เป็นปัญหาที่พบบ่อยและส่งผลต่อการดูแลผู้ป่วยและความถูกต้องของงานวิจัย การศึกษานี้เปรียบเทียบวิธีการแทนค่าสูญหาย 10 วิธี ได้แก่ เทคนิคแบบดั้งเดิม (Mean, Median, k-NN, MICE, MissForest), วิธีแบบผสม (HyperImpute) และวิธีแบบการเรียนรู้เชิงลึกหรือ DL (MLPRegressor, AEImputer, MIWAE, GAIN) โดยใช้ข้อมูลจำลอง 63 ตัวแปร ภายใต้เงื่อนไขควบคุม ได้แก่ ขนาดตัวอย่าง 3 ระดับ (5,000, 25,000, 50,000), กลไกการสูญหาย 3 รูปแบบ (MCAR, MAR, MNAR) และอัตราการสูญหาย 6 ระดับ (10% ถึง 60%) โดยประเมินผลด้วย RMSE, MAPE, เวลาในการประมวลผล และการใช้หน่วยความจำ ผลการศึกษาพบว่า Mean และ Median ยังให้ผลลัพธ์ที่ดี พร้อมความเร็วและใช้ทรัพยากรต่ำ MissForest และ HyperImpute ให้ความแม่นยำที่สมดุลกับประสิทธิภาพ เหมาะกับกรณีที่ข้อมูลขาดในระดับปานกลาง วิธีที่นิยมอย่าง MICE กลับมีข้อจำกัดกับชุดข้อมูลที่ไม่เป็นพาราเมตริก ทำให้ผลลัพธ์ด้อยกว่าในหลายเงื่อนไข ด้าน DL แม้บางวิธีให้ผลลัพธ์ดี แต่ต้องอาศัยการปรับแต่งพารามิเตอร์อย่างละเอียด และใช้ทรัพยากรมาก โดยวิธีกลุ่ม AEs มีความเสถียรที่สุด ส่วน GAIN มีความไวต่อรูปแบบข้อมูลและขนาดตัวอย่าง ให้ผลลัพธ์ไม่สม่ำเสมอ โดยสรุป แม้ DL จะมีศักยภาพ แต่ในหลายกรณี วิธีดั้งเดิมหรือแบบผสมยังคงเป็นทางเลือกที่ใช้งานได้จริงและคุ้มค่า


โครงข่ายประสาทเทียมสำหรับการวิเคราะห์การถดถอยเชิงเส้นตามบริบท, พศุตม์ รัตนศรีมงคล 2024 คณะพาณิชยศาสตร์และการบัญชี

โครงข่ายประสาทเทียมสำหรับการวิเคราะห์การถดถอยเชิงเส้นตามบริบท, พศุตม์ รัตนศรีมงคล

Chulalongkorn University Theses and Dissertations (Chula ETD)

ปัญหาการถดถอยเชิงเส้นตามบริบท คือปัญหาที่ข้อมูลมีโครงสร้างแบ่งเป็นกลุ่ม โดยแต่ละกลุ่มถูกกำหนดโดยตัวแปรบริบท งานวิจัยนี้ศึกษาการนำ ตัวแบบ Contextual Neural Network (CtxtNN) มาใช้วิเคราะห์ปัญหาประเภทนี้ และเปรียบเทียบประสิทธิภาพกับ ตัวแบบพื้นฐานอย่าง Feedforward Neural Networks (FNN) ทั้งโครงสร้างขนาดเล็ก (FNN-Small) และขนาดใหญ่ (FNN-Large) ผ่านการทดลอง 3 กรณี โดยงานวิจัยนี้จะศึกษาเฉพาะปัญหาการถดถอยเชิงเส้นตามบริบท ที่ตัวแปรต้นไม่เกิน 8 ตัว ซึ่งมีตัวแปรเชิงบริบทไม่เกิน 3 ตัว และบริบทข้อมูลไม่เกิน 3 บริบทเท่านั้น โดยจากผลการวิจัยนี้สามารถสรุปได้ว่า ตัวแบบ CtxtNN เป็นตัวแบบที่สามารถเรียนรู้ความสัมพันธ์เชิงเส้นตามบริบทได้อย่างมีประสิทธิภาพ โดยให้ผลลัพธ์ที่มีประสิทธิภาพสูงที่สุดในการแก้ปัญหาการถดถอยเชิงเส้นตามบริบทเมื่อเทียบกับตัวแบบ FNN-Small และตัวแบบ FNN-Large แม้ใช้จำนวนพารามิเตอร์ที่น้อยกว่า จึงเป็นทางเลือกที่น่าสนใจสำหรับงานวิเคราะห์ปัญหาข้อมูลที่มีความสัมพันธ์เชิงบริบทกับผลเฉลย


การประมาณค่าพารามิเตอร์ของระบบที่สามารถซ่อมแซมได้ที่มีระบบย่อยหลายระบบภายใต้ผลกระทบจากการช็อก, ปวิชญา ปรีชา 2024 คณะพาณิชยศาสตร์และการบัญชี

การประมาณค่าพารามิเตอร์ของระบบที่สามารถซ่อมแซมได้ที่มีระบบย่อยหลายระบบภายใต้ผลกระทบจากการช็อก, ปวิชญา ปรีชา

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อประมาณค่าพารามิเตอร์ของระบบที่สามารถซ่อมแซมได้ภายใต้ผลกระทบจากการช็อก (shock effect) โดยจำลองระบบที่มีองค์ประกอบย่อยสององค์ประกอบ ซึ่งอาจมีความสัมพันธ์กัน การจำลองข้อมูลใช้แบบจำลอง Power Law Process (PLP) พร้อมกำหนดระดับ shock effect แบบคูณในช่วง 1.00–1.03 เพื่อสะท้อนอิทธิพลระหว่างองค์ประกอบ จากนั้นทำการประมาณค่าพารามิเตอร์ด้วยวิธี Maximum Likelihood Estimation (MLE) ภายใต้สี่รูปแบบ ได้แก่ 1) องค์ประกอบ มีความสัมพันธ์กัน 2) มีความสัมพันธ์และทราบพารามิเตอร์ขนาด 3) ไม่มีความสัมพันธ์กัน และ 4) ไม่มีความสัมพันธ์และทราบพารามิเตอร์ขนาด โดยศึกษาในกรณีที่จำนวนการล้มเหลวสูงสุด n = 25, 35 และ 50 ผลการวิเคราะห์พบว่า พารามิเตอร์ k มีแนวโน้มถูกประมาณสูงเกินจริงเมื่อข้อมูลมีจำนวนน้อย ขณะที่ shock effect สามารถประมาณได้ใกล้เคียงค่าจริงมากขึ้นเมื่อข้อมูลเพิ่มขึ้น การประเมินความแม่นยำ ของการทำนายเวลาเกิดเหตุการณ์ถัดไปใช้เปอร์เซ็นต์ไทล์ (P5–P95) และค่าความคลาดเคลื่อนสัมบูรณ์เฉลี่ย (MAPE) พบว่า วิธี องค์ประกอบไม่มีความสัมพันธ์และทราบพารามิเตอร์ขนาด ให้ผลดีที่สุดเมื่อ shock ต่ำ (1.00–1.01) โดยเฉพาะเมื่อข้อมูลจำกัด ส่วนวิธีองค์ประกอบมีความสัมพันธ์และทราบพารามิเตอร์ขนาด เหมาะสมกว่าเมื่อ shock สูง (1.02–1.03) และข้อมูลมาก นอกจากนี้ ทุกวิธีมีแนวโน้มการกระจายเบ้ขวา โดยเฉพาะวิธีที่ไม่ทราบค่าพารามิเตอร์ขนาดซึ่งเบ้รุนแรงกว่าวิธีอื่น ในขณะที่วิธี องค์ประกอบไม่มีความสัมพันธ์และทราบพารามิเตอร์ขนาด ให้การกระจายแคบและเบ้น้อยที่สุด สะท้อนความเสถียรและความต้านทาน ต่อ outlier


Effect Of Data Visualization On Users' Running Performance On Treadmill, Thanaphon Amattayakul 2024 Faculty of Commerce and Accountancy

Effect Of Data Visualization On Users' Running Performance On Treadmill, Thanaphon Amattayakul

Chulalongkorn University Theses and Dissertations (Chula ETD)

This study investigates the effect of real-time data visualization on user performance and experience during treadmill running. Traditional treadmill displays usually present information in plain text, which may limit user engagement and motivation. To address this, the study introduced redesigned displays using data visualization techniques aligned with human perception, such as line graphs and progress bars, to make performance feedback more meaningful and easier to understand. The experiment compared three display conditions: a traditional treadmill display and two improved designs. A within-subjects design was used with 18 participants. Performance metrics such as time to exhaustion, heart rate, distance covered, and …


Enhanced Realism In Virtual Try-On Tasks Using Diffusion Methods, Saris Kiattithapanayong 2024 Faculty of Commerce and Accountancy

Enhanced Realism In Virtual Try-On Tasks Using Diffusion Methods, Saris Kiattithapanayong

Chulalongkorn University Theses and Dissertations (Chula ETD)

Virtual try-on technology is revolutionizing online retail by enabling customers to visualize garments on their bodies before purchasing. Traditional methods, often based on Generative Adversarial Networks (GANs), face challenges such as misalignment and visual artifacts, especially in complex poses. We present a virtual try-on framework leveraging diffusion models to enhance realism, accuracy, and garment detail preservation. Our approach integrates Vector Quantized Variational Autoencoders (VQ-VAEs) for precise feature matching within a diffusion U-Net architecture. By adopting image-based conditioning with the CLIP image encoder, our system utilizes visual features directly from clothing images for more faithful garment representations. Additionally, an Additional Feature …


ประสิทธิภาพการพยากรณ์ของการวิเคราะห์เชิงฟังก์ชันและการเรียนรู้เชิงลึก, บุณฑริกา พรหมสถิตย์ 2024 คณะพาณิชยศาสตร์และการบัญชี

ประสิทธิภาพการพยากรณ์ของการวิเคราะห์เชิงฟังก์ชันและการเรียนรู้เชิงลึก, บุณฑริกา พรหมสถิตย์

Chulalongkorn University Theses and Dissertations (Chula ETD)

การวิจัยนี้เปรียบเทียบประสิทธิภาพการพยากรณ์ของตัวแบบ Functional Principal Component Regression (FPCR) กับตัวแบบการเรียนรู้เชิงลึก (Recursive Neural Network: RNN, Long Short-term Memory: LSTM, Gated Recurrent Unit: GRU) ภายใต้สถานการณ์ที่มีความผันผวนแตกต่างกัน โดยใช้ชุดข้อมูลที่มีความผันผวนต่ำ (อุณหภูมิเฉลี่ยรายวัน), ปานกลาง (ปริมาณ PM 2.5 รายชั่วโมง) และสูง (อัตราการแลกเปลี่ยน Bitcoin รายนาที) ศึกษาการพยากรณ์ระยะสั้น ระยะกลาง และระยะยาว โดยใช้ตัวชี้วัด Mean Squared Error (MSE) และ Mean Integrated Squared Error (MISE) ผลการศึกษาพบว่า สำหรับข้อมูลที่มีความผันผวนต่ำ FPCR ให้ผลลัพธ์ที่แม่นยำกว่าตัวแบบการเรียนรู้เชิงลึก โดยเฉพาะในการพยากรณ์ระยะกลางและระยะยาว ในทางกลับกัน สำหรับข้อมูลที่มีความผันผวนสูง FPCR เหนือกว่าการเรียนรู้เชิงลึกเฉพาะการพยากรณ์ระยะกลางเท่านั้น สำหรับข้อมูลที่มีความผันผวนปานกลางและสูง ขนาดของชุดข้อมูลฝึกไม่มีผลกระทบอย่างชัดเจนต่อประสิทธิภาพของตัวแบบทั้งสอง อย่างไรก็ตาม ในกรณีของข้อมูลที่มีความผันผวนต่ำ เมื่อมีชุดข้อมูลขนาดใหญ่ ตัวแบบการเรียนรู้เชิงลึกให้ผลลัพธ์ที่แม่นยำกว่า FPCR ในขณะที่ FPCR มีความแม่นยำสูงกว่าหากใช้ข้อมูลจำนวนน้อยและทำการพยากรณ์ในระยะกลางถึงระยะยาว นอกจากนี้ ในการพยากรณ์ระยะสั้น FPCR มักให้ผลลัพธ์ที่ด้อยกว่าตัวแบบการเรียนรู้เชิงลึกในทุกกรณี


การเปรียบเทียบวิธีการใส่ค่าสูญหายสำหรับอนุกรมเวลาเชิงพหุ กรณีศึกษาดัชนีราคากลุ่มอุตสาหกรรมตลาดหลักทรัพย์แห่งประเทศไทย, พงษ์พล ยิ่งประทานพร 2024 คณะพาณิชยศาสตร์และการบัญชี

การเปรียบเทียบวิธีการใส่ค่าสูญหายสำหรับอนุกรมเวลาเชิงพหุ กรณีศึกษาดัชนีราคากลุ่มอุตสาหกรรมตลาดหลักทรัพย์แห่งประเทศไทย, พงษ์พล ยิ่งประทานพร

Chulalongkorn University Theses and Dissertations (Chula ETD)

การศึกษานี้มีวัตถุประสงค์เพื่อเปรียบเทียบวิธีการใส่ค่าสูญหายสำหรับอนุกรมเวลาเชิงพหุ และประเมินผลเพื่อเลือกวิธีการใส่ค่าสูญหายที่เหมาะสมที่สุดสำหรับอนุกรมเวลาเชิงพหุ โดยใช้ข้อมูลทุติยภูมิดัชนีราคากลุ่มอุตสาหกรรมของตลาดหลักทรัพย์แห่งประเทศไทย 8 กลุ่มอุตสาหกรรม จากฐานข้อมูล SETSMART ตั้งแต่วันที่ 1 มกราคม พ.ศ. 2547 ถึง 1 มกราคม พ.ศ. 2567 รวมทั้งสิ้น 4877 วัน ซึ่งได้มีการกำหนดรูปแบบการสูญหายออกเป็น 3 รูปแบบ ได้แก่ การสูญหายรูปแบบสุ่ม การสูญหายรูปแบบช่วงตามลำดับ และการสูญหายรูปแบบบล็อก และกำหนดสัดส่วนการสูญหายของข้อมูลที่ร้อยละ 5 10 20 30 40 และ 50 ตามลำดับ โดยจำแนกวิธีการใส่ค่าสูญหายออกเป็น 3 กลุ่ม ได้แก่ การใส่ค่าสูญหายด้วยวิธีการเชิงสถิติ ประกอบไปด้วย ค่าเฉลี่ย ค่ามัธยฐาน ข้อมูลสุดท้ายก่อนการสูญหาย (LOCF) ข้อมูลล่าสุดหลังการสูญหาย (NOCB) และการประมาณค่าช่วงเส้นตรง (Linear Interpolation) การใส่ค่าสูญหายด้วยวิธีการเรียนรู้ของเครื่อง ประกอบไปด้วย ค่าคาดหวังสูงที่สุด (EM) การใส่ค่าสูญหายด้วยการทดแทนแบบพหุคูณด้วยสมการลูกโซ่ (MICE) เพื่อนบ้านใกล้เคียงที่สุด (KNN) และป่าสุ่ม (Random Forest) และการใส่ค่าสูญหายด้วยวิธีการเรียนรู้เชิงลึก ประกอบไปด้วย GP-VAE USGAN และ SAITS นอกจากนี้ผู้วิจัยใช้ค่ารากที่สองของค่าความคลาดเคลื่อนกำลังสองโดยเฉลี่ย (RMSE) ค่าความคลาดเคลื่อนสัมบูรณ์โดยเฉลี่ย (MAE) และค่าร้อยละความคลาดเคลื่อนสัมบูรณ์โดยเฉลี่ย (MAPE) ในการวัดประสิทธิภาพการใส่ค่าสูญหาย ผลการศึกษาพบว่าที่รูปแบบการสูญหายทั้ง 3 รูปแบบ และสัดส่วนการสูญหายที่น้อยกว่าร้อยละ 50 การใสค่าสูญหายด้วยการประมาณค่าช่วงเส้นตรง (Linear Interpolation) มีประสิทธิภาพสูงที่สุด ในขณะที่สัดส่วนการสูญหายร้อยละ 50 ของรูปแบบการสูญหายทั้ง 3 รูปแบบ การใส่ค่าสูญหายด้วยวิธีป่าสุ่ม (Random Forest) มีประสิทธิภาพสูงที่สุด


ประสิทธิภาพของแบบจำลองแบบผสมของการเรียนรู้เชิงลึกสำหรับการพยากรณ์ราคาหุ้น, กิตติคุณ ทัดประดิษฐ 2024 คณะพาณิชยศาสตร์และการบัญชี

ประสิทธิภาพของแบบจำลองแบบผสมของการเรียนรู้เชิงลึกสำหรับการพยากรณ์ราคาหุ้น, กิตติคุณ ทัดประดิษฐ

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อวิเคราะห์และเปรียบเทียบประสิทธิภาพและระยะเวลาที่ใช้ของแบบจำลองการเรียนรู้เชิงลึกแบบผสม (Hybrid Deep Learning Models) ได้แก่ RNN, LSTM และ GRU ในรูปแบบการใช้โครงสร้างแบบผสม (Stacked layer) และเครือข่ายประสาทเทียมแบบซ้อนกัน (Cascaded neural network) ในการพยากรณ์ราคาปิดหุ้น รวมถึงศึกษาผลกระทบของการสลับลำดับของแบบจำลองภายใน Hybrid model เพื่อพิจารณาความแตกต่างของประสิทธิภาพการพยากรณ์ ทั้งในระยะสั้น 7 วัน และ ระยะยาว 30 วัน โดยมีการใช้ข้อมูลราคาหุ้นทั้งหมด 5 อุตสาหกรรม เลือกกลุ่มอุตสาหกรรมละ 3 หุ้นตามระดับความผันผวนของราคาหุ้นเมื่อเทียบกับตลาด (Beta) รวมทั้งหมด 15 ชุดข้อมูล ผลการศึกษาพบว่า ในภาพรวมการพยากรณ์ระยะสั้น 7 วันและระยะยาว 30 วัน การสร้างแบบจำลองผสมด้วยวิธี Cascaded neural network มีประสิทธิภาพที่ไม่แตกต่างกับวิธี Stacked layers อย่างมีนัยสำคัญ แต่ถ้าหากพิจารณาด้วยระยะเวลาที่ใช้ในการสร้างแบบจำลองจะพบว่า การสร้างแบบจำลองผสมด้วยวิธี Stacked layers ใช้เวลาสร้างแบบจำลองน้อยกว่าวิธี Cascaded neural network อย่างมีนัยสำคัญ ซึ่งจะช่วยประหยัดทรัพยากรในการสร้างแบบจำลองในการพยากรณ์เป็นอย่างมาก นอกจากนี้การสลับลำดับของแบบจำลองที่ใช้ในการผสมแบบจำลอง ด้วยวิธี Cascaded neural network และ Stacked layers ในการพยากรณ์ระยะสั้น 7 วันและระยะยาว 30 วัน ไม่ส่งผลต่อประสิทธิภาพของ Hybrid Deep Learning Models อย่างมีนัยสำคัญ จึงสามารถสรุปได้ว่าการสร้าง Hybrid Deep Learning Models ด้วยวิธี Stacked layers ให้ประสิทธิภาพและความคุ้มค่าที่มากกว่าการผสมด้วยวิธี Cascaded neural network ทั้งในระยะสั้นและระยะยาว


การประเมินประสิทธิภาพของการเรียนรู้เชิงรุกโดยวิธีการเลือกแบบละโมบและวิธีการสุ่มตัวอย่างของทอมป์สันกับข้อมูลรูปแบบข้อความ, อานนท์ พรหมจรรย์ 2024 คณะพาณิชยศาสตร์และการบัญชี

การประเมินประสิทธิภาพของการเรียนรู้เชิงรุกโดยวิธีการเลือกแบบละโมบและวิธีการสุ่มตัวอย่างของทอมป์สันกับข้อมูลรูปแบบข้อความ, อานนท์ พรหมจรรย์

Chulalongkorn University Theses and Dissertations (Chula ETD)

การวิจัยนี้มุ่งเน้นการประเมินประสิทธิภาพของวิธีการเรียนรู้เชิงรุกสำหรับการคัดเลือกข้อมูลที่มีประโยชน์สูงสุดในการทำป้ายกำกับ โดยเปรียบเทียบวิธีการเลือกข้อมูลแบบสุ่ม วิธีการเลือกข้อมูลโดยวิธีละโมบ และวิธีการสุ่มตัวอย่างของทอมป์สันด้วยการประมาณค่าแบบลาพลาซ ผ่านการทดลอง 100 รอบในการเลือกทวีตเกี่ยวกับการท่องเที่ยวในกรุงเทพมหานครเพื่อฝึกโมเดลการถดถอยโลจิสติก ผลการทดลองพบว่าวิธีการเลือกข้อมูลโดยวิธีละโมบให้ประสิทธิภาพสูงสุดตลอดการทดลอง เนื่องจากสามารถปรับปรุงโมเดลได้อย่างรวดเร็ว แต่ประสิทธิภาพลดลงในช่วงท้ายเมื่อจำนวนทวีตที่มีประโยชน์ลดลง ขณะที่วิธีการสุ่มตัวอย่างของทอมป์สันด้วยการประมาณค่าแบบลาพลาซใช้เวลาในการคัดเลือกข้อมูลมากที่สุดและมีประสิทธิภาพต่ำกว่าในช่วงแรก อย่างไรก็ตาม เมื่อจำนวนรอบการทดลองเพิ่มขึ้น ความแม่นยำของโมเดลก็ค่อย ๆ ดีขึ้นเมื่อเทียบกับช่วงต้น ส่วนวิธีการเลือกข้อมูลแบบสุ่มใช้เวลาน้อยที่สุด แต่ไม่มีการเรียนรู้หรือปรับปรุงโมเดล ทำให้ประสิทธิภาพไม่ดีขึ้น จากผลการทดลองสามารถสรุปได้ว่าวิธีการเลือกข้อมูลโดยวิธีละโมบเป็นทางเลือกที่มีประสิทธิภาพสูงในสถานการณ์ที่ต้องการการเรียนรู้ที่รวดเร็ว ในขณะที่วิธีการสุ่มตัวอย่างของทอมป์สันด้วยการประมาณค่าแบบลาพลาซยังคงต้องมีการศึกษาเพิ่มเติมเกี่ยวกับศักยภาพในการเรียนรู้ในระยะยาว งานวิจัยนี้สามารถนำไปประยุกต์ใช้กับการวิเคราะห์ข้อมูลข้อความ เช่น การจำแนกประเภทความรู้สึกของผู้ใช้โซเชียลมีเดีย หรือการประเมินความคิดเห็นของลูกค้าในอุตสาหกรรมต่างๆ


The Influence Of Echo Chamber On Thailand's 2023 Election, Isariyaporn Sukcharoenchaikul 2024 Faculty of Commerce and Accountancy

The Influence Of Echo Chamber On Thailand's 2023 Election, Isariyaporn Sukcharoenchaikul

Chulalongkorn University Theses and Dissertations (Chula ETD)

This research develops visual methods to explore the echo chamber effect, employing analytical approaches to investigate its dynamics through a case study on the 2023 General Election in Thailand. The study utilizes visualization techniques such as node-link diagrams, t-SNE projections, and heatmaps to analyze homophilic relationships, clustering tendencies, and polarization within online communities. To minimize inaccuracies and biases, network graphs are constructed based on contextual analysis of user-generated content, rather than relying on predefined relationship definitions (e.g., friendships, followers, retweets) or users' interpretations. The research uses the Echo Chamber Score (ECS) alongside visualizations to explore echo chambers, reveal significant variations …


A Comparative Study Of Early Fusion And Multimodal Siamese Neural Network Using Image And Text Data In Food Classification, Kanokporn Sintarasirikulchai 2024 Faculty of Commerce and Accountancy

A Comparative Study Of Early Fusion And Multimodal Siamese Neural Network Using Image And Text Data In Food Classification, Kanokporn Sintarasirikulchai

Chulalongkorn University Theses and Dissertations (Chula ETD)

The economic development under capitalism has significantly transformed people's lifestyles, resulting in a fast-paced daily life. This shift has increased the consumption of convenient food options, leading to a preference for fast food, which is often high in carbohydrates and fats. Consequently, there has been a rise in obesity and related health issues, highlighting the importance of monitoring food intake. Automated systems utilizing artificial intelligence (AI) have emerged as potent tools for providing personalized dietary advice and monitoring. With the growing volume of food-related content on social media, including images and accompanying text, leveraging multimodal data has become essential for …


Extremal Graphs For Widom–Rowlinson Colorings In K-Chromatic Graphs, John Engbers, Aysel Erey 2024 Marquette University

Extremal Graphs For Widom–Rowlinson Colorings In K-Chromatic Graphs, John Engbers, Aysel Erey

Mathematical and Statistical Science Faculty Research and Publications

The Widom–Rowlinson graph, HWR , is the fully looped path on three vertices. Let hom(G,HWR) be the number of graph homomorphisms from G to HWR or, equivalently, the number of HWR-colorings of G. We investigate extremal graphs for hom(G,HWR) for G in the family of k-chromatic graphs subject to various connectivity requirements. In particular, we determine the graphs G maximizing hom(G,HWR) in the families of n-vertex k-chromatic graphs, n-vertex connected k-chromatic graphs, n-vertex k-chromatic graphs with c components, n …


Integrating Machine Learning With Cure Models And Associated Inference, Wisdom Aselisewine 2024 University of Texas at Arlington

Integrating Machine Learning With Cure Models And Associated Inference, Wisdom Aselisewine

Mathematics Dissertations - Archive

Recent advancements in medical treatments have significantly enhanced the rates of recovery for numerous chronic illnesses. This progress has sparked growing interest in developing suitable statistical models capable of handling survival data that includes substantial cure fractions. The mixture cure model finds extensive application in analyzing survival data when there exists a cured subgroup. Standard logistic regression-based approaches for modeling the incidence part of the mixture cure model may suffer from poor predictive accuracy, especially in the presence of high dimensional covariates and/or non-linear covariate effects. To overcome this limitation, we propose the integration of distinct machine learning algorithms with …


The Performance Of Marginal Modeling Methods For Rare Events With Application To Opioid Overdose Mortality And Morbidity, Shawn Nigam 2024 University of Kentucky

The Performance Of Marginal Modeling Methods For Rare Events With Application To Opioid Overdose Mortality And Morbidity, Shawn Nigam

Theses and Dissertations--Epidemiology and Biostatistics

Opioid misuse is a nationwide epidemic, with Kentucky having one of the highest opioid overdose-related fatality rates across all US states. These rates have increased significantly over the past decade, with particularly large increases during the COVID-19 pandemic. This dissertation aims to study the behavior of these increases and the methods for the marginal modeling of count outcomes related to opioid overdose.

Opioid overdose-related fatality rates in Kentucky increased significantly during the COVID-19 pandemic. In this chapter, we characterize the changes in opioid overdose fatality rates in Kentucky and identify associations between potential factors and fatality rates. County-level opioid overdose …


Variable Selection For High-Dimensional Data With Interaction Effects: Methods, Applications, And Inferences, Leiyue Li 2024 University of Kentucky

Variable Selection For High-Dimensional Data With Interaction Effects: Methods, Applications, And Inferences, Leiyue Li

Theses and Dissertations--Statistics

For high-dimensional data where the number of variables greatly exceeds the number of observations, selecting important variables while maintaining the required heredity conditions can be challenging. This dissertation is structured into three interconnected parts. In the first part, we propose a variable selection method by implementing a well-known optimization technique, the Genetic Algorithm. An R package was developed to simplify the implementation and usage of the proposed method. We then propose another variable selection method by extending the study from the Genetic Algorithm to a different but related optimization technique, Simulated Annealing. We consider three different hierarchical structures in both …


Digital Commons powered by bepress