Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2018

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 571 - 596 of 596

Full-Text Articles in Statistics and Probability

Semiparametric Statistical Estimation And Inference With Latent Information, Qianqian Wang Jan 2018

Semiparametric Statistical Estimation And Inference With Latent Information, Qianqian Wang

Theses and Dissertations

In Chapter 1, we predicted disease risk by transformation models in the presence of missing subgroup identifiers. When a discrete covariate defining subgroup membership is missing for some of the subjects in a study, the distribution of the outcome follows a mixture distribution of the subgroup-specific distributions. Taking into account the uncertain distribution of the group membership and the covariates, we model the relation between the disease onset time and the covariates through transformation models in each sub-population, and develop a nonparametric maximum likelihood based estimation implemented through EM algorithm along with its inference procedure. We further propose methods to …


Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou Jan 2018

Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou

Theses and Dissertations

In this dissertation, we aim to address three important questions in practice, which can be solved through complex survival models. The first project focuses on studying the longitudinal fitness effect on cardiovascular disease (CVD) mortality. In the second project, we study the disease-death relation between CVD and all-cause mortality and evaluate important covariate effects on the disease or death transitions. In the third project, we compare antiretroviral treatment (ART) for HIV patients and consider both treatment effect and side effect of the drugs. The first two projects are motivated by the Aerobics Center Longitudinal Study (ACLS) datasets and the third …


Classification Of High-Dimensional Data Based On Multiple Testing Methods, Chong Ma Jan 2018

Classification Of High-Dimensional Data Based On Multiple Testing Methods, Chong Ma

Theses and Dissertations

Supervised and unsupervised classification are common topics in machine learning in both scientific and industrial fields, which usually involve three tasks: prediction, exploration, and explanation. False discovery rate (FDR) theory has a close connection to classical classification theory, which must be employed in a sophisticated way to achieve good performance in various contexts. The study aims to explore novel supervised classifiers and unsupervised classification approaches for functional data and high-dimensional data in genome study by using FDR, respectively. One work develops a novel classifier for functional data by casting the classification problem into a multiple testing task, which involves using …


การเปรียบเทียบวิธีการหาจุดเปลี่ยนแปลงแบบออฟไลน์ในข้อมูลอนุกรมเวลาที่มีความผันแปรไม่ปกติที่มีการแจกแจงปกติ 2 ตัวแปร, ปภาวิน เจริญชัยปิยกุล Jan 2018

การเปรียบเทียบวิธีการหาจุดเปลี่ยนแปลงแบบออฟไลน์ในข้อมูลอนุกรมเวลาที่มีความผันแปรไม่ปกติที่มีการแจกแจงปกติ 2 ตัวแปร, ปภาวิน เจริญชัยปิยกุล

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพในการตรวจหาจุดเปลี่ยนแปลงของวิธี E-Divisive, e-cp3o และ ks-cp3o สำหรับข้อมูลอนุกรมเวลาที่มีความผันแปรไม่ปกติที่มีการแจกแจงปกติ 2 ตัวแปรซึ่งมีการเปลี่ยนแปลงในค่าเฉลี่ย ค่าความแปรปรวน หรือค่าสหสัมพันธ์ โดยทำการจำลองและเปรียบเทียบประสิทธิภาพโดยใช้ค่า Adjusted Rand Index และเปรียบเทียบจำนวนและตำแหน่งของจุดเปลี่ยนแปลงทั้งสามวิธีในข้อมูลจริง โดยข้อมูลเป็นข้อมูลสัญญาณชีพและข้อมูลปริมาณฝุ่นละออง PM2.5 และเปรียบเทียบการเปลี่ยนแปลงของแต่ละช่วงโดยใช้การทดสอบความแตกต่างของค่าเฉลี่ยและความแปรปรวนของสองประชากร จากการศึกษาพบว่า เมื่อข้อมูลมีขนาดเล็ก (n = 90) วิธี e-cp3o และ ks-cp3o มีประสิทธิภาพในการตรวจหาจุดเปลี่ยนแปลงมากที่สุด (ค่า Adjusted Rand Index มีค่าเป็น 1) และพบว่าวิธี E-Divisive มีประสิทธิภาพสูงเฉพาะกรณีที่ข้อมูลมีการเปลี่ยนแปลงในค่าเฉลี่ย (ค่า Adjusted Rand Index มีค่าเข้าใกล้ 1) เมื่อข้อมูลมีขนาดมากขึ้น (n = 150 และ 300) วิธี E-Divisive มีประสิทธิภาพสูงที่สุดเมื่อข้อมูลมีการเปลี่ยนแปลงค่าเฉลี่ย และวิธี e-cp3o มีประสิทธิภาพสูงที่สุดในกรณีอื่น ๆ การศึกษากับข้อมูลจริงซึ่งเป็นข้อมูลสัญญาณชีพ และข้อมูลปริมาณฝุ่นละออง PM2.5 ซึ่งทำการตรวจหาจุดเปลี่ยนแปลงทั้งหมดในคราวเดียวและแบ่งข้อมูลออกเป็น 3 ส่วนก่อนแล้วจึงนำแต่ละส่วนมาตรวจหาจุดเปลี่ยนแปลง พบว่า วิธี E-Divisive และ ks-cp3o พบจำนวนจุดและตำแหน่งของจุดเปลี่ยนแปลงที่ใกล้เคียงกัน และการแบ่งข้อมูลออกเป็นช่วงย่อยก่อนจะมีประสิทธิภาพมากกว่าการตรวจข้อมูลทั้งหมดในคราวเดียวกัน


การศึกษาเปรียบเทียบการประมาณค่าจากตัวแบบการถดถอย สำหรับข้อมูลที่มีการแจกแจงแบบล็อกนอร์มอล ที่ถูกตัดปลายทางขวาแบบสุ่มที่มีการแจกแจงแบบยูนิฟอร์ม, ธนาพิพัฒน์ ทรัพย์ครองชัย Jan 2018

การศึกษาเปรียบเทียบการประมาณค่าจากตัวแบบการถดถอย สำหรับข้อมูลที่มีการแจกแจงแบบล็อกนอร์มอล ที่ถูกตัดปลายทางขวาแบบสุ่มที่มีการแจกแจงแบบยูนิฟอร์ม, ธนาพิพัฒน์ ทรัพย์ครองชัย

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาเปรียบเทียบการประมาณค่าจากตัวแบบความถดถอย เมื่อตัวแปรตามมีการแจกแจงแบบล็อกนอร์มอลและตัวแปรตามบางค่าเป็นข้อมูลที่ถูกตัดปลายทางขวาแบบที่ 1 ด้วยวิธีด้วยวิธีกำลังสองต่ำสุด (OLS), วิธีของแชตเทอร์จีและแมคลีช (CM) วิธีภาวะน่าจะเป็นสูงสุดด้วยขั้นตอนอีเอ็ม (MLE_EM) และวิธีภาวะน่าจะเป็นสูงสุดด้วยขั้นตอนอีเอ็มเมื่อมีการปรับค่าข้อมูลก่อนคำนวณ (MLE_EM_AD) ข้อมูลที่ใช้ในการศึกษาได้จากการจำลองข้อมูล 243 สถานการณ์ ขนาดตัวอย่าง (n) เท่ากับ 30, 50, 100 ร้อยละของตัวแปรตามที่ถูกตัดปลายทางขวา (r1) เท่ากับ 10, 20, 30 สัดส่วนช่วงเวลาที่เปิดรับผู้ป่วยต่อช่วงเวลาที่ติดตามการรอดชีวิต (r2) เท่ากับ 0.1, 0.2, 0.3 อัตราส่วนความแปรปรวนของตัวแปรอิสระตัวที่ 1 ต่อตัวแปรอิสระตัวที่ 2 คือ 1:1, 1:2, 1:5 และอัตราส่วนความแปรปรวนรวมของตัวแปรอิสระต่อความคลาดเคลื่อน คือ 2:1, 1:1, 1:2 จากการศึกษาพบว่า 1) วิธี MLE_EM และ MLE_EM_AD เป็นวิธีที่มีประสิทธิภาพสูงสุดเมื่อตัวอย่างมีขนาดปานกลางและใหญ่ (n = 50, 100) หรือร้อยละของข้อมูลที่ถูกตัดปลายทางขวาปานกลางและมาก (r1 = 20%, 30%) ในทางกลับกัน 2) วิธี OLS เป็นวิธีที่มีประสิทธิภาพสูงสุดเมื่อตัวอย่างมีขนาดเล็ก (n = 30) หรือตัวแปรอิสระมีการกระจายตัวน้อยกว่าความคลาดเคลื่อน แต่ CM มีประสิทธิภาพสูงสุดเมื่อตัวแปรอิสระมีการกระจายตัวมากกว่าหรือเท่ากับความคลาดเคลื่อน 3) ทุกวิธีมีประสิทธิภาพมากขึ้นเมื่อตัวอย่างมีขนาดใหญ่ขึ้น หรือตัวแปรถูกตัดปลายทางขวาน้อยลง หรือสัดส่วนช่วงเวลาที่เปิดรับผู้ป่วยต่อช่วงเวลาที่ติดตามการรอดชีวิตลดลง หรือความคลาดเคลื่อนกระจายตัวน้อยกว่าตัวแปรอิสระ


การศึกษาเปรียบเทียบการประมาณค่าจากตัวแบบการถดถอยสำหรับข้อมูลที่ถูกตัดปลายทางขวาแบบที่ 1 ที่มีการแจกแจงแบบล็อกนอร์มอล, ศิวพร ทิพย์พันธุ์ Jan 2018

การศึกษาเปรียบเทียบการประมาณค่าจากตัวแบบการถดถอยสำหรับข้อมูลที่ถูกตัดปลายทางขวาแบบที่ 1 ที่มีการแจกแจงแบบล็อกนอร์มอล, ศิวพร ทิพย์พันธุ์

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาและเปรียบเทียบวิธีการประมาณค่าจากตัวแบบ การถดถอย เมื่อตัวแปรตามมีการแจกแจงแบบล็อกนอร์มอลและตัวแปรตามบางค่าเป็นข้อมูล ที่ถูกตัดปลายทางขวาแบบที่ 1 ด้วยวิธีกำลังสองต่ำสุด (OLS) วิธีภาวะน่าจะเป็นสูงสุด (MLE) วิธีของแชตเทอร์จีและแมคลีช (CM) และวิธีภาวะน่าจะเป็นสูงสุดด้วยขั้นตอนวิธีอีเอ็ม (MLE_EM) ข้อมูลในการศึกษาได้จากการจำลองข้อมูลจำนวน 81 สถานการณ์ สถานการณ์ละ 10,000 รอบ ขนาดตัวอย่าง (n) เท่ากับ 30, 50, 100 และเปอร์เซ็นต์การถูกตัดปลายทางขวาของตัวแปรตาม (r) เท่ากับ 10%, 20%, 30% และอัตราส่วนความแปรปรวนของตัวแปรอิสระตัวที่ 1 ต่อตัวแปรอิสระ ตัวที่ 2 คือ 1:1, 1:2, 1:5 และอัตราส่วนความแปรปรวนรวมของตัวแปรอิสระต่อความคลาดเคลื่อน คือ 2:1, 1:1, 1:2 จากการศึกษาพบว่า 1) วิธี MLE และวิธี MLE_EM มีประสิทธิภาพสูงสุดเมื่อตัวอย่างมีขนาดใหญ่ (n=100) หรือตัวแปรตามถูกตัดปลายทางขวามาก (r=30%) ในทางกลับกัน 2) วิธี OLS เป็นวิธีที่มีประสิทธิภาพสูงสุดเมื่อตัวอย่างมีขนาดเล็ก (n=30) หรือตัวแปรตามถูกตัดปลายทางขวาน้อย (r=10%) และ 3) วิธี CM เป็นวิธีที่มีประสิทธิภาพสูงสุดในสถานการณ์ที่เหลือ กล่าวคือ เมื่อตัวอย่างมีขนาดปานกลาง (n=50) หรือตัวแปรตามถูกตัดปลายทางขวาปานกลาง (r=20%) นอกจากนั้นพบว่า 4) ทุกวิธีมีประสิทธิภาพมากขึ้นเมื่อตัวอย่างมีขนาดใหญ่ขึ้นหรือตัวแปรตาม ถูกตัดปลายทางขวาน้อยลงหรือความคลาดเคลื่อนกระจายตัวน้อยกว่าตัวแปรอิสระ


ความเหลื่อมล้ำทางการศึกษาของผู้เรียนที่มีความต้องการพิเศษ ในโรงเรียนสังกัดสำนักงานคณะกรรมการการศึกษาขั้นพื้นฐาน, ณปภัช บรรณาการ Jan 2018

ความเหลื่อมล้ำทางการศึกษาของผู้เรียนที่มีความต้องการพิเศษ ในโรงเรียนสังกัดสำนักงานคณะกรรมการการศึกษาขั้นพื้นฐาน, ณปภัช บรรณาการ

Chulalongkorn University Theses and Dissertations (Chula ETD)

การวิจัยครั้งนี้มีวัตถุประสงค์เพื่อ 1) เปรียบเทียบภูมิหลังทางเศรษฐกิจและสังคม ระหว่างผู้เรียนที่มีความต้องการพิเศษกับนักเรียนทั่วไปในโรงเรียนสังกัดคณะกรรมการการศึกษาขั้นพื้นฐาน 2) ศึกษาปัจจัยนำเข้าสำหรับการจัดการเรียนรู้ให้ผู้เรียนที่มีความต้องการพิเศษและระบบช่วยเหลือผู้เรียนที่มีความต้องการพิเศษในโรงเรียนสังกัดคณะกรรมการการศึกษาขั้นพื้นฐาน และ 3) วิเคราะห์ความเหลื่อมล้ำทางการศึกษาของผู้เรียนที่มีความต้องการพิเศษในโรงเรียนอันเนื่องมาจากภูมิหลังทางเศรษฐกิจและสังคม ปัจจัยนำเข้าสำหรับการจัดการเรียนรู้แก่ผู้เรียนที่มีความต้องการพิเศษ และระบบช่วยเหลือผู้เรียนที่มีความต้องการพิเศษในโรงเรียนสังกัดสำนักงานคณะกรรมการการศึกษาขั้นพื้นฐาน การวิจัยนี้เก็บรวบรวมข้อมูลจาก 2 แหล่งประกอบด้วย 1) ข้อมูลปฐมภูมิจากโรงเรียนสังกัดสำนักงานเขตพื้นที่การศึกษามัธยมศึกษาเขต 1 และ 2 ทั้งหมด 43 โรงเรียน 2) แหล่งข้อมูลทุติยภูมิโดยใช้ฐานข้อมูลของผู้สอบ O-NET รายบุคคล ระดับชั้นมัธยมศึกษาปีที่ 3 ที่ศึกษาในโรงเรียนสังกัดคณะกรรมการการศึกษาขั้นพื้นฐาน (สพฐ.) ปีการศึกษา 2560 ของสถาบันทดสอบทางการศึกษาแห่งชาติ (สทศ.) ตัวอย่างที่ใช้ในการวิจัยในครั้งนี้คือ โรงเรียนสังกัดสำนักงานเขตพื้นที่การศึกษามัธยมศึกษาเขต 1 และ 2 จำนวน 39 โรงเรียน และนักเรียนจำนวน 11,534 คน การวิเคราะห์ข้อมูลใช้สถิติบรรยาย การวิเคราะห์กลุ่มแฝง (latent class analysis) และการวิเคราะห์ความเหลื่อมล้ำ (Inequality Index) ด้วยโปรแกรม R ผลการวิจัยสรุปได้ดังนี้ 1) ภูมิหลังทางเศรษฐกิจและสังคมระหว่างผู้เรียนที่มีความต้องการพิเศษกับนักเรียนทั่วไปไม่แตกต่างกันที่ระดับนัยสำคัญ .05 (ꭓ2df = 2 = 2.408, p = 0.300) 2) ผลการศึกษาปัจจัยนำเข้าของการจัดการเรียนรู้ให้ผู้เรียนที่มีความต้องการพิเศษ ด้านความเพียงพอและคุณภาพครู พบว่าอัตราส่วนระหว่างจำนวนผู้เรียนที่มีความต้องการพิเศษต่อจำนวนครูที่รับผิดชอบงานการศึกษาพิเศษในโรงเรียน มีค่าเท่ากับ 3.77 (SD = 5.27) ค่าเฉลี่ยร้อยละของครูการศึกษาพิเศษเมื่อเปรียบเทียบกับครูที่รับผิดชอบงานการศึกษาพิเศษในโรงเรียนมีค่าเท่ากับ 16.89 (SD = 30.97) จะเห็นว่าโรงเรียนส่วนใหญ่มีครูที่รับผิดชอบงานการศึกษาพิเศษที่เพียงพอต่อความต้องการ แต่อย่างไรก็ตามครูที่รับผิดชอบงานการศึกษาพิเศษของโรงเรียนดังกล่าวนั้นมีส่วนน้อยที่มีวุฒิการศึกษาพิเศษโดยตรงซึ่งสะท้อนความขาดแคลนคุณภาพของปัจจัยนำเข้าด้านบุคลากรครูที่รับผิดชอบงานการศึกษาพิเศษของโรงเรียน นอกจากนี้ยังพบว่าสิ่งอำนวยความสะดวกและวัสดุในการผลิตสื่อของผู้เรียนที่มีความต้องการพิเศษไม่เพียงพอ สื่อการจัดกิจกรรมการเรียนรู้ของครูและอุปกรณ์นันทนาการเพียงพอระดับน้อย ด้านระบบช่วยเหลือผู้เรียนที่มีความต้องการพิเศษมีผลการประเมินดังนี้ ด้านการวางแผนการจัดการศึกษา (plan), การปฏิบัติแผนการศึกษาเฉพาะบุคคล (do) และ ด้านการวิเคราะห์และการตรวจสอบผลการจัดการเรียนรู้ (check) ปฏิบัติเป็นส่วนใหญ่ ส่วนด้านการปรับปรุงแผนและการปฏิบัติงาน (act) ปฏิบัติเป็นบางครั้ง 3) ผลการวิเคราะห์ความเหลื่อมล้ำทางการศึกษาของผู้เรียนที่มีความต้องการพิเศษ พบว่า …


โมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศ: การวิเคราะห์ความไม่แปรเปลี่ยนแบบเบส์, ชนินันท์ พฤกษ์ประมูล Jan 2018

โมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศ: การวิเคราะห์ความไม่แปรเปลี่ยนแบบเบส์, ชนินันท์ พฤกษ์ประมูล

Chulalongkorn University Theses and Dissertations (Chula ETD)

การวิจัยนี้มีวัตถุประสงค์ 1) เพื่อพัฒนาโมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศ 2) เพื่อตรวจสอบความสอดคล้องของโมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศที่พัฒนาขึ้นกับข้อมูลเชิงประจักษ์ และ 3) เพื่อทดสอบความไม่แปรเปลี่ยนของโมเดลคุณลักษณะที่พึงประสงค์ฯ ที่พัฒนาขึ้นตามกลุ่มนักศึกษาไทยและต่างประเทศด้วยวิธีการประมาณค่าความน่าจะเป็นสูงสุด และการวิเคราะห์แบบเบส์ ตัวอย่างที่ใช้ในการวิจัยครั้งนี้ คือ นักศึกษาระดับปริญญาโทและเอกจำนวน 716 คน จาก 11 มหาวิทยาลัยชั้นนำของไทย จำนวน 459 คนและจาก 11 มหาวิทยาลัยชั้นนำของโลกจำนวน 257 คน เครื่องมือที่ใช้ในงานวิจัยในครั้งนี้ ได้แก่ แบบประเมินตนเองของนักศึกษาระดับปริญญาโทและเอกต่อคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาฉบับภาษาไทยและภาษาอังกฤษ การวิเคราะห์ข้อมูลเพื่อประมาณค่าพารามิเตอร์ในโมเดลด้วยวิธีการประมาณค่าความน่าจะเป็นสูงสุด (Maximum-likelihood estimation) จากโปรแกรม LISREL 8.72 และการวิเคราะห์แบบเบส์ (Bayesian analysis) จากโปรแกรม R 3.6.1 ผลการวิจัยสรุปได้ดังนี้ 1. โมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศ ประกอบด้วย 3 องค์ประกอบ 10 ตัวบ่งชี้ ประกอบด้วย องค์ประกอบที่ 1 ด้านความรู้ แบ่งออกเป็น 3 ตัวบ่งชี้ ได้แก่ ความรู้ในสาขาวิชาชีพ ความรู้ในศาสตร์อื่น ๆ และ ความรู้เท่าทันต่อการเปลี่ยนแปลง องค์ประกอบที่ 2 ด้านทักษะการเรียนรู้และการทำงาน แบ่งออกเป็น 4 ตัวบ่งชี้ ได้แก่ การคิดเชิงนวัตกรรม การจัดการตนเองและองค์กร การทำงานเป็นทีมแบบท้าทาย และ การเรียนรู้เทคโนโลยีสม่ำเสมอ และองค์ประกอบที่ 3 ด้านคุณธรรมจริยธรรม แบ่งออกเป็น 3 ตัวบ่งชี้ ได้แก่ คุณธรรมการอยู่ร่วมกับผู้อื่นในสังคม จรรยาบรรณทางวิชาชีพและวิชาการ และจิตอาสาและสำนึกสาธารณะ 2. โมเดลคุณลักษณะที่พึงประสงค์ของบัณฑิตระดับบัณฑิตศึกษาของไทยและต่างประเทศสอดคล้องกับข้อมูลเชิงประจักษ์ (χ2 = 29.17, df = 19, χ2/df = 1.535, p= .06339, RMSEA = .027, GFI= …


Introductory Statistics, Barbara Illowsky, Susan Dean Jan 2018

Introductory Statistics, Barbara Illowsky, Susan Dean

Open Access Textbooks

Introductory Statistics follows scope and sequence requirements of a one-semester introduction to statistics course and is geared toward students majoring in fields other than math or engineering. The text assumes some knowledge of intermediate algebra and focuses on statistics application over theory. Introductory Statistics includes innovative practical applications that make the text relevant and accessible, as well as collaborative exercises, technology integration problems, and statistics labs.


Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu Jan 2018

Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu

Theses and Dissertations

Structural Equation Modeling (SEM) is a framework of statistical methods that allows us to represent complex relationships between variables. SEM is widely used in economics, genetics and the behavioral sciences (e.g. psychology, psychobiology, sociology and medicine). Model complexity is defined as a model’s ability to fit different data patterns and it plays an important role in model selection when applying SEM. As in linear regression, the number of free model parameters is typically used in traditional SEM model fit indices as a measure of the model complexity. However, only using number of free model parameters to indicate SEM model complexity …


Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang Jan 2018

Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang

Theses and Dissertations

Modern big data often emerge as tensors. Standard statistical methods are inadequate to deal with datasets of large volume, high dimensionality, and complex structure. Therefore, it is important to develop algorithms such as low-rank tensor decomposition for data compression, dimensionality reduction, and approximation.

With the advancement in technology, high-dimensional images are becoming ubiquitous in the medical field. In lung radiation therapy, the respiratory motion of the lung introduces variabilities during treatment as the tumor inside the lung is moving, which brings challenges to the precise delivery of radiation to the tumor. Several approaches to quantifying this uncertainty propose using a …


An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker Jan 2018

An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker

Computer Science Faculty Publications

Cryo-electron microscopy (cryo-EM) is a structure determination method for large molecular complexes. As more and more atomic structures are determined using this technique, it is becoming possible to perform statistical characterization of side-chain conformations. Two data sets were involved to characterize block lengths for each of the 18 types of amino acids. One set contains 9131 structures resolved using X-ray crystallography from density maps with better than or equal to 1.5 Å resolutions, and the other contains 237 protein structures derived from cryo-EM density maps with 2-4 Å resolutions. The results show that the normalized probability density function of block …


Spatial Modelling And Wildlife Health Surveillance: A Case Study Of White Nose Syndrome In Ontario, Lauren Yee Jan 2018

Spatial Modelling And Wildlife Health Surveillance: A Case Study Of White Nose Syndrome In Ontario, Lauren Yee

Theses and Dissertations (Comprehensive)

Wildlife data is often limited by survey effort, small sample sizes, and spatial biases associated with collection and missing data. These factors can create unique challenges from a surveillance perspective when trying to extract spatial patterns of habitat suitability and disease distributions for conservation and management purposes. This thesis examined data quality from a wildlife health database in the context of spatial analysis of wildlife disease. Spatial analysis of the data to predict habitat suitability of bats and white nose syndrome afflicted bats was examined by using the MaxEnt modelling method. Methods to reduce spatial bias were examined and specific …


Particle Filters For State Estimation Of Confined Aquifers, Graeme Field Jan 2018

Particle Filters For State Estimation Of Confined Aquifers, Graeme Field

UNF Graduate Theses and Dissertations

Mathematical models are used in engineering and the sciences to estimate properties of systems of interest, increasing our understanding of the surrounding world and driving technological innovation. Unfortunately, as the systems of interest grow in complexity, so to do the models necessary to accurately describe them. Analytic solutions for problems with such models are provably intractable, motivating the use of approximate yet still accurate estimation techniques. Particle filtering methods have emerged as a popular tool in the presence of such models, spreading from its origins in signal processing to a diverse set of fields throughout engineering and the sciences including …


Unmasking Cost Growth Behavior: A Longitudinal Study, Cory N. D'Amico, Edward D. White, Jonathan D. Ritschel, Scott R. Kozlak Jan 2018

Unmasking Cost Growth Behavior: A Longitudinal Study, Cory N. D'Amico, Edward D. White, Jonathan D. Ritschel, Scott R. Kozlak

Faculty Publications

This article examines how cost growth factors (CGF) change over a program’s acquisition life cycle for 36 Department of Defense aircraft programs. Starting from Milestone B, the authors examine CGFs at five gateways: Critical Design Review, First Flight (FF), the end of Developmental Test and Evaluation (DT&E), Initial Operational Capability, and Full Operational Capability. Each CGF is assigned a color rating based upon the program’s cost growth: Green (low), Amber (moderate), or Red (high). Significant findings include dependencies among similar CGF color ratings and cost growth occurring primarily between FF and the end of DT&E during a program’s life cycle.


Netnographic Slog: Creative Elicitation Strategies To Encourage Participation In An Online Community Of Practice For Early Education And Care, Ruth Wallace Jan 2018

Netnographic Slog: Creative Elicitation Strategies To Encourage Participation In An Online Community Of Practice For Early Education And Care, Ruth Wallace

Research outputs 2014 to 2021

Active, participatory netnography, in contrast to passive netnography, is essential if researchers are to gain rich rewards from the rigorous collection of qualitative data. However, researchers should be aware of the ‘netnographic slog’; “the blood, sweat and tears” associated with eliciting quality data and encouraging active participation in online communities.

This article examines the – Supporting Nutrition for Australian Childcare (SNAC) – online community of practice, established to support healthy eating practices in early childhood education and care settings. To ensure research rigour, Kozinets’ netnographic steps were employed. Garnering member participation in this online community was a slog; most community …


Statistical Methods For Detecting Causal Rare Variants And Analyzing Multiple Phenotypes, Xinlan Yang Jan 2018

Statistical Methods For Detecting Causal Rare Variants And Analyzing Multiple Phenotypes, Xinlan Yang

Dissertations, Master's Theses and Master's Reports

This dissertation includes two papers with each distributed in one chapter. To date, genome-wide association studies (GWAS) have identified a large number of common variants that are associated with complex diseases successfully. However, the common variants identified by GWAS only account for a small proportion of trait heritability. Many studies showed that rare variants could explain parts of the missing heritability. Since the well-developed common variant detecting methods are underpowered for rare variant association tests unless sample sizes or effect sizes are very large, investigation the roles of rare variants in complex diseases presents substantial challenges. In chapter 1, we …


Statistical Methods For Analyzing Multivariate Phenotypes And Detecting Rare Variant Associations, Huanhuan Zhu Jan 2018

Statistical Methods For Analyzing Multivariate Phenotypes And Detecting Rare Variant Associations, Huanhuan Zhu

Dissertations, Master's Theses and Master's Reports

This dissertation includes four papers with each distributed in one chapter.

In chapter 1, I compared the performance of eight multivariate phenotype association tests. The motivation to conduct this power comparison paper is as follows. For nearly 15 years, genome-wide association studies (GWAS) have been widely used to identify genetic variants associated with human diseases and traits. GWAS typically investigate genetic variants for a predefined phenotype, thus fail to identify weak but important effects. In recent years, many multivariate association tests have been developed. However, there is a lack of comprehensive summary of such kinds of approaches. To fill this …


Offline And Online Density Estimation For Large High-Dimensional Data, Aref Majdara Jan 2018

Offline And Online Density Estimation For Large High-Dimensional Data, Aref Majdara

Dissertations, Master's Theses and Master's Reports

Density estimation has wide applications in machine learning and data analysis techniques including clustering, classification, multimodality analysis, bump hunting and anomaly detection. In high-dimensional space, sparsity of data in local neighborhood makes many of parametric and nonparametric density estimation methods mostly inefficient.

This work presents development of computationally efficient algorithms for high-dimensional density estimation, based on Bayesian sequential partitioning (BSP). Copula transform is used to separate the estimation of marginal and joint densities, with the purpose of reducing the computational complexity and estimation error. Using this separation, a parallel implementation of the density estimation algorithm on a 4-core CPU is …


Application Of Remote Sensing And Machine Learning Modeling To Post-Wildfire Debris Flow Risks, Priscilla Addison Jan 2018

Application Of Remote Sensing And Machine Learning Modeling To Post-Wildfire Debris Flow Risks, Priscilla Addison

Dissertations, Master's Theses and Master's Reports

Historically, post-fire debris flows (DFs) have been mostly more deadly than the fires that preceded them. Fires can transform a location that had no history of DFs to one that is primed for it. Studies have found that the higher the severity of the fire, the higher the probability of DF occurrence. Due to high fatalities associated with these events, several statistical models have been developed for use as emergency decision support tools. These previous models used linear modeling approaches that produced subpar results. Our study therefore investigated the application of nonlinear machine learning modeling as an alternative. Existing models …


Wildfire Emissions In The Context Of Global Change And The Implications For Mercury Pollution, Aditya Kumar Jan 2018

Wildfire Emissions In The Context Of Global Change And The Implications For Mercury Pollution, Aditya Kumar

Dissertations, Master's Theses and Master's Reports

Wildfires are episodic disturbances that exert a significant influence on the Earth system. They emit substantial amounts of atmospheric pollutants, which can impact atmospheric chemistry/composition and the Earth’s climate at the global and regional scales. This work presents a collection of studies aimed at better estimating wildfire emissions of atmospheric pollutants, quantifying their impacts on remote ecosystems and determining the implications of 2000s-2050s global environmental change (land use/land cover, climate) for wildfire emissions following the Intergovernmental Panel on Climate Change (IPCC) A1B socioeconomic scenario.

A global fire emissions model is developed to compile global wildfire emission inventories for major atmospheric …


Complex-Valued Time Series Modeling For Improved Activation Detection In Fmri Studies, Daniel W. Adrian, Ranjan Maitra, Daniel B. Rowe Jan 2018

Complex-Valued Time Series Modeling For Improved Activation Detection In Fmri Studies, Daniel W. Adrian, Ranjan Maitra, Daniel B. Rowe

Mathematical and Statistical Science Faculty Research and Publications

A complex-valued data-based model with th order autoregressive errors and general real/imaginary error covariance structure is proposed as an alternative to the commonly used magnitude-only data-based autoregressive model for fMRI time series. Likelihood-ratio-test-based activation statistics are derived for both models and compared for experimental and simulated data. For a dataset from a right-hand finger-tapping experiment, the activation map obtained using complex-valued modeling more clearly identifies the primary activation region (left functional central sulcus) than the magnitude-only model. Such improved accuracy in mapping the left functional central sulcus has important implications in neurosurgical planning for tumor and epilepsy patients. Additionally, we …


On Extended Quadratic Hazard Rate Distribution: Development, Properties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein G. Hamedani, Wenhui Sheng, Munir Ahmad Jan 2018

On Extended Quadratic Hazard Rate Distribution: Development, Properties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein G. Hamedani, Wenhui Sheng, Munir Ahmad

Mathematical and Statistical Science Faculty Research and Publications

In this paper, we propose a flexible extended quadratic hazard rate (EQHR) distribution with increasing, decreasing, bathtub and upside-down bathtub hazard rate function. The EQHR density is arc, right-skewed and symmetrical shaped. This distribution is also obtained from compounding mixture distributions. Stochastic orderings, descriptive measures on the basis of quantiles, order statistics and reliability measures are theoretically established. Characterizations of the EQHR distribution are studied via different techniques. Parameters of the EQHR distribution are estimated using the maximum likelihood method. Goodness of fit of this distribution through different methods is studied.


An Analysis Of Equity-Linked Insurance Pricing, Clara C. Ortgies Jan 2018

An Analysis Of Equity-Linked Insurance Pricing, Clara C. Ortgies

Honors Program Theses

This comprehensive study of equity-linked insurance options will explore the pricing of certificates of deposit and life insurance options using a present value method. With this study, I will be able to construct and price various equity-linked insurance products, with a focus on life insurance, that insurance companies could then sell to prospective customers. I will use concepts and formulas based in actuarial math, probability theory, and financial engineering in order to construct, price, and analyze new equity-linked insurance products. The fundamental methodology I will use involves applying pricing theory based on the expected value of the insurance payoff present …


A Land Use Regression Model For Explaining Spatial Variation In Air Pollution Levels Using A Wind Sector Based Approach, Owen Naughton, Aoife Donnelly, Paul Nolan, Francesco Pilla, Bruce Misstear, Brian Broderick Jan 2018

A Land Use Regression Model For Explaining Spatial Variation In Air Pollution Levels Using A Wind Sector Based Approach, Owen Naughton, Aoife Donnelly, Paul Nolan, Francesco Pilla, Bruce Misstear, Brian Broderick

Articles

Estimating pollutant concentrations at a local and regional scale is essential for good ambient air quality information in environmental and health policy decision making. Here we present a land use regression (LUR) modelling methodology that exploits the high temporal resolution of fixed-site monitoring (FSM) to produce viable air quality maps. The methodology partitions concentration time series from a national FSM network into wind-dependent sectors or “wedges”. A LUR model is derived using predictor variables calculated within the directional wind sectors, and compared against the long-term average concentrations within each sector. This study demonstrates the value of incorporating the relative position …


On Modified Burr Xii-Inverse Exponential Distribution: Prop¬Erties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein Hamedani, Haitham M. Yousof, Azeem Ali, Munir Ahmad Jan 2018

On Modified Burr Xii-Inverse Exponential Distribution: Prop¬Erties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein Hamedani, Haitham M. Yousof, Azeem Ali, Munir Ahmad

Mathematical and Statistical Science Faculty Research and Publications

In this paper, a flexible lifetime distribution with increasing, increasing and decreasing and modified bathtub hazard rate called Modified Burr XII-Inverse Exponential (MBXII-IE) is introduced. The density function of MBXII-IE has exponential, left-skewed, right-skewed and symmetrical shapes. Descriptive measures such as moments, moments of order statistics, incomplete moments, inequality measures, residual life function and reliability measures are theoretically established. The MBXII-IE distribution is characterized via different techniques. Parameters of MBXII-IE distribution are estimated using maximum likelihood method. The simulation study is performed to illustrate the performance of the Maximum Likelihood Estimates (MLEs) of the parameters of the MBXII-IE distribution. The …