Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 3871 - 3900 of 12817

Full-Text Articles in Statistics and Probability

Bayesian Tail Probability Estimation And Model Selection, Nan Shen Jan 2021

Bayesian Tail Probability Estimation And Model Selection, Nan Shen

Graduate Research Theses & Dissertations

Bayesian statistics is a prevalent and important field in statistics that assigns Bayesian probabilities, which represent a state of knowledge, to unknown quantities. We study Bayesian statistics with its applications through two projects in this report.

In the first project, we investigate the reasons that the Bayesian estimator of the tail probability is always higher than the frequentist estimator. Sufficient conditions for this phenomenon are established by looking at Taylor series approximations about the tail and by using Jensen's Inequality, both of which point to the convexity of the distribution function.

The second project is about redefining the Bayesian information …


Gene-Based Disease Classification Using Bayesian Self-Organizing Map Neural Networks, Guangting Zhou Jan 2021

Gene-Based Disease Classification Using Bayesian Self-Organizing Map Neural Networks, Guangting Zhou

Graduate Research Theses & Dissertations

Genes perform vital roles in living beings. By taking charges of protein synthesis, genes are able to take control of the expression of living traits. There are a lot of diseases associated closely to our genes. By analyzing genetic information, we are able to detect or classify gene based diseases. Among genetic disease information technologies, microarray can be one of the widely used ones. Usually, microarray data records thousands of gene expression features from a small number of samples including both normal and abnormal expressed tissues. It provides standardized comparison information between normal and diseased tissues, so as to provide …


Robust Determinants Of Happiness: High-Dimensional Bayesian Treatment Of Model Uncertainty, Milivoje Davidovic Jan 2021

Robust Determinants Of Happiness: High-Dimensional Bayesian Treatment Of Model Uncertainty, Milivoje Davidovic

Graduate Research Theses & Dissertations

The thesis investigates the most relevant economic and institutional determinants of happiness in some 93 countries worldwide, covering the period 2006-2019. We employ the Bayesian Model Averaging (BMA) fixed effect model (country demeaned and time demeaned) using a working panel data set with 651 observations. Our initial goal is to address the problem of model uncertainty in panel data models of happiness, aiming at selecting a set variables that are likely to be included in as "true" model of happiness. In addition, we aim to investigate the causal relationship running from selected economic and institutional variable to index of happiness. …


Frequentist Methods In Handling Misrepresentation Risk, Rexford Mawunyegah Akakpo Jan 2021

Frequentist Methods In Handling Misrepresentation Risk, Rexford Mawunyegah Akakpo

Graduate Research Theses & Dissertations

A commonly encountered risk in insurance business is misrepresentation risk. Misrepresentation is a type of insurance fraud where a policyholder or a policy applicant falsifies his or her risk status in order to pay cheaper premiums for more expensive future risks. It is difficult and expensive for insurance companies to detect this kind of risk. With high cost of sophisticated underwriting, it becomes a norm for insurance companies to regularly rely on the policy applicant to self-report most of their risk statuses. We employ a frequentist approach by using expectation-maximization (EM) algorithm to carry out maximum likelihood estimation of the …


Addressing The Ecological Fallacy With Lagrangian Inference, Michael Schwob Jan 2021

Addressing The Ecological Fallacy With Lagrangian Inference, Michael Schwob

Calvert Undergraduate Research Awards

Most epidemiologists elect to use statistical models that use population-level data to make inference on the spread of some virus or disease. This has become commonplace in the fields of epidemiology and biostatistics since most data used to construct and verify epidemic models are recorded at the population-level. Obtaining inference from a population-level model may be beneficial in studying the spread of disease in a homogeneous population, but the use of such models to describe a heterogeneous population results in inadequate inference. The inaccuracy of these models is further amplified when one tries to make individual-level inference from these population-level …


Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis Jan 2021

Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis

Theses and Dissertations

High-throughput chromosome conformation capture technology (Hi-C) has revealed extensive DNA looping and folding into discrete 3D domains. These include Topologically Associating Domains (TADs) and chromatin loops, the 3D domains critical for cellular processes like gene regulation and cell differentiation. The relatively low resolution of Hi-C data (regions of several kilobases in size) prevents precise mapping of domain boundaries by conventional TAD/loop-callers. However, high resolution genomic annotations associated with boundaries, such as CTCF and members of cohesin complex, suggest a computational approach for precise location of domain boundaries.

We developed preciseTAD, an optimized machine learning framework that leverages a random …


Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu Jan 2021

Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu

Theses and Dissertations

For analyzing large electronic health records (EHR) with time-to-event endpoints, such as in kidney transplantation, a major challenge is to provide an accurate risk analyses, while accounting for a multitude of epidemiological and statistical complexities. Motivated by a right-censored kidney transplantation EHR dataset derived from the United Network of Organ Sharing (UNOS), this dissertation, through a culmination of two interrelated yet distinctly different projects, focuses on developments of novel statistical procedures and methodologies to address some pressing issues arising in EHR-based research. In the first project, we aim to decouple the causal effects of treatments (here, studying subgroups, such as …


The Relevance Of Credit Risk In The Determination Of Commercial Banks’ Profitability: Evidence From Ghana, Godwin Kwabla Ekpe Jan 2021

The Relevance Of Credit Risk In The Determination Of Commercial Banks’ Profitability: Evidence From Ghana, Godwin Kwabla Ekpe

Graduate Research Theses & Dissertations

Existing empirical literature on the relationship between credit risk and bank’s profitability is replete with mixed results. This research investigates the probable effect of credit risk on banks’ profitability by examining the nature of the relationship between two measures of credit risk (Loss provisioning rate and Actual provisioning charge rate) and two measures ofprofitability (Return on assets and Return on Equity). The investigation is conducted using data on the Ghanaian banking industry. Various modeling techniques are used to fit the data, including frequentist beta regression and Bayesian beta regression models. The results across all models suggest negative linear relationship between …


Traffic Fatality Rate Prediction Based On Deep Neural Network And Bayesian Neural Network, Yiqun Hu Jan 2021

Traffic Fatality Rate Prediction Based On Deep Neural Network And Bayesian Neural Network, Yiqun Hu

Graduate Research Theses & Dissertations

There have been numerous studies on traffic accidents and their fatality rate. For this challenging machine learning regression problem, Neural Networks (NNs) have produced state-of-the-art data. Despite their success, they are often used in a fre- quentist scheme, which means they cannot account for uncertainty in their forecasts. BNNs are comprised of a Probabilistic Model and a Neural Network. The aim of such a design is to bring together the benefits of Neural Networks and stochastic modeling. Neural networks have the ability to approximate continuous functions uni- versally. Statistical models allow for the direct definition of a model with known …


A Time Series Analysis Approach To Forecasting Covid-19 Cases And Deaths: An Analysis Of Covid-19 Data In Colombia, Andrea Jackson-Sagredo Jan 2021

A Time Series Analysis Approach To Forecasting Covid-19 Cases And Deaths: An Analysis Of Covid-19 Data In Colombia, Andrea Jackson-Sagredo

Graduate Research Theses & Dissertations

The novel Coronavirus, known as COVID-19 is a highly contagious and transmissible infectious disease that has taken a toll throughout the entire world for over a year. The inner workings and long term effects of COVID-19 continue to be misunderstood. While COVID-19 has impacted all countries tremendously, Latin American countries and specifically Colombia have been impacted significantly by the virus. This thesis investigates the potential to forecast COVID-19 cases and deaths using Time Series Analysis methods and models for the South American country of Colombia. Time series analysis on Colombian COVID-19 data begins with data processing on a data set …


Stochastic Infection In Network Models With Applications To Pollution Analysis, Alexander Thor Wold Jan 2021

Stochastic Infection In Network Models With Applications To Pollution Analysis, Alexander Thor Wold

Graduate Research Theses & Dissertations

The continued adoption and escalation of commercial surface extraction techniques threatens to contaminate adjacent river networks across the coal mining landscape. We look to simulate the movement of this pollution across connected graph structures using stochastic block model methodologies, fueled from work and theory derived from exponential random graph models. We begin our study by applying our virtual experiment to the motivating material, later offering an exhaustive walk through of the simulation itself. Afterwards, we present our findings and their implications to water pollution analysis, emphasizing a need for this research and expanding on some inferential statistics left for later …


The Simulation Extrapolation Method With Differential Measurement Error, Dominic Partipilo Jan 2021

The Simulation Extrapolation Method With Differential Measurement Error, Dominic Partipilo

Graduate Research Theses & Dissertations

Most of statistical theory operates under the assumption that the true values of covariates have been measured correctly, but it is not always possible to obtain the true values of these covariates. A common issue, specifically in regression models, is that predictors are misclassified or measured with systematic measurement error. There have been many methods developed for handling measurement error, specifically in the case where measurement error is nondifferential, where the measurement error can be treated as independent from the covariates. The frequentist method known as simulation extrapolation (SIMEX) is one of these methods that specifically handles the case for …


Optimization Of Dynamic Objective Functions Using Path Integrals, Paramahansa Pramanik Jan 2021

Optimization Of Dynamic Objective Functions Using Path Integrals, Paramahansa Pramanik

Graduate Research Theses & Dissertations

Path integrals are used to find an optimal strategy for a firm under a Walrasian system. We define dynamic optimal strategies and develop an integration method to capture all non-additive non-convex strategies. We also show that the method can solve the non-linear case, for example Merton-Garman-Hamiltonian system, which the traditional Pontryagin maximum principle cannot solve in closed form. Furthermore, we assume that the strategy space and time are inseparable with respect to a contract. Under this assumption we show that the strategy spacetime is a dynamic curved Liouville-like 2-brane quantum gravity surface under asymmetric information and that traditional Euclidean geometry …


Methods For High-Dimensional Spatial Data: Dimension Reduction And Covariance Approximation, Paul May Jan 2021

Methods For High-Dimensional Spatial Data: Dimension Reduction And Covariance Approximation, Paul May

Electronic Theses and Dissertations

In spatial statistics, because quantities are correlated based on their relative positions in space, data is modeled as a single realization of a multivariate stochastic process. Spatial data can be high-dimensional either through a large number of observed variables per location, or through a large number of observed locations. The two are often handled differently, with the former addressed through dimension reduction and the latter addressed through appropriate modeling of the spatial correlation between locations. The main body of this dissertation is a three-part work. Parts 2 and 3 pertain to the "many variables" problem, proposing novel methods of dimension …


Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee Jan 2021

Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee

Senior Independent Study Theses

According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.

We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …


An Evaluation Of The Performance Of Proc Arima's Identify Statement: A Data-Driven Approach Using Covid-19 Cases And Deaths In Florida, Fahmida Akter Shahela Jan 2021

An Evaluation Of The Performance Of Proc Arima's Identify Statement: A Data-Driven Approach Using Covid-19 Cases And Deaths In Florida, Fahmida Akter Shahela

Electronic Theses and Dissertations, 2020-2023

Understanding data on novel coronavirus (COVID-19) pandemic, and modeling such data over time are crucial for decision making at managing, fighting, and controlling the spread of this emerging disease. This thesis work looks at some aspects of exploratory analysis and modeling of COVID-19 data obtained from the Florida Department of Health (FDOH). In particular, the present work is devoted to data collection, preparation, description, and modeling of COVID-19 cases and deaths reported by FDOH between March 12, 2020, and April 30, 2021. For modeling data on both cases and deaths, this thesis utilized an autoregressive integrated moving average (ARIMA) times …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke Jan 2021

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Impact Of Case Management On Childhood Lead Exposure In Marion County, Indiana, Maliki Yacouba Jan 2021

Impact Of Case Management On Childhood Lead Exposure In Marion County, Indiana, Maliki Yacouba

Walden Dissertations and Doctoral Studies

The Centers for Disease Control and Prevention recently declared that no amount of childhood blood lead level (BLL) is safe. The purpose of this quantitative study with a retrospective cohort design was to evaluate the effectiveness of case management intervention on children diagnosed with elevated BLL (EBLL; ≥ 5 μg/dL) in Marion, County, Indiana. The health belief model was used as the theoretical foundation for the study. A data set of 160 lead exposure case management records was analyzed to find whether: (a) BLL at post-case-management time significantly differ from BLL at baseline (b) BLL at post-case-management time is affected …


A New Generalized Modified Weibull Distribution, Morad Alizadeh, Muhammad Nauman Khan, Mahdi Rasekhi, Gholamhossein Hamedani Jan 2021

A New Generalized Modified Weibull Distribution, Morad Alizadeh, Muhammad Nauman Khan, Mahdi Rasekhi, Gholamhossein Hamedani

Mathematical and Statistical Science Faculty Research and Publications

We introduce a new distribution, so called A new generalized modified Weibull (NGMW) distribution. Various structural properties of the distribution are obtained in terms of Meijer’s G–function, such as moments, moment generating function, conditional moments, mean deviations, order statistics and maximum likelihood estimators. The distribution exhibits a wide range of shapes with varying skewness and assumes all possible forms of hazard rate function. The NGMW distribution along with other distributions are fitted to two sets of data, arising in hydrology and in reliability. It is shown that the proposed distribution has a superior performance among the compared distributions as …


Upper-Sided Ewma-Based Distribution-Specific Tolerance Limits, Owen Visser Jan 2021

Upper-Sided Ewma-Based Distribution-Specific Tolerance Limits, Owen Visser

UNF Graduate Theses and Dissertations

Tolerance limits are constructed from sample data to ascertain if a proportion of a process is within specification limits. There exists multiple methods of calculating the sample size requirements for tolerance limits under various assumptions. In this research, a distribution-specific algorithm that utilizes the exponentially weighted moving average technique (EWMA), first introduced by Sa and Razaila (2004), is reconstructed. The algorithm is used to calculate the required sample sizes for continuous construction of upper-sided tolerance limits. The sample sizes and intervals constructed from them are compared to three existing methods for various distributions. The distribution-specific algorithm was observed to reduce …


Prediction Intervals For Fractionally Integrated Time Series And Volatility Models, Rukman Ekanayake Jan 2021

Prediction Intervals For Fractionally Integrated Time Series And Volatility Models, Rukman Ekanayake

Doctoral Dissertations

"The two of the main formulations for modeling long range dependence in volatilities associated with financial time series are fractionally integrated generalized autoregressive conditional heteroscedastic (FIGARCH) and hyperbolic generalized autoregressive conditional heteroscedastic (HYGARCH) models. The traditional methods of constructing prediction intervals for volatility models, either employ a Gaussian error assumption or are based on asymptotic theory. However, many empirical studies show that the distribution of errors exhibit leptokurtic behavior. Therefore, the traditional prediction intervals developed for conditional volatility models yield poor coverage. An alternative is to employ residual bootstrap-based prediction intervals. One goal of this dissertation research is to develop …


Modeling Time Series With Conditional Heteroscedastic Structure, Ratnayake Mudiyanselage Isuru Panduka Ratnayake Jan 2021

Modeling Time Series With Conditional Heteroscedastic Structure, Ratnayake Mudiyanselage Isuru Panduka Ratnayake

Doctoral Dissertations

"Models with a conditional heteroscedastic variance structure play a vital role in many applications, including modeling financial volatility. In this dissertation several existing formulations, motivated by the Generalized Autoregressive Conditional Heteroscedastic model, are further generalized to provide more effective modeling of price range data well as count data. First, the Conditional Autoregressive Range (CARR) model is generalized by introducing a composite range-based multiplicative component formulation named the Composite CARR model. This formulation enables a more effective modeling of the long and short-term volatility components present in price range data. It treats the long-term volatility as a stochastic component that in …


Depicting Bivariate Relationship With A Gaussian Ellipse, Mamunur Rashid, Jyotirmoy Sarkar Jan 2021

Depicting Bivariate Relationship With A Gaussian Ellipse, Mamunur Rashid, Jyotirmoy Sarkar

Mathematics Faculty Publications

For data on two continuous variables, how should one depict the summary statistics (means, SDs, correlation coefficient, coefficient of determination, regression lines) so that their values can be read off easily from the depiction and potential outliers can be flagged also? We propose the Gaussian covariance ellipse as an answer that will benefit all users of statistics.


A Model For Inhalation Of Infectious Aerosol Contaminants In An Aircraft Passenger Cabin, Bert A. Silich Jan 2021

A Model For Inhalation Of Infectious Aerosol Contaminants In An Aircraft Passenger Cabin, Bert A. Silich

International Journal of Aviation, Aeronautics, and Aerospace

Aerosol contamination of an aircraft cabin by infectious passengers is a concern of passengers, aircrew and the aviation industry. This may be especially important during a pandemic, such as COVID-19, where the full extent of aerosol transmission is not well understood. A statistical method to determine the number of infectious passengers on board along with a mathematical model estimating the contaminant concentration of aerosols in the cabin and the number of inhaled infectious particles by passengers is presented. An example is used to demonstrated how the results can be estimated during normal operations and emergency conditions with malfunctions of the …


Option Implied Volatility's Predictability On Monthly Stock Returns, Hung T. Dao Jan 2021

Option Implied Volatility's Predictability On Monthly Stock Returns, Hung T. Dao

Senior Independent Study Theses

Since the trading of options is based on underlying stocks, it is reasonable to assume that information from the options market can be used to explain the returns in the stock market. Our independent study investigates the relationship between options implied volatility and stock returns. Previous studies have found significant results in using implied volatility in predicting stock returns. This paper provides a discussion of such studies, the theoretical framework for the research topic, and the Black-Scholes model, which is famous for its application in implied volatility calculation. Monthly returns of 20 large US firms are regressed against implied volatility …


การเปรียบเทียบวิธีในการพยากรณ์ราคาหุ้นด้วยแบบจำลองอารีม่า, โครงข่ายประสาทเทียม และตัวแบบผสม, กาญจน์ภิวรรณ จงศิริวิโรจ Jan 2021

การเปรียบเทียบวิธีในการพยากรณ์ราคาหุ้นด้วยแบบจำลองอารีม่า, โครงข่ายประสาทเทียม และตัวแบบผสม, กาญจน์ภิวรรณ จงศิริวิโรจ

Chulalongkorn University Theses and Dissertations (Chula ETD)

การวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบวิธีการพยากรณ์ราคาปิดหุ้นรายวันในอนาคต โดยใช้ตัวแบบอารีม่าซึ่งสร้างจากวิธีการค้นหาแบบกริด โครงข่ายประสาทเทียมและตัวแบบผสมในการพยากรณ์ราคาของหุ้น ภายใต้ตัวอย่างหุ้นที่ถูกเลือกมาตามระดับความผันผวนจากสูงไปต่ำ ในกลุ่มอุตสาหกรรมเทคโนโลยีและชิ้นส่วนอิเล็กทรอนิกส์ ได้แก่ HANA, DELTA และ SVI ตามลำดับ โดยเก็บข้อมูลราคาปิดรายวันของหุ้นตั้งแต่เดือนตุลาคม พ.ศ. 2559 ถึงเดือนตุลาคม พ.ศ. 2564 ( 5 ปีย้อนหลัง ) ซึ่งอาศัยการแบ่งชุดข้อมูลฝึกสอนด้วยวิธี ตรวจสอบไขว้ (rolling forward validation) ทั้งวิธีตรวจสอบไขว้แบบสะสม และวิธีตรวจสอบไขว้แบบ moving window ซึ่งผลการวิจัยพบว่า เมื่อใช้ค่าเฉลี่ยของร้อยละความผิดพลาดสัมบูรณ์เป็นเกณฑ์ในการคัดเลือกตัวแบบ ทั้งสองวิธีการแบ่งชุดข้อมูลย่อยนั้น โครงข่ายประสาทเทียมมีความแม่นยำมากที่สุดในการพยากรณ์ราคาปิดของหุ้น HANA, DELTA และ SVI รวมถึงตัวแบบผสมดังกล่าวไม่จำเป็นต้องมีประสิทธิภาพดีกว่าการใช้แต่ละตัวแบบเพียงลำพังเสมอไป ตัวแบบอารีม่าซึ่งสร้างจากวิธีการค้นหาแบบกริดสามารถพยากรณ์ได้ดีกว่าในหุ้นที่มีระดับความผันผวนกลางและระดับต่ำ ในขณะที่โครงข่ายประสาทเทียมสามารถพยากรณ์ได้ดีในทุกระดับความผันผวนราคาหุ้น


โครงข่ายประสาทเทียมสำหรับการวิเคราะห์การถดถอยเชิงเส้นตามบริบทนัยทั่วไป, ชยานนท์ ขัตติยาภิรักษ์ Jan 2021

โครงข่ายประสาทเทียมสำหรับการวิเคราะห์การถดถอยเชิงเส้นตามบริบทนัยทั่วไป, ชยานนท์ ขัตติยาภิรักษ์

Chulalongkorn University Theses and Dissertations (Chula ETD)

ปัญหาความสัมพันธ์เชิงเส้นตามบริบท คือปัญหาที่มีตัวแปรต้นที่แบ่งข้อมูลออกเป็นกลุ่มต่าง ๆ โดยในแต่ละกลุ่มจะมีความสัมพันธ์กับผลเฉลยในลักษณะเชิงเส้นที่แตกต่างกัน ทางผู้วิจัยได้สนใจที่จะนำวิธีโครงข่ายประสาทเทียม (Neural Networks) มาแก้ไขปัญหาประเภทดังกล่าว โดยพัฒนาโครงสร้างที่ชื่อว่า Generalized Contextual Regression (GCR) และเปรียบเทียบกับโครงสร้างที่เคยมีมาก่อน ได้แก่ Feedforward Neural Networks (FNN) ซึ่งเป็นโครงสร้างพื้่นฐาน และ Contextual Regression (CR) ซึ่งนำเสนอโดย Liu และ Wang (2017) งานวิจัยนี้จะศึกษาเฉพาะปัญหาการถดถอยเชิงเส้น ที่ตัวแปรต้นไม่เกิน 10 ตัว ซึ่งมีตัวแปรเชิงบริบทไม่เกิน 3 ตัวเท่านั้น โดยจากผลการวิจัยพบว่าวิธี GCR มีประสิทธิภาพสูงที่สุดในการแก้ไขปัญหาความสัมพันธ์เชิงเส้นตามบริบทเมื่อเปรียบเทียบกับวิธี FNN และ CR


การเปรียบเทียบประสิทธิภาพของวิธีการสร้างช่วงความเชื่อมั่นสำหรับสัมประสิทธิ์การถดถอยลอจิสติกในข้อมูลที่มีมิติสูง โดยใช้การประมาณสองขั้นตอนด้วยวิธี Lasso + Mle And A Bootstrap Lasso + Partial Ridge, ณิชากร ไทยวงษ์ Jan 2021

การเปรียบเทียบประสิทธิภาพของวิธีการสร้างช่วงความเชื่อมั่นสำหรับสัมประสิทธิ์การถดถอยลอจิสติกในข้อมูลที่มีมิติสูง โดยใช้การประมาณสองขั้นตอนด้วยวิธี Lasso + Mle And A Bootstrap Lasso + Partial Ridge, ณิชากร ไทยวงษ์

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบวิธีการสร้างช่วงความเชื่อมั่นสำหรับสัมประสิทธิ์การถดถอยลอจิสติกในข้อมูลที่มีมิติสูง โดยใช้การประมาณสองขั้นตอนด้วยวิธี Lasso+MLE และวิธี Lasso+ Partial Ridge ซึ่งในการศึกษานี้จะจำลองข้อมูลทั้งหมด 8 ชุด และเปรียบเทียบประสิทธิภาพของช่วงความเชื่อมั่นที่ได้จากการสร้างช่วงความเชื่อมั่นทั้งหมด 4 วิธี ได้แก่ วิธี Parametric Bootstrap Lasso+MLE, วิธี Parametric Bootstrap Lasso+Partial Ridge, วิธี Paired Bootstrap Lasso+MLE และวิธี Paired Bootstrap Lasso+Partial Ridge โดยใช้เกณฑ์ในการเปรียบเทียบประสิทธิภาพของช่วงความเชื่อมั่น คือ ความกว้างเฉลี่ยของช่วงความเชื่อมั่น ค่าความน่าจะเป็นครอบคลุม ค่าความแม่นยำ และค่าความไว จากการศึกษาภายใต้ขอบเขตดังกล่าวผลปรากฏว่า วิธี Parametric Bootstrap Lasso+Partial Ridge มีประสิทธิภาพในการสร้างช่วงความเชื่อมั่นมากที่สุด รองลงมาคือ วิธี Paired Bootstrap Lasso+Partial Ridge และวิธี Paired Bootstrap Lasso+MLE ตามลำดับ และวิธีที่มีประสิทธิภาพในการสร้างช่วงความเชื่อมั่นน้อยที่สุด ก็คือ วิธี Parametric Bootstrap Lasso+MLE ดังนั้นจึงสรุปได้ว่า การสร้างช่วงความเชื่อมั่นสำหรับสัมประสิทธิ์การถดถอยลอจิสติกโดยใช้การประมาณสองขั้นตอนด้วยวิธี Lasso+Partial Ridge มีประสิทธิภาพมากกว่าวิธี Lasso+MLE


การแบ่งส่วนรูปภาพดอกไม้ด้วยการใช้ซาเลียนซีแมปร่วมกับการประยุกต์ใช้ปริภูมิสีเอชเอสวีและหน้ากากสี, ธนณัฏฐ์ หงษ์ทอง Jan 2021

การแบ่งส่วนรูปภาพดอกไม้ด้วยการใช้ซาเลียนซีแมปร่วมกับการประยุกต์ใช้ปริภูมิสีเอชเอสวีและหน้ากากสี, ธนณัฏฐ์ หงษ์ทอง

Chulalongkorn University Theses and Dissertations (Chula ETD)

การจำแนกประเภทรูปภาพดอกไม้เป็นสิ่งที่ท้าทาย เนื่องจากความคล้ายคลึงกันทางกายภาพของดอกไม้ เทคนิคการแบ่งส่วนรูปภาพ (Image segmentation) สามารถลดความซับซ้อนขององค์ประกอบภายในพื้นหลังภาพ ทำให้การจำแนกประเภทรูปภาพดอกไม้มีประสิทธิภาพมากขึ้น งานวิจัยชิ้นนี้ได้นำเสนอแนวคิดการแบ่งส่วนรูปภาพ โดยอิงการใช้ประโยชน์จากซาเลียนซีแมป (Saliency map) ในการเลือกบริเวณที่สนใจภายในภาพ และการใช้ปริภูมิสีเอชเอสวี (HSV) ผนวกกับการใช้หน้ากากสี (Color mask) ในการช่วยลดรายละเอียดที่ไม่สำคัญภายในพื้นหลังของรูปภาพ ผลการทดลองแสดงให้เห็นว่าวิธีการที่นำเสนอให้ผลลัพธ์การแบ่งส่วนรูปภาพโดยวัดจากค่าเฉลี่ย IoU เท่ากับ 54% (ซึ่งมากกว่างานวิจัยก่อนหน้า 13 %) ในขณะที่ค่าความถูกต้อง ความแม่นยำ ค่าความครบถ้วน และค่า F1 เมื่อจำแนกประเภทดอกไม้ด้วยแบบจำลอง VGG16 ที่ผ่านการปรับโครงสร้างเท่ากับ 87 %


การวิเคราะห์ข้อมูลด้วยภาพเพื่อการจัดซื้อหนังสือด้วยข้อมูลบรรณานุกรมของสำนักงานวิทยทรัพยากร จุฬาลงกรณ์มหาวิทยาลัย, ธนศาสตร์ ทักษิณ Jan 2021

การวิเคราะห์ข้อมูลด้วยภาพเพื่อการจัดซื้อหนังสือด้วยข้อมูลบรรณานุกรมของสำนักงานวิทยทรัพยากร จุฬาลงกรณ์มหาวิทยาลัย, ธนศาสตร์ ทักษิณ

Chulalongkorn University Theses and Dissertations (Chula ETD)

ปัจจุบันการจัดซื้อหนังสือของสำนักงานวิทยทรัพยากร จุฬาลงกรณ์มหาวิทยาลัยจะจัดซื้อตามคำแนะนำของผู้ใช้งานและประสบการณ์ของบรรณารักษ์ โดยส่วนมากจะจัดซื้อหนังสือที่สอดคล้องกับหลักสูตรการเรียนการสอนซึ่งยังไม่ตรงตามความต้องการของผู้ใช้งาน การวิจัยนี้เป็นการเปรียบเทียบประสิทธิภาพการจัดซื้อหนังสือก่อนและหลังการใช้โปรแกรมเพื่อตัดสินใจซื้อ ซึ่งสามารถวางแผนการจัดซื้อหนังสือได้อย่างมีประสิทธิภาพ โดยแสดงภาพปริมาณและราคาที่เหมาะสมของหนังสือแต่ละเล่มที่ตรงกับความต้องการของผู้ใช้จริง ผู้วิจัยได้คัดเลือกบรรณารักษ์ของสำนักงานวิทยทรัพยากรฯ แบบเจาะจงในการทําวิจัยและศึกษาความต้องการของผู้ใช้งานในการซื้อหนังสือของสำนักงานวิทยทรัพยากรฯ และพัฒนาโปรแกรมการแนะนำหนังสือโดยให้บรรณารักษ์เป็นผู้ทดสอบคุณภาพโปรแกรม การทดสอบใช้ข้อมูลหนังสือจากสำนักงานวิทยทรัพยากรฯ ที่ตีพิมพ์ในช่วงปี ค.ศ. 2010-2019 โดยใช้ค่าดัชนีแจ็คการ์ดวัดประสิทธิภาพการแนะนำหนังสือของโปรแกรมซึ่งคือค่าความคล้ายคลึงของการเลือกหนังสือก่อนและหลังใช้ภาพแสดงข้อมูลจากโปรแกรมในสถานการณ์ต่าง ๆ 8 สถานการณ์ ผลการทดสอบคือบรรณารักษ์สามารถเลือกหนังสือคล้ายคลึงกับภาพที่โปรแกรมแนะนำคือการใช้ภาพแสดงข้อมูลสามารถบอกข้อดีและข้อเสียของการเลือกซื้อหนังสือด้วยวิธีปัจจุบันและสามารถแนะนำเงื่อนไขเพิ่มเติมเพื่อให้วิธีการเลือกซื้อหนังสือในปัจจุบันมีประสิทธิภาพมากขึ้น