Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2023

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 781 - 810 of 841

Full-Text Articles in Statistics and Probability

Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede Jan 2023

Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede

Statistical and Data Sciences: Faculty Publications

During the emergence of Data Science as a distinct discipline, discussions of what exactly constitutes Data Science have been a source of contention, with no clear resolution. These disagreements have been exacerbated by the lack of a clear single disciplinary 'parent.' Many early efforts at defining curricula and courses exist, with the EDISON Project's Data Science Framework (EDISON-DSF) from the European Union being the most complete. The EDISON-DSF includes both a Data Science Body of Knowledge (DS-BoK) and Competency Framework (CF-DS). This paper takes a critical look at how EDISON's CF-DS compares to recent work and other published curricular or …


Deep Learning-Based Technique For The Perception Of The Cervical Cancer, Aya Haraz, Hossam El-Din Moustafa, Abeer Twakol Khaleel, Ahmed H. Eltanboly Jan 2023

Deep Learning-Based Technique For The Perception Of The Cervical Cancer, Aya Haraz, Hossam El-Din Moustafa, Abeer Twakol Khaleel, Ahmed H. Eltanboly

Mansoura Engineering Journal

In third-world countries, cervical cancer is the most prevalent and leading cause of death. It is affected by a variety of factors, including smoking, poor nutritional status, immunological inadequacy, and prolonged use of contraception. The Pap smear test, which is intended to prevent cervical cancer, finds preneoplastic changes in cervical epithelial cells. This study framework classified cervical cancer cells from Pap smears into five specified cell types using machine learning-based classification algorithms. The SIPaKMeD database is used in this investigation. This public dataset, which was manually cropped from 966 cluster cell images taken from Pap smear slides, has 4045 isolated …


Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi Jan 2023

Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi

Graduate Research Posters

Background: Head and neck cancer is the 6th most common cancer worldwide with an expected 1.08 million new cases each year. Such cancer data are ultra-high dimensional with thousands of clinical features and gene expressions, making it challenging for the traditional analytical tools to extract the potential biomarker for the cancer survival and control false discoveries. In addition, presence of heavy censoring can affect the screening procedures based on Kaplan-Meier (K-M) survival estimates.

Aim: To propose a model free, ultra-high dimensional feature screening method with two-dimensional survival outcome allowing false discovery rate (FDR) control.

Method: 516 primary tumor patients with …


Automated Machine Learning: Intellient Binning Data Preparation And Regularized Regression Classfier, Jianbin Zhu Jan 2023

Automated Machine Learning: Intellient Binning Data Preparation And Regularized Regression Classfier, Jianbin Zhu

Electronic Theses and Dissertations, 2020-2023

Automated machine learning (AutoML) has become a new trend which is the process of automating the complete pipeline from the raw dataset to the development of machine learning model. It not only can relief data scientists' works but also allows non-experts to finish the jobs without solid knowledge and understanding of statistical inference and machine learning. One limitation of AutoML framework is the data quality differs significantly batch by batch. Consequently, fitted model quality for some batches of data can be very poor due to distribution shift for some numerical predictors. In this dissertation, we develop an intelligent binning to …


Meta-Analysis Of Scent Detection Canines And Potential Factors Influencing Their Success Rates, Molly Marie Jaskinia Jan 2023

Meta-Analysis Of Scent Detection Canines And Potential Factors Influencing Their Success Rates, Molly Marie Jaskinia

Graduate Student Theses, Dissertations, & Professional Papers

Objective: This is a meta-analysis focused on the success rates of scent detection canines and potential factors that could influence their accuracy. A series of statistical analyses were conducted to determine if certain demographic factors, such as the dog’s gender, age, and breed, have an effect on a scent dog’s accuracy during a search. Or if more circumstantial factors, like the dog’s level of experience in scent work, the type of target scent, and their handler’s awareness of the target’s location, affect the outcome of the search.

Materials and Methods: A dataset was created from 37 different articles consisting of …


Applications Of Transfer Learning From Malicious To Vulnerable Binaries, Sean Patrick Mcnulty Jan 2023

Applications Of Transfer Learning From Malicious To Vulnerable Binaries, Sean Patrick Mcnulty

Graduate Student Theses, Dissertations, & Professional Papers

Malware detection and vulnerability detection are important cybersecurity tasks. Previous research has successfully applied a variety of machine learning methods to both. However, despite their potential synergies, previous research has yet to unite these two tasks. Given the recent success of transfer learning in many domains, such as language modeling and image recognition, this thesis investigated the use of transfer learning to improve vulnerability detection. Specifically, we pre-trained a series of models to detect malicious binaries and used the weights from those models to kickstart the detection of vulnerable binaries. In our study, we also investigated five different data representations …


Poisson Regression Model With Application To Wastewater Surveillance Under A Threshold Linear Mixed Model For Covid-19 Sensitivity Rates, Norou Diawara, Hueiwang Anna Jeng, Kyle Curtis, Raul Gonzalez, Nancy Welch, Cynthia Jackson, Rekha Singh, David Jurgens, Sasanka Adikari, Omotomilola Jegede Jan 2023

Poisson Regression Model With Application To Wastewater Surveillance Under A Threshold Linear Mixed Model For Covid-19 Sensitivity Rates, Norou Diawara, Hueiwang Anna Jeng, Kyle Curtis, Raul Gonzalez, Nancy Welch, Cynthia Jackson, Rekha Singh, David Jurgens, Sasanka Adikari, Omotomilola Jegede

Mathematics & Statistics Faculty Publications

A Threshold Linear Mixed Model (TLMM) has been developed to identify specific thresholds based on wastewater SARS-CoV-2 viral concentrations, which reflect COVID-19 cases. The thresholds can guide decisions regarding public health responses and prevention measures. To assess the practical application of TLMM, a simple simulation was conducted using a sample size of 100 and 500 replications. The simulation allowed for comparing parameter estimators by assessing bias and standard deviation and the root of the mean square error. The model and estimation procedures were applied to reported wastewater and clinic data to test its application for real-world scenarios. Our results demonstrated …


Utilizing Markov Chains To Estimate Allele Progression Through Generations, Ronit Gandhi Jan 2023

Utilizing Markov Chains To Estimate Allele Progression Through Generations, Ronit Gandhi

Honors Program: Senior Projects (Public)

All populations display patterns in allele frequencies over time. Some alleles cease to exist, while some grow to become the norm. These frequencies can shift or stay constant based on the conditions the population lives in. If in Hardy-Weinberg equilibrium, the allele frequencies stay constant. Most populations, however, have bias from environmental factors, sexual preferences, other organisms, etc. We propose a stochastic Markov chain model to study allele progression across generations. In such a model, the allele frequencies in the next generation depend only on the frequencies in the current one.

We use this model to track a recessive allele …


Evaluating Ai Sentiment Analysis, Aakriti Shah Jan 2023

Evaluating Ai Sentiment Analysis, Aakriti Shah

Honors Program Theses

This paper presents a comparative analysis of human and AI performance on a sentiment analysis task involving the coding of qualitative data from community program transcripts. The results demonstrate promising but imperfect agreement between two AI models, Claude and Bing, versus three human annotators and one expert annotator using the Community Capitals framework categories. While both models achieved fair alignment with human judgment, confusion patterns emerged involving metaphorical language and text overlapping multiple categories. The findings provide a case study for benchmarking conversational AI systems against human baselines to reveal limitations and target improvements. Key gaps center around distinguishing between …


A Qualitative Analysis Of Construct Measurement Techniques Used In Industrial/Organizational Research, Benjamin Michael, Andrea F. Snell, Katie Rosneck Jan 2023

A Qualitative Analysis Of Construct Measurement Techniques Used In Industrial/Organizational Research, Benjamin Michael, Andrea F. Snell, Katie Rosneck

Williams Honors College, Honors Research Projects

This project aims to challenge the appropriateness of the methodological strategies and tools utilized within psychological research. We will look at the types of statistical modeling used and the context in which they are used, such as measurement modeling, confirmatory factor analysis, and bifactor analysis within survey development, as well as the use of psychological constructs such as extraversion and leadership. The objective of this research is to search for and recognize patterns from the content of some of the top journal articles in the field of industrial and organizational psychology. The information gained from analyzing the content of the …


Applications For Functional Data Analysis, Kacy D. Kane Jan 2023

Applications For Functional Data Analysis, Kacy D. Kane

Graduate Research Theses & Dissertations

Functional Data Analysis is often used in the study of data that exists over a continuum, such as time. There are two datasets that will be considered here. For the first study we have a dataset on the efficacy of a lobectomy in reduction or elimination of epileptic seizures in patients. After an initial analysis of the dataset from a multinomial model perspective, we found that there were outliers in our dataset. From there, we considered a Multinomial Mixture Model to aid in the detection of outliers. In our second dataset we are considering a social robotics dataset where the …


Macroeconomic Factors Influencing Foreign Direct Investment In Some Selected Countries In Africa, Richard Essel Mensah Jan 2023

Macroeconomic Factors Influencing Foreign Direct Investment In Some Selected Countries In Africa, Richard Essel Mensah

Graduate Research Theses & Dissertations

This paper investigates the possible factors that influence foreign direct investment inflow rate to Africa after controlling for other macroeconomic factors. Using the heterogenous Toeplitz mixed method on a sample of 23 countries from 1998 – 2020, we find evidence of the statistical significance of a relationship between the amount of trade done in Africa and the FDI inflow rate in Africa. We also find a statistical relationship between the labor force participation rate and the FDI inflow rate to Africa. Although the Fixed effect and GLM method did not find the relationship between LFP rate and FDI inflow to …


Impacts Of Covid-19 On Industrial Growth In The United States, Emily G. Warthman, Charles J. Landis Jan 2023

Impacts Of Covid-19 On Industrial Growth In The United States, Emily G. Warthman, Charles J. Landis

Williams Honors College, Honors Research Projects

COVID-19 has caused massive ramifications on all parts of life in the world and industry growth/decline is not immune to it. This report will analyze nine different industries’ profit and revenue from quarterly data during the years 2009-2022. Forecast models will be generated using various methods and different techniques of validating to predict the values from Q2 2020- Q4 2022 based on historical data. After which, a comparison will be conducted between those predicted values to the actual average revenue and profit generated by order of greatest error percentage made. Thorough research will then be completed to determine if there …


Modeling The Bidirectional Relationship Between Shared-Patient Physician Networks And Patient Longitudinal Treatment Patterns: Application To Physician Risky-Prescribing, Xin Ran Jan 2023

Modeling The Bidirectional Relationship Between Shared-Patient Physician Networks And Patient Longitudinal Treatment Patterns: Application To Physician Risky-Prescribing, Xin Ran

Dartmouth College Ph.D Dissertations

Risky-prescribing is a pressing public health concern in the United States. Opioids, benzodiazepines, and non-benzodiazepine sedative-hypnotics (sedative-hypnotics) are three commonly-prescribed but potentially risky drug groups, prescribed alone or in combination. Physician shared-patient networks provide a unique perspective in studying physician network characteristics and structures, as well as their association with the delivery of health care. Understanding how physician shared-patient networks are related to their prescribing may inform network-based interventions targeting risky-prescribing, which is yet to be fully studied.

We investigated patient receipt of risky prescriptions and physician risky-prescribing intensity through the scope of shared-patient networks. We used retrospective Medicare insurance …


การเปรียบเทียบวิธีการใส่ค่าสูญหาย ในการวิเคราะห์การถดถอยโลจิสติก เมื่อตัวแปรตามมีการสูญหายแบบนอนอิกนอร์เรเบิล, อภิชาติ ฉัตรเรืองเลิศ Jan 2023

การเปรียบเทียบวิธีการใส่ค่าสูญหาย ในการวิเคราะห์การถดถอยโลจิสติก เมื่อตัวแปรตามมีการสูญหายแบบนอนอิกนอร์เรเบิล, อภิชาติ ฉัตรเรืองเลิศ

Chulalongkorn University Theses and Dissertations (Chula ETD)

งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบวิธีการใส่ค่าสูญหาย ในการวิเคราะห์การถดถอยโลจิสติก เมื่อตัวแปรตามมีการสูญหายแบบนอนอิกนอร์เรเบิล วิธีการที่ใช้ศึกษา คือ วิธี Complete Case Analysis (CC) วิธี Mode Imputation (MODE) วิธี Expectation Maximization Algorithm (EM) วิธี Multiple Imputation (MI) วิธี Hard Cutoff Augmentation (HARDCUT) วิธี Parceling Augmentation (PARCELING) และวิธี Fuzzy Augmentation (FUZZY) งานวิจัยนี้ใช้การจำลองข้อมูลในการศึกษาตามขนาดของตัวอย่าง ร้อยละของการสูญหายของข้อมูล และระดับของการสูญหายแบบนอนอิกนอร์เรเบิล การจำลองข้อมูลในแต่ละสถานการณ์จะกระทำ 5,000 รอบ โดยมีเกณฑ์ที่ใช้เปรียบเทียบประสิทธิภาพของวิธีการใส่ค่าสูญหาย ได้แก่ ค่าเฉลี่ยของค่าเฉลี่ยความคลาดเคลื่อนกำลังสอง (Average Mean Squared Error: AMSE) ของค่าประมาณความน่าจะเป็นของการเกิดเหตุการณ์ที่สนใจ (P(Y = 1)) และค่าประสิทธิภาพสัมพัทธ์ (Relative Efficiency: RE) จากผลการทดลองสรุปได้ว่า ค่า AMSE จะลดลงเมื่อขนาดของตัวอย่างใหญ่ขึ้น และจะมีค่ามากขึ้นเมื่อร้อยละของการสูญหายของข้อมูลเพิ่มขึ้น เมื่อพิจารณาผลของระดับของการสูญหายแบบนอนอิกนอร์เรเบิลต่อค่า AMSE พบว่ามีเพียง AMSE ของวิธี MODE เท่านั้นที่มีแนวโน้มเพิ่มขึ้น เมื่อระดับของการสูญหายแบบนอนอิกนอร์เรเบิลเพิ่มขึ้น และเมื่อพิจารณาค่า RE โดยเปรียบเทียบ AMSE ของวิธี CC กับวิธีการใส่ค่าสูญหายวิธีอื่น พบว่า วิธี EM และวิธี FUZZY ให้ค่า AMSE เท่ากับ AMSE ของวิธี CC ในขณะที่ AMSE ของวิธีอื่น ๆ มีค่าน้อยกว่า AMSE ของวิธี CC


Exploring Information Leakage In Historical Stock Market Data, Edison Hua Jan 2023

Exploring Information Leakage In Historical Stock Market Data, Edison Hua

Dissertations and Theses

Information leakage is a major concern for traders who want to execute large orders without affecting the market price. In this paper, we explore the sources and effects of information leakage in historical stock market data using various methods and metrics. We first define information leakage as a pattern caused by a trader that would otherwise not occur without the trader’s activity. Using historical data, the direct impact of a potential large trade cannot be measured, but we consider a minimal impact large trade to be one that minimizes changes to the established trading data. We then analyze how information …


Shallow Water Coral Distribution And Its Response To Climate Change, Amaury De Jesus Jan 2023

Shallow Water Coral Distribution And Its Response To Climate Change, Amaury De Jesus

Dissertations and Theses

Shallow water corals are one of the main reef-building organisms that secrete carbonates as their skeletons, and therefore, are one of the major sinks of CO2 in the ocean. These reef builders are also very crucial to marine environments and human society. As the global energy demand continues rising, fossil fuel burning increases at a faster pace despite the increase in energy supply using clean and renewable energy. The increase of CO2 in the atmosphere has been shown to exacerbate global warming and may cause ocean acidification, threatening the habitat of shallow-water corals. Many recent observations show alarming signs of …


Application Of Sentiment Analysis And Machine Learning Techniques To Predict Daily Cryptocurrency Price Returns, Edward Wu Jan 2023

Application Of Sentiment Analysis And Machine Learning Techniques To Predict Daily Cryptocurrency Price Returns, Edward Wu

CMC Senior Theses

This paper examines the effects of social media sentiment relating to Bitcoin on the daily price returns of Bitcoin and other popular cryptocurrencies by utilizing sentiment analysis and machine learning techniques to predict daily price returns. Many investors think that social media sentiment affects cryptocurrency prices. However, the results of this paper find that social media sentiment relating to Bitcoin does not add significant predictive value to forecasting daily price returns for each of the six cryptocurrencies used for analysis and that machine learning models that do not assume linearity between the current day price return and previous daily price …


Modeling Growth And Stress Factors For Converted Silvopasture Systems In The Missouri Ozarks, Bailee N. Suedmeyer Jan 2023

Modeling Growth And Stress Factors For Converted Silvopasture Systems In The Missouri Ozarks, Bailee N. Suedmeyer

Graduate Theses/Dissertations

Silvopasture systems are becoming increasingly popular among sustainable agriculture ranchers, due to the increase in knowledge of benefits to the cattle and ability to grow cool season grasses beneath the canopy. This project focuses on the forest crop aspect of silvopasture systems from monitoring of the health of the trees over time to recommendations for thinning management to keep it functioning as viable silvopasture. The study site consists of five acres of upland hardwood forest area in Southern Missouri with 18 monumented fixed area plots. Arial and ground data was collected at each plot throughout the growing season, along with …


The Influence Of Instrumental Sources Of Variance On Mass Spectral Comparison Algorithms, Isabel Cristina Galvez Valencia Jan 2023

The Influence Of Instrumental Sources Of Variance On Mass Spectral Comparison Algorithms, Isabel Cristina Galvez Valencia

Graduate Theses, Dissertations, and Problem Reports (ETD)

Current search algorithms for the identification of substances based only on their electron ionization mass spectra provide the correct compound as their top result approximately 80% of the time. One contributing factor to the ~20% deviation in the first-hit recognition rate is that traditional algorithms work by comparing the unknown spectrum to an ‘ideal’ or consensus spectrum of each reference compound. The inclusion of replicate reference spectra in a database has been shown to improve the probability of ranking the correct identity in the number one position, but the variance in ion abundances caused by different conditions or different instruments …


Investigating Collaborative Explainable Ai (Cxai)/Social Forum As An Explainable Ai (Xai) Method In Autonomous Driving (Ad), Tauseef Ibne Mamun Jan 2023

Investigating Collaborative Explainable Ai (Cxai)/Social Forum As An Explainable Ai (Xai) Method In Autonomous Driving (Ad), Tauseef Ibne Mamun

Dissertations, Master's Theses and Master's Reports

Explainable AI (XAI) systems primarily focus on algorithms, integrating additional information into AI decisions and classifications to enhance user or developer comprehension of the system's behavior. These systems often incorporate untested concepts of explainability, lacking grounding in the cognitive and educational psychology literature (S. T. Mueller et al., 2021). Consequently, their effectiveness may be limited, as they may address problems that real users don't encounter or provide information that users do not seek.

In contrast, an alternative approach called Collaborative XAI (CXAI), as proposed by S. Mueller et al (2021), emphasizes generating explanations without relying solely on algorithms. CXAI centers …


Additive P-Value Combination Test, Xing Ling Jan 2023

Additive P-Value Combination Test, Xing Ling

Dissertations, Master's Theses and Master's Reports

This dissertation includes four Chapters. A brief description of each chapter is organized as follows.

In Chapter 1, some developments on multiple hypotheses tests are introduced. Some preliminaries about the definition and the assumption are included.

In Chapter 2, a Stable Combination Test is proposed to combine $p$-values from multiple hypotheses tests. We show the proposed method controls the family-wise error rate at the target level and maintains asymptotically optimal power even when the elementary p-values from the individual hypotheses are dependent.

In Chapter 3, a deeper dig into the additive p-value combination test is performed. A common idea behind …


Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun Jan 2023

Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun

Theses and Dissertations

Noncompliance to treatment assignment is widespread in randomized trials and presents challenges in causal inference. In the presence of noncompliance, the most commonly estimated effect of treatment assignment, also known as intent-to-treat (ITT) effect, is biased. Of interest in this setting is the complier average causal effect (CACE), the ITT effect among compliers. Further complication arises when the outcome variable is partially observed.

My research focuses on estimating the distribution of a site-specific CACE in a multisite randomized controlled trial (MRCT) by maximum likelihood (ML). Assuming compliance missing at random (MAR). We express the likelihood as an integral with respect …


Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd Jan 2023

Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd

Theses and Dissertations

In 2015, Open Science Framework directly replicated 100 psychology studies and found astonishingly low replication rates. Since, researchers have suggested factors that may have influenced the low rates, including the metrics used to assess replications. The definitions used to decide whether a replication study was successful all suffer from flaws. Therefore, we propose a new metric for assessing replication that can estimate the likelihood a study successfully replicated rather than forcing a binary choice and accounts for study design limitations.

Using equivalence study techniques, we first propose a new metric to assess replication, defining a successful replication as one where …


Early Termination In Phase Ii Clinical Trials: Admissible Designs Using Decreasingly Informative Priors, Chen Wang Jan 2023

Early Termination In Phase Ii Clinical Trials: Admissible Designs Using Decreasingly Informative Priors, Chen Wang

Theses and Dissertations

In Phase II clinical trials, Thall and Simon’s Bayesian posterior probability design is commonly implemented to allow for an early termination to determine whether a new treatment warrants further investigation in a larger-scale Phase III trial; this in turn requires a pre-selected prior distribution based on known clinical opinion or historical information. Moreover, this Bayesian approach can result in an issue of inflating type I error rate by monitoring interim data to inform early termination decisions. Alternatively, a Bayesian approach with the decreasingly informative prior (DIP), which is an informative yet skeptical prior, can be implemented to overcome the contentious …


Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli Jan 2023

Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli

Theses and Dissertations

Investigations into the association between chemical exposure and health outcomes are increasingly focused on the role of chemical mixtures, as opposed to individual chemicals. The analysis of chemical mixture data required the development of novel statistical methods, one of these being Bayesian group index regression. A statistical challenge common to all chemical mixture analyses is the ubiquitous presence of below detection limit (BDL) data. We propose an extension of Bayesian group index regression that treats both regression effects and missing BDL observations as parameters in a model estimated through a Markov Chain Monte Carlo algorithm that we refer to as …


Integrative Post-Gwas Analyses Of Psychiatric Disorders: Identifying Putative Risk Genes And Gene Sets Using Transcriptome, Proteome And Methylome Information, Huseyin Gedik Jan 2023

Integrative Post-Gwas Analyses Of Psychiatric Disorders: Identifying Putative Risk Genes And Gene Sets Using Transcriptome, Proteome And Methylome Information, Huseyin Gedik

Theses and Dissertations

Genome-wide association studies (GWAS) of psychiatric disorders (PD) yield numerous loci with significant signals, but often they do not implicate specific protein coding genes. Because GWAS risk loci are enriched in expression/protein/methylation quantitative loci (e/p/mQTL, hereafter xQTL), transcriptome/proteome/methylome-wide association studies (T/P/MWAS, hereafter XWAS), which integrate information from GWAS and x-level (mRNA, protein or DNA methylation levels) coming from largest xQTL studies, can link GWAS signals to effects on specific genes. For gene level analyses, researchers use mendelian randomization (MR) methods to fine-map the association between x-levels and trait. However, none of the previous studies ever jointly analyzed XWAS of multiple …


Dynamic Overnight Effect On Next Day Stock Market Forecasting, Thomas J. Lee Jan 2023

Dynamic Overnight Effect On Next Day Stock Market Forecasting, Thomas J. Lee

Graduate Research Theses & Dissertations

Using a cross section of stocks that have high frequency trading data from 2007 to 2018, we document whether various intraday momentum patterns found in the financial literature over the years continue to hold over time. The first half hour return on the market is often seen as having predictive power over the last half hour of trading, or overnight returns are thought to reverse in the next day's first half hour of trading. We find that while there is some evidence for these patterns, especially in the earlier years, these patterns tend to weaken over time as investors take …


The Impact Of Faculty Composition On Cost Per Student: A Mixed Model Approach, Arun Sleeba Jan 2023

The Impact Of Faculty Composition On Cost Per Student: A Mixed Model Approach, Arun Sleeba

Graduate Research Theses & Dissertations

This thesis aims to explore whether research universities in the United States, specifically those classified as Carnegie I or II institutions, utilize part time contingent faculty(rPTF) as a cost saving strategy. Additionally, it sought to determine if there was a differential impact on total costs when comparing public and private universities. Employing a linear mixed effects model with random intercept and slopes, this study analyzed the relationship between rPTF (ratio of part-time to total faculty) and total cost. This study did not provide substantial evidence to support the notion despite observing a negative correlation between rPTF and total cost. Regarding …


Enhanced Maximum Likelihood Models For Underreported Variables: Extending To Multiple Claims Dimension, Shalaka Sudhanshu Sarpotdar Jan 2023

Enhanced Maximum Likelihood Models For Underreported Variables: Extending To Multiple Claims Dimension, Shalaka Sudhanshu Sarpotdar

Graduate Research Theses & Dissertations

This thesis builds upon the foundations laid out in Xia et al. [2023], which explored the utilizationof Maximum Likelihood approach to model misrepresentation data in Generalized Linear Models (GLM) ratemaking models. We introduce the concept of “underreported variables”, a form of insurance misrepresentation where insured individuals provide inaccurate information about risk factors that influence insurance eligibility, premiums, and insured amounts. Unlike fraudulent misrepresentation, underreported variables arise from a lack of awareness regarding the insured’s mental and physical health conditions, rather than fraudulent intent. The study rigorously tests the proposed model using health insurance data and extends its applicability to other …