Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Smith College (49)
- Southern Methodist University (37)
- Kennesaw State University (27)
- Central Bank of Nigeria (23)
- Old Dominion University (23)
-
- City University of New York (CUNY) (19)
- University of Central Florida (16)
- Chapman University (13)
- West Virginia University (13)
- Department of Primary Industries and Regional Development, Western Australia (12)
- Illinois State University (12)
- California Polytechnic State University, San Luis Obispo (11)
- East Tennessee State University (11)
- LSU Health New Orleans (11)
- Claremont Colleges (10)
- Georgia Southern University (10)
- University of Arkansas, Fayetteville (10)
- Embry-Riddle Aeronautical University (9)
- University of Kentucky (9)
- Virginia Commonwealth University (9)
- Rochester Institute of Technology (7)
- Clemson University (6)
- Dartmouth College (6)
- Purdue University (6)
- Binghamton University (5)
- Murray State University (5)
- The University of Akron (5)
- University of Louisville (5)
- University of New Mexico (5)
- University of South Florida (5)
- Keyword
-
- Machine Learning (36)
- Machine learning (36)
- Statistics (27)
- Deep learning (15)
- Data Science (14)
-
- Data science (13)
- Classification (12)
- Time series (11)
- COVID-19 (10)
- Deep Learning (9)
- Artificial Intelligence (8)
- Forecasting (8)
- Regression (8)
- Prediction (7)
- Western Australia (7)
- Logistic regression (6)
- Neural Network (6)
- Clustering (5)
- Data analysis (5)
- Natural language processing (5)
- Sentiment analysis (5)
- Simulation (5)
- Time Series (5)
- Analysis (4)
- Analytics (4)
- Baseball (4)
- Bayesian (4)
- Bioinformatics (4)
- Biostatistics (4)
- CNN (4)
- Publication Year
- Publication
-
- Statistical and Data Sciences: Faculty Publications (45)
- SMU Data Science Review (27)
- CBN Journal of Applied Statistics (JAS) (20)
- Symposium of Student Scholars (20)
- Electronic Theses and Dissertations (18)
-
- Theses and Dissertations (17)
- Master's Theses (11)
- Annual Symposium on Biomathematics and Ecology Education and Research (10)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (10)
- Mathematics & Statistics Faculty Publications (10)
- School of Public Health Faculty Publications (10)
- College of Graduate Studies: Theses & Dissertations (9)
- Data Science and Data Mining (9)
- Dissertations, Theses, and Capstone Projects (9)
- Articles (7)
- Computational and Data Sciences (PhD) Dissertations (7)
- Dissertations (7)
- CMC Senior Theses (6)
- All Dissertations (5)
- Honors College Theses (5)
- Northeast Journal of Complex Systems (NEJCS) (5)
- Publications and Research (5)
- Statistical Science Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Dartmouth College Ph.D Dissertations (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Doctor of Data Science and Analytics Dissertations (4)
- Electronic Theses & Dissertations (2024 - present) (4)
- Fisheries Research Articles (4)
- Honors Projects (4)
- Publication Type
- File Type
Articles 451 - 480 of 550
Full-Text Articles in Data Science
Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee
Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee
Statistical and Data Sciences: Faculty Publications
No abstract provided.
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr
Statistical and Data Sciences: Faculty Publications
We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.
A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity, Adewale Aremu Adesanya
A Transdisciplinary Analysis Of Just Transition Pathways To 100% Renewable Electricity, Adewale Aremu Adesanya
Dissertations, Master's Theses and Master's Reports
The transition to using clean, affordable, and reliable electrical energy is critical for enhancing human opportunities and capabilities. In the United States, many states and localities are engaging in this transition despite the lack of ambitious federal policy support. This research builds on the theoretical framework of the multilevel perspective (MLP) of sociotechnical transitions as well as the concept of energy justice to investigate potential pathways to 100 percent renewable energy (RE) for electricity provision in the U.S. This research seeks to answer the question: what are the technical, policy, and perceptual pathways, barriers, and opportunities for just transition to …
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee
CMC Senior Theses
This paper attempts to quantify predictive power of social media sentiment and financial data in stock prediction by utilizing a comprehensive set of stock-related fundamental and technical variables and social media sentiments. For conducting sentiment analysis, this study employs a pretrained finBERT model that provides three different sentiment classifications and respective softmax scores. Hence, the significance of these variables is evaluated with XGBoost regression and Shapley Additive exPlanations (SHAP) frameworks. Through investigating feature importance, this study finds that statistical properties of sentiment variables provide a stronger predictive power than a weighted sentiment score and that it is possible to quantify …
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard
CMC Senior Theses
Using a Twitter API, I gather and analyze tweets by performing sentiment analysis to solve the GOAT debate among professional athletes with the primary focus on comparing Michael Jordan and LeBron James. Athletes from the National Football League (NFL), the National Basketball Association (NBA), Major League Baseball (MLB), and the National Collegiate Athletic Association (NCAA) Division 1 Men's and Women's Basketball were selected to compare how sentiment polarity varies across sports. Sentiment polarity is measured by labeling text as "positive", "neutral", or "negative" which allows us to determine which athlete/sport is highly favored among the Twitter community when it comes …
Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington
Review Of Forecasting Univariate Time-Series Data With Application To Water-Energy Nexus Studies & Proposal Of Parallel Hybrid Sarima-Ann Model, Cory Sumner Yarrington
Graduate Theses, Dissertations, and Problem Reports (ETD)
The necessary materials for most human activities are water and energy. Integrated analysis to accurately forecast water and energy consumption enables the implementation of efficient short and long-term resource management planning as well as expanding policy and research possibilities for the supportive infrastructure. However, the integral relationship between water and energy (water-energy nexus) poses a difficult problem for modeling. The accessibility and physical overlay of data sets related to water-energy nexus is another main issue for a reliable water-energy consumption forecast. The framework of urban metabolism (UM) uses several types of data to build a global view and highlight issues …
Statistical And Machine Learning Approaches To Depressive Disorders Among Adults In The United States: From Factor Discovery To Prediction Evaluation, Minhwa Lee
Senior Independent Study Theses
According to the National Institutes of Mental Health (NIMH), depressive disorders (or major depression) are considered one of the most common and serious health risks in the United States. Our study focuses on extracting non-medical factors of depressive disorders diagnosis, such as overall health states, health risk behaviors, demography, and healthcare access, using the Behavioral Risk Factor Surveillance System (BRFSS) data set collected by the Centers for Disease Control and Prevention (CDC) in 2018.
We set the two objectives of our study about depressive disorders diagnosis in the United States as follows. First, we aim to utilize machine learning algorithms …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang
Self-Exciting Point Process For Modelling Terror Attack Data, Siyi Wang
Theses and Dissertations (Comprehensive)
Terrorism becomes more rampant in recent years because of separatism and extreme nationalism, which brings a serious threat to the national security of many countries in the world. The analysis of spatial and temporal patterns of terror data is significant in containing terrorism. This thesis focuses on building and applying a temporal point process called self-exciting point process to fit the terror data from 1970 to 2018 of 10 countries. The data come from the Global Terrorism database. Further, an application in predicting the number of terror events based on the self-exciting model is another main innovative idea, in which …
Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani
Modeling Multivariate Hopfield-Transformer Hawkes Process: Application To Sovereign Credit Default Swaps, Mohsen Bahremani
Theses and Dissertations (Comprehensive)
Hawkes process was evolved so that the past events contribute to the occurrence time of future events by self-exciting or mutually exciting. However, many real-world data do not follow the Hawkes process's assumptions (i.e., positivity, additivity, and exponential decay) and become more complex to be modeled by the traditional Hawkes processes, so the neural Hawkes process was developed to tackle the challenges. However, Recurrent Neural Networks (RNN) fail to capture long-term dependencies among multiple point processes, and Transformer Hawkes processes only address temporal characteristics of Hawkes processes. In this thesis, we proposed a combination of neural networks and Hawkes processes …
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Statistical Science Theses and Dissertations
This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.
In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …
Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi
Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi
Publications and Research
In Spring 2020, I did a project, "Decision Tree Predicting the Party of Legislators," and construct a decision tree model to predict legislators' parties' based on their votes. We also use this model to identify legislators who frequently voted against their parties. We used the legislators' roll call votes, Office of Clerk U.S. House of Representatives Data Sets (Categorical values) collected in 2018 and 2019. In this new project, We study the 2018 and 2019 vote data using Principal Component Analysis (PCA). The goal is to find a (compressed) model using unsupervised learning to distinguish the legislators' parties, and PCA …
Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman
Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman
Master's Theses
Several regions of the Western United States utilize statistical binary classification models to predict and manage debris flow initiation probability after wildfires. As the occurrence of wildfires and large intensity rainfall events increase, so has the frequency in which development occurs in the steep and mountainous terrain where these events arise. This resulting intersection brings with it an increasing need to derive improved results from existing models, or develop new models, to reduce the economic and human impacts that debris flows may bring. Any development or change to these models could also theoretically increase the ease of collection, processing, and …
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Statistical and Data Sciences: Faculty Publications
The advancement of scientific knowledge increasingly depends on ensuring that data-driven research is reproducible: that two people with the same data obtain the same results. However, while the necessity of reproducibility is clear, there are significant behavioral and technical challenges that impede its widespread implementation and no clear consensus on standards of what constitutes reproducibility in published research. We present fertile, an R package that focuses on a series of common mistakes programmers make while conducting data science projects in R, primarily through the RStudio integrated development environment. fertile operates in two modes: proactively, to prevent reproducibility mistakes from happening …
Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano
Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano
Honors College Theses
The work performed in this research aims to evaluate the performance of commercially available breathing masks in filtering airborne nanoparticles at manufacturing sites. Nanoparticles are found virtually anywhere, from dust in a worksite to a simple sneeze. Therefore, they pose a substantial threat to human health as their velocity and volatility are high. This research analyzes if current efforts of breathing masks to hinder nanoparticles are effective, especially at manufacturing sites. Data has been collected in order to analyze the behavior of nanoparticles and to measure nanoparticle levels at manufacturing sites and its working environment. Data is statistical in nature …
A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann
A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li
Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour
Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour
Journal of the Arab American University مجلة الجامعة العربية الامريكية للبحوث
This paper aimed to compare the modern methods of cash flow forecasting with the traditional ones. In other words, the researcher compared between the Probabilistic Neural Networks and Transfer Function. It is worth mentioning that cash flow forecasting , nowadays, is very important and helps the upper management plan, control, assess the performance and make decisions. More specifically, in this paper, the Artificial Neural networks were used to diagnose the nature of the cash flow for the next period of time and then forecast the cash flow. The experiment was conducted in The General company for Electricity Distribution in Baghdad. …
Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite
Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite
SMU Data Science Review
In this paper, modeling techniques for the forecasting of wind speed using historical values observed by Light Detection and Ranging (LIDAR) sensors in an offshore context are described. Both univariate time series and multivariate time series modeling techniques leveraging meteorological data collected simultaneously with the LIDAR data are evaluated for potential contributions to predictive ability. Accurate and timely ability to predict wind values is essential to the effective integration of wind power into existing power grid systems. It allows for both the management of rapid ramp-up / down of base production capacity due to highly variable wind power inputs and …
Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird
Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird
Statistical and Data Sciences: Faculty Publications
This paper presents an undergraduate machine learning course that emphasizes algorithmic understanding and programming skills while assuming no statistical training. Emphasizing the development of good habits of mind, this course trains students to be independent machine learning practitioners through an iterative, cyclical framework for teaching concepts while adding increasing depth and nuance. Beginning with unsupervised learning, this course is sequenced as a series of machine learning ideas and concepts with specific algorithms acting as concrete examples. This paper also details course organization including evaluation practices and logistics.
Machine Learning Applications For Drug Repurposing, Hansaim Lim
Machine Learning Applications For Drug Repurposing, Hansaim Lim
Dissertations, Theses, and Capstone Projects
The cost of bringing a drug to market is astounding and the failure rate is intimidating. Drug discovery has been of limited success under the conventional reductionist model of one-drug-one-gene-one-disease paradigm, where a single disease-associated gene is identified and a molecular binder to the specific target is subsequently designed. Under the simplistic paradigm of drug discovery, a drug molecule is assumed to interact only with the intended on-target. However, small molecular drugs often interact with multiple targets, and those off-target interactions are not considered under the conventional paradigm. As a result, drug-induced side effects and adverse reactions are often neglected …
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
SMU Data Science Review
In this paper, we explore a representation methodology for the compression of DNA isolates. Using lossless string compression via tokenization of frequently repeated segments of DNA, we reduce the length of the isolates to be counted as k-mers for classification. With this new representation, we apply a previously established feature sampling method to dramatically reduce the feature space. In understanding the genetic diversity, we also look at conserving biological function across these spaces. Using a random forest model we were able to predict the resistance or susceptibility of bacteria with 85-90\% accuracy, with a 30-50\% reduction in overall isolate length, …
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Mathematics Theses and Dissertations
Cell assemblies, defined as groups of neurons forming temporal spike coordination, are thought to be fundamental units supporting major cognitive functions. However, detecting cell assemblies is challenging since they can occur at a range of time scales and with a range of precisions, from synchronous spikes to co-variations in firing rate. In this dissertation, we use a recently published cell assembly detection (CAD) algorithm that is capable of detecting assemblies at a range of time scales and precisions. We first showed that the CAD method can be applied to sparser spike train data than what have previously been reported. This …
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Dissertations
In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.
First, to improve the prediction accuracy of learning …
A Novel Correction For The Adjusted Box-Pierce Test — New Risk Factors For Emergency Department Return Visits Within 72 Hours For Children With Respiratory Conditions — General Pediatric Model For Understanding And Predicting Prolonged Length Of Stay, Sidy Danioko
Computational and Data Sciences (PhD) Dissertations
This thesis represents the results of three research projects that underline the breadth and depth of my interests.
Firstly, I devoted some efforts to the well-known Box-Pierce goodness-of-fit tests for time series models which has been an important research topic over the last few decades. All previously proposed tests are focused on changes of the test statistics. Instead, I adopted a different approach that takes the best performing test and modifying the rejection region. Thus, I developed a semiparametric correction of the Adjusted Box-Pierce test that attains the best I error rates for all sample sizes and lags and outperforms …
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Student Works (2020-2029)
One of the approaches for structural health monitoring (SHM) consists of two major components, i.e. a network of sensors to collect the response data and an extraction method to obtain information on the structural health condition. Data mining (DM) is a novel data extraction technology which can employ for development of inverse analysis. Implementation of DM techniques in different areas of civil engineering has recently given very good results. However, application of DM in SHM is not used as much as expected, thus, many challenges are still ahead. Therefore, it is necessary to develop the applicability of DM in SHM. …
Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis
Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis
Dissertations and Theses (Open Access)
Tumor cells have heterogeneous genotypes, which drives progression and treatment resistance. Such genetic intratumor heterogeneity plays a role in the process of clonal evolution that underlies tumor progression and treatment resistance. Single-cell DNA sequencing is a promising experimental method for studying intratumor heterogeneity, but brings unique statistical challenges in interpreting the resulting data. Researchers lack methods to determine whether sufficiently many cells have been sampled from a tumor. In addition, there are no proven computational methods for determining the ploidy of a cell, a necessary step in the determination of copy number. In this work, software for calculating probabilities from …
Hoop Dreams: An Empirical Analysis Of The Gender Wage Gap In Professional Basketball, Hailey Dicicco
Hoop Dreams: An Empirical Analysis Of The Gender Wage Gap In Professional Basketball, Hailey Dicicco
Business and Economics Presentations
The gender wage gap is a very prominent point of discussion in the professional world, but in the sports world, it has taken the spotlight in recent years. One sport that has seen discussion and debate over salary differences is the National Basketball Association and Women’s National Basketball Association. In 2018, the average salary in the NBA was 6.4 million dollars, while the average salary in the WNBA was 71,635 dollars. A reason why these salaries are so differently is due to the amount of revenue that each league brings in. The NBA brings in roughly 7.4 billion dollars a …
Quantitatively Motivated Model Development Framework: Downstream Analysis Effects Of Normalization Strategies, Jessica M. Rudd
Quantitatively Motivated Model Development Framework: Downstream Analysis Effects Of Normalization Strategies, Jessica M. Rudd
Doctor of Data Science and Analytics Dissertations
Through a review of epistemological frameworks in social sciences, history of frameworks in statistics, as well as the current state of research, we establish that there appears to be no consistent, quantitatively motivated model development framework in data science, and the downstream analysis effects of various modeling choices are not uniformly documented. Examples are provided which illustrate that analytic choices, even if justifiable and statistically valid, have a downstream analysis effect on model results. This study proposes a unified model development framework that allows researchers to make statistically motivated modeling choices within the development pipeline. Additionally, a simulation study is …
Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen
Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen
Master's Theses
Current methods of production forecasting such as decline curve analysis (DCA) or numerical simulation require years of historical production data, and their accuracy is limited by the choice of model parameters. Unconventional resources have proven challenging to apply traditional methods of production forecasting because they lack long production histories and have extremely variable model parameters. This research proposes a data-driven alternative to reservoir simulation and production forecasting techniques. We create a proxy-well model for predicting cumulative oil production by selecting statistically significant well completion parameters and reservoir information as independent predictor variables in regression-based models. Then, principal component analysis (PCA) …