Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (12)
- Artificial Intelligence and Robotics (7)
- Applied Statistics (5)
- Statistics and Probability (5)
- Social and Behavioral Sciences (4)
-
- Business (3)
- Engineering (3)
- Statistical Models (3)
- Theory and Algorithms (3)
- Business Analytics (2)
- Econometrics (2)
- Economics (2)
- Sports Studies (2)
- Statistical Methodology (2)
- Applied Mathematics (1)
- Arts and Humanities (1)
- Astrophysics and Astronomy (1)
- Atmospheric Sciences (1)
- Bioinformatics (1)
- Biostatistics (1)
- Business Intelligence (1)
- Categorical Data Analysis (1)
- Communication (1)
- Communication Technology and New Media (1)
- Computational Linguistics (1)
- Computer Engineering (1)
- Data Storage Systems (1)
- Earth Sciences (1)
- Institution
-
- Southern Methodist University (4)
- California Polytechnic State University, San Luis Obispo (3)
- Universitas Negeri Yogyakarta (3)
- Chapman University (2)
- Syracuse University (2)
-
- Belmont University (1)
- CCT College Dublin (1)
- Central Washington University (1)
- City University of New York (CUNY) (1)
- Clemson University (1)
- Dartmouth College (1)
- East Tennessee State University (1)
- Louisiana State University (1)
- Michigan Technological University (1)
- Old Dominion University (1)
- Purdue University (1)
- Southern Adventist University (1)
- The University of Akron (1)
- Universitas Negeri Malang (1)
- University of Arkansas, Fayetteville (1)
- University of Central Florida (1)
- University of Missouri, St. Louis (1)
- West Virginia University (1)
- Publication
-
- SMU Data Science Review (4)
- Elinvo (Electronics, Informatics, and Vocational Education) (3)
- Master's Theses (3)
- Sport Management - All Scholarship (2)
- All Dissertations (1)
-
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Science Faculty Scholarship (1)
- Computer Science Faculty Works (1)
- Dartmouth College Undergraduate Theses (1)
- Data Science Undergraduate Honors Theses (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Dissertations, Theses, and Capstone Projects (1)
- Engineering Faculty Articles and Research (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors Undergraduate Theses (1)
- ICT (1)
- Knowledge Engineering and Data Science (1)
- LSU Doctoral Dissertations (1)
- MS in Computer Science Project Reports (1)
- Mathematics & Statistics Faculty Publications (1)
- SPARK Symposium Presentations (1)
- The Journal of Purdue Undergraduate Research (1)
- Undergraduate Honors Theses (1)
- Williams Honors College, Honors Research Projects (1)
- Publication Type
Articles 1 - 30 of 32
Full-Text Articles in Data Science
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Master's Theses
Unsupervised clustering often faces the challenge of determining the correct number of clusters in the absence of a true target variable. Traditional methods such as the Elbow Method and the Silhouette Score can produce ambiguous results and rely on assumptions about cluster shape or separation. To address this, we created the Clustering Rivals and Buddies (CRAB) algorithm which evaluates clusters based on stability across multiple subsamples. CRAB uses pairwise classifications to identify points that consistently group together called “Buddies” and points that remain separated called “Rivals.” Applied with K-means, CRAB accurately recovers underlying cluster structures in both spherical and non-spherical …
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Master's Theses
Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …
Logistic-T Multinomial Mixture Model For Clustering For Microbiome Data, Wenshu Dai, Yuan Fang, Sanjeena Subedi
Logistic-T Multinomial Mixture Model For Clustering For Microbiome Data, Wenshu Dai, Yuan Fang, Sanjeena Subedi
Mathematics & Statistics Faculty Publications
The logistic-normal multinomial distribution has been used for modelling microbiome data obtained from high-throughput sequencing technologies, which are compositional in nature. A logistic-normal multinomial distribution is a hierarchical multinomial distribution that assumes the latent variable which are the additive log-ratio (ALR) transformed proportions in a multinomial distribution follows a Gaussian distribution. Model-based clustering algorithms have also been developed for clustering microbiome data based on the logistic-normal models. However, the Gaussian assumption may violated when the ALR transformed variable exhibit heavy-tailed distributions or has outliers. Our study introduces a novel mixture of logistic-t multinomial models that effectively address these challenges. Utilizing …
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
Honors Undergraduate Theses
With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Sport Management - All Scholarship
The purpose of this study was to identify player types that exist in the modern National Basketball Association (NBA), test whether player types are paid differently controlling for performance and other factors and construct successful rosters with cheaper payrolls.
We collected performance statistics and salary data for players and teams across five seasons (2018-19 to 2022-23). Cluster analysis is leveraged to group together player-seasons to identify the player types that exist in the NBA. Linear regression models are run to test for differences in pay by cluster membership while controlling for performance, age, and contractual details. Linear programming simulation models …
Frequent Itemset Mining With Tidyclust In R, Andrew D. Kerr
Frequent Itemset Mining With Tidyclust In R, Andrew D. Kerr
Master's Theses
Unsupervised learning is closely associated with clustering, however other methods fall under this umbrella such as data mining. In R, the tidyclust package provides a unified interface for clustering models, yet lacks support for data mining. This thesis addresses this gap by introducing the Apriori and ECLAT algorithms into tidyclust, with a focus on frequent itemset mining. Unlike traditional clustering models, frequent itemsets produce groupings of column variables, rather than cluster labels or partitions of observations. To address this, a novel clustering approach is proposed: items (columns) are grouped based on their ”dominant” frequent itemset. A key contribution is a …
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster
Undergraduate Honors Theses
Machine learning is a method that employs statistical algorithms to identify patterns and make predictions from data. This study applies machine learning techniques to analyze data from Major League Baseball (MLB) teams between 1998 and 2024, with the goal of determining which factors strongly influence a team's likelihood of reaching the postseason and in accurately predicting the teams that do and do not qualify for the postseason. Data exploration and unsupervised machine learning methods such as clustering were used to identify underlying patterns in team performance metrics and determine potential significant contributors to team success. Many different supervised learning methods …
Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova
Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova
Engineering Faculty Articles and Research
Fire Influence on Regional to Global Environments and Air Quality (FIREX-AQ) was a field campaign aimed at better understanding the impact of wildfires and agricultural fires on air quality and climate. The FIREX-AQ campaign took place in August 2019 and involved two aircraft and multiple coordinated satellite observations. This study applied and evaluated a self-supervised machine learning (ML) method for the active fire and smoke plume identification and tracking in the satellite and sub-orbital remote sensing datasets collected during the campaign. Our unique methodology combines remote sensing observations with different spatial and spectral resolutions. With as much as a 10% …
Spotify Recommender Using Content Filtering, Colin M. Anderson
Spotify Recommender Using Content Filtering, Colin M. Anderson
SPARK Symposium Presentations
This project is a music recommender system that analyzes Spotify data to establish relationships between songs by analyzing their musical components. A user can receive recommendations through two methods. Firstly, the user can select a song from the database that they enjoy, and the system will provide them with a list of recommendations, as well as predict the genre of the song that they have entered. Secondly, the user can set personalized values for various musical features (i.e. energy, danceability). In either case, by selecting a number 1 to 5 recommendations that the user would like, they will get that …
Discovering Latent Themes: Mixed-Methods Comparative Analysis Of Topic Extraction And Clustering., Laura Byrne
Discovering Latent Themes: Mixed-Methods Comparative Analysis Of Topic Extraction And Clustering., Laura Byrne
ICT
This study investigates whether clustering and topic modeling can uncover themes within the 167 verses of the King James Version of the Book of Esther. A standardized preprocessing pipeline was applied, and TF-IDF and sentence-embedding feature spaces were used to evaluate topic extraction (NMF, GSDMM, BERTopic, Top2Vec) and clustering (HDBSCAN, DBSCAN, Gaussian Mixtures, Agglomerative) using coherence, cluster validity, lexical distinctiveness, and cross-model similarity metrics. NMF provided the most interpretable topics, GSDMM yielded compact low-overlap topics, BERTopic offered entity-centered groupings, and Top2Vec highlighted broad thematic regions. Cross-model analysis revealed overlapping motifs, notably decree and banquet themes, but clusters and topics did …
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Dissertations, Master's Theses and Master's Reports
Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …
Where To Build Food Banks: A Machine Learning Approach, Gavin Ruan
Where To Build Food Banks: A Machine Learning Approach, Gavin Ruan
The Journal of Purdue Undergraduate Research
Over 44 million Americans currently suffer from food insecurity, of whom 13 million are children. Food insecurity has been shown to cause a wide range of both physical and developmental issues. Across the United States, thousands of food banks and pantries serve as vital sources of food and other forms of aid for food-insecure families. By optimizing food bank locations, food banks and their resources would become more accessible to families who desperately require it. The aim of this paper is to build a machine learning framework that is able to optimize food bank locations and to consider factors such …
The Importance Of Data Preparation In A Data Science Problem, Sophia Beard
The Importance Of Data Preparation In A Data Science Problem, Sophia Beard
Data Science Undergraduate Honors Theses
This study is going to be based on an inventory outlier automation data science problem that is being solved to identify and prescribe inventory level outliers to help keep shelves stocked in terms of beverages. The objective of this paper will address why it is so important to understand the data that is involved in a particular data science problem and how planning ahead ensures a successful outcome in the data science world. In this data science project, Spatiotemporal Outlier Analysis for Inventory Intervention Automation, it was crucial for the team to understand, research, and visualize the data we were …
Optimizing Nba Roster Construction, Nick R. Riccardi
Optimizing Nba Roster Construction, Nick R. Riccardi
Sport Management - All Scholarship
This study aims to quantify the effect that complementary player types have on team success in the National Basketball Association. Using cluster analysis, player-seasons are redefined from their traditional basketball positions to better encompass the roles that players play. For the 10 seasons of data, the best player for each of the 30 teams in the league is determined and teams are grouped based on the cluster of their best player. Ordinary Least Squares regressions are performed to test what player types fit together best. The results of this study show the importance of complementary workers to a firm’s success.
Static Malware Family Clustering Via Structural And Functional Characteristics, David George, Andre Mauldin, Josh Mitchell, Sufiyan Mohammed, Robert Slater
Static Malware Family Clustering Via Structural And Functional Characteristics, David George, Andre Mauldin, Josh Mitchell, Sufiyan Mohammed, Robert Slater
SMU Data Science Review
Static and dynamic analyses are the two primary approaches to analyzing malicious applications. The primary distinction between the two is that the application is analyzed without execution in static analysis, whereas the dynamic approach executes the malware and records the behavior exhibited during execution. Although each approach has advantages and disadvantages, dynamic analysis has been more widely accepted and utilized by the research community whereas static analysis has not seen the same attention. This study aims to apply advancements in static analysis techniques to demonstrate the identification of fine-grained functionality, and show, through clustering, how malicious applications may be grouped …
Comparing Igneous Geochemical Data From Hawaii And Southern California Via Machine Learning, Miro Manestar
Comparing Igneous Geochemical Data From Hawaii And Southern California Via Machine Learning, Miro Manestar
MS in Computer Science Project Reports
Bi-plots are commonly used in geochemical analyses. However, their use can become cumbersome in the case of multi-variate analyses. Therefore, this thesis explores the application of unsupervised machine learning techniques, specifically PCA and K-Means, to analyze large geochemical data sets from two distinct regions, Hawaii and the \acrfull{prb} in Southern California. The IBM Foundational Methodology for Data Science was utilized to ensure proper data preparation and analysis. PCA provided dimensionality reduction, revealing which features correlated most strongly with variances within the data. K-Means clustering allowed for deeper interpretation of the data. The analysis yielded valuable insights into the composition and …
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War, Yucheol Shin, Yvan Sojdehei, Limin Zheng, Brad Blanchard
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War, Yucheol Shin, Yvan Sojdehei, Limin Zheng, Brad Blanchard
SMU Data Science Review
The Ukrainian-Russian war has garnered significant attention worldwide, with fake news obstructing the formation of public opinion and disseminating false information. This scholarly paper explores the use of unsupervised learning methods and the Bidirectional Encoder Representations from Transformers (BERT) to detect fake news in news articles from various sources. BERT topic modeling is applied to cluster news articles by their respective topics, followed by summarization to measure the similarity scores. The hypothesis posits that topics with larger variances are more likely to contain fake news. The proposed method was evaluated using a dataset of approximately 1000 labeled news articles related …
Market Segmentation And Recency Frequency Monetary Value Analysis For A Freemium Mobile Game, Satvik Ajmera, Taylor Bonar, Dylan Scott, Carol Miu, Alana Manuel
Market Segmentation And Recency Frequency Monetary Value Analysis For A Freemium Mobile Game, Satvik Ajmera, Taylor Bonar, Dylan Scott, Carol Miu, Alana Manuel
SMU Data Science Review
Bricks ‘N Balls is a freemium game that relies on in-app purchases and ad monetization from users to be profitable at no upfront cost to the players. This study explores how in-game data analytics and purchase data can be used to segment players. Features taken into consideration for segmentation include past purchasing habits along with the players interactions within the missions. This study uses the Recency Frequency Monetary Value (RFM) framework to extract insights on player purchasing behavior to segment players into clusters and predict how much users will spend in the future.
Unsupervised Contrastive Representation Learning For Knowledge Distillation And Clustering, Fei Ding
Unsupervised Contrastive Representation Learning For Knowledge Distillation And Clustering, Fei Ding
All Dissertations
Unsupervised contrastive learning has emerged as an important training strategy to learn representation by pulling positive samples closer and pushing negative samples apart in low-dimensional latent space. Usually, positive samples are the augmented versions of the same input and negative samples are from different inputs. Once the low-dimensional representations are learned, further analysis, such as clustering, and classification can be performed using the representations. Currently, there are two challenges in this framework. First, the empirical studies reveal that even though contrastive learning methods show great progress in representation learning on large model training, they do not work well for small …
A Machine Learning Approach To Revenue Generation Within The Professional Hair Care Industry, Alexander K. Sepenu, Linda Eliasen
A Machine Learning Approach To Revenue Generation Within The Professional Hair Care Industry, Alexander K. Sepenu, Linda Eliasen
SMU Data Science Review
The cosmetic and beauty industry continues to grow and evolve to satisfy its patrons. In the United States, the industry is heavily science-driven, innovative, and fast-paced, suggesting that to remain productive and profitable, companies must seek smart alternatives to their current modus operandi or risk losing out on this multi-billion-dollar industry to fierce competition. In this paper, the authors seek to utilize machine learning models such as clustering and regression to improve the efficiency of current sales and customer segmentation models to help HairCo (pseudonym for confidentiality), a professional hair products manufacturer, strategize their marketing and sales efforts for revenue …
A Comparison Of K-Means And Agglomerative Clustering For Users Segmentation Based On Question Answerer Reputation In Brainly Platform, Puji Winar Cahyo, Landung Sudarmana
A Comparison Of K-Means And Agglomerative Clustering For Users Segmentation Based On Question Answerer Reputation In Brainly Platform, Puji Winar Cahyo, Landung Sudarmana
Elinvo (Electronics, Informatics, and Vocational Education)
Brainly is a question and answer (Q&A) site that students can use as a media for questions and answers. Students can also use Brainly to find and share educational information that helps students solve their homework problems. In Brainly, users can answer questions according to their interests. However, it could be that the interest is not necessarily following the competencies possessed. It causes many answers to the questions given not to have a high rating because the answers given are of low quality to be prioritized as the main answer. This study aims to apply the K-Means and Agglomerative Clustering …
Connecting The Dots: The Boons And Banes Of Network Modeling, Sharlee Climer
Connecting The Dots: The Boons And Banes Of Network Modeling, Sharlee Climer
Computer Science Faculty Works
Network modeling transforms data into a structure of nodes and edges such that edges represent relationships between pairs of objects, then extracts clusters of densely connected nodes in order to capture high-dimensional relationships hidden in the data. This efficient and flexible strategy holds potential for unveiling complex patterns concealed within massive datasets, but standard implementations overlook several key issues that can undermine research efforts. These issues range from data imputation and discretization to correlation metrics, clustering methods, and validation of results. Here, we enumerate these pitfalls and provide practical strategies for alleviating their negative effects. These guidelines increase …
Piecewise Linear Manifold Clustering, Artyom Diky
Piecewise Linear Manifold Clustering, Artyom Diky
Dissertations, Theses, and Capstone Projects
This work studies the application of topological analysis to non-linear manifold clustering. A novel method, that exploits the data clustering structure, allows to generate a topological representation of the point dataset. An analysis of topological construction under different simulated conditions is performed to explore the capabilities and limitations of the method, and demonstrated statistically significant improvements in performance. Furthermore, we introduce a new information-theoretical validation measure for clustering, that exploits geometrical properties of clusters to estimate clustering compressibility, for evaluation of the clustering goodness-of-fit without any prior information about true class assignments. We show how the new validation measure, when …
Exploring The Long Tail, Joseph H. Hajjar
Exploring The Long Tail, Joseph H. Hajjar
Dartmouth College Undergraduate Theses
The migration of datasets online has created a near-infinite inventory for big name retailers such as Amazon and Netflix, giving rise to recommendation systems to assist users in navigating the massive catalog. This has also allowed for the possibility of retailers storing much less popular, uncommon items which would not appear in a more traditional brick-and-mortar setting due to the cost of storage. Nevertheless, previous work has highlighted the profit potential which lies in the so-called "long tail'' of niche, unpopular items. Unfortunately, due to the limited amount of data in this subset of the inventory, recommendation systems often struggle …
Novel Applications Of Statistical And Machine Learning Methods To Analyze Trial-Level Data From Cognitive Measures, Chelsea Parlett
Novel Applications Of Statistical And Machine Learning Methods To Analyze Trial-Level Data From Cognitive Measures, Chelsea Parlett
Computational and Data Sciences (PhD) Dissertations
Many cognitive tasks and measures can benefit from trial-level analyses including Item Response Theory models as well as other Bayesian and Machine Learning models. Specifically, this dissertation focuses mainly on task-based measures of metamemory and how within-set variability as well as item-level characteristics can improve the inferences researchers make about these measures.First, a clustering analysis of judgements of learning across a task is examined in order to detect different participant strategies on a metamemory task and whether strategy use differs by age. Second, the benefits of using item response theory models to analyze both individual and item-level differences in metamemory …
Clustering Data To Classify Hearthstone Decks, Tim Inzitari
Clustering Data To Classify Hearthstone Decks, Tim Inzitari
Williams Honors College, Honors Research Projects
The esports game of "Hearthstone" is a collectible card game with a competitive format that has every team submit 4 decks of 30 cards each. Using K-Means clustering an adaptable way to group data for classifying can be made that works well in every update of the game. This system will take in a list of decks and cluster them to easily classify large amounts of information in a timely fashion. This system will be able to be used by the Universities esports department for years to come to aid the preparation of "Hearthstone" matches. This model uses qualities about …
Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang
Identification And Classification Of Radio Pulsar Signals Using Machine Learning, Di Pang
Graduate Theses, Dissertations, and Problem Reports (ETD)
Automated single-pulse search approaches are necessary as ever-increasing amount of observed data makes the manual inspection impractical. Detecting radio pulsars using single-pulse searches, however, is a challenging problem for machine learning because pul- sar signals often vary significantly in brightness, width, and shape and are only detected in a small fraction of observed data.
The research work presented in this dissertation is focused on development of ma- chine learning algorithms and approaches for single-pulse searches in the time domain. Specifically, (1) We developed a two-stage single-pulse search approach, named Single- Pulse Event Group IDentification (SPEGID), which automatically identifies and clas- …
Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott
Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott
LSU Doctoral Dissertations
Modern human-machine systems such as microservices rely upon agile engineering practices which require changes to be tested and released more frequently than classically engineered systems. A critical step in the testing of such systems is the generation of realistic workloads or load testing. Generated workload emulates the expected behaviors of users and machines within a system under test in order to find potentially unknown failure states. Typical testing tools rely on static testing artifacts to generate realistic workload conditions. Such artifacts can be cumbersome and costly to maintain; however, even model-based alternatives can prevent adaptation to changes in a system …
Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski
Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski
Knowledge Engineering and Data Science
Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …
Topik Modeling Penelitian Dosen Jptei Uny Pada Google Scholar Menggunakan Latent Dirichlet Allocation, Akhsin Nurlayli, Moch. Ari Nasichuddin
Topik Modeling Penelitian Dosen Jptei Uny Pada Google Scholar Menggunakan Latent Dirichlet Allocation, Akhsin Nurlayli, Moch. Ari Nasichuddin
Elinvo (Electronics, Informatics, and Vocational Education)
The mapping of research topics for lecturers is necessary to determine the research tendencies in a department or study program. This study aims to implement topic modeling in the publication titles of the Department of Electronics and Informatics Education Engineering of Universitas Negeri Yogyakarta (JPTEI UNY) lecturers taken from Google Scholar. The method used for topic modeling is the Latent Dirichlet Allocation (LDA). LDA is a generative probabilistic model for finding the semantic structure of a corpus collection based on the hierarchical bayesian analysis. After the topic modeling process, the results showed that JPTEI UNY lecturers tend to have …