Open Access. Powered by Scholars. Published by Universities.®

Computer Engineering Commons

Open Access. Powered by Scholars. Published by Universities.®

Clustering

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 91

Full-Text Articles in Computer Engineering

Improving Approach Of Evolutionary Strategies For Clustering Technique Enhancement, Duaa Mahde Saleh, Hasanen S. Abdullah, Ahmad Zamsuri Jun 2026

Improving Approach Of Evolutionary Strategies For Clustering Technique Enhancement, Duaa Mahde Saleh, Hasanen S. Abdullah, Ahmad Zamsuri

Journal of Soft Computing and Computer Applications

The existence of the information has been the essential aspect of the whole society. Information is concentrated in all forms to be effectively utilized. Clustering — an unsupervised learning technique. It is based on data similarity that gives rise to issues in collection, challenges and instability in data structure. It proposes an advanced evolutionary method by combining two approaches. Firstly, it adopts the evolutionary approach and integrates the advantages between two methods to design one. Among them are Differential Evolution (DE) and Genetic Algorithm (GA), Evolutionary Strategy (ES) and Genetic Programming (GP), and Evolutionary Programming (EP) and Particle Swarm Optimization …


Computational Modeling For Automatic Superconducting Cavity Fault Prediction And Classification Using Time Series Signals, Md Monibor Rahman Aug 2025

Computational Modeling For Automatic Superconducting Cavity Fault Prediction And Classification Using Time Series Signals, Md Monibor Rahman

Electrical & Computer Engineering Theses & Dissertations

Processing multivariate time series signals collected from sensor networks is challenging because of complex temporal dependencies and non-stationarity. With the advent of artificial intelligence (AI) like machine learning and deep learning, it has become possible to process sensor-driven time series data more effectively than traditional statistical methods.

This dissertation aims to develop machine learning and deep learning models to address machine fault diagnosis using multivariate time series signals collected from the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. The first goal of the proposed work is to develop deep learning–based classification models and an unsupervised fault clustering approach …


Toward Strategy Identification And Subtask Decomposition In Task Exploration, Tom Odem Jan 2025

Toward Strategy Identification And Subtask Decomposition In Task Exploration, Tom Odem

Master's Projects

This research builds on work in anticipatory human-machine interaction, a subfield of human-machine interaction where machines can facilitate advantageous interactions by anticipating a user’s future state. The aim of this research is to further a machine’s understanding of user knowledge, skill, and behavior in pursuit of implicit coordination. A task explorer pipeline was developed that uses clustering techniques, paired with factor analysis and string edit distance, to automatically identify key global and local strategies that are used to complete tasks. Global strategies identify generalized sets of actions used to complete tasks, while local strategies identify sequences that used those sets …


Comparative Analysis Of Embedding Techniques With Clustering Algorithms For Malware Opcodes, Ayush Koul Jan 2025

Comparative Analysis Of Embedding Techniques With Clustering Algorithms For Malware Opcodes, Ayush Koul

Master's Projects

Malware detection and classification remain critical challenges in cybersecurity, especially as malicious software becomes increasingly sophisticated and prevalent. While much of the work involving embeddings has traditionally relied on supervised learning approaches, there is significant potential in leveraging unsupervised learning techniques to discern hidden structures in malware data. By employing embedding techniques to convert malware samples into high-dimensional vector representations, we can capture the subtle and complex patterns inherent in malicious code without relying on pre-labeled data. This unsupervised approach helps categorize malware into predefined malware families, greatly aiding in developing cybersecurity solutions. In contrast to traditional supervised models that …


A Method For Battlefield Situation Information Ontology Construction Based On Top-Down And Bottom-Up Integration, Cong Zhou, Sihang Zhou, Jian Huang, Dong Wang Oct 2024

A Method For Battlefield Situation Information Ontology Construction Based On Top-Down And Bottom-Up Integration, Cong Zhou, Sihang Zhou, Jian Huang, Dong Wang

Journal of System Simulation

Abstract: The construction of the unified expression model of battlefield situational information is challenging due to the complexity of data sources and the significant differences in data structures and expression methods. Ontologies, as semantic conceptual models, are often used to describe concepts, relationships, and attributes within knowledge domains. An ontology construction method for the battlefield situational information domain based on a top-down and bottom-top integration is proposed. The top-down method is used to construct the upper ontology, in which a conceptual hierarchy model with a clear top-down structure is designed to establish the hierarchical relationships and semantic associations. A bottom-up …


Just-In-Time Learning Energy Consumption Predictive Modeling Method In Multi-Condition Production Process, Sheng Wei, Yan Wang, Zhicheng Ji Jun 2024

Just-In-Time Learning Energy Consumption Predictive Modeling Method In Multi-Condition Production Process, Sheng Wei, Yan Wang, Zhicheng Ji

Journal of System Simulation

Abstract: Aiming at the problem that the global energy consumption prediction model is only suitable for part of the prediction sample and the model is computationally intensive, the idea of just-in-time learning is introduced, and the local weighted partial least squares method combined with the energy consumption model is used to establish a temporary local energy consumption prediction model. The inertia weights of the particle swarm algorithm are improved, considering the effects of particle fitness, number of iterations and population size on the convergence speed and convergence accuracy of the particle swarm algorithm, a nonlinear change adaptive inertia weight strategy …


Unsupervised Complex Condition Recognition Based On Stochastic Neighborhood Embedding, Lin Huang, Shanjun Liu, Wei Wang, Li Gong Jun 2024

Unsupervised Complex Condition Recognition Based On Stochastic Neighborhood Embedding, Lin Huang, Shanjun Liu, Wei Wang, Li Gong

Journal of System Simulation

Abstract: Modern industrial production equipment usually has a complex structure and runs alternately in different working conditions. Accurate working conditions identification based on monitoring data is the basis of health monitoring of the system, but the monitoring data of the system usually has a high dimension and a large data volume. To identify the complex equipment operating conditions, an unsupervised operating condition identification method based on stochastic neighborhood embedding is proposed. The stochastic neighborhood embedding algorithm can simultaneously preserve the local and global structural characteristics of the data, and also calculate the probability similarity of data points in high-dimensional and …


Hyper-Heuristic Approach With K-Means Clustering For Inter-Cell Scheduling, Yanlin Zhao, Yunna Tian Apr 2024

Hyper-Heuristic Approach With K-Means Clustering For Inter-Cell Scheduling, Yanlin Zhao, Yunna Tian

Journal of System Simulation

Abstract: According to the actual production situation of China's manufacturing industry, a hyperheuristic algorithm based on K-means clustering is proposed for inter-cell scheduling problem of flexible job-shop. K-means clustering is applied to group entities with similar attributes into the corresponding work cluster decision blocks, and the ant colony algorithm is used to select heuristic rules for each decision block. The optimal scheduling solutions are generated by using corresponding heuristic rules for scheduling of entities in each decision block. Computational results show that, the computational granularity is properly increased by the form of decision blocks, and the computational efficiency of the …


Multi Base Station Energy Efficient Cluster-Aware Routing For Wireless Sensor Networks With Realtime Data Backup, Martinaa M Apr 2024

Multi Base Station Energy Efficient Cluster-Aware Routing For Wireless Sensor Networks With Realtime Data Backup, Martinaa M

Theses and Dissertations

Wireless Sensor Networks (WSNs) is created, stemming from their applications in distinct areas. This research focuses on implementing an efficient clustering and routing protocols to maximize the lifespan of the WSN by proposing a novel method known as the Energy Efficient Cluster-aware Routing Protocol (EECR). The proposed method comprises of three steps: cluster formation, cluster head (CH) selection, and multi-hop data transmission. The factors needed are residual energy, the minimum distance to the base station (BS), and the minimum Load Count as given in the Energy and Distance CH selection algorithm. The shortest pathway is estimated by the Energy Route …


Exploring Human Aging Proteins Based On Deep Autoencoders And K-Means Clustering, Sondos M. Hammad, Mohamed Talaat Saidahmed, Elsayed A. Sallam, Reda Elbasiony Mar 2024

Exploring Human Aging Proteins Based On Deep Autoencoders And K-Means Clustering, Sondos M. Hammad, Mohamed Talaat Saidahmed, Elsayed A. Sallam, Reda Elbasiony

Journal of Engineering Research

Aging significantly affects human health and the overall economy, yet understanding of the underlying molecular mechanisms remains limited. Among all human genes, almost three hundred and five have been linked to human aging. While certain subsets of these genes or specific aging-related genes have been extensively studied. There has been a lack of comprehensive examination encompassing the entire set of aging-related genes. Here, the main objective is to overcome understanding based on an innovative approach that combines the capabilities of deep learning. Particularly using One-Dimensional Deep AutoEncoder (1D-DAE). Followed by the K-means clustering technique as a means of unsupervised learning. …


Understanding Patient Profiles In Sickle Cell Disease Using Unsupervised Machine Learning, Raj Kamal Somavarapu Jan 2024

Understanding Patient Profiles In Sickle Cell Disease Using Unsupervised Machine Learning, Raj Kamal Somavarapu

Browse all Theses and Dissertations

Sickle Cell Disease (SCD) is one of the most prevalent genetic blood disorders affecting millions of people worldwide. It is often accompanied by acute and/or chronic pain leading to increased healthcare costs and adverse outcomes. Effective management of SCD requires an understanding of the diverse physiological profiles. This study employs unsupervised machine learning, specifically K-means clustering to categorize the patients suffering with SCD into different clusters based on their vital signs. The main aim is to identify the groups that reflect similarities in physiological and pain profiles, allowing an in-depth analysis to reveal distinctive features distinguishing patient clusters. The project …


Opinion Graphs Construction For Reviews Using Transfer Learning And Large Language Models, Yichen Lin Jan 2024

Opinion Graphs Construction For Reviews Using Transfer Learning And Large Language Models, Yichen Lin

Master's Projects

With the rapid development of the Internet, reading online reviews before making a purchase, booking a hotel, or making a restaurant reservation has become a part of daily life. Customers often consider reviews as crucial supplementary information before making decisions on how to spend their money. However, reading many reviews to gain helpful information takes time and effort. This project proposes a new method OpinionGraphGenerator that aims to create opinion graphs from hotel reviews to reduce the high volume of text in reviews while preserving essential insights. In an opinion graph, vertices are semantically similar opinions, where each opinion consists …


Cluster Analysis For Concept Drift Detection In Malware, Aniket Mishra Jan 2024

Cluster Analysis For Concept Drift Detection In Malware, Aniket Mishra

Master's Projects

The rapid evolution of malware presents significant challenges for detection systems. This is due to malware families adapting through feature manipulation and obfuscation, which causes concept drift. A clustering based approach is used to detect and adapt to these shifts. The KronoDroid dataset is segmented into batch sizes of 50 and analyzed with MiniBatch K-Means clustering. The silhouette coefficient is used to evaluate clustering quality, and help identify drift by detecting significant changes in cluster patterns. Concept drift will cause retraining of supervised classifiers, including Linear SVM, RF, MLP, and XGBoost. Three scenarios are used: static models, periodic retraining, and …


Ki-Cook: Clustering Multimodal Cooking Representations Through Knowledge-Infused Learning, Revathy Venkataramanan, Swati Padhee, Saini Rohan Rao, Ronak Kaoshik, Anirudh Sundara Rajan, Amit Sheth Jul 2023

Ki-Cook: Clustering Multimodal Cooking Representations Through Knowledge-Infused Learning, Revathy Venkataramanan, Swati Padhee, Saini Rohan Rao, Ronak Kaoshik, Anirudh Sundara Rajan, Amit Sheth

Publications

Cross-modal recipe retrieval has gained prominence due to its ability to retrieve a text representation given an image representation and vice versa. Clustering these recipe representations based on similarity is essential to retrieve relevant information about unknown food images. Existing studies cluster similar recipe representations in the latent space based on class names. Due to inter-class similarity and intraclass variation, associating a recipe with a class name does not provide sufficient knowledge about recipes to determine similarity. However, recipe title, ingredients, and cooking actions provide detailed knowledge about recipes and are a better determinant of similar recipes. In this study, …


Analyzing Ground Motion Records With Cvi Fuzzy Art, Dustin Tanksley, Xinzhe Yuan, Genda Chen, Donald C. Wunsch Jan 2023

Analyzing Ground Motion Records With Cvi Fuzzy Art, Dustin Tanksley, Xinzhe Yuan, Genda Chen, Donald C. Wunsch

Civil, Architectural and Environmental Engineering Faculty Research & Creative Works

This paper explores using Cluster Validity Indices Fuzzy Adaptative Resonance Theory (CVI Fuzzy ART) to cluster ground motion records (GMRs). Clustering the features extracted from a supervised network trained for predicting the structure damage results in less overfitting from the trained network. Using Cluster Validity Indices (CVIs) to evaluate the clustering gives feedback to how well the data is being classified, allowing further separation of the data. By using CVI Fuzzy ART in combination with features extracted from a trained Convolutional Neural Network (CNN), we were able to form additional clusters in the data. Within the primary clusters, accuracy was …


A Cascade Framework For Privacy-Preserving Point-Of-Interest Recommender System, Longyin Cui, Xiwei Wang Apr 2022

A Cascade Framework For Privacy-Preserving Point-Of-Interest Recommender System, Longyin Cui, Xiwei Wang

Computer Science Faculty Publications

Point-of-interest (POI) recommender systems (RSes) have gained significant popularity in recent years due to the prosperity of location-based social networks (LBSN). However, in the interest of personalization services, various sensitive contextual information is collected, causing potential privacy concerns. This paper proposes a cascaded privacy-preserving POI recommendation (CRS) framework that protects contextual information such as user comments and locations. We demonstrate a minimized trade-off between the privacy-preserving feature and prediction accuracy by applying a semi-decentralized model to real-world datasets.


K-Means Clustering Using Gravity Distance, Ajinkya Vishwas Indulkar Apr 2022

K-Means Clustering Using Gravity Distance, Ajinkya Vishwas Indulkar

Masters Theses & Specialist Projects

Clustering is an important topic in data modeling. K-means Clustering is a well-known partitional clustering algorithm, where a dataset is separated into groups sharing similar properties. Clustering an unbalanced dataset is a challenging problem in data modeling, where some group has a much larger number of data points than others. When a K-means clustering algorithm with Euclidean distance is applied to such data, the algorithm fails to form good clusters. The standard K-means tends to split data into smaller clusters during a clustering process evenly.

We propose a new K-means clustering algorithm to overcome the disadvantage by introducing a different …


Topological Hierarchies And Decomposition: From Clustering To Persistence, Kyle A. Brown Jan 2022

Topological Hierarchies And Decomposition: From Clustering To Persistence, Kyle A. Brown

Browse all Theses and Dissertations

Hierarchical clustering is a class of algorithms commonly used in exploratory data analysis (EDA) and supervised learning. However, they suffer from some drawbacks, including the difficulty of interpreting the resulting dendrogram, arbitrariness in the choice of cut to obtain a flat clustering, and the lack of an obvious way of comparing individual clusters. In this dissertation, we develop the notion of a topological hierarchy on recursively-defined subsets of a metric space. We look to the field of topological data analysis (TDA) for the mathematical background to associate topological structures such as simplicial complexes and maps of covers to clusters in …


Efficient Yet Robust Privacy For Video Streaming, Luke Cranfill, Junggab Son Jul 2021

Efficient Yet Robust Privacy For Video Streaming, Luke Cranfill, Junggab Son

Master of Science in Computer Science Theses

MPEG-DASH is a video streaming standard that outlines protocols for sending audio and video content from a server to a client over HTTP. The standard has been widely utilized by the video streaming industry. However, it creates an opportunity for an adversary to invade users’ privacy. While a user is watching a video, information is leaked in the form of meta-data, the size and time that the server sent data to the user. This information is not protected by encryption and can be used to create a fingerprint for a video. Once the fingerprint is created, the adversary can use …


A Quantitative Validation Of Multi-Modal Image Fusion And Segmentation For Object Detection And Tracking, Nicholas Lahaye, Michael J. Garay, Brian D. Bue, Hesham El-Askary, Erik Linstead Jun 2021

A Quantitative Validation Of Multi-Modal Image Fusion And Segmentation For Object Detection And Tracking, Nicholas Lahaye, Michael J. Garay, Brian D. Bue, Hesham El-Askary, Erik Linstead

Mathematics, Physics, and Computer Science Faculty Articles and Research

In previous works, we have shown the efficacy of using Deep Belief Networks, paired with clustering, to identify distinct classes of objects within remotely sensed data via cluster analysis and qualitative analysis of the output data in comparison with reference data. In this paper, we quantitatively validate the methodology against datasets currently being generated and used within the remote sensing community, as well as show the capabilities and benefits of the data fusion methodologies used. The experiments run take the output of our unsupervised fusion and segmentation methodology and map them to various labeled datasets at different levels of global …


Performance Analysis Of Whale Optimization Based Data Clustering, Ahamed Shafeeq B M, Zahid Ahmed Ansari, Shyam Karanth May 2021

Performance Analysis Of Whale Optimization Based Data Clustering, Ahamed Shafeeq B M, Zahid Ahmed Ansari, Shyam Karanth

Future Computing and Informatics Journal

Data clustering is the method of gathering of data points so that the more similar points will be in the same group. It is a key role in exploratory data mining and a popular technique used in many fields to analyze statistical data. Quality clusters are the key requirement of the cluster analysis result. There will be tradeoffs between the speed of the clustering algorithm and the quality of clusters it produces. Both the quality and speed criteria must be considered for the state-of-the-art clustering algorithm for applications. The Bio-inspired technique has ensured that the process is not trapped in …


Learning Discriminative And Efficient Attention For Person Re-Identification Using Agglomerative Clustering Frameworks, Kshitij Nikhal Apr 2021

Learning Discriminative And Efficient Attention For Person Re-Identification Using Agglomerative Clustering Frameworks, Kshitij Nikhal

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

Recent advancements like multiple contextual analysis, attention mechanisms, distance-aware optimization, and multi-task guidance have been widely used for supervised person re-identification (ReID), but the implementation and effects of such methods in unsupervised person ReID frameworks are non-trivial and unclear, respectively. Moreover, with increasing size and complexity of image- and video-based ReID datasets, manual or semi-automated annotation procedures for supervised ReID are becoming labor intensive and cost prohibitive, which is undesirable especially considering the likelihood of annotation errors increase with scale/complexity of data collections. Therefore, this thesis proposes a new iterative clustering framework that incorporates (a) two attention architectures that learn …


Can Generative Adversarial Networks Help Us Fight Financial Fraud?, Sean Mciver Jan 2021

Can Generative Adversarial Networks Help Us Fight Financial Fraud?, Sean Mciver

Dissertations

Transactional fraud datasets exhibit extreme class imbalance. Learners cannot make accurate generalizations without sufficient data. Researchers can account for imbalance at the data level, algorithmic level or both. This paper focuses on techniques at the data level. We evaluate the evidence of the optimal technique and potential enhancements. Global fraud losses totalled more than 80 % of the UK’s GDP in 2019. The improvement of preprocessing is inherently valuable in fighting these losses. Synthetic minority oversampling technique (SMOTE) and extensions of SMOTE are currently the most common preprocessing strategies. SMOTE oversamples the minority classes by randomly generating a point between …


Texture-Driven Image Clustering In Laser Powder Bed Fusion, Alexander H. Groeger Jan 2021

Texture-Driven Image Clustering In Laser Powder Bed Fusion, Alexander H. Groeger

Browse all Theses and Dissertations

The additive manufacturing (AM) field is striving to identify anomalies in laser powder bed fusion (LPBF) using multi-sensor in-process monitoring paired with machine learning (ML). In-process monitoring can reveal the presence of anomalies but creating a ML classifier requires labeled data. The present work approaches this problem by printing hundreds of Inconel-718 coupons with different processing parameters to capture a wide range of process monitoring imagery with multiple sensor types. Afterwards, the process monitoring images are encoded into feature vectors and clustered to isolate groups in each sensor modality. Four texture representations were learned by training two convolutional neural network …


Texture-Driven Image Clustering In Laser Powder Bed Fusion, Alexander H. Groeger Jan 2021

Texture-Driven Image Clustering In Laser Powder Bed Fusion, Alexander H. Groeger

Browse all Theses and Dissertations

The additive manufacturing (AM) field is striving to identify anomalies in laser powder bed fusion (LPBF) using multi-sensor in-process monitoring paired with machine learning (ML). In-process monitoring can reveal the presence of anomalies but creating a ML classifier requires labeled data. The present work approaches this problem by printing hundreds of Inconel-718 coupons with different processing parameters to capture a wide range of process monitoring imagery with multiple sensor types. Afterwards, the process monitoring images are encoded into feature vectors and clustered to isolate groups in each sensor modality. Four texture representations were learned by training two convolutional neural network …


Clustered Mobile Data Collection In Wsns: An Energy-Delay Trade-Of, İzzet Fati̇h Şentürk Jan 2021

Clustered Mobile Data Collection In Wsns: An Energy-Delay Trade-Of, İzzet Fati̇h Şentürk

Turkish Journal of Electrical Engineering and Computer Sciences

Wireless sensor networks enable monitoring remote areas with limited human intervention. However, the network connectivity between sensor nodes and the base station (BS) may not be always possible due to the limited transmission range of the nodes. In such a case, one or more mobile data collectors (MDCs) can be employed to visit nodes for data collection. If multiple MDCs are available, it is desirable to minimize the energy cost of mobility while distributing the cost among the MDCs in a fair manner. Despite availability of various clustering algorithms, there is no single fits all clustering solution when different requirements …


Proportional Voting Based Semi-Unsupervised Machine Learning Intrusion Detection System, Yang G. Kim, Ohbong Kwon, John Yoon Dec 2020

Proportional Voting Based Semi-Unsupervised Machine Learning Intrusion Detection System, Yang G. Kim, Ohbong Kwon, John Yoon

Publications and Research

Feature selection of NSL-KDD data set is usually done by finding co-relationships among features, irrespective of target prediction. We aim to determine the relationship between features and target goals to facilitate different target detection goals regardless of the correlated feature selection. The unbalanced data structure in NSL-KDD data can be relaxed by Proportional Representation (PR). However, adopting PR would deny the notion of winner-take-all by attracting a majority of the vote and also provide a fairly proportional share for any grouping of like-minded data. Furthermore, minorities and majorities would get a fair share of power and representation in data structure …


Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski Dec 2020

Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski

Knowledge Engineering and Data Science

Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …


Slashing Quality Index Modeling And Simulation Based On Data Dispersion Clustering, Yuxian Zhang, Xiaoyi Qian, Dong Xiao, Jianhui Wang Aug 2020

Slashing Quality Index Modeling And Simulation Based On Data Dispersion Clustering, Yuxian Zhang, Xiaoyi Qian, Dong Xiao, Jianhui Wang

Journal of System Simulation

Abstract: For the sensitivity of noise and outliers data in the typical partitioning clustering algorithm, a clustering algorithm based on data dispersion was proposed. The data dispersion was defined and introduced to a non-Euclidean distance. The similarity metric was established, and the data clustering was realized. The optimal clustering number was obtained by the validity function based on improved partition coefficient. Then the proposed clustering algorithm was applied to quality index model in slashing process. A size add-on quality index model was built by radial basis function neural networks. The node number of hidden layer was determined and the center …


Key Technologies Of Precaution And Prediction Of Abnormal Spatial-Temporal Trajectory: A Review Of Recent Advances, Gongda Qiu, He Ming, Yang Jie, Yuting Cao, Jihong Sun Jun 2020

Key Technologies Of Precaution And Prediction Of Abnormal Spatial-Temporal Trajectory: A Review Of Recent Advances, Gongda Qiu, He Ming, Yang Jie, Yuting Cao, Jihong Sun

Journal of System Simulation

Abstract: The ex-post disposition of a major incident, which is expected to transform into prediction and precaution of abnormal behavior, is increasingly unable to meet the urgent needs of the society.Therapid development and popularization of sensor network and positioning technology lay the foundation for mining spatial-temporal trajectory data. With the key objective of prediction and precaution of abnormal trajectory based on big data mining, the future research directions and prospects on trajectory clustering and recognitionareanalyzed, discussed and elaboratedinthis paper.Temporal trajectory prediction applied in prediction and precaution of abnormal spatial-temporal trajectory is also presented, providing a reference for further research on …