Open Access. Powered by Scholars. Published by Universities.®

Computer Engineering Commons

Open Access. Powered by Scholars. Published by Universities.®

Deep Learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 170

Full-Text Articles in Computer Engineering

Large Language Models For Bacterial Genomic Analysis, Manvendra Chavan Jan 2025

Large Language Models For Bacterial Genomic Analysis, Manvendra Chavan

Master's Projects

Identification of bacterial gene sequences with agricultural applications has the potential to transform agricultural biotechnology. These genes can be used in environmentally friendly pest control strategies. One such use case is identifying genes with potential insecticidal properties. With an increasing number of genomic information and decreasing numbers of available annotated sequences, finding new insecticidal genes has become more challenging.The traditional methods relying on sequence alignment and annotated databases are not effective in detecting functionally relevant genes lacking close homology to known cases. This project investigates the data-driven classification of genes by sequence modeling. This research is focused on learning DNA …


Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo Jan 2025

Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo

Dissertations, Master's Theses and Master's Reports

Medical Image Segmentation is a critical task in the field of medical imaging, playing a crucial role in diagnostics, treatment planning, and disease monitoring. The emergence of Deep Learning (DL) has ushered in a new era in Artificial Intelligence (AI), propelling remarkable advancements in key domains like language translation, object recognition, and recommendation systems. This evolution has been accompanied by continuous enhancements in computational efficiency and improvements in predictive accuracy. The introduction of sophisticated algorithms, such as convolutional neural networks (CNNs) and transformers, exemplifies these advancements. DL algorithms have demonstrated exceptional efficacy in medical image segmentation tasks, showcasing the potential …


Enhancing Image Classification Using A Convolutional Neural Network Model, Zena M. Saadi, Ahmed T. Sadiq, Omar Z. Akif, Marwa M. Eid Dec 2024

Enhancing Image Classification Using A Convolutional Neural Network Model, Zena M. Saadi, Ahmed T. Sadiq, Omar Z. Akif, Marwa M. Eid

Journal of Soft Computing and Computer Applications

In recent years, with the rapid development of the current classification system in digital content identification, automatic classification of images has become the most challenging task in the field of computer vision. As can be seen, vision is quite challenging for a system to automatically understand and analyze images, as compared to the vision of humans. Some research papers have been done to address the issue in the low-level current classification system, but the output was restricted only to basic image features. However, similarly, the approaches fail to accurately classify images. For the results expected in this field, such as …


Heterogeneous Collaborative Robotics: Multi-Robot Navigation In Dynamic Environments, Tyler Nicholas Raettig Dec 2024

Heterogeneous Collaborative Robotics: Multi-Robot Navigation In Dynamic Environments, Tyler Nicholas Raettig

Theses and Dissertations

Abstract: The challenges of multi-robot navigation in dynamic environments, focusing on uncertainties in obstacle complexities, partial observation, and the transition of policies from simulations to the real world. The proposed approach utilizes a deep reinforcement learning (DRL) framework enabling a Light Detection and Ranging (LiDAR)-equipped robot to communicate with a camera-equipped robot to achieve optimal paths despite their different sensors. The key contributions include the development of a cooperative architecture for information exchange between robots, a DRL-based framework for learning navigation policies, and a training mechanism based on dynamic randomization for enhanced real-world adaptability. Experimental validation using Gazebo simulations demonstrates …


End-To-End Learning For A Low-Cost Robotics Arm, Abhishek Chothani Dec 2024

End-To-End Learning For A Low-Cost Robotics Arm, Abhishek Chothani

Theses and Dissertations

Robotic manipulation is a cornerstone of automation, with the ultimate goal of developing versatile systems capable of executing a wide range of real-world tasks autonomously. Traditional robotics approaches, while reliable and widely adopted in industrial settings, often struggle with adaptability, perception, and dynamic task execution. This thesis explores the evolution from classical robotics techniques to modern learning-based approaches, leveraging advancements in artificial intelligence to overcome these limitations.

Initially, the thesis presents a pick-and-place pipeline built using the Drake robotics framework and the KUKA iiwa robotic arm. This system employs a pseudoinverse controller for inverse kinematics to perform structured tasks like …


Detecting Data Poisoning Attacks In Federated Learning For Healthcare Applications Using Deep Learning, Mohammed Aljanabi, Sahar Yousif Mohammed, Alaa Hamza Omran Oct 2024

Detecting Data Poisoning Attacks In Federated Learning For Healthcare Applications Using Deep Learning, Mohammed Aljanabi, Sahar Yousif Mohammed, Alaa Hamza Omran

Iraqi Journal for Computer Science and Mathematics

This work introduces a new approach to protecting the data in the healthcare applications of federated learning based on the classification of skin cancer. The recommended solution established and prevents the data poisoning attacks by using deep learning and CNN architectures namely VGG16. In a federated learning system which comprises of ten healthcare facilities, the approach enables the training of models in a collaborative way without compromising the medical data or the patients’ information. Data is meticulously prepared and preprocessed using the Skin Cancer MNIST: According to the HAM10000 dataset. As for the federated learning approach, VGG16’s feature extraction capability …


Interpreting Black-Box Time Series Classifiers Using Parameterised Event Primitives, Ephrem Tibebe Mekonnen, Luca Longo, Pierpaolo Dondio Oct 2024

Interpreting Black-Box Time Series Classifiers Using Parameterised Event Primitives, Ephrem Tibebe Mekonnen, Luca Longo, Pierpaolo Dondio

Conference papers

Amidst the remarkable performance of deep learning models in time series classification, there is a pressing demand for methods that unveil their prediction rationale. Existing feature importance techniques often neglect the temporal nature of time series data, focusing solely on segment importance. Addressing this gap, this paper introduces a local model-agnostic method akin to LIME, which generates neighbouring samples by randomly perturbing segments of the original instance. Subsequently, weights are computed for each neighbouring instance based on its distance from the original, elucidating its influence. Parameterised event primitives (PEPs) are then extracted from these perturbed samples, encompassing increasing and decreasing …


Deep-Learning Based Microstructure Reconstruction And Generation, Cameron J. Maloney, Lucas Taliaferro Oct 2024

Deep-Learning Based Microstructure Reconstruction And Generation, Cameron J. Maloney, Lucas Taliaferro

College of Engineering Summer Undergraduate Research Program

Characterizing the microstructural behavior of materials is crucial for understanding their properties and performance. Traditional imaging methods, such as optical microscopy and electron microscopy, are effective but costly and time-consuming. Computational approaches can reduce costs and time while expanding the accessibility of microstructural analysis through the generation of new microstructure images. Traditional computational approaches, namely descriptor-based approaches, are slow but effective in low-data scenarios. Modern approaches use machine learning (ML), which is faster but often requires a lot of data to approach the performance of descriptor-based methods. This research leverages a special data-efficient Generative Adversarial Network (GAN) architecture to artificially …


An Adaptive Hybrid Deep Learning Architecture For Providing Guaranteed Qos In 5g Cellular Networks, Rajilal Mv Ms Aug 2024

An Adaptive Hybrid Deep Learning Architecture For Providing Guaranteed Qos In 5g Cellular Networks, Rajilal Mv Ms

Theses and Dissertations

Wireless network systems must have effective resource allocation, particularly in the context of 5G networks when flexibility is needed to meet a range of network requirements. Resource allocation is essential in cellular network contexts to guarantee equitable access to customers, partners, and cellular service users. Since resource distribution determines network performance, it offers significant advantages when executed well. One of the biggest issues with 5G technology is resource allocation, particularly when it comes to the Quality of Service (QoS) for various applications. Resources in wireless networks include items like channels, power, and spectrum; these must all be apportioned according to …


Real-Time Gun Detection In Video Streams Using Yolo V8, Harish Kumar Reddy Kunchala Aug 2024

Real-Time Gun Detection In Video Streams Using Yolo V8, Harish Kumar Reddy Kunchala

Electronic Theses, Projects, and Dissertations

In this research, we advance the domain of public safety by developing a machine learning model that utilizes the YOLO v8 architecture for real-time detection of firearms in video streams. A diverse and extensive dataset, capturing a range of firearms in varying lighting and backgrounds, was meticulously assembled and preprocessed to enhance the model's adaptability to real-world scenarios. Leveraging the YOLO v8 framework, known for its real-time object detection accuracy, the model was fine-tuned to accurately identify firearms across different shapes and orientations.

The training phase capitalized on GPU computing and transfer learning to expedite the learning process while preserving …


Ensemble Learning For Accurate Prediction Of Heart Sounds Using Gammatonegram Images, Sinam Ashinikumar Singh, Sinam Ajitkumar Singh, Aheibam Dinamani Singh Jul 2024

Ensemble Learning For Accurate Prediction Of Heart Sounds Using Gammatonegram Images, Sinam Ashinikumar Singh, Sinam Ajitkumar Singh, Aheibam Dinamani Singh

Turkish Journal of Electrical Engineering and Computer Sciences

The analysis of heart sound signals constitutes a pivotal domain in healthcare, with the prediction of imbalanced heart sounds offering critical diagnostic insights. However, the inherent diversity in cardiac sound patterns presents a substantial challenge in predicting imbalanced signals. Many scientific disciplines have focused a great deal of emphasis on the problem of class inequality. We introduce an ensemble learning approach employing a convolutional neural network model-based deep learning algorithm to effectively tackle the challenges associated with predicting imbalanced heart sound signals. We use a Gammatone filter bank to extract relevant features from the heard sound signal. Our approach leverages …


Autonomous Microgrid System, Xavier Kuehn, Brian Xiong Jun 2024

Autonomous Microgrid System, Xavier Kuehn, Brian Xiong

Computer Science and Engineering Senior Theses

Microgrids have made a revolutionary change in the realm of energy distribution due to the features that they offer, including localized, resilient, and sustainable energy solutions. Operating renewable resources in a microgrid while maintaining generation-load balance and acceptable voltage-frequency limits has been an open research problem. This thesis presents smart python agents for microgrid systems to automate the operations and control of microgrid renewable resources in an effort to provide resilient solutions to the intermittence issues that could potentially arise within the microgrid energy system. The smart agents operate the microgrids by not only integrating the use of renewable energy …


Automated Brain Tumor Classifier With Deep Learning, Venkata Sai Krishna Chaitanya Kandula May 2024

Automated Brain Tumor Classifier With Deep Learning, Venkata Sai Krishna Chaitanya Kandula

Electronic Theses, Projects, and Dissertations

Brain Tumors are abnormal growth of cells within the brain that can be categorized as benign (non-cancerous) or malignant (cancerous). Accurate and timely classification of brain tumors is crucial for effective treatment planning and patient care. Medical imaging techniques like Magnetic Resonance Imaging (MRI) provide detailed visualizations of brain structures, aiding in diagnosis and tumor classification[8].

In this project, we propose a brain tumor classifier applying deep learning methodologies to automatically classify brain tumor images without any manual intervention. The classifier uses deep learning architectures to extract and classify brain MRI images. Specifically, a Convolutional Neural Network (CNN) …


Deep Learning Using Vision And Lidar For Global Robot Localization, Brett E. Gowling May 2024

Deep Learning Using Vision And Lidar For Global Robot Localization, Brett E. Gowling

Master's Theses

As the field of mobile robotics rapidly expands, precise understanding of a robot’s position and orientation becomes critical for autonomous navigation and efficient task performance. In this thesis, we present a snapshot-based global localization machine learning model for a mobile robot, the e-puck, in a simulated environment. Our model uses multimodal data to predict both position and orientation using the robot’s on-board cameras and LiDAR sensor. In an effort to minimize localization error, we explore different sensor configurations by varying the number of cameras and LiDAR layers used. Additionally, we investigate the performance benefits of different multimodal fusion strategies while …


Development Of Deep Neural Architecture For Continuous Sign Language Video Generation, Natarajan B Apr 2024

Development Of Deep Neural Architecture For Continuous Sign Language Video Generation, Natarajan B

Theses and Dissertations

This dissertation presents a deep neural network based sign language video generation framework for translating the multilingual sentences into sign videos. This thesis addresses the challenges persist with the sign language video generation such as (i) Handling longer sequences of input sentences and new words (ii) Pose estimation with higher accuracy (iii) High quality photo realistic sign gesture video generation (iv) Improving realism in sign video generation. Hence, the thesis focuses four contributions to address the above issues.

The first contribution of this thesis automates the translation of multilingual sentences into sign glosses without manual intervention by incorporating Hybrid Neural …


Enhancing Mobile App User Experience: A Deep Learning Approach For System Design And Optimization, Deepesh Haryani Apr 2024

Enhancing Mobile App User Experience: A Deep Learning Approach For System Design And Optimization, Deepesh Haryani

Harrisburg University Dissertations and Theses

This paper presents a comprehensive framework for enhancing user experience in mobile applications through the integration of deep learning systems. The proposed system design encompasses various components, including data collection and preprocessing, model development and training, integration with mobile applications, dataset management service, model training service, model serving, hyperparameter optimization, metadata and artifact store, and workflow orchestration. Each component is meticulously designed with a focus on scalability, efficiency, isolation, and critical analysis. Innovative design principles are employed to ensure seamless integration, usability, and automation. Additionally, the paper discusses distributed training service design, advanced optimization techniques, and decision criteria for hyperparameter …


A Spatial Data Framework For Indoor Positioning Using Machine Learning Techniques, Venkateswari P Apr 2024

A Spatial Data Framework For Indoor Positioning Using Machine Learning Techniques, Venkateswari P

Theses and Dissertations

The last few years have seen an increase in interest in indoor positioning and localization as potential research and development areas. WiFi is a strong substitute that supports positioning based on indoor floor plans. In this thesis, the Principal Featured - Kohonen Deep Structure (PF-KDS) model is developed to position WiFi devices more accurately and efficiently for indoor floor planning. Initially, spatial data analysis is conducted using the Principal Feature Enhanced Auto-Encoder algorithm, extracting principal features for dimensionality reduction.

Following this, the Kohonen Self- Organizing Deep Structured Learning technique is devised for precise position estimation by considering a new path …


Exploring The Use Of Enhanced Swad Towards Building Learned Models That Generalize Better To Unseen Sources, Brandon M. Weinhofer Mar 2024

Exploring The Use Of Enhanced Swad Towards Building Learned Models That Generalize Better To Unseen Sources, Brandon M. Weinhofer

USF Tampa Graduate Theses and Dissertations

Deep learning models, typically, take significant time to train. Classifier ensembles are areliable way to increase classifier accuracy and perhaps generalizability to unseen sources of data. These classifiers can be combined with a simple voting scheme. The problem is that having multiple models can very heavily increase training time. Snapshot ensembles have been shown to provide a boost in performance by creating an ensemble of classifiers with different weights during the training of a single deep learned model. This can somewhat solve the problem of the increased training time as you do not have to train separate models. As Machine …


Insights Into Cellular Evolution: Temporal Deep Learning Models And Analysis For Cell Image Classification, Xinran Zhao Mar 2024

Insights Into Cellular Evolution: Temporal Deep Learning Models And Analysis For Cell Image Classification, Xinran Zhao

Master's Theses

Understanding the temporal evolution of cells poses a significant challenge in developmental biology. This study embarks on a comparative analysis of various machine-learning techniques to classify cell colony images across different timestamps, thereby aiming to capture dynamic transitions of cellular states. By performing Transfer Learning with state-of-the-art classification networks, we achieve high accuracy in categorizing single-timestamp images. Furthermore, this research introduces the integration of temporal models, notably LSTM (Long Short Term Memory Network), R-Transformer (Recurrent Neural Network enhanced Transformer) and ViViT (Video Vision Transformer), to undertake this classification task to verify the effectiveness of incorporating temporal features into the classification …


An Integrative Computational Intelligence For Robust Anomaly Detection In Social Networks, Helina Rajini Suresh, K R. Harsavarthini, R Mageswaran, Hirald Dwaraka Praveena, C Gnanaprakasam, C.Sakthi Lakshmi Priya Jan 2024

An Integrative Computational Intelligence For Robust Anomaly Detection In Social Networks, Helina Rajini Suresh, K R. Harsavarthini, R Mageswaran, Hirald Dwaraka Praveena, C Gnanaprakasam, C.Sakthi Lakshmi Priya

Iraqi Journal for Computer Science and Mathematics

Anomaly detection is one of the most important tasks for maintaining the integrity, security, and trustworthiness of online communities in a social network. This paper proposes AdaptoDetect, which represents a new framework; it discusses a new anomaly detection approach called Pufferfish Optimization Technique for feature selection, together with a Graph Embedding Autoencoder for identifying anomalies. What makes AdaptoDetect special is that, with the use of POT, it has a distinctive capability in dynamic adaptation against network changes by selecting only the most relevant features in social network data. The technique for optimization underlines the important attributes for anomaly detection so …


Applications Of Predictive And Generative Ai Algorithms: Regression Modeling, Customized Large Language Models, And Text-To-Image Generative Diffusion Models, Suhaima Jamal Jan 2024

Applications Of Predictive And Generative Ai Algorithms: Regression Modeling, Customized Large Language Models, And Text-To-Image Generative Diffusion Models, Suhaima Jamal

College of Graduate Studies: Theses & Dissertations

The integration of Machine Learning (ML) and Artificial Intelligence (AI) algorithms has radically changed predictive modeling and classification tasks, enhancing a multitude of domains with unprecedented analytical capabilities. Predictive modeling leverages ML and AI to forecast future trends or behaviors based on historical data, while classification tasks categorize data into distinct classes, from email filtering to medical diagnosis. Concurrently, text-to-image generation has emerged as a transformative potential, allowing visual content creation directly from textual descriptions. These advancements are pivotal in design, art, entertainment, and visual communication, as well as enhancing creativity and productivity. This work explores three significant studies in …


Birdsong Classification Using Deep Learning And Mixit, Sasanka Kosuru Jan 2024

Birdsong Classification Using Deep Learning And Mixit, Sasanka Kosuru

Master's Projects

The identification of bird species using deep learning techniques presents a novel approach in bioacoustics, by significantly advancing our understanding and enhancing our capabilities in bird species recognition from audio recordings. The value of audio over visual data for monitoring ecological patterns in birds can be highlighted with the deployment of automated recording devices in remote wildlife sensing, offering a more cost-effective, non-invasive, and practical solution. However, the methods of processing and classifying the audio remain challenging due to the complexity of bird audio, characterized by diverse vocalizations and imminent environmental noise, which poses difficult challenges to perform effective classification. …


Emotion Detection Using Ensemble Learning, Priya Harika Yerapothu Jan 2024

Emotion Detection Using Ensemble Learning, Priya Harika Yerapothu

Master's Projects

Emotion detection is gaining exponential necessity in today’s technological age. This research seeks to delve into ways conversational AI could be enhanced by integrating emotional intelligence using an ensemble learning approach. Traditional machine learning along with advanced neural network architectures are implemented to improve the understanding and intricacies of emotion detection from textual data. The dataset we use is GoEmotions dataset, annotated with 27 emotional labels, to conduct a detailed analysis of emotion recognition. Various machine learning models, such as HistGradientBoosting, LightGBM, CatBoost, and MLP, will be evaluated side by side with advanced models of Bidirectional Long Short-Term Memory (BiLSTM) …


Community Detection Using Deep Learning: Variational Graph Autoencoder Enhanced With Leiden And K-Truss Techniques, Jyotika Hariom Patil Jan 2024

Community Detection Using Deep Learning: Variational Graph Autoencoder Enhanced With Leiden And K-Truss Techniques, Jyotika Hariom Patil

Master's Projects

Community detection in networks is essential for understanding the complex structures of connected systems. Traditional deep learning (DL) methods such as Graph Neural Networks (GNNs) and Graph Convolutional Networks (GCNs) have shown promised results in supervised tasks, like classification, but often fail in unsupervised tasks like community detection because of the lack of labels. Self- supervised approaches where we integrate crucial community information offer a solution. This project seeks to explore DL methods for community detection, focusing specifically on using Graph Variational Autoencoders (VGAEs). While classical approaches can efficiently handle small to medium-sized networks, they typically struggle with larger-sized structures. …


Deciphering Speech Through Vision: A Deep Learning Lip Reading System, Srujith Rao Ambati Jan 2024

Deciphering Speech Through Vision: A Deep Learning Lip Reading System, Srujith Rao Ambati

Master's Projects

Lip-reading, a ubiquitous field between computer vision and speech processing, focuses on identifying what spoken words a person generates depending on their uttering lip movements. This paper presents a streamlined lip-reading solution that employs machine learning and deep learning. First, Our work utilizes the Multi-Task Cascaded Convo- lutional Networks to detect facial “landmarks,” including the face and lips region, and the aligns the face. The aligned faces are segmented to get the lip images. Lip images are preprocessed using the Real-Enhanced Super Resolution Generative Adversarial Network to enhance image resolution to identify subtle lip movement in video images: a critical …


Security Information And Event Management Optimization Using Deep Federated Learning In Cloud-Based Autonomous Cyber-Physical Systems, Mohamed Mounir Moussa Jan 2024

Security Information And Event Management Optimization Using Deep Federated Learning In Cloud-Based Autonomous Cyber-Physical Systems, Mohamed Mounir Moussa

Wayne State University Dissertations

The integration of cloud-based technologies into Connected and Autonomous Vehicles (CAVs) is reshaping the field by combining Deep Federated Learning (DFL), Security Information and Event Management (SIEM), and cloud-dew computing. This solution leverages cloud-based resource provisioning, which is crucial for allocating scalable and efficient computational resources in a dynamic manner. These resources are essential for managing the intricate data and computing requirements of distributed systems, especially in the intelligent vehicle sector. This provisioning facilitates the efficient control of route mapping and cybersecurity in Connected Autonomous Vehicles (CAVs), guaranteeing the ability to process and make decisions in real-time.The research evaluates the …


Multi-Semantic-Stage Neural Networks For Robust And Interpretable Deep Learning, Christopher J. Menart Jan 2024

Multi-Semantic-Stage Neural Networks For Robust And Interpretable Deep Learning, Christopher J. Menart

Browse all Theses and Dissertations

Deep neural networks have great representational power. However, most deep neural nets today optimize directly for performance on a single task defined only by labeled training data. This excludes potential sources of knowledge and ways of learning which could improve their performance, and address challenges, such as explainability, which are pressing to the field. We propose a framework for neural network architecture which generalizes it to a graph of many semantically-meaningful variables. We call it the Multi-Semantic-Stage Neural Network (MSSNN). An MSSNN models its domain as a web of conditional probabilities, i.e. a collection of inter-related tasks which can learn …


Deep Learning Techniques For Image Segmentation In Dermoscopic Skin Cancer Images, Norsang Lama Jan 2024

Deep Learning Techniques For Image Segmentation In Dermoscopic Skin Cancer Images, Norsang Lama

Doctoral Dissertations

"Melanoma is recognized as the most lethal type of skin cancer, responsible for a significant proportion of skin cancer-related deaths. However, early detection of melanoma is essential for successful treatment outcomes. Computer-aided skin cancer diagnosis tools can save lives by enabling earlier detection of skin cancer. Image segmentation is a crucial step in computer-aided diagnosis as it allows the detection of critical features or regions in an image. Thus, an accurate image segmentation method is necessary to create a more precise computer-aided diagnostic tool for skin cancer diagnosis. This dissertation includes investigating and developing deep learning techniques to improve image …


Robotic Gas Source Localization And Distribution Mapping Via Deep Reinforcement Learning, Iliya Kulbaka Jan 2024

Robotic Gas Source Localization And Distribution Mapping Via Deep Reinforcement Learning, Iliya Kulbaka

UNF Graduate Theses and Dissertations

This research aims to advance the fields of Gas Source Localization (GSL) and Gas Distribution Mapping (GDM) by developing deep reinforcement learning (DRL) methodologies suitable for complex, real-world environments. GSL and GDM are crucial for applications such as environmental monitoring, hazardous material detection, and search-and-rescue missions, where safe and efficient exploration is essential. Traditional methods often fall short in dynamic settings influenced by factors like wind and obstacles. To address these limitations, this study proposes novel neural network architectures and learning frameworks for adaptive exploration and mapping, integrating Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) layers, and Deep Q-Networks …


Hindi Image Captioning Using Indictrans2 And Encoder-Decoder Architecture, Anahita Vayalombrone Dinesh Jan 2024

Hindi Image Captioning Using Indictrans2 And Encoder-Decoder Architecture, Anahita Vayalombrone Dinesh

Master's Projects

One of the most prominent tasks that lie on the conjunction of Natural Language Processing (NLP) and computer vision, is image captioning. Image captioning is the generative task of achieving textual descriptions from images. Its application finds use in many real-world scenarios like aiding the visually impaired, editing applications, recommendation systems, and medical imaging. This research focus lies in Hindi image captioning, the official language of India, as it has not been explored as far as its need. Several challenges such as the lack of substantial Hindi text data for training models, the need for human annotators to verify the …