Open Access. Powered by Scholars. Published by Universities.®

Computer vision

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 157

Full-Text Articles in Artificial Intelligence and Robotics

Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane Jun 2026

Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane

Beyond: Undergraduate Research Journal

Autonomous tracking of agile unmanned aerial vehicles (UAVs) presents significant challenges for real-time perception and control systems. This work presents AIRHOUND (Autonomous Intelligent Rotorcraft for Hostile Object Unified Navigation and Detection), a UAV platform implementing vision-based yaw tracking through a modular ROS2 software architecture. The system employs YOLOv8 object detection optimized with NVIDIA TensorRT for embedded deployment on an NVIDIA Jetson Orin companion computer. Detected targets are processed through a geometric tracking module that converts pixel coordinates to angular yaw errors using pinhole camera intrinsics, with a proportional controller generating rate-limited yaw commands. These commands are streamed to a PX4 …


Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud Jun 2026

Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud

Dissertations

Deep learning has become the dominant approach for solving vision-centric problems; however, its successful deployment in real-world applications remains limited by high computational cost, data dependency, and insufficient integration with practical and human-centered environments. While state-of-the art deep learning models often achieve impressive performance in controlled settings, they frequently fail to generalize or operate efficiently under deployment constraints such as limited resources, complex data modalities, and real-time interaction requirements. These limitations motivate the need for a deployment-oriented deep learning framework that balances accuracy, efficiency, and practical usability.

This dissertation investigates the design and deployment of efficient deep learning architectures for …


Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan May 2026

Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan

Dissertations

Digital three-dimensional (3D) models are central to engineering design, analysis, and manufacturing, but learning pipelines for man-made geometry often operate on sampled carriers that do not preserve all of the structure present in exact CAD representations. This dissertation studies learning-based reconstruction and part decomposition for structured man-made 3D geometry, from general object benchmarks to CAD-derived datasets, with a focus on neural implicit representations trained from signed-distance samples, point clouds, and tessellated meshes. The goal is to make these models more accurate, more part-aware, and more consistently supervised.

First, signed distance function (SDF) reconstruction with implicit neural representations is improved through …


Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella May 2026

Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella

Capstone Projects

Cataloging and digitizing the objects inside a building manually is a task that is often impractical at scale. This project therefore automates the process, using a custom-made system. Using a photogrammetry-based 3D reconstruction of a room, this system is applied to sequences of 2D images used to make the 3D models. The system applies object detection, image segmentation, and image-text models to identify and describe objects, using CNN based models such as YOLO and OpenCLIP. Each analyzed object is then stored in a structured database with spatial coordinates from the 3D scanning, descriptive attributes from the image-text models, and other …


Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun Mar 2026

Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun

Electronic Theses and Dissertations

The performance of deep neural networks (DNNs) is strongly influenced by the characteristics and quality of the underlying datasets. This Ph.D. dissertation addresses three pervasive data challenges-imbalance, quality degradation, and scarcity-that commonly hinder the effectiveness of DNNs in computer vision (CV) and natural language processing (NLP) applications.

Class imbalance remains one of the most frequent causes of degraded model generalization. While Focal Loss effectively mitigates inter-class imbalance by assigning higher weights to minority classes, it struggles with intra-class imbalance, particularly in video datasets where longer clips dominate feature representation. To address this, I implement and utilize …


Video Generation Techniques For Novel View Synthesis With Flow-Matching Transformers, Xiuyuan Qiu Mar 2026

Video Generation Techniques For Novel View Synthesis With Flow-Matching Transformers, Xiuyuan Qiu

Master's Theses

Novel view synthesis (NVS) aims to generate images of a scene from unseen camera viewpoints. Recent work, such as Stable Virtual Camera, shows that large-scale image diffusion models like Stable Diffusion can be adapted for pose-conditioned view synthesis by incorporating video-generation techniques with camera conditioning. In this thesis, we introduce MVFlow, a new NVS model that extends this approach to a different image generation architecture: a flow-matching diffusion transformer, specifically FLUX.1, which has demonstrated strong performance in image synthesis. We evaluate MVFlow under varying input view counts and pose distance settings. Our results show that this architectural transfer is feasible; …


Pl-Mamba: A 3d Point Cloud Semantic Segmentation Network Based On Bimodal Fusion, He Zhu, Feng Zhou, Mengxiao Zhu, Ju Dai Jan 2026

Pl-Mamba: A 3d Point Cloud Semantic Segmentation Network Based On Bimodal Fusion, He Zhu, Feng Zhou, Mengxiao Zhu, Ju Dai

Journal of System Simulation

Abstract: To enhance the semantic discrimination capability in point cloud semantic segmentation, a 3D point cloud semantic segmentation network named PL-Mamba is proposed, which is centered on the fusion of point cloud (P) and language (L) dual modalities. This method takes PointMamba as the backbone network, leveraging its excellent long-sequence modeling and global perception capabilities. It introduces a language prompt mechanism and uses a pretrained language model BERT to encode the context of category labels, obtaining semantically rich text features. The text information serves as a language guided token and is deeply integrated with point cloud features through cross modal …


Towards Multimodal Guideline-Aligned Agentic Systems, Wenliang Zhong Jan 2026

Towards Multimodal Guideline-Aligned Agentic Systems, Wenliang Zhong

Computer Science and Engineering Dissertations - Archive

I present my work on building multimodal guideline-aligned agentic systems designed to enable AI agents to solve complex real-world tasks. My research addresses two critical perspectives: (1) Instruction-Aware Embedding Models for flexible and universal embedding tasks, and (2) Guideline-Driven LLM Agents that leverage domain-specific guidelines to perform expert-level tasks. These components address embedding and generation tasks, respectively, and lay the foundation for a hybrid agent capable of tackling challenging real-world applications.

From the embedding perspective, I first address the instruction-following capabilities of embedding models. While Large Language Models (LLMs) excel at instruction following, they are primarily designed for generation rather …


Reliable And Label-Efficient Learning For Open-World Visual Perception And Robot Learning Under Uncertainty, Zongyao Lyu Jan 2026

Reliable And Label-Efficient Learning For Open-World Visual Perception And Robot Learning Under Uncertainty, Zongyao Lyu

Computer Science and Engineering Dissertations

Modern learning systems deployed in open-world environments must make reliable decisions despite predictive uncertainty, previously unseen classes, limited annotations, and distribution shifts. This dissertation develops methods for reliable and label-efficient learning in visual perception and robot control.

First, this work studies uncertainty in object detection by representing semantic and spatial predictions probabilistically. A deep-ensemble framework aggregates detections into class-probability distributions and probabilistic bounding boxes, while a subsequent extension combines deep ensembles with Monte Carlo dropout to further investigate predictive uncertainty. Second, this dissertation addresses open-set recognition, where classes absent during training may appear at inference time. An empirical study shows …


Unsafe2safe: Controllable Image Anonymization For Downstream Utility, Minh Dinh Jan 2026

Unsafe2safe: Controllable Image Anonymization For Downstream Utility, Minh Dinh

Dartmouth College Master’s Theses

Large-scale image datasets frequently contain identifiable or sensitive content, raising privacy risks when training models that may memorize and leak such information. We present Unsafe2Safe, a fully automated pipeline that detects privacy-prone images and rewrites only their sensitive regions using multimodally guided diffusion editing. Unsafe2Safe operates in two stages. Stage 1 uses a vision--language model to (i) inspect images for privacy risks, (ii) generate paired private and public captions that respectively include and omit sensitive attributes, and (iii) prompt a large language model to produce structured, identity-neutral edit instructions conditioned on the public caption. Stage 2 employs instruction-driven diffusion editors …


Pixel-Perfect Segmentation Of Solar Filaments, Jamie Harris Dec 2025

Pixel-Perfect Segmentation Of Solar Filaments, Jamie Harris

Undergraduate Research Symposium

The observation and classification of solar filaments has a drastic impact on the ability to predict solar-magnetic weather phenomena that threatens to put both satellite infrastructure and astronauts at risk. Using the Hɑ filter provided by the Global Oscillations Network Group (GONG), a network of six telescopes around the world dedicated to 24/7 surveillance of the sun, we are able to get images that clearly and prominently display filament activity. With the vast amount of images the GONG takes, it is not possible to manually analyze every image. Using the U-Net model for computer vision, we were able to train …


Intelligence Architectures And Machine Learning Applications In Contemporary Spine Care, Rahul Kumar, Conor Dougherty, Kyle Sporn, Akshay Khanna, Puja Ravi, Pranay Prabhakar, Nasif Zaman Sep 2025

Intelligence Architectures And Machine Learning Applications In Contemporary Spine Care, Rahul Kumar, Conor Dougherty, Kyle Sporn, Akshay Khanna, Puja Ravi, Pranay Prabhakar, Nasif Zaman

SKMC Student Presentations and Publications

The rapid evolution of artificial intelligence (AI) and machine learning (ML) technologies has initiated a paradigm shift in contemporary spine care. This narrative review synthesizes advances across imaging-based diagnostics, surgical planning, genomic risk stratification, and post-operative outcome prediction. We critically assess high-performing AI tools, such as convolutional neural networks for vertebral fracture detection, robotic guidance platforms like Mazor X and ExcelsiusGPS, and deep learning-based morphometric analysis systems. In parallel, we examine the emergence of ambient clinical intelligence and precision pharmacogenomics as enablers of personalized spine care. Notably, genome-wide association studies (GWAS) and polygenic risk scores are enabling a shift from …


Intelligent Motion Tracking: A Low-Cost Surveillance Framework Using Soc Devices And Sensor Fusion, Aung Myat Khaung, Nicholas Michael Stiffler Aug 2025

Intelligent Motion Tracking: A Low-Cost Surveillance Framework Using Soc Devices And Sensor Fusion, Aung Myat Khaung, Nicholas Michael Stiffler

Research from the Berry Summer Thesis Institute, 2025

This thesis presents the design and implementation of a lightweight surveillance system capable of realtime motion detection, object tracking, and behavioral history reconstruction in controlled environments. The system uses System-on-Chip devices such as Raspberry Pi boards equipped with NOIR cameras, monocular cameras, and break-beam sensors that work together to detect and track single or multiple moving objects like colored balls. The prototype is validated in structured settings with the goal of eventual deployment in more dynamic environments, addressing the challenge of reliably tracking visually similar objects with minimal distinguishing features. The architecture integrates computer vision with sensor fusion by combining …


Developing Deep Learning Methods For Cognitive Impairment Detection Based On Non-Invasive Data, Muath Alsuhaibani Aug 2025

Developing Deep Learning Methods For Cognitive Impairment Detection Based On Non-Invasive Data, Muath Alsuhaibani

Electronic Theses and Dissertations

Cognitive impairment detection is on the rise to help reduce the burden of healthcare costs on institutions and individuals. Mild Cognitive Impairment (MCI) is an early stage of cognitive decline progressing to Alzheimer’s disease (AD) or AD-related Dementia (ADRD). Detecting the early stages of AD/ADRD is crucial for early interventions among older adults to mitigate cognitive decline over time. However, the current diagnostic methods are often costly and/or invasive, such as MRI and PET scans. Thus, the search for non-invasive and cost-effective screening tools for the early detection of cognitive impairment using speech, language, visual, and motor data is growing. …


Analysis Of Vision Transformers And Domain Adaptation In Long-Range Facial Recognition, Zachary Michael Swanson Jul 2025

Analysis Of Vision Transformers And Domain Adaptation In Long-Range Facial Recognition, Zachary Michael Swanson

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

Atmospheric turbulence presents a significant barrier to long-range facial recognition, introducing severe geometric distortions and blur that degrade image quality. This thesis investigates deep learning approaches for mitigating these effects, with a focus on transformer based architectures and domain adaptation strategies.

An in-depth benchmarking study was performed using convolutional neural networks (CNNs) and vision transformers (ViTs) on the Husker BRIAR Research Collection from up to 500m (HBRC-500) face dataset. The results demonstrated that vision transformers, particularly hierarchical vision transformers like the shifted-window (Swin) transformer, outperform CNN-based models at long distances due to their ability to model global spatial relationships and …


Enriching Vision Representation By Deep Neural Networks And Self-Supervised Learning, Yucong Shen May 2025

Enriching Vision Representation By Deep Neural Networks And Self-Supervised Learning, Yucong Shen

Dissertations

Nowadays, more and more interesting computer vision tasks are tackled by deep learning approaches. However, the increasing model complexity imposes significant computational and storage costs. To address this challenge, this dissertation explores efficient deep learning techniques, proposing morphological layer, an efficient feature extraction layer. It achieves competitive image classification accuracy with significantly decreased model parameters. Another attempt at efficient deep learning is a proposed channel pruning approach that compresses deep neural networks by identifying and removing redundant channels using optimal transport theory. This approach achieves significant reductions in model size and computational cost while maintaining or even improving performance across …


Automated Solar Pv Analysis With Machine Learning And Computer Vision: Dataset And Methodology, Malachi Massey May 2025

Automated Solar Pv Analysis With Machine Learning And Computer Vision: Dataset And Methodology, Malachi Massey

Electrical Engineering and Computer Science Undergraduate Honors Theses

Solar power is a vital resource in a world being threatened with the ever-evolving impacts of climate change. A combination of new and developing technologies have allowed solar photovoltaic installation to increase at an exponential rate. With this rapid and unprecedented growth comes the task of maintaining tens of thousands of square miles of solar photovoltaic panels. Manually observing and testing solar PV panels for defects or obstructions is costly and time-consuming, distracting valuable resources from the continued installation of new units. This research aims to (i) firstly, introduce a novel dataset on solar PV obstruction, named De-Solar dataset; (ii) …


Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova Apr 2025

Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova

Engineering Faculty Articles and Research

Fire Influence on Regional to Global Environments and Air Quality (FIREX-AQ) was a field campaign aimed at better understanding the impact of wildfires and agricultural fires on air quality and climate. The FIREX-AQ campaign took place in August 2019 and involved two aircraft and multiple coordinated satellite observations. This study applied and evaluated a self-supervised machine learning (ML) method for the active fire and smoke plume identification and tracking in the satellite and sub-orbital remote sensing datasets collected during the campaign. Our unique methodology combines remote sensing observations with different spatial and spectral resolutions. With as much as a 10% …


From Image Enhancement To Model Protection Integrating Generative Ai And Secure Learning In Computer Vision, Mohammad Shahab Uddin Apr 2025

From Image Enhancement To Model Protection Integrating Generative Ai And Secure Learning In Computer Vision, Mohammad Shahab Uddin

Electrical & Computer Engineering Theses & Dissertations

This dissertation aims to address critical challenges in the field of computer vision and machine learning, focusing on three key areas: image translation, denoising, and model security. The research encompasses novel methodologies and models that significantly advance existing techniques. This dissertation will not only provide valuable contributions to the academic community but also hold significant potential for practical applications in domains ranging from surveillance to autonomous systems.

Consequently, this dissertation proposes three goals. First, we present new approaches for converting optical videos to infrared videos using deep learning. To apply powerful deep learning based algorithms for object detection and classification …


Accelerated Multiobjective Calibration Of Fused Deposition Modeling 3d Printers Using Multitask Bayesian Optimization And Computer Vision, Craig S. Ganitano, Benji Maruyama, Gilbert L. Peterson Feb 2025

Accelerated Multiobjective Calibration Of Fused Deposition Modeling 3d Printers Using Multitask Bayesian Optimization And Computer Vision, Craig S. Ganitano, Benji Maruyama, Gilbert L. Peterson

Faculty Publications

Proper process parameter calibration is critical to the success of fused deposition modeling (FDM) three-dimensional (3D) printing, but is time-consuming and requires expertise. While existing systems for autonomous calibration have demonstrated success in calibrating for a single objective, users may need to balance multiple conflicting objectives. Herein, an easily deployable, camera-based system for autonomous calibration of FDM printers that optimizes for both part quality and completion time is presented. Autonomous calibration is achieved through a novel, multifaceted computer vision characterization and a multitask learning extension to Bayesian optimization. The system is demonstrated on four popular filament types using two distinct …


Scalable, Secure, And Adaptable Perception Systems Through Adversarial Analysis And Federated Fine-Tuning, Arkajyoti Mitra Jan 2025

Scalable, Secure, And Adaptable Perception Systems Through Adversarial Analysis And Federated Fine-Tuning, Arkajyoti Mitra

Computer Science and Engineering Dissertations - Archive

Perception systems are fundamental to intelligent machines, enabling them to sense, understand, and interpret complex environments. However, as perception increasingly underpins critical applications such as autonomous vehicles, IoT healthcare devices, and smart trading platforms, challenges related to security, scalability, and environmental understanding have become more pressing. This work addresses three core research questions: (1) How can we identify, analyze, and mitigate adversarial vulnerabilities in perception systems to ensure reliable operation under adversarial conditions? (2.1) How can AVPS models be efficiently scaled and fine-tuned across decentralized and resource-constrained environments while preserving privacy and performance? (2.2) How can we scale generative models …


Comparative Evaluation Of Traditional Machine Learning And Deep Cnn Models For Static Hand Gesture Recognition, Anamol Khadka, Prit Desai Jan 2025

Comparative Evaluation Of Traditional Machine Learning And Deep Cnn Models For Static Hand Gesture Recognition, Anamol Khadka, Prit Desai

Computer Science and Engineering Student Research - Archive

Hand gesture recognition plays a vital role in facilitating natural and intuitive human-computer interaction, with applications ranging from sign language translation to touchless control systems. This study presents a comparative evaluation of traditional machine learning models and a deep convolutional neural network (CNN) for static hand gesture classification. The experimental dataset comprises 24,000 training images and 6,000 testing images, spanning 20 gesture classes. Traditional models, including k-Nearest Neighbors (KNN) and Support Vector Machines (SVM), utilize handcrafted features such as convex hull, convexity defects, and Hu moments. In contrast, the deep learning approach fine-tunes a ResNet18 architecture to learn features directly …


Human Perception Of Ai Capabilities At Classifying Perturbed Roadway Signs, Katherine R. Garcia, Jing Chen, Yanru Xiao, Scott Mishler, Cong Wang, Bin Hu Jan 2025

Human Perception Of Ai Capabilities At Classifying Perturbed Roadway Signs, Katherine R. Garcia, Jing Chen, Yanru Xiao, Scott Mishler, Cong Wang, Bin Hu

Psychology Faculty Publications

Artificial Intelligence (AI) is crucial to numerous functions required for driving automation systems, including the computer vision techniques used to detect the roadway environment and make real-time decisions. However, the images used as inputs to the AI system may be maliciously perturbed, or manipulated, causing the AI system to make an incorrect classification. In this study, we examined humans’ perception of the AI’s computer vision capability of classifying various road sign images, including the original images, images with two different types of malicious attacks, and images that are scrambled randomly at the pixel level. Our results showed that participants rated …


Information-Theoretic Methods For Efficient Training And Robust Evaluations In Self-Supervised Learning, Oscar Skean Jan 2025

Information-Theoretic Methods For Efficient Training And Robust Evaluations In Self-Supervised Learning, Oscar Skean

Theses and Dissertations--Computer Science

Self-supervised learning (SSL) has become a cornerstone of modern machine learning, offering a scalable alternative to costly human annotation by constructing pretext tasks directly from raw data. While SSL has delivered strong results across vision, language, and multimodal domains, two major limitations persist: (1) SSL methods are often significantly slower to train than supervised counterparts, and (2) evaluation protocols remain narrow, with most studies relying on linear probing accuracy on the pretraining dataset. . These challenges are particularly acute for large language models (LLMs), where training costs and interpretability of intermediate representations are critical concerns.

In this work, we propose …


Using Quanser Platform To Introduce Engineering Technology Students To Autonomous Vehicles, Otilia Popescu, Logan Beaver, Murat Kuzlu, Krishnanand Kaipa Jan 2025

Using Quanser Platform To Introduce Engineering Technology Students To Autonomous Vehicles, Otilia Popescu, Logan Beaver, Murat Kuzlu, Krishnanand Kaipa

Engineering Technology Faculty Publications

The area of autonomous vehicles is not new, but the latest advances in various technologies gave it a new boost in the last decade and it keeps growing in interest. However, undergraduate curricula rarely include courses specific to this area, which is considered mostly an interdisciplinary graduate field. While various programs introduce students to the background needed to understand and approach the field, specific work on autonomous vehicle projects is left for extra curriculum activities or student clubs, and eventually for senior (capstone) projects. This paper presents the work of a team of electrical engineering technology students on an autonomous …


Human Perception Of Ai Capabilities At Classifying Perturbed Roadway Signs, Katherine R. Garcia, Jing Chen, Yanru Xiao, Scott Mishler, Cong Wang, Bin Hu Jan 2025

Human Perception Of Ai Capabilities At Classifying Perturbed Roadway Signs, Katherine R. Garcia, Jing Chen, Yanru Xiao, Scott Mishler, Cong Wang, Bin Hu

Computer Science Faculty Publications

Artificial Intelligence (AI) is crucial to numerous functions required for driving automation systems, including the computer vision techniques used to detect the roadway environment and make real-time decisions. However, the images used as inputs to the AI system may be maliciously perturbed, or manipulated, causing the AI system to make an incorrect classification. In this study, we examined humans’ perception of the AI’s computer vision capability of classifying various road sign images, including the original images, images with two different types of malicious attacks, and images that are scrambled randomly at the pixel level. Our results showed that participants rated …


Embodied Ai For Challenging Rearrangement Tasks In The Context Of Service And Assistive Robots, Mariia Khan Jan 2025

Embodied Ai For Challenging Rearrangement Tasks In The Context Of Service And Assistive Robots, Mariia Khan

Theses: Doctorates and Masters

Embodied AI explores intelligent agents that learn through interaction with their environment, aiming to replicate human-like learning processes. Achieving this requires agents capable of understanding a scene via various sensors, reasoning about their actions, and reacting accordingly. These abilities are necessary for service domestic robots to assist humans in their day-to-day activities. Embodied AI tasks can include but are not limited to: visual exploration, visual navigation, instruction following and embodied question answering, which typically consider static (unchanging) environments, where objects do not move over time. This thesis addresses one of the most challenging Embodied AI tasks – visual room rearrangement, …


Toward Embodied Navigation Through Vision And Language, Muraleekrishna Gopinathan Jan 2025

Toward Embodied Navigation Through Vision And Language, Muraleekrishna Gopinathan

Theses: Doctorates and Masters

Embodied AI is a challenging but exciting field in which a robot learns to interact with human-living spaces to perform various tasks. This thesis studies the embodied navigation problem in which a robotic agent navigates in a previously unseen indoor environment based on a challenging task. In particular, the Vision-and-Language Navigation (VLN) task requires a robot to navigate based on a descriptive human-language instruction. This thesis aims to improve VLN agents on four key aspects - their understanding of the environment, training via additional data, correcting navigational errors, and predicting the layout of the environment for better planning.

First, we …


Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni Dec 2024

Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni

Theses and Dissertations

This dissertation addresses critical challenges in neural network design by leveraging entropy-based techniques to improve model efficiency, interpretability, and bias reduction. Focusing on the unique demands of computer vision applications, particularly object detection and classification for real-time systems, this work introduces a series of innovative methods centered on information theory. At the core of these methods is the Probabilistic Explanations of Entropic Knowledge (PEEK) framework, a tool developed to analyze and visualize entropy distributions across feature maps. PEEK offers insights into information flow within neural networks, making it possible to pinpoint layers that contribute meaningfully to decision-making or identify those …


Collectively Advancing Deep Learning For Animal Detection In Drone Imagery: Successes, Challenges, And Research Gaps, Daniel Axford, Ferdous Sohel, Mathew A. Vanderklift, Amanda J. Hodgson Nov 2024

Collectively Advancing Deep Learning For Animal Detection In Drone Imagery: Successes, Challenges, And Research Gaps, Daniel Axford, Ferdous Sohel, Mathew A. Vanderklift, Amanda J. Hodgson

Research outputs 2022 to 2026

Drones have emerged as a powerful tool in animal detection, significantly advancing wildlife monitoring, conservation, and management by capturing high-resolution, real-time imagery over areas often inaccessible or challenging for human observers to reach. However, manual analysis of drone imagery for animal detection is labour-intensive and time-consuming. The application of deep learning methods, particularly convolutional neural networks, in automating animal detection from drone imagery has the potential to revolutionise wildlife monitoring, conservation, and management protocols. This review provides a comprehensive overview of the increasing use and prospects of deep learning in animal detection using drone imagery. It explores successful applications of …