Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer vision

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 349

Full-Text Articles in Computer Sciences

Dbssnet: Dual-Branch Spectral-Spatial Network With Data-Driven And Knowledge-Guided Band Selection For Uav Hyperspectral Wheat Rust Detection, Subin Kim Aug 2026

Dbssnet: Dual-Branch Spectral-Spatial Network With Data-Driven And Knowledge-Guided Band Selection For Uav Hyperspectral Wheat Rust Detection, Subin Kim

All Graduate Theses and Dissertations, Fall 2023 to Present

Wheat rust is a serious plant disease that can reduce crop yield and quality. In practice, the disease is often noticed only after visible symptoms appear, when some damage may already be difficult to reverse. This thesis studies whether drone-based imaging can help detect wheat rust earlier and more reliably in field environments. 

Unlike an ordinary color photograph, a hyperspectral image records reflected light at many narrow wavelengths. These measurements can reveal useful information about plant condition, but they are also high dimensional, noisy, and difficult to analyze when only a limited number of labeled field samples are available. To …


Scalefusion: Hierarchical Feature Aggregation For Unified Multiscale Object Detection, Ayşenur Yaylaci, Berru Kaya, Mehmet Kiliçarslan Jul 2026

Scalefusion: Hierarchical Feature Aggregation For Unified Multiscale Object Detection, Ayşenur Yaylaci, Berru Kaya, Mehmet Kiliçarslan

Turkish Journal of Electrical Engineering and Computer Sciences

Detecting objects across a wide range of scales, particularly small ones, remains a significant challenge in computer vision. Existing methods often improve small object detection at the cost of performance on larger objects or introduce significant computational overhead through external techniques like image slicing. This paper introduces ScaleFusion, a novel, unified, end-to-end object detection architecture designed to provide robust performance across all scales within a single model. The core of our approach is a hierarchical feature aggregation strategy structured like a tree. ScaleFusion processes an image by running a shared backbone network only on fine-grained patches at the lowest level …


Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane Jun 2026

Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane

Beyond: Undergraduate Research Journal

Autonomous tracking of agile unmanned aerial vehicles (UAVs) presents significant challenges for real-time perception and control systems. This work presents AIRHOUND (Autonomous Intelligent Rotorcraft for Hostile Object Unified Navigation and Detection), a UAV platform implementing vision-based yaw tracking through a modular ROS2 software architecture. The system employs YOLOv8 object detection optimized with NVIDIA TensorRT for embedded deployment on an NVIDIA Jetson Orin companion computer. Detected targets are processed through a geometric tracking module that converts pixel coordinates to angular yaw errors using pinhole camera intrinsics, with a proportional controller generating rate-limited yaw commands. These commands are streamed to a PX4 …


Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud Jun 2026

Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud

Dissertations

Deep learning has become the dominant approach for solving vision-centric problems; however, its successful deployment in real-world applications remains limited by high computational cost, data dependency, and insufficient integration with practical and human-centered environments. While state-of-the art deep learning models often achieve impressive performance in controlled settings, they frequently fail to generalize or operate efficiently under deployment constraints such as limited resources, complex data modalities, and real-time interaction requirements. These limitations motivate the need for a deployment-oriented deep learning framework that balances accuracy, efficiency, and practical usability.

This dissertation investigates the design and deployment of efficient deep learning architectures for …


Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan May 2026

Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan

Dissertations

Digital three-dimensional (3D) models are central to engineering design, analysis, and manufacturing, but learning pipelines for man-made geometry often operate on sampled carriers that do not preserve all of the structure present in exact CAD representations. This dissertation studies learning-based reconstruction and part decomposition for structured man-made 3D geometry, from general object benchmarks to CAD-derived datasets, with a focus on neural implicit representations trained from signed-distance samples, point clouds, and tessellated meshes. The goal is to make these models more accurate, more part-aware, and more consistently supervised.

First, signed distance function (SDF) reconstruction with implicit neural representations is improved through …


Three-Dimensional Shape Cues Affect Human And Artificial Recognition Systems Differently, Mikayla Cutler, Luke D. Baumel, Joseph Tocco, William Friebel, George K. Thiruvathukal, Nicholas Baker Dr. May 2026

Three-Dimensional Shape Cues Affect Human And Artificial Recognition Systems Differently, Mikayla Cutler, Luke D. Baumel, Joseph Tocco, William Friebel, George K. Thiruvathukal, Nicholas Baker Dr.

Computer Science: Faculty Publications and Other Works

Humans and neural networks use shape and texture information differently. While humans weigh shape heavily in their ultimate classification decision, neural networks are more biased towards texture cues. Many tests of shape vs. texture bias have focused on shape recognition from an object’s external contour. However, shape information is also conveyed through internal contours, shading, and attached shadows, especially when an object is viewed from noncanonical perspectives. Using models from ShapeNet, we created datasets of 120,000 texture-substituted images of objects from many viewpoints with and without shading and attached shadows. We tested humans’ and several neural networks’ ability to classify …


Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella May 2026

Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella

Capstone Projects

Cataloging and digitizing the objects inside a building manually is a task that is often impractical at scale. This project therefore automates the process, using a custom-made system. Using a photogrammetry-based 3D reconstruction of a room, this system is applied to sequences of 2D images used to make the 3D models. The system applies object detection, image segmentation, and image-text models to identify and describe objects, using CNN based models such as YOLO and OpenCLIP. Each analyzed object is then stored in a structured database with spatial coordinates from the 3D scanning, descriptive attributes from the image-text models, and other …


Spaceforge: Spatial Reconstruction For Signal Simulations, Compton Ross May 2026

Spaceforge: Spatial Reconstruction For Signal Simulations, Compton Ross

Honors Theses

Signal simulation environments require accurate three dimensional representations of physical spaces, yet current methods for generating these representations, including Light Detection and Ranging (LiDAR) scanning, manual 3D modeling, and commercial photogrammetry, are both costly and time intensive. SpaceForge addresses this gap with a prompt guided pipeline that takes an ordinary indoor photograph and a configurable set of simulation relevant object categories as input and produces a voxelized 3D scene compatible with downstream signal simulation workflows. The pipeline proceeds through five major stages: open set object detection and segmentation, object level preprocessing, single image 3D mesh reconstruction, heuristic pose estimation and …


Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun Mar 2026

Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun

Electronic Theses and Dissertations

The performance of deep neural networks (DNNs) is strongly influenced by the characteristics and quality of the underlying datasets. This Ph.D. dissertation addresses three pervasive data challenges-imbalance, quality degradation, and scarcity-that commonly hinder the effectiveness of DNNs in computer vision (CV) and natural language processing (NLP) applications.

Class imbalance remains one of the most frequent causes of degraded model generalization. While Focal Loss effectively mitigates inter-class imbalance by assigning higher weights to minority classes, it struggles with intra-class imbalance, particularly in video datasets where longer clips dominate feature representation. To address this, I implement and utilize …


Video Generation Techniques For Novel View Synthesis With Flow-Matching Transformers, Xiuyuan Qiu Mar 2026

Video Generation Techniques For Novel View Synthesis With Flow-Matching Transformers, Xiuyuan Qiu

Master's Theses

Novel view synthesis (NVS) aims to generate images of a scene from unseen camera viewpoints. Recent work, such as Stable Virtual Camera, shows that large-scale image diffusion models like Stable Diffusion can be adapted for pose-conditioned view synthesis by incorporating video-generation techniques with camera conditioning. In this thesis, we introduce MVFlow, a new NVS model that extends this approach to a different image generation architecture: a flow-matching diffusion transformer, specifically FLUX.1, which has demonstrated strong performance in image synthesis. We evaluate MVFlow under varying input view counts and pose distance settings. Our results show that this architectural transfer is feasible; …


Pl-Mamba: A 3d Point Cloud Semantic Segmentation Network Based On Bimodal Fusion, He Zhu, Feng Zhou, Mengxiao Zhu, Ju Dai Jan 2026

Pl-Mamba: A 3d Point Cloud Semantic Segmentation Network Based On Bimodal Fusion, He Zhu, Feng Zhou, Mengxiao Zhu, Ju Dai

Journal of System Simulation

Abstract: To enhance the semantic discrimination capability in point cloud semantic segmentation, a 3D point cloud semantic segmentation network named PL-Mamba is proposed, which is centered on the fusion of point cloud (P) and language (L) dual modalities. This method takes PointMamba as the backbone network, leveraging its excellent long-sequence modeling and global perception capabilities. It introduces a language prompt mechanism and uses a pretrained language model BERT to encode the context of category labels, obtaining semantically rich text features. The text information serves as a language guided token and is deeply integrated with point cloud features through cross modal …


Zero-Shot Segmentation Of Estuary Mudflats Using The Segment Anything Model, Jaren Unzen Jan 2026

Zero-Shot Segmentation Of Estuary Mudflats Using The Segment Anything Model, Jaren Unzen

Honors Theses and Capstones

Estuary mudflats are ecologically sensitive environments that require consistent monitoring. Traditional satellite-based classification workflows are often constrained by the high cost and labor-intensive nature of manual data annotation. This study evaluates the utility of Segment Anything Model 3 (SAM 3), a foundational computer vision model, to automate mudflat segmentation without domain-specific fine-tuning. By leveraging the model’s text-prompting capabilities alongside specialized pre- and post-processing techniques, we generated segmentation masks in a zero-shot framework. Our approach achieved an F1- score of 0.51, demonstrating the inherent challenges of spectrally complex coastal features. Despite this, the results highlight a promising pathway for adapting large-scale …


Towards Multimodal Guideline-Aligned Agentic Systems, Wenliang Zhong Jan 2026

Towards Multimodal Guideline-Aligned Agentic Systems, Wenliang Zhong

Computer Science and Engineering Dissertations - Archive

I present my work on building multimodal guideline-aligned agentic systems designed to enable AI agents to solve complex real-world tasks. My research addresses two critical perspectives: (1) Instruction-Aware Embedding Models for flexible and universal embedding tasks, and (2) Guideline-Driven LLM Agents that leverage domain-specific guidelines to perform expert-level tasks. These components address embedding and generation tasks, respectively, and lay the foundation for a hybrid agent capable of tackling challenging real-world applications.

From the embedding perspective, I first address the instruction-following capabilities of embedding models. While Large Language Models (LLMs) excel at instruction following, they are primarily designed for generation rather …


Unsafe2safe: Controllable Image Anonymization For Downstream Utility, Minh Dinh Jan 2026

Unsafe2safe: Controllable Image Anonymization For Downstream Utility, Minh Dinh

Dartmouth College Master’s Theses

Large-scale image datasets frequently contain identifiable or sensitive content, raising privacy risks when training models that may memorize and leak such information. We present Unsafe2Safe, a fully automated pipeline that detects privacy-prone images and rewrites only their sensitive regions using multimodally guided diffusion editing. Unsafe2Safe operates in two stages. Stage 1 uses a vision--language model to (i) inspect images for privacy risks, (ii) generate paired private and public captions that respectively include and omit sensitive attributes, and (iii) prompt a large language model to produce structured, identity-neutral edit instructions conditioned on the public caption. Stage 2 employs instruction-driven diffusion editors …


Reliable And Label-Efficient Learning For Open-World Visual Perception And Robot Learning Under Uncertainty, Zongyao Lyu Jan 2026

Reliable And Label-Efficient Learning For Open-World Visual Perception And Robot Learning Under Uncertainty, Zongyao Lyu

Computer Science and Engineering Dissertations

Modern learning systems deployed in open-world environments must make reliable decisions despite predictive uncertainty, previously unseen classes, limited annotations, and distribution shifts. This dissertation develops methods for reliable and label-efficient learning in visual perception and robot control.

First, this work studies uncertainty in object detection by representing semantic and spatial predictions probabilistically. A deep-ensemble framework aggregates detections into class-probability distributions and probabilistic bounding boxes, while a subsequent extension combines deep ensembles with Monte Carlo dropout to further investigate predictive uncertainty. Second, this dissertation addresses open-set recognition, where classes absent during training may appear at inference time. An empirical study shows …


Pixel-Perfect Segmentation Of Solar Filaments, Jamie Harris Dec 2025

Pixel-Perfect Segmentation Of Solar Filaments, Jamie Harris

Undergraduate Research Symposium

The observation and classification of solar filaments has a drastic impact on the ability to predict solar-magnetic weather phenomena that threatens to put both satellite infrastructure and astronauts at risk. Using the Hɑ filter provided by the Global Oscillations Network Group (GONG), a network of six telescopes around the world dedicated to 24/7 surveillance of the sun, we are able to get images that clearly and prominently display filament activity. With the vast amount of images the GONG takes, it is not possible to manually analyze every image. Using the U-Net model for computer vision, we were able to train …


Precision Agriculture In The Age Of Ai: A Systematic Review Of Machine Learning Methods For Crop Disease Detection, Munir Majdalawieh, Carla Martins, Mohammed Radi, Maher Alaraj, Shafaq Khan Dec 2025

Precision Agriculture In The Age Of Ai: A Systematic Review Of Machine Learning Methods For Crop Disease Detection, Munir Majdalawieh, Carla Martins, Mohammed Radi, Maher Alaraj, Shafaq Khan

All Works

Artificial Intelligence (AI) has become a critical tool in modern precision agriculture, particularly in the detection of plant diseases and pests. This study provides a comprehensive review of current AI methodologies applied to crop disease detection, with a focus on machine learning models, dataset availability, and performance metrics. Our findings indicate that Convolutional Neural Networks (CNNs) are the most widely used and cost-effective approach, while Vision Transformers (ViTs) exhibit superior accuracy but require significantly higher computational resources. We identify key research gaps, including the geographic bias in dataset origins, the trade-off between data quality and quantity, and the limited exploration …


Towards Vision-Brain Understanding At Scales: From Classical To Quantum Machine Learning Approaches, Xuan-Bac Nguyen Dec 2025

Towards Vision-Brain Understanding At Scales: From Classical To Quantum Machine Learning Approaches, Xuan-Bac Nguyen

Graduate Theses and Dissertations

In recent years, large-scale learning approaches such as unsupervised and self-supervised learning have revolutionized artificial intelligence. These methods enable machines to learn high-level representations without explicit human supervision, achieving remarkable success across vision, language, and multimodal tasks. However, such advances come at a cost—they rely on massive datasets, billions of parameters, and extensive computational resources. Despite these achievements, artificial systems still fall short of the remarkable learning efficiency of the human brain, which can infer, adapt, and generalize from limited experiences. This gap motivates a deeper exploration of how biological intelligence acquires knowledge and how these principles can inspire the …


Deconstructing The Black Box: An Explainability Analysis Of Deep Learning Architectures In Cytopathology, Chase A. Garrett Nov 2025

Deconstructing The Black Box: An Explainability Analysis Of Deep Learning Architectures In Cytopathology, Chase A. Garrett

Master's Theses or Doctor of Nursing Practice

Deep learning shows strong potential in medical-image analysis, yet adoption in cyptopathology remains limited. Cytopathology could benefit from deep learning applications by improving diagnostic efficiency and accuracy. However deep learning comes with a notorious “black box” that keeps the models from being transparent and trustworthy for widespread clinical adoption. We conducted a comprehensive and comparative analysis of several deep learning architectures for multi-class classification of acute leukemia types, ALL, AML, and normal healthy cells from peripheral blood smear images. The models in this research include a Vision Transformer (ViT) and a diverse selection of Convolutional Neural Network (CNN) models. The …


Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh Nov 2025

Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh

Doctoral

Image captioning models enable us to automatically generate natural language image descriptions for previously unseen images. It combines the two fields of computer vision and natural language generation, allowing models to interpret the con tent of an image and communicate that knowledge through natural language text.

Research into image captioning has the potential benefit of reducing the gap in digital information availability between fully sighted individuals and those who are visually impaired. However, automatically generated captions often fail to provide the required level of detail and specificity to achieve this goal. Furthermore, current standard evaluation methods are insufficient at measuring …


Intelligence Architectures And Machine Learning Applications In Contemporary Spine Care, Rahul Kumar, Conor Dougherty, Kyle Sporn, Akshay Khanna, Puja Ravi, Pranay Prabhakar, Nasif Zaman Sep 2025

Intelligence Architectures And Machine Learning Applications In Contemporary Spine Care, Rahul Kumar, Conor Dougherty, Kyle Sporn, Akshay Khanna, Puja Ravi, Pranay Prabhakar, Nasif Zaman

SKMC Student Presentations and Publications

The rapid evolution of artificial intelligence (AI) and machine learning (ML) technologies has initiated a paradigm shift in contemporary spine care. This narrative review synthesizes advances across imaging-based diagnostics, surgical planning, genomic risk stratification, and post-operative outcome prediction. We critically assess high-performing AI tools, such as convolutional neural networks for vertebral fracture detection, robotic guidance platforms like Mazor X and ExcelsiusGPS, and deep learning-based morphometric analysis systems. In parallel, we examine the emergence of ambient clinical intelligence and precision pharmacogenomics as enablers of personalized spine care. Notably, genome-wide association studies (GWAS) and polygenic risk scores are enabling a shift from …


Intelligent Motion Tracking: A Low-Cost Surveillance Framework Using Soc Devices And Sensor Fusion, Aung Myat Khaung, Nicholas Michael Stiffler Aug 2025

Intelligent Motion Tracking: A Low-Cost Surveillance Framework Using Soc Devices And Sensor Fusion, Aung Myat Khaung, Nicholas Michael Stiffler

Research from the Berry Summer Thesis Institute, 2025

This thesis presents the design and implementation of a lightweight surveillance system capable of realtime motion detection, object tracking, and behavioral history reconstruction in controlled environments. The system uses System-on-Chip devices such as Raspberry Pi boards equipped with NOIR cameras, monocular cameras, and break-beam sensors that work together to detect and track single or multiple moving objects like colored balls. The prototype is validated in structured settings with the goal of eventual deployment in more dynamic environments, addressing the challenge of reliably tracking visually similar objects with minimal distinguishing features. The architecture integrates computer vision with sensor fusion by combining …


Developing Deep Learning Methods For Cognitive Impairment Detection Based On Non-Invasive Data, Muath Alsuhaibani Aug 2025

Developing Deep Learning Methods For Cognitive Impairment Detection Based On Non-Invasive Data, Muath Alsuhaibani

Electronic Theses and Dissertations

Cognitive impairment detection is on the rise to help reduce the burden of healthcare costs on institutions and individuals. Mild Cognitive Impairment (MCI) is an early stage of cognitive decline progressing to Alzheimer’s disease (AD) or AD-related Dementia (ADRD). Detecting the early stages of AD/ADRD is crucial for early interventions among older adults to mitigate cognitive decline over time. However, the current diagnostic methods are often costly and/or invasive, such as MRI and PET scans. Thus, the search for non-invasive and cost-effective screening tools for the early detection of cognitive impairment using speech, language, visual, and motor data is growing. …


Alignment Of Perceptual Similarity Metrics With Human Perception, Abhijay Ghildyal Jul 2025

Alignment Of Perceptual Similarity Metrics With Human Perception, Abhijay Ghildyal

Dissertations and Theses

Perceptual similarity metrics are used for quantitatively evaluating the similarity between two images as it would appear to human perception. These metrics aim to mimic the human visual system, providing a more accurate assessment of visual similarity. Such visual assessments are considered to be more advanced than simple pixel-wise comparisons such as ℓp norm distances. Thus, a human-like assessment of visual similarity, makes the metrics valuable for applications in image compression, restoration, and enhancement, where evaluating perceptual quality is crucial. Perceptual similarity metrics have progressively become more correlated with human judgments on perceptual similarity; however, despite recent advances, the …


Analysis Of Vision Transformers And Domain Adaptation In Long-Range Facial Recognition, Zachary Michael Swanson Jul 2025

Analysis Of Vision Transformers And Domain Adaptation In Long-Range Facial Recognition, Zachary Michael Swanson

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

Atmospheric turbulence presents a significant barrier to long-range facial recognition, introducing severe geometric distortions and blur that degrade image quality. This thesis investigates deep learning approaches for mitigating these effects, with a focus on transformer based architectures and domain adaptation strategies.

An in-depth benchmarking study was performed using convolutional neural networks (CNNs) and vision transformers (ViTs) on the Husker BRIAR Research Collection from up to 500m (HBRC-500) face dataset. The results demonstrated that vision transformers, particularly hierarchical vision transformers like the shifted-window (Swin) transformer, outperform CNN-based models at long distances due to their ability to model global spatial relationships and …


Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He Jul 2025

Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He

Research Collection School Of Computing and Information Systems

Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenges, we propose Instruct2See, a novel zero-shot framework capable of handling both seen and unseen obstacles. The core idea of our approach is to unify obstruction removal by treating it as a soft-hard mask restoration problem, where any obstruction can be represented using multi-modal prompts, such as visual semantics and textual instructions, …


Enriching Vision Representation By Deep Neural Networks And Self-Supervised Learning, Yucong Shen May 2025

Enriching Vision Representation By Deep Neural Networks And Self-Supervised Learning, Yucong Shen

Dissertations

Nowadays, more and more interesting computer vision tasks are tackled by deep learning approaches. However, the increasing model complexity imposes significant computational and storage costs. To address this challenge, this dissertation explores efficient deep learning techniques, proposing morphological layer, an efficient feature extraction layer. It achieves competitive image classification accuracy with significantly decreased model parameters. Another attempt at efficient deep learning is a proposed channel pruning approach that compresses deep neural networks by identifying and removing redundant channels using optimal transport theory. This approach achieves significant reductions in model size and computational cost while maintaining or even improving performance across …


An Edge Computing Device Optimized And Transfer Learning Enhanced Deep Learning Model For Detecting Wildfire Flame And Smoke, Giovanny Vazquez May 2025

An Edge Computing Device Optimized And Transfer Learning Enhanced Deep Learning Model For Detecting Wildfire Flame And Smoke, Giovanny Vazquez

UNLV Theses, Dissertations, Professional Papers, and Capstones

The integration of autonomous unmanned aerial vehicles (UAVs) with edge computing technology and deep learning (DL)-based object detection offers a groundbreaking solution for real-time wildfire detection, enabling rapid data processing directly on devices and minimizing response delays in critical scenarios. However, although showing early promise, performance is often constrained by limited training data and edge computing devices that lack graphics processing unit (GPU) acceleration. This thesis seeks to address these limitations in two stages.First, this work explores the transformative potential of Transfer Learning (TL) to enhance wildfire object detection model accuracy while also investigating TL’s impact, for DL-based object detection …


Automated Solar Pv Analysis With Machine Learning And Computer Vision: Dataset And Methodology, Malachi Massey May 2025

Automated Solar Pv Analysis With Machine Learning And Computer Vision: Dataset And Methodology, Malachi Massey

Electrical Engineering and Computer Science Undergraduate Honors Theses

Solar power is a vital resource in a world being threatened with the ever-evolving impacts of climate change. A combination of new and developing technologies have allowed solar photovoltaic installation to increase at an exponential rate. With this rapid and unprecedented growth comes the task of maintaining tens of thousands of square miles of solar photovoltaic panels. Manually observing and testing solar PV panels for defects or obstructions is costly and time-consuming, distracting valuable resources from the continued installation of new units. This research aims to (i) firstly, introduce a novel dataset on solar PV obstruction, named De-Solar dataset; (ii) …


Modeling Language And Vision At Human Scales, Clayton Fields May 2025

Modeling Language And Vision At Human Scales, Clayton Fields

Boise State University Theses and Dissertations

The impressive results that have recently been achieved in natural language processing and artificial intelligence have been primarily driven by the introduction of the transformer deep learning architecture, increasingly large models with many parameters and using enormous datasets. The size of models and their training data requirements present costly demands that freeze many researchers out of training with cutting edge models. Beyond these practical implications, current methods learn from text alone, without the rich array of sensory information that human beings use in learning language. This means that language models are often incapable of reasoning about the concrete world that …