Open Access. Powered by Scholars. Published by Universities.®

Computer Vision

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 76

Full-Text Articles in Artificial Intelligence and Robotics

Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega May 2026

Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega

Honors Theses

Running gait analysis plays a critical role in injury prevention and performance optimization, however, existing approaches often rely on specialized laboratory equipment or wearable sensors with limited interpretability. Recent advances in computer vision, particularly 2D human pose estimation, enable markerless motion analysis from standard video. However, progress remains constrained by the lack of publicly available datasets designed for running form analysis.

In this work, we introduce a preliminary dataset and benchmark for stride-level running gait analysis. The dataset consists of 73 treadmill running videos from 15 participants with varying experience levels, annotated with over 4,600 stride-level labels across multiple biomechanical …


Machine Learning For Handwritten Character Recognition, Hannah Freitag May 2026

Machine Learning For Handwritten Character Recognition, Hannah Freitag

Honors Capstones

Handwritten character recognition remains a challenging problem in machine learning due to the high variability of handwriting across individuals and the visual similarity between certain character classes. This project explores whether Singular Value Decomposition (SVD)-based dimensionality reduction can serve as an effective preprocessing step for a fully connected neural network trained on the EMNIST Balanced dataset, a 47-class benchmark of handwritten digits and letters. By projecting 784- dimensional pixel inputs onto the top 70 principal components, approximately 90% of the total variance is preserved while reducing input dimensionality by 91%. The resulting SVD-based model achieves approximately 94% test accuracy, outperforming …


Enhancing Traffic Safety Through Ai-Driven, Privacy-Preserving, And Secure Impaired Driving Detection Systems, Razan Alsulieman May 2026

Enhancing Traffic Safety Through Ai-Driven, Privacy-Preserving, And Secure Impaired Driving Detection Systems, Razan Alsulieman

Dissertations

Drunk driving remains a major threat to road safety worldwide, contributing significantly to traffic injuries and fatalities each year. Traditional detection approaches are largely reactive and vehicle-centric, relying on in-vehicle sensors, breathalyzers, or post-incident enforcement. These methods often depend on driver cooperation, intrusive hardware installations, or limited monitoring environments, restricting their scalability and effectiveness in large transportation systems. At the same time, modern cities increasingly deploy roadside cameras, surveillance networks, and drone- based monitoring systems, creating new opportunities for proactive intoxication detection at the infrastructure level. However, leveraging such external monitoring introduces challenges related to secure data collection, reliable AI-based …


Pediatric Inflammatory Bowel Disease Tissue Classification From Pathology Slide Images: Detecting Phenotypes Using Computer Vision, Chloe Martin-King, Ali Nael, Louis Ehwerhemuepha, Blake Calvo, Quinn Gates, Jamie Janchoi, Elisa Ornelas, Melissa Perez, Andrea Venderby, John Miklavcic, Peter Chang, Aaron Sassoon, Brian Rubio, Ghislaine Barrigan, Kenneth Grant Feb 2026

Pediatric Inflammatory Bowel Disease Tissue Classification From Pathology Slide Images: Detecting Phenotypes Using Computer Vision, Chloe Martin-King, Ali Nael, Louis Ehwerhemuepha, Blake Calvo, Quinn Gates, Jamie Janchoi, Elisa Ornelas, Melissa Perez, Andrea Venderby, John Miklavcic, Peter Chang, Aaron Sassoon, Brian Rubio, Ghislaine Barrigan, Kenneth Grant

Food Science Faculty Articles and Research

Background and Aims

With the advent of computer vision algorithms, we hypothesize that histopathology images from endoscopic biopsies may be utilized for automated classification of histologic phenotypes, thus guiding Crohn’s disease and ulcerative colitis diagnosis and treatment. The aim of our study is to assess whether artificial intelligence can be used to improve pediatric inflammatory bowel disease outcomes by aiding pathologists with accurate detection of abnormal tissue sections.

Methods

Three two-dimensional (2D) convolutional neural networks with multiple instance learning were developed to classify histopathology tissue sections as normal vs abnormal and as containing active inflammation and/or chronic changes/architectural distortion.

Results …


Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci Jan 2026

Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci

Theses and Dissertations--Electrical and Computer Engineering

Fine-grained Temporal Action Segmentation (TAS) has become a cornerstone of video understanding, offering dense frame-level predictions essential for clinical assessment, surgical skill evaluation, and human-computer interaction. While TAS methods have delivered strong results on coarse-grained benchmarks, two fundamental challenges persist: (1) global attention mechanisms dilute boundary information critical for subsecond precision, a phenomenon we term the temporal granularity bottleneck, and (2) dense frame-level annotation remains prohibitively expensive, with most datasets requiring exhaustive labeling of lengthy untrimmed videos. These challenges are particularly pronounced in medical domains, where sub-second primitives define clinical outcomes while expert annotation remains scarce. In this dissertation, we …


Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail Jan 2026

Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail

Theses and Dissertations (Comprehensive)

Deploying deep learning models for medical image analysis on mobile devices requires a balance between inference latency, memory footprint, and delineating anatomical boundaries with high accuracy. While Convolutional Neural Networks (CNNs) and mobile Vision Transformers (ViTs) offer efficiency, they often struggle to model the irregular, non-local geometric structures inherent in biological tissues without incurring prohibitive computational costs. In this thesis, we introduce GeoViG (Geometric Vision Graph), an architecture that bridges the gap between efficient grid-based processing and explicit Geometric Deep Learning. GeoViG introduces a novel transition from high-resolution pixel grids to low-resolution dynamic graphs via a SpreadEdgePool operator, a geometry-aware …


Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou Dec 2025

Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou

Undergraduate Research Symposium

Various computational models of first impressions have been developed to uncover the mechanisms driving these judgments. However, the implicit notion of a singular ``human'' often overlooks meaningful individual differences in beliefs, attitudes, and associations, as well as culturally grounded group-level constructs. In this paper, we extend Cultural Consensus Theory (CCT) to estimate culturally shared beliefs about faces by incorporating latent constructs structured around interpretable facial features extracted via computer vision algorithms. We apply our model to a large-scale dataset of people’s first impressions of faces. Our approach reveals a robust mapping between facial features and culturally constructed impressions, allowing us …


Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga Dec 2025

Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga

MS in Computer Science Project Reports

We present a grading system that accelerates evaluation of open-ended student work across scanned and digital workflows. The system crops answer regions from PDFs, assigns submissions via OCR on identity regions only, and groups answers by visual semantics using a vision LLM. Instructors review and edit groups, apply rubric items once per group, and export grades from an on-screen table. The solution integrates Ghostscript rasterization, PdfPig page orchestration, SkiaSharp region extraction, Tesseract identity OCR, and GPT-4o Vision for grouping. We detail the architecture, token-budgeted batching strategy, and persistence design, then describe testing results for grouping quality, time-on-task, and usability. The …


Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui Dec 2025

Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui

Graduate Theses and Dissertations

Negation is a fundamental linguistic concept used by humans to convey information that they do not desire. Despite this, minimal research has focused on negation within text-guided image editing. This lack of research means that vision-language models (VLMs) for image editing may struggle to understand negation, implying that they struggle to provide accurate results. One barrier to achieving human-level intelligence is the lack of a standard collection by which research into negation can be evaluated. This thesis presents the first large-scale dataset, Negative Instruction (NeIn), for studying negation within instruction-based image editing. Our dataset comprises 366,957 quintuplets, i.e., source image, …


Correcting Class Imbalance Through Synthetic Training Data And 3d Modeling For Carabid Pitfall Trap Sampling, Blair Mirka Nov 2025

Correcting Class Imbalance Through Synthetic Training Data And 3d Modeling For Carabid Pitfall Trap Sampling, Blair Mirka

Geography ETDs

Crowdsourced biodiversity data provide an accessible foundation for large-scale ecological monitoring, but class imbalance limits automated species identification, particularly for rare taxa. This research explores the use of synthetic training data generated from 3D models of carabid beetle museum specimens to improve detection and classification performance for underrepresented species in crowdsourced datasets. High-resolution 3D models were created to simulate variation in lighting, orientation, and background. These synthetic images were incorporated into convolutional neural network training datasets at varying synthetic-to-real ratios to assess their impact on classification accuracy. Models were evaluated using controlled pitfall-trap imagery to examine the influence of scene …


Mosquito Classification And Explainability From Image Data Via Deep Learning Techniques, Farhat Binte Azam Oct 2025

Mosquito Classification And Explainability From Image Data Via Deep Learning Techniques, Farhat Binte Azam

USF Tampa Graduate Theses and Dissertations

According to the World Health Organization (WHO), mosquitoes are the deadliest animals on Earth, responsible for more human deaths annually than any other species. Mosquito-borne illnesses continue to pose severe risks to global health. In 2015 alone, there were an estimated 214 million malaria cases worldwide. Similarly, a 2016 report from the Centers for Disease Control and Prevention (CDC) revealed that Puerto Rico’s Department of Health received over 62,500 suspected cases of Zika, with 29,345 confirmed positive cases. In 2019, Southeast Asia experienced its worst dengue outbreak in recorded history. Of the approximately 4,500 mosquito species distributed across 34 genera, …


Robustness Investigation, Detection, And Defense Of Deep Learning Models Against False Data Injection, Amirhossein Nazeri Aug 2025

Robustness Investigation, Detection, And Defense Of Deep Learning Models Against False Data Injection, Amirhossein Nazeri

All Dissertations

This dissertation addresses the critical challenge of adversarial robustness in deep learning systems, focusing on two fundamental domains: time-series prediction and object detection. As these AI systems become increasingly deployed in safety-critical applications from power grid management to autonomous vehicles their vulnerability to adversarial attacks poses significant risks to infrastructure and human safety.

The first contribution introduces a novel stealthy black-box False Data Injection (FDI) attack specifically designed for quasi-periodic time-series data. Unlike existing attacks that produce easily detectable anomalies, our method generates adversarial perturbations that preserve the underlying periodicity and statistical properties of the data, effectively bypassing traditional anomaly …


A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater Jun 2025

A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater

SMU Data Science Review

Paleography, the study of historical handwriting, is essential for preserving societal understanding of cultural, social, and legal frameworks from the past. Medieval manuscripts, often exhibiting refined craftsmanship, present unique challenges to modern readers due to differences in handwriting conventions and the absence of standardized punctuation and spaces. These texts hold valuable insights into the evolution of written communication, literacy, and language development. However, interpreting them requires specialized knowledge and technological solutions. Convolutional Neural Networks (CNNs) can be leveraged to classify scripts, an important step in Historical Document analysis. These models extract and analyze hierarchical features from images, addressing inconsistencies in …


Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim Jun 2025

Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim

Master's Theses

The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Defend: A 1m Dataset Foundation Model For Tobacco Analysis, Matthew J. Shepard May 2025

Defend: A 1m Dataset Foundation Model For Tobacco Analysis, Matthew J. Shepard

Electrical Engineering and Computer Science Undergraduate Honors Theses

The study of tobacco imagery and marketing is a complex challenge that involves extremely large datasets. It also demands a detailed analysis of the so- cial context and specific types of tobacco being marketed. Despite major recent advances in computer vision and foundation model technology, this still poses a substantial challenge. Through the DEFEND model, we aspire to address these obstacles by integrating features such as multimodal learning, hierarchical under- standing, and feature extraction to develop a foundation model designed to handle the unique challenges of tobacco image analysis. One of the core elements of DE- FEND is the Tobacco …


Multimodal Learning For Visual Perception And Robotic Action, Taisei Hanyu May 2025

Multimodal Learning For Visual Perception And Robotic Action, Taisei Hanyu

Electrical Engineering and Computer Science Undergraduate Honors Theses

Multimodal learning aims to weave information from images, language, depth, and other sensors into one coherent representation, much as people naturally combine sight, speech, and sound. Progress toward that goal is slowed by three gaps: vision encoders that cannot balance crisp object boundaries with global context, 3-D semantic maps that are computationally prohibitive for real-time, open-vocabulary queries, and vision-language-action pipelines that depend on large token pools with weak relational grounding.

We first introduce AerialFormer, a lightweight hybrid of convolutional and Transformer layers that captures long-range structure without sacrificing fine detail. On the large-scale iSAID benchmark it reaches 69.3% mean IoU, …


Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira Apr 2025

Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira

Doctoral Dissertations and Master's Theses

This dissertation proposes researching an approach to incorporate and align Software black-box testing methods into Machine Learning (ML) applications, specifically in the context of computer vision models. Typically, testing methods within Software Engineering (SE) encompass a range of test types that assess levels of a software system, such as Unit, Integration, Functional, and System testing [1]. The testing spectrum offers two perspectives on the system: black-box, where the system’s code is hidden, and white-box, where the system's code is exposed for testing. Software Quality pairs testing with requirements, in a many-to-one relationship, to ensure proper validation of the software system. …


The Asset Management Optimization Engine: An Ai And Machine Learning Model Approach To Pavement Asset Management, Matt Versdahl Mar 2025

The Asset Management Optimization Engine: An Ai And Machine Learning Model Approach To Pavement Asset Management, Matt Versdahl

USF Tampa Graduate Theses and Dissertations

While state Departments of Transportation (DOT) face major funding challenges, the need to find optimal ways to preserve and maintain pavement assets remains. Asset management employs a lowest cost lifecycle method to analyze asset costs and determine the best investment strategies to preserve it throughout its lifecycle. As new technology emerges, so do opportunities to leverage it. DOTs collect a significant amount of performance data on pavement and use it to decide how to keep it in a state of good repair. The literature in this area focuses on engineering techniques applied to treatment strategies. This dissertation research focuses on …


Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago Jan 2025

Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago

Dartmouth College Ph.D Dissertations

Natural language describes entities in the world, some real and some abstract. It is also common practice to complement human learning of natural language with visual cues. This is evident in the heavily graphical nature of children’s literature which underscores the importance of visual cues in language acquisition. Similarly, the notion of “visual learners” is well recognized, reflecting the understanding that visual signals such as illustrations, gestures, and depictions effectively supplement language. In machine learning, two primary paradigms have emerged for training systems involving natural language. The first paradigm encompasses setups where pre-training and downstream tasks are exclusively in natural …


Using Satellite Image Segmentation To Detect Trails, Jeremy Reynolds Jan 2025

Using Satellite Image Segmentation To Detect Trails, Jeremy Reynolds

Theses and Dissertations

This masters thesis proposes an innovative approach to satellite image segmentation by focusing on the detection and mapping of walking, hiking, and biking trails. The motivation behind this project comes from the underexplored area in segmentation techniques for trail identification and offers potential benefits for urban planning, environmental monitoring, and public health. The problem statement addresses the need for a model that can differentiate between various trail types and other natural or man-made elements. The project aims for efficiency and scalability in processing satellite imagery across different compute hardware. The work details several stages: researching existing segmentation techniques, specifically road …


Design And Analysis Of Facial Recognition Algorithms For Home Monitoring, Nathaniel F. Bernich Jan 2025

Design And Analysis Of Facial Recognition Algorithms For Home Monitoring, Nathaniel F. Bernich

Honors Theses and Capstones

Facial recognition "in the wild" has posed a challenge in the field of computer vision. Though facial recognition algorithms are generally proficient at recognizing faces up close, subjects at awkward angles and greater distances from the camera make monitoring areas with this software a practical challenge. At UNH's Cognitive Assistive Robotics Lab (CARL), overcoming the weak areas of face recognition is essential to the task of home monitoring. The CARL research team is implementing a suite of robotics and computer vision technologies to monitor patients with Alzheimer's dementia in their homes. This necessitates a reliable and effective facial recognition pipeline …


Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo Jan 2025

Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo

Dissertations, Master's Theses and Master's Reports

Medical Image Segmentation is a critical task in the field of medical imaging, playing a crucial role in diagnostics, treatment planning, and disease monitoring. The emergence of Deep Learning (DL) has ushered in a new era in Artificial Intelligence (AI), propelling remarkable advancements in key domains like language translation, object recognition, and recommendation systems. This evolution has been accompanied by continuous enhancements in computational efficiency and improvements in predictive accuracy. The introduction of sophisticated algorithms, such as convolutional neural networks (CNNs) and transformers, exemplifies these advancements. DL algorithms have demonstrated exceptional efficacy in medical image segmentation tasks, showcasing the potential …


Cropsync: Ai-Powered Sustainable Crop Management, Ziad Doughan, Ibrahim Mneimneh, Zouheir Nakouzi, Noor Al Khaib, Samer Damaj, Jamal Chaaban, Hamza Mrad, Sari Itani Dec 2024

Cropsync: Ai-Powered Sustainable Crop Management, Ziad Doughan, Ibrahim Mneimneh, Zouheir Nakouzi, Noor Al Khaib, Samer Damaj, Jamal Chaaban, Hamza Mrad, Sari Itani

BAU Journal - Science and Technology

CropSync is a smart agriculture system that uses AI and IoT technologies to enable sustain- able crop management and precision farming. The system aims to address the challenges faced by the agriculture sector, such as increasing food production to meet global population demands while minimizing environmental impact. CropSync integrates sensors, cameras, and cloud-based analytics to provide farmers with real-time insights and recommendations for optimizing crop cul- tivation. The system upholds engineering professional and ethical standards, considering broader social, environmental, and economic implications. From a social perspective, CropSync improves food security and enhances farmers’ livelihoods through increased productivity and efficient re- …


Towards Comprehensive And Interpretable Video Understanding, Khoa Vo Dec 2024

Towards Comprehensive And Interpretable Video Understanding, Khoa Vo

Graduate Theses and Dissertations

Video understanding is a critical domain in computer vision, focusing on analysis of sequential visual data to extract meaningful spatiotemporal information for tasks such as action recognition, video captioning, video retrieval, and temporal action localization, etc. Despite significant advancements with spatio-temporal convolutional neural networks and attention-based video models, current methods face limitations, including inadequate representation of main actors, lack of fine-grained modeling of relevant objects, and limited interpretability.
This thesis addresses these challenges by proposing novel approaches that enhance video understanding through modeling interactions among entities (actors and objects) and between entities and the environment, while improving interpretability in the …


Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi Nov 2024

Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi

LSU Master's Theses

Wearable exoskeletons offer significant potential in enhancing human mobility in industrial environments. However, their adaptability to dynamic, task-intensive settings presents challenges, especially in accurately predicting locomotion modes such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This research proposes a multimodal framework that integrates visual data and speech commands to improve locomotion mode prediction in unpredictable environments. Multimodal data was collected using smart glasses, capturing both the user’s perspective (field-of-view, FOV) and voice during locomotion tasks. State-of-the-art models—CLIP, ImageBind, and GPT-4o—process these visual and linguistic inputs to predict locomotion activities. The models were evaluated in zero-shot and fine-tuned …


Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa Oct 2024

Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa

Research Collection School Of Computing and Information Systems

Industry 4.0, the digitalization of manufacturing promises to lead to lowered cost, efficient processes and even discovery of new business models. However, many of the enterprises have huge investments in legacy machines which are not 'smart'. In this study, we thus designed a cost-efficient solution to retrofit a legacy conveyor belt-based cutlery washing machine with a commodity web camera. We then applied computer vision (using both traditional image processing and deep learning techniques) to infer the speed and utilization of the machine. We detailed the algorithms that we designed for computing both speed andutilization. With the existing operational constraints of …


The Impact Of Model Variations On The Robustness Of Deep Learning Models In Adversarial Settings, Firuz Juraev, Mohammed Abuhamad, Simon S. Woo, George K. Thiruvathukal, Tamer Abuhmed Aug 2024

The Impact Of Model Variations On The Robustness Of Deep Learning Models In Adversarial Settings, Firuz Juraev, Mohammed Abuhamad, Simon S. Woo, George K. Thiruvathukal, Tamer Abuhmed

Computer Science: Faculty Publications and Other Works

Rapid advancements of deep learning are accelerating adoption in a wide variety of applications, including safety-critical applications such as self-driving vehicles, drones, robots, and surveillance systems. These advancements include applying variations of sophisticated techniques that improve the performance of models. However, such models are not immune to adversarial manipulations, which can cause the system to misbehave and remain unnoticed by experts. The frequency of modifications to existing deep learning models necessitates thorough analysis to determine the impact on models’ robustness. In this work, we present an experimental evaluation of the effects of model modifications on deep learning model robustness using …


Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim Aug 2024

Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim

Research Collection School Of Computing and Information Systems

Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessment by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes unstructured data (i.e. an image frame with facial line segments) and structured data (i.e. features of facial expressions) to detect facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 21 facial palsy patients. Our experimental results show that among various data modalities (i.e. unstructured data - RGB images …


Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller Jul 2024

Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller

2024 Symposium

Vision loss presents significant challenges in daily life. Existing solutions for blind and visually impaired individuals are often limited in functionality, expensive, or complex to use. Vysion Software addresses this gap by developing a user-friendly, all-in-one AI companion app that provides features including text summarization, real-time audio descriptions, and AI-enhanced navigation. This project details the development plan, initial functionalities, and future vision for Vysion Software.