Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (21)
- Computer Engineering (13)
- Data Science (11)
- Software Engineering (9)
- Life Sciences (7)
-
- Medicine and Health Sciences (7)
- Numerical Analysis and Scientific Computing (6)
- Social and Behavioral Sciences (6)
- Other Computer Sciences (5)
- Databases and Information Systems (4)
- Graphics and Human Computer Interfaces (4)
- Mathematics (4)
- Robotics (4)
- Theory and Algorithms (4)
- Applied Mathematics (3)
- Electrical and Computer Engineering (3)
- Statistics and Probability (3)
- Algebraic Geometry (2)
- Bioinformatics (2)
- Biomedical Engineering and Bioengineering (2)
- Disability Studies (2)
- Earth Sciences (2)
- Geometry and Topology (2)
- Information Security (2)
- OS and Networks (2)
- Other Computer Engineering (2)
- Statistical Models (2)
- Institution
-
- California Polytechnic State University, San Luis Obispo (10)
- University of Arkansas, Fayetteville (6)
- Singapore Management University (5)
- San Jose State University (4)
- University of South Florida (4)
-
- University of Kentucky (3)
- Chapman University (2)
- City University of New York (CUNY) (2)
- Fordham University (2)
- Georgia Southern University (2)
- Northern Illinois University (2)
- Arkansas State University (1)
- Beirut Arab University (1)
- Bucknell University (1)
- Clemson University (1)
- Dartmouth College (1)
- Eastern Washington University (1)
- Embry-Riddle Aeronautical University (1)
- Indian Statistical Institute (1)
- Kennesaw State University (1)
- LSU New Orleans (1)
- Louisiana State University (1)
- Loyola University Chicago (1)
- MBZUAI (1)
- Michigan Technological University (1)
- Purdue University (1)
- Southern Adventist University (1)
- Southern Methodist University (1)
- The College of Wooster (1)
- University at Albany, State University of New York (1)
- Publication Year
- Publication
-
- Master's Theses (10)
- Master's Projects (4)
- Research Collection School Of Computing and Information Systems (4)
- USF Tampa Graduate Theses and Dissertations (4)
- Graduate Theses and Dissertations (3)
-
- Dissertations, Theses, and Capstone Projects (2)
- Electrical Engineering and Computer Science Undergraduate Honors Theses (2)
- Faculty Publications (2)
- Honors Theses (2)
- Theses and Dissertations--Computer Science (2)
- 2024 Symposium (1)
- All Dissertations (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- BAU Journal - Science and Technology (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Science: Faculty Publications and Other Works (1)
- Dartmouth College Ph.D Dissertations (1)
- Data Science Undergraduate Honors Theses (1)
- Dissertations and Theses (1)
- Dissertations and Theses Collection (Open Access) (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Doctoral Dissertations and Master's Theses (1)
- Electronic Theses & Dissertations (2024 - present) (1)
- Food Science Faculty Articles and Research (1)
- Geography ETDs (1)
- Graduate Research Theses & Dissertations (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors Capstones (1)
- Honors College Theses (1)
- Publication Type
Articles 1 - 30 of 75
Full-Text Articles in Artificial Intelligence and Robotics
Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega
Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega
Honors Theses
Running gait analysis plays a critical role in injury prevention and performance optimization, however, existing approaches often rely on specialized laboratory equipment or wearable sensors with limited interpretability. Recent advances in computer vision, particularly 2D human pose estimation, enable markerless motion analysis from standard video. However, progress remains constrained by the lack of publicly available datasets designed for running form analysis.
In this work, we introduce a preliminary dataset and benchmark for stride-level running gait analysis. The dataset consists of 73 treadmill running videos from 15 participants with varying experience levels, annotated with over 4,600 stride-level labels across multiple biomechanical …
Machine Learning For Handwritten Character Recognition, Hannah Freitag
Machine Learning For Handwritten Character Recognition, Hannah Freitag
Honors Capstones
Handwritten character recognition remains a challenging problem in machine learning due to the high variability of handwriting across individuals and the visual similarity between certain character classes. This project explores whether Singular Value Decomposition (SVD)-based dimensionality reduction can serve as an effective preprocessing step for a fully connected neural network trained on the EMNIST Balanced dataset, a 47-class benchmark of handwritten digits and letters. By projecting 784- dimensional pixel inputs onto the top 70 principal components, approximately 90% of the total variance is preserved while reducing input dimensionality by 91%. The resulting SVD-based model achieves approximately 94% test accuracy, outperforming …
Pediatric Inflammatory Bowel Disease Tissue Classification From Pathology Slide Images: Detecting Phenotypes Using Computer Vision, Chloe Martin-King, Ali Nael, Louis Ehwerhemuepha, Blake Calvo, Quinn Gates, Jamie Janchoi, Elisa Ornelas, Melissa Perez, Andrea Venderby, John Miklavcic, Peter Chang, Aaron Sassoon, Brian Rubio, Ghislaine Barrigan, Kenneth Grant
Pediatric Inflammatory Bowel Disease Tissue Classification From Pathology Slide Images: Detecting Phenotypes Using Computer Vision, Chloe Martin-King, Ali Nael, Louis Ehwerhemuepha, Blake Calvo, Quinn Gates, Jamie Janchoi, Elisa Ornelas, Melissa Perez, Andrea Venderby, John Miklavcic, Peter Chang, Aaron Sassoon, Brian Rubio, Ghislaine Barrigan, Kenneth Grant
Food Science Faculty Articles and Research
Background and Aims
With the advent of computer vision algorithms, we hypothesize that histopathology images from endoscopic biopsies may be utilized for automated classification of histologic phenotypes, thus guiding Crohn’s disease and ulcerative colitis diagnosis and treatment. The aim of our study is to assess whether artificial intelligence can be used to improve pediatric inflammatory bowel disease outcomes by aiding pathologists with accurate detection of abnormal tissue sections.Methods
Three two-dimensional (2D) convolutional neural networks with multiple instance learning were developed to classify histopathology tissue sections as normal vs abnormal and as containing active inflammation and/or chronic changes/architectural distortion.Results …
Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci
Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci
Theses and Dissertations--Electrical and Computer Engineering
Fine-grained Temporal Action Segmentation (TAS) has become a cornerstone of video understanding, offering dense frame-level predictions essential for clinical assessment, surgical skill evaluation, and human-computer interaction. While TAS methods have delivered strong results on coarse-grained benchmarks, two fundamental challenges persist: (1) global attention mechanisms dilute boundary information critical for subsecond precision, a phenomenon we term the temporal granularity bottleneck, and (2) dense frame-level annotation remains prohibitively expensive, with most datasets requiring exhaustive labeling of lengthy untrimmed videos. These challenges are particularly pronounced in medical domains, where sub-second primitives define clinical outcomes while expert annotation remains scarce. In this dissertation, we …
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Theses and Dissertations (Comprehensive)
Deploying deep learning models for medical image analysis on mobile devices requires a balance between inference latency, memory footprint, and delineating anatomical boundaries with high accuracy. While Convolutional Neural Networks (CNNs) and mobile Vision Transformers (ViTs) offer efficiency, they often struggle to model the irregular, non-local geometric structures inherent in biological tissues without incurring prohibitive computational costs. In this thesis, we introduce GeoViG (Geometric Vision Graph), an architecture that bridges the gap between efficient grid-based processing and explicit Geometric Deep Learning. GeoViG introduces a novel transition from high-resolution pixel grids to low-resolution dynamic graphs via a SpreadEdgePool operator, a geometry-aware …
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Undergraduate Research Symposium
Various computational models of first impressions have been developed to uncover the mechanisms driving these judgments. However, the implicit notion of a singular ``human'' often overlooks meaningful individual differences in beliefs, attitudes, and associations, as well as culturally grounded group-level constructs. In this paper, we extend Cultural Consensus Theory (CCT) to estimate culturally shared beliefs about faces by incorporating latent constructs structured around interpretable facial features extracted via computer vision algorithms. We apply our model to a large-scale dataset of people’s first impressions of faces. Our approach reveals a robust mapping between facial features and culturally constructed impressions, allowing us …
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
MS in Computer Science Project Reports
We present a grading system that accelerates evaluation of open-ended student work across scanned and digital workflows. The system crops answer regions from PDFs, assigns submissions via OCR on identity regions only, and groups answers by visual semantics using a vision LLM. Instructors review and edit groups, apply rubric items once per group, and export grades from an on-screen table. The solution integrates Ghostscript rasterization, PdfPig page orchestration, SkiaSharp region extraction, Tesseract identity OCR, and GPT-4o Vision for grouping. We detail the architecture, token-budgeted batching strategy, and persistence design, then describe testing results for grouping quality, time-on-task, and usability. The …
Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui
Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui
Graduate Theses and Dissertations
Negation is a fundamental linguistic concept used by humans to convey information that they do not desire. Despite this, minimal research has focused on negation within text-guided image editing. This lack of research means that vision-language models (VLMs) for image editing may struggle to understand negation, implying that they struggle to provide accurate results. One barrier to achieving human-level intelligence is the lack of a standard collection by which research into negation can be evaluated. This thesis presents the first large-scale dataset, Negative Instruction (NeIn), for studying negation within instruction-based image editing. Our dataset comprises 366,957 quintuplets, i.e., source image, …
Correcting Class Imbalance Through Synthetic Training Data And 3d Modeling For Carabid Pitfall Trap Sampling, Blair Mirka
Correcting Class Imbalance Through Synthetic Training Data And 3d Modeling For Carabid Pitfall Trap Sampling, Blair Mirka
Geography ETDs
Crowdsourced biodiversity data provide an accessible foundation for large-scale ecological monitoring, but class imbalance limits automated species identification, particularly for rare taxa. This research explores the use of synthetic training data generated from 3D models of carabid beetle museum specimens to improve detection and classification performance for underrepresented species in crowdsourced datasets. High-resolution 3D models were created to simulate variation in lighting, orientation, and background. These synthetic images were incorporated into convolutional neural network training datasets at varying synthetic-to-real ratios to assess their impact on classification accuracy. Models were evaluated using controlled pitfall-trap imagery to examine the influence of scene …
Mosquito Classification And Explainability From Image Data Via Deep Learning Techniques, Farhat Binte Azam
Mosquito Classification And Explainability From Image Data Via Deep Learning Techniques, Farhat Binte Azam
USF Tampa Graduate Theses and Dissertations
According to the World Health Organization (WHO), mosquitoes are the deadliest animals on Earth, responsible for more human deaths annually than any other species. Mosquito-borne illnesses continue to pose severe risks to global health. In 2015 alone, there were an estimated 214 million malaria cases worldwide. Similarly, a 2016 report from the Centers for Disease Control and Prevention (CDC) revealed that Puerto Rico’s Department of Health received over 62,500 suspected cases of Zika, with 29,345 confirmed positive cases. In 2019, Southeast Asia experienced its worst dengue outbreak in recorded history. Of the approximately 4,500 mosquito species distributed across 34 genera, …
Robustness Investigation, Detection, And Defense Of Deep Learning Models Against False Data Injection, Amirhossein Nazeri
Robustness Investigation, Detection, And Defense Of Deep Learning Models Against False Data Injection, Amirhossein Nazeri
All Dissertations
This dissertation addresses the critical challenge of adversarial robustness in deep learning systems, focusing on two fundamental domains: time-series prediction and object detection. As these AI systems become increasingly deployed in safety-critical applications from power grid management to autonomous vehicles their vulnerability to adversarial attacks poses significant risks to infrastructure and human safety.
The first contribution introduces a novel stealthy black-box False Data Injection (FDI) attack specifically designed for quasi-periodic time-series data. Unlike existing attacks that produce easily detectable anomalies, our method generates adversarial perturbations that preserve the underlying periodicity and statistical properties of the data, effectively bypassing traditional anomaly …
A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater
A Modern Approach To Classifying Medieval Latin Scripts, Robert L. Lane Jr, Rafia Mirza, Robert Slater
SMU Data Science Review
Paleography, the study of historical handwriting, is essential for preserving societal understanding of cultural, social, and legal frameworks from the past. Medieval manuscripts, often exhibiting refined craftsmanship, present unique challenges to modern readers due to differences in handwriting conventions and the absence of standardized punctuation and spaces. These texts hold valuable insights into the evolution of written communication, literacy, and language development. However, interpreting them requires specialized knowledge and technological solutions. Convolutional Neural Networks (CNNs) can be leveraged to classify scripts, an important step in Historical Document analysis. These models extract and analyze hierarchical features from images, addressing inconsistencies in …
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Master's Theses
The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Defend: A 1m Dataset Foundation Model For Tobacco Analysis, Matthew J. Shepard
Defend: A 1m Dataset Foundation Model For Tobacco Analysis, Matthew J. Shepard
Electrical Engineering and Computer Science Undergraduate Honors Theses
The study of tobacco imagery and marketing is a complex challenge that involves extremely large datasets. It also demands a detailed analysis of the so- cial context and specific types of tobacco being marketed. Despite major recent advances in computer vision and foundation model technology, this still poses a substantial challenge. Through the DEFEND model, we aspire to address these obstacles by integrating features such as multimodal learning, hierarchical under- standing, and feature extraction to develop a foundation model designed to handle the unique challenges of tobacco image analysis. One of the core elements of DE- FEND is the Tobacco …
Multimodal Learning For Visual Perception And Robotic Action, Taisei Hanyu
Multimodal Learning For Visual Perception And Robotic Action, Taisei Hanyu
Electrical Engineering and Computer Science Undergraduate Honors Theses
Multimodal learning aims to weave information from images, language, depth, and other sensors into one coherent representation, much as people naturally combine sight, speech, and sound. Progress toward that goal is slowed by three gaps: vision encoders that cannot balance crisp object boundaries with global context, 3-D semantic maps that are computationally prohibitive for real-time, open-vocabulary queries, and vision-language-action pipelines that depend on large token pools with weak relational grounding.
We first introduce AerialFormer, a lightweight hybrid of convolutional and Transformer layers that captures long-range structure without sacrificing fine detail. On the large-scale iSAID benchmark it reaches 69.3% mean IoU, …
Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira
Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira
Doctoral Dissertations and Master's Theses
This dissertation proposes researching an approach to incorporate and align Software black-box testing methods into Machine Learning (ML) applications, specifically in the context of computer vision models. Typically, testing methods within Software Engineering (SE) encompass a range of test types that assess levels of a software system, such as Unit, Integration, Functional, and System testing [1]. The testing spectrum offers two perspectives on the system: black-box, where the system’s code is hidden, and white-box, where the system's code is exposed for testing. Software Quality pairs testing with requirements, in a many-to-one relationship, to ensure proper validation of the software system. …
The Asset Management Optimization Engine: An Ai And Machine Learning Model Approach To Pavement Asset Management, Matt Versdahl
The Asset Management Optimization Engine: An Ai And Machine Learning Model Approach To Pavement Asset Management, Matt Versdahl
USF Tampa Graduate Theses and Dissertations
While state Departments of Transportation (DOT) face major funding challenges, the need to find optimal ways to preserve and maintain pavement assets remains. Asset management employs a lowest cost lifecycle method to analyze asset costs and determine the best investment strategies to preserve it throughout its lifecycle. As new technology emerges, so do opportunities to leverage it. DOTs collect a significant amount of performance data on pavement and use it to decide how to keep it in a state of good repair. The literature in this area focuses on engineering techniques applied to treatment strategies. This dissertation research focuses on …
Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago
Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago
Dartmouth College Ph.D Dissertations
Natural language describes entities in the world, some real and some abstract. It is also common practice to complement human learning of natural language with visual cues. This is evident in the heavily graphical nature of children’s literature which underscores the importance of visual cues in language acquisition. Similarly, the notion of “visual learners” is well recognized, reflecting the understanding that visual signals such as illustrations, gestures, and depictions effectively supplement language. In machine learning, two primary paradigms have emerged for training systems involving natural language. The first paradigm encompasses setups where pre-training and downstream tasks are exclusively in natural …
Using Satellite Image Segmentation To Detect Trails, Jeremy Reynolds
Using Satellite Image Segmentation To Detect Trails, Jeremy Reynolds
Theses and Dissertations
This masters thesis proposes an innovative approach to satellite image segmentation by focusing on the detection and mapping of walking, hiking, and biking trails. The motivation behind this project comes from the underexplored area in segmentation techniques for trail identification and offers potential benefits for urban planning, environmental monitoring, and public health. The problem statement addresses the need for a model that can differentiate between various trail types and other natural or man-made elements. The project aims for efficiency and scalability in processing satellite imagery across different compute hardware. The work details several stages: researching existing segmentation techniques, specifically road …
Design And Analysis Of Facial Recognition Algorithms For Home Monitoring, Nathaniel F. Bernich
Design And Analysis Of Facial Recognition Algorithms For Home Monitoring, Nathaniel F. Bernich
Honors Theses and Capstones
Facial recognition "in the wild" has posed a challenge in the field of computer vision. Though facial recognition algorithms are generally proficient at recognizing faces up close, subjects at awkward angles and greater distances from the camera make monitoring areas with this software a practical challenge. At UNH's Cognitive Assistive Robotics Lab (CARL), overcoming the weak areas of face recognition is essential to the task of home monitoring. The CARL research team is implementing a suite of robotics and computer vision technologies to monitor patients with Alzheimer's dementia in their homes. This necessitates a reliable and effective facial recognition pipeline …
Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo
Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo
Dissertations, Master's Theses and Master's Reports
Medical Image Segmentation is a critical task in the field of medical imaging, playing a crucial role in diagnostics, treatment planning, and disease monitoring. The emergence of Deep Learning (DL) has ushered in a new era in Artificial Intelligence (AI), propelling remarkable advancements in key domains like language translation, object recognition, and recommendation systems. This evolution has been accompanied by continuous enhancements in computational efficiency and improvements in predictive accuracy. The introduction of sophisticated algorithms, such as convolutional neural networks (CNNs) and transformers, exemplifies these advancements. DL algorithms have demonstrated exceptional efficacy in medical image segmentation tasks, showcasing the potential …
Cropsync: Ai-Powered Sustainable Crop Management, Ziad Doughan, Ibrahim Mneimneh, Zouheir Nakouzi, Noor Al Khaib, Samer Damaj, Jamal Chaaban, Hamza Mrad, Sari Itani
Cropsync: Ai-Powered Sustainable Crop Management, Ziad Doughan, Ibrahim Mneimneh, Zouheir Nakouzi, Noor Al Khaib, Samer Damaj, Jamal Chaaban, Hamza Mrad, Sari Itani
BAU Journal - Science and Technology
CropSync is a smart agriculture system that uses AI and IoT technologies to enable sustain- able crop management and precision farming. The system aims to address the challenges faced by the agriculture sector, such as increasing food production to meet global population demands while minimizing environmental impact. CropSync integrates sensors, cameras, and cloud-based analytics to provide farmers with real-time insights and recommendations for optimizing crop cul- tivation. The system upholds engineering professional and ethical standards, considering broader social, environmental, and economic implications. From a social perspective, CropSync improves food security and enhances farmers’ livelihoods through increased productivity and efficient re- …
Towards Comprehensive And Interpretable Video Understanding, Khoa Vo
Towards Comprehensive And Interpretable Video Understanding, Khoa Vo
Graduate Theses and Dissertations
Video understanding is a critical domain in computer vision, focusing on analysis of sequential visual data to extract meaningful spatiotemporal information for tasks such as action recognition, video captioning, video retrieval, and temporal action localization, etc. Despite significant advancements with spatio-temporal convolutional neural networks and attention-based video models, current methods face limitations, including inadequate representation of main actors, lack of fine-grained modeling of relevant objects, and limited interpretability.
This thesis addresses these challenges by proposing novel approaches that enhance video understanding through modeling interactions among entities (actors and objects) and between entities and the environment, while improving interpretability in the …
Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi
Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi
LSU Master's Theses
Wearable exoskeletons offer significant potential in enhancing human mobility in industrial environments. However, their adaptability to dynamic, task-intensive settings presents challenges, especially in accurately predicting locomotion modes such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This research proposes a multimodal framework that integrates visual data and speech commands to improve locomotion mode prediction in unpredictable environments. Multimodal data was collected using smart glasses, capturing both the user’s perspective (field-of-view, FOV) and voice during locomotion tasks. State-of-the-art models—CLIP, ImageBind, and GPT-4o—process these visual and linguistic inputs to predict locomotion activities. The models were evaluated in zero-shot and fine-tuned …
Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa
Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa
Research Collection School Of Computing and Information Systems
Industry 4.0, the digitalization of manufacturing promises to lead to lowered cost, efficient processes and even discovery of new business models. However, many of the enterprises have huge investments in legacy machines which are not 'smart'. In this study, we thus designed a cost-efficient solution to retrofit a legacy conveyor belt-based cutlery washing machine with a commodity web camera. We then applied computer vision (using both traditional image processing and deep learning techniques) to infer the speed and utilization of the machine. We detailed the algorithms that we designed for computing both speed andutilization. With the existing operational constraints of …
The Impact Of Model Variations On The Robustness Of Deep Learning Models In Adversarial Settings, Firuz Juraev, Mohammed Abuhamad, Simon S. Woo, George K. Thiruvathukal, Tamer Abuhmed
The Impact Of Model Variations On The Robustness Of Deep Learning Models In Adversarial Settings, Firuz Juraev, Mohammed Abuhamad, Simon S. Woo, George K. Thiruvathukal, Tamer Abuhmed
Computer Science: Faculty Publications and Other Works
Rapid advancements of deep learning are accelerating adoption in a wide variety of applications, including safety-critical applications such as self-driving vehicles, drones, robots, and surveillance systems. These advancements include applying variations of sophisticated techniques that improve the performance of models. However, such models are not immune to adversarial manipulations, which can cause the system to misbehave and remain unnoticed by experts. The frequency of modifications to existing deep learning models necessitates thorough analysis to determine the impact on models’ robustness. In this work, we present an experimental evaluation of the effects of model modifications on deep learning model robustness using …
Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim
Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim
Research Collection School Of Computing and Information Systems
Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessment by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes unstructured data (i.e. an image frame with facial line segments) and structured data (i.e. features of facial expressions) to detect facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 21 facial palsy patients. Our experimental results show that among various data modalities (i.e. unstructured data - RGB images …
Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller
Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller
2024 Symposium
Vision loss presents significant challenges in daily life. Existing solutions for blind and visually impaired individuals are often limited in functionality, expensive, or complex to use. Vysion Software addresses this gap by developing a user-friendly, all-in-one AI companion app that provides features including text summarization, real-time audio descriptions, and AI-enhanced navigation. This project details the development plan, initial functionalities, and future vision for Vysion Software.
Context In Computer Vision: A Taxonomy, Multi-Stage Integration, And A General Framework, Xuan Wang
Context In Computer Vision: A Taxonomy, Multi-Stage Integration, And A General Framework, Xuan Wang
Dissertations, Theses, and Capstone Projects
Contextual information has been widely used in many computer vision tasks, such as object detection, video action detection, image classification, etc. Recognizing a single object or action out of context could be sometimes very challenging, and context information may help improve the understanding of a scene or an event greatly. However, existing approaches design specific contextual information mechanisms for different detection tasks.
In this research, we first present a comprehensive survey of context understanding in computer vision, with a taxonomy to describe context in different types and levels. Then we proposed MultiCLU, a new multi-stage context learning and utilization framework, …