Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Deep Learning

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 432

Full-Text Articles in Computer Sciences

Viability Assessment Of Bovine Embryos: A Public Dataset And Deep Learning Baselines, Erfan Khayyati Aug 2026

Viability Assessment Of Bovine Embryos: A Public Dataset And Deep Learning Baselines, Erfan Khayyati

All Graduate Theses and Dissertations, Fall 2023 to Present

Improving the success rates of cattle breeding is essential for sustainable agriculture, global food security, and high-quality livestock production. Currently, determining whether a lab-grown bovine embryo is healthy enough for a successful pregnancy requires highly trained experts to manually evaluate days of continuous time-lapse video footage. This process is not only incredibly time-consuming but also highly subjective; human reviewers often suffer from visual fatigue when tracking subtle, microscopic cellular changes over a seven-day period, leading to significant disagreement among even top experts on an embryo’s true potential. Furthermore, assessing bovine embryos is notoriously difficult due to their dark, lipid-dense cellular …


Edge Ai And Iot-Based Intelligent Smart Irrigation System For Sustainable Agriculture, Afra Alshamsi Jun 2026

Edge Ai And Iot-Based Intelligent Smart Irrigation System For Sustainable Agriculture, Afra Alshamsi

Thesis/ Dissertation Defenses

Traditional agriculture faces several challenges, including excessive water consumption, inefficient irrigation practices, and the increasing demand for sustainable food production. The rapid advancement of the Internet of Things (IoT) and Artificial Intelligence (AI) provides opportunities to develop intelligent agricultural systems that improve resource management and support precision farming. This thesis presents an Edge AI and IoT-based smart irrigation system designed for real-time environmental monitoring, automated irrigation control, and intelligent irrigation prediction to support the optimal growth of arugula plants. The proposed system continuously monitors key cultivation parameters, including soil moisture, temperature, humidity, light intensity, pH, nutrient levels (NPK), and water …


Streamlined Biomedical Image Processing Pipelines, Jiehyun Kim May 2026

Streamlined Biomedical Image Processing Pipelines, Jiehyun Kim

Graduate Doctoral Dissertations

This dissertation focuses on advancing carotid artery analysis through a series of visualizations and deep learning tools for calcified plaque assessment and related biomedical imaging tasks. Accurate plaque evaluation is essential, but current workflows depend on slow, clinician-dependent manual review. To address these limitations, this work introduces the CACTAS framework, a set of tools and methods that enable fast and reliable plaque segmentation for clinicians.

The first study, the CACTAS-Tool, provides a web-based labeling tool that enables clinicians to label plaque directly in three dimensions through a streamlined one-click interface. This tool significantly reduces the effort required to generate high-quality …


Mri-Based Deep Learning Radiomics Model For Automated Classification Of Disc Degeneration In The Lumbar Spine, Shiv Patil, Om Gandhi, Mert Karabacak, Matthew Carr, Konstantinos Margetis May 2026

Mri-Based Deep Learning Radiomics Model For Automated Classification Of Disc Degeneration In The Lumbar Spine, Shiv Patil, Om Gandhi, Mert Karabacak, Matthew Carr, Konstantinos Margetis

Student Papers, Posters & Projects

Disc degeneration in the lumbar spine is a major cause of low back pain (LBP). The accurate grading of disc degeneration on magnetic resonance imaging (MRI) is critical for clinical management and patient selection for spine surgery. This study aims to develop and evaluate machine learning (ML) models that combine features from deep learning (DL) and radiomics for the automated prediction of Pfirrmann grade (PG), a measure of disc degeneration, using multi-parametric lumbar spine MRI. Sagittal T1, T2, and T2 SPACE MRIs of 218 patients with LBP were acquired from the SPIDER dataset. For each intervertebral disc and available sequence, …


Visual Interpretability Of Multimodal Tissue Perfusion Classification Using Grad-Cam And Saliency Maps, Metehan Zorluoglu May 2026

Visual Interpretability Of Multimodal Tissue Perfusion Classification Using Grad-Cam And Saliency Maps, Metehan Zorluoglu

UNLV Theses, Dissertations, Professional Papers, and Capstones

Accurate identification of the tissue perfusion phase from hand images can aid doctors in decision-making with non-invasive techniques. The present study proposes a multimodal deep learning model for classifying the tissue perfusion phase using infrared, thermal, and visible spectrum images of the human hand. The proposed model consists of various preprocessing techniques such as manipulation, homography alignments, and masking. The significant contribution of this thesis is the interpretability analysis of deep learning models, achieved through the analysis of saliency maps and the Gradient-weighted Class Activation Mapping (Grad-CAM) methods. The purpose of this method is to find out how the convolutional …


Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega May 2026

Stridevision: Automated Detection Of Running Form Deviations From 2d Pose Estimation And Machine Learning, Paulina Eguibar Ortega

Honors Theses

Running gait analysis plays a critical role in injury prevention and performance optimization, however, existing approaches often rely on specialized laboratory equipment or wearable sensors with limited interpretability. Recent advances in computer vision, particularly 2D human pose estimation, enable markerless motion analysis from standard video. However, progress remains constrained by the lack of publicly available datasets designed for running form analysis.

In this work, we introduce a preliminary dataset and benchmark for stride-level running gait analysis. The dataset consists of 73 treadmill running videos from 15 participants with varying experience levels, annotated with over 4,600 stride-level labels across multiple biomechanical …


Leveraging Convolutional Neural Networks For Through-The-Wall Radar Imaging: Challenges, Impacts, And Future Directions, Tumaini Edgar, Abdulla F. Ally, Abdi T. Abdalla Apr 2026

Leveraging Convolutional Neural Networks For Through-The-Wall Radar Imaging: Challenges, Impacts, And Future Directions, Tumaini Edgar, Abdulla F. Ally, Abdi T. Abdalla

Tanzania Journal of Engineering and Technology (TJET)

Through-the-wall radar imaging (TWRI) is an essential technology for military and rescue applications; however, its performance in detecting and visualizing high-quality images of targets behind walls is significantly degraded by multipath reflections and signal attenuation. This paper reviews the current state of TWRI and its challenges, and explores the transformative potential of deep learning, particularly convolutional neural networks (CNNs), in addressing these challenges. Peer-reviewed articles published from 2018 to 2024 were analysed to examine CNN applications in addressing TWRI challenges. The analysis reveals that using CNNs, TWRI systems can be more effective by filtering wall distortions, reducing noise, lowering computational …


Intelligent Deep Learning-Based Sign Language Translation System, Nada Rasem Shahin Apr 2026

Intelligent Deep Learning-Based Sign Language Translation System, Nada Rasem Shahin

Dissertations

The Deaf and Hard of Hearing (DHH) community uses sign language as a primary means of communication. However, the shortage of sign language interpreters and the existence of hundreds of sign languages limit accessibility and inclusion. Sign Language Machine Translation (SLMT) systems present a promising solution for bridging the communication gap between the DHH and the hearing individuals, supporting inclusive societies. In smart cities, such systems play an essential role in improving the quality of life on a community level. In particular, as the population’s well-being is critical, developing intelligent assistive technologies, such as SLMT systems, is necessary to provide …


Development Of Deep Fused Neural Architecture For Ancient Tamil Palm-Leaf Manuscript Recognition, Hariharan P Mr Mar 2026

Development Of Deep Fused Neural Architecture For Ancient Tamil Palm-Leaf Manuscript Recognition, Hariharan P Mr

Theses and Dissertations

Digitizing Tamil palm-leaf manuscripts is important for education, communication, and the preservation of cultural heritage. The complex structure of the Tamil script, the wide range of handwriting styles, and the degradation seen in ancient Tamil palm-leaf manuscripts make these texts very difficult to read and understand. Digital Image Processing (DIP), document analysis techniques, and traditional Optical Character Recognition (OCR) are unable to handle noise, background interference, faded ink, and limited labelled data, motivating the need for robust, effective Deep Learning (DL)- based solutions.

As a prerequisite to understanding and designing effective recognition systems for ancient manuscripts, this thesis first examines …


Developing Deep Neural Network Based Brain Computational Models From Psychophysics Data Of Some Simple Perceptual Phenomena: Visual As Well As Auditory, Chandran Keerthi S Feb 2026

Developing Deep Neural Network Based Brain Computational Models From Psychophysics Data Of Some Simple Perceptual Phenomena: Visual As Well As Auditory, Chandran Keerthi S

Doctoral Theses

The thesis, consisting of nine chapters, explores the methodology of building testable brain computational models using Deep Neural Networks (DNN), which are trained by psychophysics data. Psychophysics is the quantitative study of perception of physical stimuli. In psychophysics experiments, one or more parameters associated with the stimuli are changed, and the human subject’s responses to the stimuli are recorded. This thesis encompasses both experimental psychophysics works, as well as computational models of the phenomena involved. The contributory chapters of the thesis start in Chapter 2 with the perspective building of the novel methodology followed throughout this research involving psychophysics on …


Artificial Intelligence Methods For Circadian Rhythm Recovery And Disease Identification From High-Dimensional Omics Data, Aram Ansary Ogholbake Jan 2026

Artificial Intelligence Methods For Circadian Rhythm Recovery And Disease Identification From High-Dimensional Omics Data, Aram Ansary Ogholbake

Theses and Dissertations--Computer Science

High-throughput transcriptomic and proteomic technologies have enabled opportunities for studying biological processes and disease mechanisms. However, extracting meaningful biological information from these high-dimensional datasets remains challenging due to limited sample sizes, biological heterogeneity, measurement noise, and the absence of biological annotations. In particular, many molecular datasets lack temporal information required for circadian analysis, making the study of circadian rhythms difficult. Moreover, disease diagnosis and biomarker discovery from transcriptomic data often rely on complex machine learning models whose predictions are difficult to interpret and may not generalize well across independent datasets and experimental platforms. These limitations motivate the development of artificial …


On-Device Artificial Intelligence Solutions With Applications To Smart Environments, Fabrizio De Vita, Dario Bruneo, Sajal K. Das Jan 2026

On-Device Artificial Intelligence Solutions With Applications To Smart Environments, Fabrizio De Vita, Dario Bruneo, Sajal K. Das

Computer Science Faculty Research & Creative Works

Recent advances in Artificial Intelligence (AI) and the increasing availability of computational power have accelerated the diffusion of Intelligent Cyber-Physical Systems (ICPSs), enabling smart applications with reasoning capabilities. However, the limited resources of embedded and Edge devices significantly constrain the complexity of deep learning models that can be effectively deployed. Traditional approaches rely on cloud-based training and edge-only inference, a paradigm that becomes inadequate when low latency, privacy, security, and high customization are required. In this context, On-device AI is emerging as a new paradigm in which both training and inference are performed directly on the device, avoiding data transfer …


Robust Deep Learning One-Class Classification, Shahd Alnofaie Jan 2026

Robust Deep Learning One-Class Classification, Shahd Alnofaie

Graduate Studies Theses and Dissertations 2026

One-Class Classification (OCC) focuses on learning the characteristics of normal data and identifying observations that deviate from this learned pattern as anomalies. It is commonly used in applications such as medical diagnosis, cybersecurity, industrial monitoring, and fraud detection, where abnormal examples are often rare or unavailable during training. Classical approaches such as SVDD and LS-SVDD describe normal data using a hypersphere. While effective in some settings, these methods rely on shallow representations and can be sensitive to noise and contaminated observations. To address these limitations, this dissertation introduces a Deep LS-SVDD framework that combines hypersphere-based data description with deep neural …


Effective Deep Learning Architectures For Structured Data Analysis And Generation, Md Atik Ahamed Jan 2026

Effective Deep Learning Architectures For Structured Data Analysis And Generation, Md Atik Ahamed

Theses and Dissertations--Computer Science

The effective utilization of structured data is fundamental to modern machine learning, yet it presents distinct challenges in both predictive analysis and generative modeling. Traditional deep learning architectures, particularly Transformers, often suffer from quadratic computational complexity when processing long sequences. This dissertation addresses these limitations by introducing novel architectures based on State-Space Models (SSMs) and Diffusion Models. In the area of predictive analysis, we focus on overcoming the computational bottlenecks of attention mechanisms for tabular and time-series data. First, we introduce MambaTab, a selective state-space architecture designed for efficient tabular classification. By leveraging the linear complexity of SSMs, MambaTab significantly …


Deep Learning Approaches For Voltammetric Analysis Of Coffee, Ryan Koes Jan 2026

Deep Learning Approaches For Voltammetric Analysis Of Coffee, Ryan Koes

Honors Theses

This thesis investigates deep learning approaches for voltammetric analysis of brewed coffee using a low-cost electrochemical system and screen-printed electrodes (SPEs). Traditional analytical methods, such as high-performance liquid chromatography (HPLC) and gas chromatography-mass spectrometry (GC-MS), provide precise quantification of key compounds but require expensive instrumentation and specialized expertise, limiting accessibility. While SPEs offer a more accessible alternative, they yielded poor results with traditional processing; however, when combined with a neural network, the system proved more effective. In experiments with 132 coffee samples, mean errors for caffeine, CGA, and TDS predictions were 52.98 ppm, 70.48 ppm, and 0.08%, respectively. These findings …


Multimodal Emotion Detection System, Shubhankar Sameer Munshi Jan 2026

Multimodal Emotion Detection System, Shubhankar Sameer Munshi

Master's Projects

Trying to understand emotion from speech is a problem that is present in human computer interaction. Nevertheless, there are still some shortcomings in current SER methods. Text-based systems may miss vital vocal cues, such as sarcasm, tone changes, and delivery. On the other hand, purely audio-based systems are prone to noise and unstable acoustic features. The combination of linguistic and acoustic features in multimodal approaches partially solves this problem, but many existing approaches use inflexible multimodal fusion techniques that cannot adjust their behaviors according to the quality of input signals. In this work, we propose a multimodal approach based on …


A Unified Framework For Evaluating Training Efficiency In Deep (Bayesian) Neural Networks: Metrics, Overtraining, Stopping Criteria, And Grokking Computer Science, Eduardo Cueto Mendoza Jan 2026

A Unified Framework For Evaluating Training Efficiency In Deep (Bayesian) Neural Networks: Metrics, Overtraining, Stopping Criteria, And Grokking Computer Science, Eduardo Cueto Mendoza

Doctoral

Measuring training efficiency for artificial neural networks is an open research problem, current literature reports several attempts to define measures or create reporting frameworks. Current methods lack generality as they require measurements of the hardware or software thus, comparing efficiency between different systems can be difficult. Similarly, current metrics or frameworks generally do not propose the use of the metrics to directly improve training efficiency. This thesis presents three main contributions: (1) a novel framework that quantifies the training efficiency of a neural architecture on a learning task as the average ratio of model accuracy to total energy consumption during …


Real-Time Isolated Asl Recognition: Evaluating Spatial-Temporal Networks And Multimodal Llms, Raga Mouni Batchu Jan 2026

Real-Time Isolated Asl Recognition: Evaluating Spatial-Temporal Networks And Multimodal Llms, Raga Mouni Batchu

West Chester University Graduate Theses, Dissertations, and Final Projects

This thesis investigates the deployment of high-accuracy Isolated ASL Recognition (ISLR) in resource-constrained edge environments. We train a lightweight Spatio-Temporal Attention Network (SSTAN,∼2.7 M parameters,∼10 MB) on the WLASL-100 benchmark, achieving 75.25% Top-1 and 88.24% Top-5 accuracy with 139 ms CPU-only inference. A systematic comparison against frontier multimodal LLMs (Gemini 3 Flash, Gemini 3.1 Pro, Qwen 3 VL) shows SSTAN outperforms the best LLM baseline by∼1.85×in accuracy while being 22–230×faster and up to 40×cheaper annually. The LLMs’ core limitation is a lack of fine-grained temporal perception; they impose English-language semantic priors rather than learning the articulatory distinctions that define ASL …


Modeling And Mitigating Atmospheric Degradation In Computer Vision With Application In Renewable Energy Prediction, Sumit Laha Jan 2026

Modeling And Mitigating Atmospheric Degradation In Computer Vision With Application In Renewable Energy Prediction, Sumit Laha

Graduate Studies Theses and Dissertations 2026

Weather-induced variability poses significant challenges to the reliability and performance of modern computational systems, particularly those relying on visual perception and environmental prediction. This dissertation focuses on enhancing computer vision and machine learning based predictive models that operate under varying atmospheric conditions. Two representative weather-impacted applications are investigated: image dehazing and solar photovoltaic (PV) power output forecasting. Image dehazing focuses on the restoration of clear, unobstructed visuals from hazy or foggy images, a task that is vital for various applications. On the other hand, photovoltaic (PV) power forecasting aims to predict future solar energy generation based on historical sky images …


Adaptive Deep Learning In Physical Layer Applications, Ali Owfi Dec 2025

Adaptive Deep Learning In Physical Layer Applications, Ali Owfi

All Dissertations

Traditionally, signal processing models in communication systems have been designed based on solid foundations in statistics and information theory, often assuming linearity and optimizing for simplified models. However, real-world communication systems exhibit numerous imperfections and non-linearities that traditional linear models struggle to capture accurately. Deep Learning (DL)-based approaches, unconstrained by rigid mathematical models, have shown promise in optimizing system performance by accommodating specific hardware configurations and dynamic channel conditions as an alternative to the traditional methods. Despite all the recent research efforts on DL-based methods for physical layer applications, DL models have still not been widely applied to physical layer …


Exploiting The In-Distribution Embedding Space With Deep Learning And Gaussian Discriminant Analysis For An Out-Of-Distribution Malware Attach Detection, Tosin Olusola Ige Dec 2025

Exploiting The In-Distribution Embedding Space With Deep Learning And Gaussian Discriminant Analysis For An Out-Of-Distribution Malware Attach Detection, Tosin Olusola Ige

Open Access Theses & Dissertations

State-of-the-art machine and deep learning models generally perform well on previously seen data, albeit with wrong close world assumption that all real-world data are from previously seen train and validation samples, hence there poor performance when exposed to data which deviates from previously seen training and validation set. This is clearly evident in the domain of cybersecurity where the world continues to experience several high profile malware attacks despite advancement in state-of-the-art research. The reason being that the constant evolvement of innovation in the development of tools and method deployed to carry out various attacks had given hackers and other …


Artificial Intelligence For Reliability: Predictive Health Maintenance And Geolocation In Gps-Denied Environments, Rafael Toche Pizano Dec 2025

Artificial Intelligence For Reliability: Predictive Health Maintenance And Geolocation In Gps-Denied Environments, Rafael Toche Pizano

Graduate Theses and Dissertations

In this dissertation, we explore the potential of machine learning and deep learning techniques to enhance the performance and robustness of applications across two major domains. By addressing the challenges within these fields, we demonstrate that we can leverage learning algorithms to obtain substantial improvements in accuracy and robustness. First, we tackle a problem in the field of predictive health maintenance. We propose a novel auto encoder and neural network based methodology to predict failure times in complex aviation systems to learn to distinguish between normal and abnormal operational behavior, and use this information to inform the neural network to …


Detection Of Phase-Binning And Interpolation Artifacts In 4-Dimensional Computed Tomography Imaging Using Deep Learning And Rule-Based Approaches, Jorge Cisneros, Nathan H. Feldt, Yevgeniy Vinogradskiy, Richard Castillo, Edward Castillo Dec 2025

Detection Of Phase-Binning And Interpolation Artifacts In 4-Dimensional Computed Tomography Imaging Using Deep Learning And Rule-Based Approaches, Jorge Cisneros, Nathan H. Feldt, Yevgeniy Vinogradskiy, Richard Castillo, Edward Castillo

Department of Radiation Oncology Faculty Papers

BACKGROUND: Four-dimensional computed tomography (4DCT) imaging is a crucial component to lung cancer radiotherapy planning and enables CT-ventilation-based functional avoidance planning to mitigate radiation toxicity. However, 4DCT scans are frequently impaired by acquisition artifacts that corrupt downstream analyses that depend on lung segmentation and deformable image registration, such as CT-ventilation and dose accumulation.

PURPOSE: This study develops 3D deep learning models to identify phase-binning artifacts at the voxel level and a heuristic, rule-based method to identify interpolation slices within 4DCT images.

METHODS: We introduce a generator that systematically inserts synthetic phase-binning and interpolation artifacts into any artifact-free breathing phase obtained …


Cnn Based Deep Learning Modeling With Explainability Analysis For Detecting Fraudulent Blockchain Transactions, Mohammad Hasan, Mohammad Shahriar Rahman, Mohammad Jabed Morshed Chowdhury, Iqbal H. Sarker Dec 2025

Cnn Based Deep Learning Modeling With Explainability Analysis For Detecting Fraudulent Blockchain Transactions, Mohammad Hasan, Mohammad Shahriar Rahman, Mohammad Jabed Morshed Chowdhury, Iqbal H. Sarker

Research outputs 2022 to 2026

In the era of growing cryptocurrency adoption, Blockchain has emerged as a leading player in the digital payment landscape. However, this widespread popularity also brings forth various security challenges, including the need to safeguard against fraudulent activities. One of the paramount challenges in this regard is the detection of fraudulent transactions within the realm of Bitcoin data. This task significantly influences the trust and security of digital payments. Yet, it's a formidable challenge given the relatively low occurrence of fraudulent Bitcoin transactions. While deep learning techniques have demonstrated their prowess in fraud detection, there remains a scarcity of studies exploring …


Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian Dec 2025

Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian

Research Collection School Of Computing and Information Systems

Infrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving …


Spatio-Temporal Gcn With Softmax Classifier For Skeleton-Based Human Action Recognition, Kabul Khudaybergenov, Avazjon Marakhimov, Zahriddin Muminov Nov 2025

Spatio-Temporal Gcn With Softmax Classifier For Skeleton-Based Human Action Recognition, Kabul Khudaybergenov, Avazjon Marakhimov, Zahriddin Muminov

Chemical Technology, Control and Management

Skeleton-based human action recognition is an important research area with many practical applications. Most existing methods rely on single representations of skeletal sequences, which cannot totally obtain all the complex features of human movements. This paper presents LFHAR (Latent Features for Human Action Recognition), a new framework that uses multiple spatio-temporal latent representations to improve the extraction of action features. Our method captures how skeletal poses change over time and combines motion information from both individual joints and connected body parts. The proposed approach applies graph-based processing to each skeleton frame in a sequence, then arranges the resulting graph features …


Skin Cancer Image Classification Using Deep Learning With Data Segmentation Technique, Akhilesh Kumar Shrivas, Hema Vastrakar Oct 2025

Skin Cancer Image Classification Using Deep Learning With Data Segmentation Technique, Akhilesh Kumar Shrivas, Hema Vastrakar

Karbala International Journal of Modern Science

The human skin is an impressive organ and structural element often impacted by a diverse range of recognized and unknown diseases. Diagnosing disorders that affect the outermost layer of the body is the most uncertain and difficult component in the scientific field. Dermatological diseases are one of the most significant health concerns in the 21st century since their identification is challenging and costly, plagued with challenges and the subjectivity that comes with human interpretation. The main objective of this piece of research work is to develop a robust model for the classification of skin cancer diseases using deep convolution neural …


Bifusenet: A Multimodal Network For Estimating Blood Alcohol Concentration Via Bidirectional Hierarchical Fusion, Abdullah Tariq, Arooba Maqsood, Martin Masek, Syed Zulqarnain Gilani Oct 2025

Bifusenet: A Multimodal Network For Estimating Blood Alcohol Concentration Via Bidirectional Hierarchical Fusion, Abdullah Tariq, Arooba Maqsood, Martin Masek, Syed Zulqarnain Gilani

Research outputs 2022 to 2026

Drunk driving remains a significant public safety challenge, demanding innovative alternatives to conventional methods such as field sobriety tests and breathalysers. Estimating a driver's level of intoxication through facial cues is particularly challenging due to the subtle and person-specific nature of alcohol-induced behaviours. In this paper, we present BiFuseNet, a 3D spatio-temporal multi-modal network designed to classify alcohol impairment levels into three categories: sober, moderate, and severe. Unlike prior approaches that rely on either uni-modal RGB video or hand-crafted facial features, our method exploits complementary physiological cues from RGB and infrared (IR) facial videos. We introduce a Bi-directional Hierarchical Fusion …


Detecting Polar Ring Galaxies Via Deep Learning, Fawad Kirmani, Anathavishnu S. Unnii, Varsha P. Kulkarni, Kyle Lackey, John R. Rose Oct 2025

Detecting Polar Ring Galaxies Via Deep Learning, Fawad Kirmani, Anathavishnu S. Unnii, Varsha P. Kulkarni, Kyle Lackey, John R. Rose

Faculty Publications

Polar ring galaxies (PRGs) are peculiar galaxies that show a ring of stars, gas, and dust oriented roughly over the poles of the central ‘host’ galaxy (i.e. roughly orthogonal to the disc of the host galaxy). The formation models for these rings involve mergers or tidal interactions of the host galaxy with another galaxy. Although the identified PRGs look different from each other, they all have a ring that is not in the same plane as the disc of the host galaxy. Unlike in galaxies such as our Milky Way, where stars form in spiral arms, the rings exemplify an …


Topoimages: Incorporating Local Topology Encoding Into Deep Learning Models For Medical Image Classification, Pengfei Gu, Hongxiao Wang, Yejia Zhang, Huimin Li, Chaoli Wang, Danny Z. Chen Oct 2025

Topoimages: Incorporating Local Topology Encoding Into Deep Learning Models For Medical Image Classification, Pengfei Gu, Hongxiao Wang, Yejia Zhang, Huimin Li, Chaoli Wang, Danny Z. Chen

Computer Science Faculty Publications

Topological structures in image data, such as connected components and loops, play a crucial role in understanding image content (e.g., biomedical objects). Despite remarkable successes of numerous image processing methods that rely on appearance information, these methods often lack sensitivity to topological structures when used in general deep learning (DL) frameworks. In this paper, we introduce a new general approach, called TopoImages (for Topology Images), which computes a new representation of input images by encoding local topology of patches. In TopoImages, we leverage persistent homology (PH) to encode geometric and topological features inherent in image patches. Our main objective is …