Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer vision

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 241 - 270 of 349

Full-Text Articles in Computer Sciences

Generative Adversarial Networks For Online Visual Object Tracking Systems, Ghsoun Zin Jan 2019

Generative Adversarial Networks For Online Visual Object Tracking Systems, Ghsoun Zin

Theses and Dissertations (Comprehensive)

Object Tracking is one of the essential tasks in computer vision domain as it has numerous applications in various fields, such as human-computer interaction, video surveillance, augmented reality, and robotics. Object Tracking refers to the process of detecting and locating the target object in a series of frames in a video. The state-of-the-art for tracking-by-detection framework is typically made up of two steps to track the target object. The first step is drawing multiple samples near the target region of the previous frame. The second step is classifying each sample as either the target object or the background. Visual object …


Learning About Large Scale Image Search: Lessons From Global Scale Hotel Recognition To Fight Sex Trafficking, Abby Stylianou Dec 2018

Learning About Large Scale Image Search: Lessons From Global Scale Hotel Recognition To Fight Sex Trafficking, Abby Stylianou

McKelvey School of Engineering Graduate Student Theses & Dissertations

Hotel recognition is a sub-domain of scene recognition that involves determining what hotel is seen in a photograph taken in a hotel. The hotel recognition task is a challenging computer vision task due to the properties of hotel rooms, including low visual similarity between rooms in the same hotel and high visual similarity between rooms in different hotels, particularly those from the same chain. Building accurate approaches for hotel recognition is important to investigations of human trafficking. Images of human trafficking victims are often shared by traffickers among criminal networks and posted in online advertisements. These images are often taken …


Entity-Grounded Image Captioning, Annika Lindh, Robert J. Ross, John D. Kelleher Sep 2018

Entity-Grounded Image Captioning, Annika Lindh, Robert J. Ross, John D. Kelleher

Conference papers

An urgent limitation in current Image Captioning models is their tendency to produce generic captions that avoid the interesting detail which makes each image unique. To address this limitation, we propose an approach that enforces a stronger alignment between image regions and specific segments of text. The model architecture is composed of a visual region proposer, a region-order planner and a region-guided caption generator. The region-guided caption generator incorporates a novel information gate which allows visual and textual input of different frequencies and dimensionalities in a Recurrent Neural Network.


Enhancing 3d Visual Odometry With Single-Camera Stereo Omnidirectional Systems, Carlos A. Jaramillo Sep 2018

Enhancing 3d Visual Odometry With Single-Camera Stereo Omnidirectional Systems, Carlos A. Jaramillo

Dissertations, Theses, and Capstone Projects

We explore low-cost solutions for efficiently improving the 3D pose estimation problem of a single camera moving in an unfamiliar environment. The visual odometry (VO) task -- as it is called when using computer vision to estimate egomotion -- is of particular interest to mobile robots as well as humans with visual impairments. The payload capacity of small robots like micro-aerial vehicles (drones) requires the use of portable perception equipment, which is constrained by size, weight, energy consumption, and processing power. Using a single camera as the passive sensor for the VO task satisfies these requirements, and it motivates the …


Artificial Intelligence For Cognitive Behavior Assessment In Children, Srujana Gattupalli Aug 2018

Artificial Intelligence For Cognitive Behavior Assessment In Children, Srujana Gattupalli

Computer Science and Engineering Dissertations - Archive

Cognitive impairments in early childhood can lead to poor academic performance and require proper remedial intervention at the appropriate time. ADHD a?ects about 6-7% of children and is a psychiatric neurodevelopmental disorder that is very hard to diagnose or tell apart from other disorders. Cognitive insu?ciencies hinder the development of working memory and can a?ect school success and even have long term e?ects that can result in low self-esteem and self-acceptance. The main aim of this research is to investigate development of an automated and non-intrusive system for assessing physical exercises related to the treatment and diagnosis of Attention De?cit …


Bounding Box Improvement With Reinforcement Learning, Andrew Lewis Cleland Jun 2018

Bounding Box Improvement With Reinforcement Learning, Andrew Lewis Cleland

Dissertations and Theses

In this thesis, I explore a reinforcement learning technique for improving bounding box localizations of objects in images. The model takes as input a bounding box already known to overlap an object and aims to improve the fit of the box through a series of transformations that shift the location of the box by translation, or change its size or aspect ratio. Over the course of these actions, the model adapts to new information extracted from the image. This active localization approach contrasts with existing bounding-box regression methods, which extract information from the image only once. I implement, train, and …


Face Detection And Recognition Using Moving Window Accumulator With Various Deep Learning Architecture, Anil Kumar Nayak May 2018

Face Detection And Recognition Using Moving Window Accumulator With Various Deep Learning Architecture, Anil Kumar Nayak

Computer Science and Engineering Theses - Archive

Recent advancement in the field of Computer Vision and Deep Learning is making object detection and recognition easier. Hence, growing research activities in the field of deep learning are enabling researchers to find new ideas in the area of face detection and recognition. Implementation of such systems has a number of challenges when it comes to the current approaches. In this paper, we have presented a system of Face Detection and Recognition with newly designed deep learning classification models like CNN, Inception and various state of art models like SVM and we also compared the result with FaceNet. Multiple approaches …


From Text Classification To Image Clustering, Problems Less Optimized, Amirhossein Herandi May 2018

From Text Classification To Image Clustering, Problems Less Optimized, Amirhossein Herandi

Computer Science and Engineering Theses - Archive

Machine Learning is thriving. Every industry is using its techniques in some way to improve their efficiency and revenue. However, the focus on research is not divided equally between all of the different areas and problems that this field can tackle and analyze. Currently, Computer Vision is the one area that is being focused very extensively by researchers and companies alike, and as a result has seen an amazing boost in the recent years. This ranges from the well-known problems of classification that use discriminative models all the way to more novel problems that use generative models such as style …


Applications Of Deep Learning In Large-Scale Object Detection And Semantic Segmentation, Wei Xiang May 2018

Applications Of Deep Learning In Large-Scale Object Detection And Semantic Segmentation, Wei Xiang

Computer Science and Engineering Dissertations - Archive

With the massive storage of multimedia data and increasing computational power of mobile devices, developing scalable computer vision applications has become the primary motivation for both research and industrial community. Among these applications, object detection and semantic segmentation are two of the most popular topics which, in addition, serve as the fundamental features for many computer vision systems under platforms like mobile, healthcare, autonomous driving, etc. Inspired by the current and foreseeable trend, this thesis focuses on developing both effective and efficient object detection and semantic segmentation models, with the large-scale, publicly available data sets sourced for various applications. In …


Integrity Monitoring For Automated Aerial Refueling: A Stereo Vision Approach, Thomas R. Stuart Mar 2018

Integrity Monitoring For Automated Aerial Refueling: A Stereo Vision Approach, Thomas R. Stuart

Theses and Dissertations

Unmanned aerial vehicles (UAVs) increasingly require the capability to y autonomously in close formation including to facilitate automated aerial refueling (AAR). The availability of relative navigation measurements and navigation integrity are essential to autonomous relative navigation. Due to the potential non-availability of the global positioning system (GPS) during military operations, it is highly desirable that relative navigation can be accomplished without the use of GPS. This paper develops two algorithms designed to provide relative navigation measurements solely from a stereo image pair. These algorithms were developed and analyzed in the context of AAR using a stereo camera system modeling that …


Object Localization, Segmentation, And Classification In 3d Images, Allan Zelener Feb 2018

Object Localization, Segmentation, And Classification In 3d Images, Allan Zelener

Dissertations, Theses, and Capstone Projects

We address the problem of identifying objects of interest in 3D images as a set of related tasks involving localization of objects within a scene, segmentation of observed object instances from other scene elements, classifying detected objects into semantic categories, and estimating the 3D pose of detected objects within the scene. The increasing availability of 3D sensors motivates us to leverage large amounts of 3D data to train machine learning models to address these tasks in 3D images. Leveraging recent advances in deep learning has allowed us to develop models capable of addressing these tasks and optimizing these tasks jointly …


Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger Jan 2018

Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger

Theses and Dissertations--Computer Science

Human appearance is highly variable and depends on individual preferences, such as fashion, facial expression, and makeup. These preferences depend on many factors including a person's sense of style, what they are doing, and the weather. These factors, in turn, are dependent upon geographic location and time. In our work, we build computational models to learn the relationship between human appearance, geographic location, and time. The primary contributions are a framework for collecting and processing geotagged imagery of people, a large dataset collected by our framework, and several generative and discriminative models that use our dataset to learn the relationship …


Deep Probabilistic Models For Camera Geo-Calibration, Menghua Zhai Jan 2018

Deep Probabilistic Models For Camera Geo-Calibration, Menghua Zhai

Theses and Dissertations--Computer Science

The ultimate goal of image understanding is to transfer visual images into numerical or symbolic descriptions of the scene that are helpful for decision making. Knowing when, where, and in which direction a picture was taken, the task of geo-calibration makes it possible to use imagery to understand the world and how it changes in time. Current models for geo-calibration are mostly deterministic, which in many cases fails to model the inherent uncertainties when the image content is ambiguous. Furthermore, without a proper modeling of the uncertainty, subsequent processing can yield overly confident predictions. To address these limitations, we propose …


Leveraging Overhead Imagery For Localization, Mapping, And Understanding, Scott Workman Jan 2018

Leveraging Overhead Imagery For Localization, Mapping, And Understanding, Scott Workman

Theses and Dissertations--Computer Science

Ground-level and overhead images provide complementary viewpoints of the world. This thesis proposes methods which leverage dense overhead imagery, in addition to sparsely distributed ground-level imagery, to advance traditional computer vision problems, such as ground-level image localization and fine-grained urban mapping. Our work focuses on three primary research areas: learning a joint feature representation between ground-level and overhead imagery to enable direct comparison for the task of image geolocalization, incorporating unlabeled overhead images by inferring labels from nearby ground-level images to improve image-driven mapping, and fusing ground-level imagery with overhead imagery to enhance understanding. The ultimate contribution of this thesis …


Estimating Meteorological Visibility Range Under Foggy Weather Conditions: A Deep Learning Approach, Hazar Chaabani, Naoufel Werghi, Faouzi Kamoun, Bilal Taha, Fatma Outay, Ansar Ul Haque Yasar Jan 2018

Estimating Meteorological Visibility Range Under Foggy Weather Conditions: A Deep Learning Approach, Hazar Chaabani, Naoufel Werghi, Faouzi Kamoun, Bilal Taha, Fatma Outay, Ansar Ul Haque Yasar

All Works

© 2018 The Authors. Published by Elsevier Ltd. Systems capable of estimating visibility distances under foggy weather conditions are extremely useful for next-generation cooperative situational awareness and collision avoidance systems. In this paper, we present a brief review of noticeable approaches for determining visibility distance under foggy weather conditions. We then propose a novel approach based on the combination of a deep learning method for feature extraction and an SVM classifier. We present a quantitative evaluation of the proposed solution and show that our approach provides better performance results compared to an earlier approach that was based on the combination …


Shadow Patching: Exemplar-Based Shadow Removal, Ryan Sears Hintze Dec 2017

Shadow Patching: Exemplar-Based Shadow Removal, Ryan Sears Hintze

Theses and Dissertations

Shadow removal is an important problem for both artists and algorithms. Previous methods handle some shadows well but, because they rely on the shadowed data, perform poorly in cases with severe degradation. Image-completion algorithms can completely replace severely degraded shadowed regions, and perform well with smaller-scale textures, but often fail to reproduce larger-scale macrostructure that may still be visible in the shadowed region. This paper provides a general framework that leverages degraded (e.g., shadowed) data to guide the image completion process by extending the objective function commonly used in current state-of-the-art image completion energy-minimization methods. This approach achieves realistic shadow …


Delving Into Salient Object Subitizing And Detection, Shengfeng He, Jianbo Jiao, Xiaodan Zhang, Guoqiang Han, Rynson W.H Lau Oct 2017

Delving Into Salient Object Subitizing And Detection, Shengfeng He, Jianbo Jiao, Xiaodan Zhang, Guoqiang Han, Rynson W.H Lau

Research Collection School Of Computing and Information Systems

Subitizing (i.e., instant judgement on the number) and detection of salient objects are human inborn abilities. These two tasks influence each other in the human visual system. In this paper, we delve into the complementarity of these two tasks. We propose a multi-task deep neural network with weight prediction for salient object detection, where the parameters of an adaptive weight layer are dynamically determined by an auxiliary subitizing network. The numerical representation of salient objects is therefore embedded into the spatial representation. The proposed joint network can be trained end-to-end using backpropagation. Experiments show the proposed multi-task network outperforms existing …


Improved Scoring Models For Semantic Image Retrieval Using Scene Graphs, Erik Timothy Conser Sep 2017

Improved Scoring Models For Semantic Image Retrieval Using Scene Graphs, Erik Timothy Conser

Dissertations and Theses

Image retrieval via a structured query is explored in Johnson, et al. [7]. The query is structured as a scene graph and a graphical model is generated from the scene graph's object, attribute, and relationship structure. Inference is performed on the graphical model with candidate images and the energy results are used to rank the best matches. In [7], scene graph objects that are not in the set of recognized objects are not represented in the graphical model. This work proposes and tests two approaches for modeling the unrecognized objects in order to leverage the attribute and relationship models to …


Refining Bounding-Box Regression For Object Localization, Naomi Lynn Dickerson Sep 2017

Refining Bounding-Box Regression For Object Localization, Naomi Lynn Dickerson

Dissertations and Theses

For the last several years, convolutional neural network (CNN) based object detection systems have used a regression technique to predict improved object bounding boxes based on an initial proposal using low-level image features extracted from the CNN. In spite of its prevalence, there is little critical analysis of bounding-box regression or in-depth performance evaluation. This thesis surveys an array of techniques and parameter settings in order to further optimize bounding-box regression and provide guidance for its implementation. I refute a claim regarding training procedure, and demonstrate the effectiveness of using principal component analysis to handle unwieldy numbers of features produced …


Cross-Correlation-Based Structural System Identification Using Unmanned Aerial Vehicles., Hyungchul Yoon, Vedhus Hoskere, Jong-Woong Park, Billie F. Spencer Sep 2017

Cross-Correlation-Based Structural System Identification Using Unmanned Aerial Vehicles., Hyungchul Yoon, Vedhus Hoskere, Jong-Woong Park, Billie F. Spencer

Michigan Tech Publications, Part 1

Computer vision techniques have been employed to characterize dynamic properties of structures, as well as to capture structural motion for system identification purposes. All of these methods leverage image-processing techniques using a stationary camera. This requirement makes finding an effective location for camera installation difficult, because civil infrastructure (i.e., bridges, buildings, etc.) are often difficult to access, being constructed over rivers, roads, or other obstacles. This paper seeks to use video from Unmanned Aerial Vehicles (UAVs) to address this problem. As opposed to the traditional way of using stationary cameras, the use of UAVs brings the issue of the camera …


An Intelligent Multimodal Upper-Limb Rehabilitation Robotic System, Alexandros Lioulemes Aug 2017

An Intelligent Multimodal Upper-Limb Rehabilitation Robotic System, Alexandros Lioulemes

Computer Science and Engineering Dissertations - Archive

A traffic accident, a battlefield injury, or a stroke can lead to brain or musculoskeletal injuries that impact motor and cognitive functions and can drastically change a person's life. In such situations, rehabilitation plays a critical role in the ability of the patient to partially or totally regain motor function, but the optimal training approach remains unclear. Robotic technologies are recognized as powerful tools to promote neuroplasticity and stimulate motor re-learning. Moreover, they deliver high-intensity, repetitive, active and task-oriented training; in addition, they provide objective measurements for patient evaluation. The primary focus of this research is to investigate the development …


Formresnet: Formatted Residual Learning For Image Restoration, Jianbo Jiao, Wei-Chih Tu, Shengfeng He Aug 2017

Formresnet: Formatted Residual Learning For Image Restoration, Jianbo Jiao, Wei-Chih Tu, Shengfeng He

Research Collection School Of Computing and Information Systems

In this paper, we propose a deep CNN to tackle the image restoration problem by learning the structured residual. Previous deep learning based methods directly learn the mapping from corrupted images to clean images, and may suffer from the gradient exploding/vanishing problems of deep neural networks. We propose to address the image restoration problem by learning the structured details and recovering the latent clean image together, from the shared information between the corrupted image and the latent image. In addition, instead of learning the pure difference (corruption), we propose to add a 'residual formatting layer' to format the residual to …


Deshadownet: A Multi-Context Embedding Deep Network For Shadow Removal, Liangqiong Qu, Jiandong Tian, Shengfeng He, Yandong Tang, Rynson W. H. Lau Jul 2017

Deshadownet: A Multi-Context Embedding Deep Network For Shadow Removal, Liangqiong Qu, Jiandong Tian, Shengfeng He, Yandong Tang, Rynson W. H. Lau

Research Collection School Of Computing and Information Systems

Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a multi-context architecture, where the output shadow matte is predicted by embedding information from three different perspectives. The first global network extracts shadow features from a global view. Two levels of features are derived from the global network and transferred to two parallel networks. While one extracts the appearance of the input image, the …


Computer Vision Based Route Mapping, Ryan S. Kehlenbeck, Zachary Cody Jun 2017

Computer Vision Based Route Mapping, Ryan S. Kehlenbeck, Zachary Cody

Computer Science and Software Engineering

The problem our project solves is the integration of edge detection techniques with mapping libraries to display routes based on images. To do this, we used the OpenCV library within an Android application. This application lets a user import an image from their device, and uses edge detection to pull out a path from the image. The application can find the user's location and uses it alongside the path data from the image to create a route using the physical roads near the location. The shape of the route matches the edges from the given image and the user can …


Bayesian Optimization For Refining Object Proposals, With An Application To Pedestrian Detection, Anthony D. Rhodes May 2017

Bayesian Optimization For Refining Object Proposals, With An Application To Pedestrian Detection, Anthony D. Rhodes

Student Research Symposium

We devise an algorithm using a Bayesian optimization framework in conjunction with contextual visual data for the efficient localization of objects in still images. Recent research has demonstrated substantial progress in object localization and related tasks for computer vision. However, many current state-of-the-art object localization procedures still suffer from inaccuracy and inefficiency, in addition to failing to successfully leverage contextual data. We address these issues with the current research.

Our method encompasses an active search procedure that uses contextual data to generate initial bounding-box proposals for a target object. We train a convolutional neural network to approximate an offset distance …


Tandem 2.0: Image And Text Data Generation Application, Christopher J. Vitale Feb 2017

Tandem 2.0: Image And Text Data Generation Application, Christopher J. Vitale

Dissertations, Theses, and Capstone Projects

First created as part of the Digital Humanities Praxis course in the spring of 2012 at the CUNY Graduate Center, Tandem explores the generation of datasets comprised of text and image data by leveraging Optical Character Recognition (OCR), Natural Language Processing (NLP) and Computer Vision (CV). This project builds upon that earlier work in a new programming framework. While other developers and digital humanities scholars have created similar tools specifically geared toward NLP (e.g. Voyant-Tools), as well as algorithms for image processing and feature extraction on the CV side, Tandem explores the process of developing a more robust and user-friendly …


A Neural Network Approach To Visibility Range Estimation Under Foggy Weather Conditions, Hazar Chaabani, Faouzi Kamoun, Hichem Bargaoui, Fatma Outay, Ansar Ul Haque Yasar Jan 2017

A Neural Network Approach To Visibility Range Estimation Under Foggy Weather Conditions, Hazar Chaabani, Faouzi Kamoun, Hichem Bargaoui, Fatma Outay, Ansar Ul Haque Yasar

All Works

© 2017 The Authors. Published by Elsevier B.V. The degradation of visibility due to foggy weather conditions is a common trigger for road accidents and, as a result, there has been a growing interest to develop intelligent fog detection and visibility range estimation systems. In this contribution, we provide a brief overview of the state-of-the-art contributions in relation to estimating visibility distance under foggy weather conditions. We then present a neural network approach for estimating visibility distances using a camera that can be fixed to a roadside unit (RSU) or mounted onboard a moving vehicle. We evaluate the proposed solution …


Assessing The Importance Of Features For Detection Of Hard Exudates In Retinal Images, Kemal Akyol, Baha Şen, Şafak Bayir, Hasan Basri̇ Çakmak Jan 2017

Assessing The Importance Of Features For Detection Of Hard Exudates In Retinal Images, Kemal Akyol, Baha Şen, Şafak Bayir, Hasan Basri̇ Çakmak

Turkish Journal of Electrical Engineering and Computer Sciences

Diabetes disrupts the operation of the eye and leads to vision loss, affecting particularly the nerve layer and capillary vessels in this layer by changes in the blood vessels of the retina.~Suddenly loss and blurred vision problems occur in the image, depending on the phase of the disease, called diabetic retinopathy. Hard exudates are one of the primary signs of diabetic retinopathy. Automatic recognition of hard exudates in retinal images can contribute to detection of the disease. We present an automatic screening system for the detection of hard exudates. This system consists of two main steps. Firstly, the features were …


Extracting And Modeling Useful Information From Videos For Supporting Continuous Queries, Manish Kumar Annappa Dec 2016

Extracting And Modeling Useful Information From Videos For Supporting Continuous Queries, Manish Kumar Annappa

Computer Science and Engineering Theses - Archive

Automating video stream processing for inferring situations of interest has been an ongoing challenge. This problem is currently exacerbated by the volume of surveillance/monitoring videos generated. Currently, manual or context-based customized techniques are used for this purpose. To the best to our knowledge the attempted work in this area use a custom query language to extract data and infer simple situations from the video streams, thus adding an additional overhead to learn their query language. Objective of the work in this thesis is to develop a framework that extracts data from video streams generating a data representation such that simple …


Indoor Scene Localization To Fight Sex Trafficking In Hotels, Abigail Stylianou Dec 2016

Indoor Scene Localization To Fight Sex Trafficking In Hotels, Abigail Stylianou

McKelvey School of Engineering Graduate Student Theses & Dissertations

Images are key to fighting sex trafficking. They are: (a) used to advertise for sex services,(b) shared among criminal networks, and (c) connect a person in an image to the place where the image was taken. This work explores the ability to link images to indoor places in order to support the investigation and prosecution of sex trafficking. We propose and develop a framework that includes a database of open-source information available on the Internet, a crowd-sourcing approach to gathering additional images, and explore a variety of matching approaches based both on hand-tuned features such as SIFT and learned features …