Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (16)
- Engineering (12)
- Computer Engineering (9)
- Data Storage Systems (7)
- Artificial Intelligence and Robotics (5)
-
- Other Computer Sciences (5)
- Social and Behavioral Sciences (5)
- Information Security (3)
- Electrical and Computer Engineering (2)
- Library and Information Science (2)
- OS and Networks (2)
- Programming Languages and Compilers (2)
- Software Engineering (2)
- Systems Architecture (2)
- Theory and Algorithms (2)
- Aerospace Engineering (1)
- Anthropology (1)
- Archaeological Anthropology (1)
- Art Practice (1)
- Arts and Humanities (1)
- Aviation (1)
- Digital Communications and Networking (1)
- Geography (1)
- Hardware Systems (1)
- Numerical Analysis and Scientific Computing (1)
- Physical and Environmental Geography (1)
- Power and Energy (1)
- Institution
- Keyword
-
- Avatars (3)
- #antcenter (2)
- Augmented reality (2)
- Classification (2)
- Computer graphics (2)
-
- Context (2)
- GUI (2)
- Intelligent virtual agents (2)
- Perception of avatars (2)
- Virtual worlds (2)
- Visualization (2)
- Web video (2)
- AI (1)
- Academic libraries (1)
- Animation (1)
- Animator (1)
- Application software (1)
- Art (1)
- Automation--Human factors (1)
- Biometric identification (1)
- Blender (1)
- CEMD matching (1)
- Categorization (1)
- Clustering (1)
- Collaboration (1)
- Color theory (1)
- Common Pattern Discovery (1)
- Computational modeling (1)
- Computer errors (1)
- Computer industry (1)
- Publication
-
- Research Collection School Of Computing and Information Systems (19)
- Theses and Dissertations (4)
- Computer Science Faculty Publications (3)
- Conference papers (3)
- Master's Theses (3)
-
- Honors Scholar Theses (2)
- International Conference on Information and Communication Technologies (2)
- Computer Engineering (1)
- Departmental Papers (CS) (1)
- Dyson College- Seidenberg School of CSIS : Collaborative Projects and Presentations (1)
- Electrical & Computer Engineering Faculty Publications (1)
- Georgia Library Quarterly (1)
- Masters Theses & Specialist Projects (1)
- VCU Libraries Faculty and Staff Presentations (1)
- Publication Type
- File Type
Articles 1 - 30 of 43
Full-Text Articles in Graphics and Human Computer Interfaces
Exercise Power Grid Display And Web Interface, Alexander (Alex) Chernetz
Exercise Power Grid Display And Web Interface, Alexander (Alex) Chernetz
Computer Engineering
The 2008-2009 expansion of the Recreation Center at Cal Poly includes three new rooms with cardiovascular fitness equipment. As part of its ongoing commitment to sustainable development, the new machines connect to the main power grid and generate power during a workout. This document explains the process of quantifying and expressing the power generated using two interfaces: an autonomous display designed for a television with a text size and amount of detail adaptable to multiple television sizes and viewing distances, and an interactive, more detailed Web interface accessible with any Java-capable computer system or browser.
Evaluating Head Gestures For Panning 2-D Spatial Information, Matthew O. Derry
Evaluating Head Gestures For Panning 2-D Spatial Information, Matthew O. Derry
Master's Theses
New, often free, spatial information applications such as mapping tools, topological imaging, and geographic information systems are becoming increasingly available to the average computer user. These systems, which were once available only to government, scholastic, and corporate institutions with highly skilled operators, are driving a need for new and innovative ways for the average user to navigate and control spatial information intuitively, accurately, and efficiently. Gestures provide a method of control that is well suited to navigating the large datasets often associated with spatial information applications. Several different types of gestures and different applications that navigate spatial data are examined. …
Vireo/Dvmm At Trecvid 2009: High-Level Feature Extraction, Automatic Video Search, And Content-Based Copy Detection, Chong-Wah Ngo, Yu-Gang Jiang, Xiao-Yong Wei, Wanlei Zhao, Yang Liu, Jun Wang, Shiai Zhu, Shih-Fu Chang
Vireo/Dvmm At Trecvid 2009: High-Level Feature Extraction, Automatic Video Search, And Content-Based Copy Detection, Chong-Wah Ngo, Yu-Gang Jiang, Xiao-Yong Wei, Wanlei Zhao, Yang Liu, Jun Wang, Shiai Zhu, Shih-Fu Chang
Research Collection School Of Computing and Information Systems
This paper presents overview and comparative analysis of our systems designed for 3 TRECVID 2009 tasks: high-level feature extraction, automatic search, and content-based copy detection.
Towards Google Challenge: Combining Contextual And Social Information For Web Video Categorization, Xiao Wu, Wan-Lei Zhao, Chong-Wah Ngo
Towards Google Challenge: Combining Contextual And Social Information For Web Video Categorization, Xiao Wu, Wan-Lei Zhao, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Web video categorization is a fundamental task for web video search. In this paper, we explore the Google challenge from a new perspective by combing contextual and social information under the scenario of social web. The semantic meaning of text (title and tags), video relevance from related videos, and user interest induced from user videos, are integrated to robustly determine the video category. Experiments on YouTube videos demonstrate the effectiveness of the proposed solution. The performance reaches 60% improvement compared to the traditional text based classifiers.
Localizing Volumetric Motion For Action Recognition In Realistic Videos, Xiao Wu, Chong-Wah Ngo, Jintao Li, Yongdong Zhang
Localizing Volumetric Motion For Action Recognition In Realistic Videos, Xiao Wu, Chong-Wah Ngo, Jintao Li, Yongdong Zhang
Research Collection School Of Computing and Information Systems
This paper presents a novel motion localization approach for recognizing actions and events in real videos. Examples include StandUp and Kiss in Hollywood movies. The challenge can be attributed to the large visual and motion variations imposed by realistic action poses. Previous works mainly focus on learning from descriptors of cuboids around space time interest points (STIP) to characterize actions. The size, shape and space-time position of cuboids are fixed without considering the underlying motion dynamics. This often results in large set of fragmentized cuboids which fail to capture long-term dynamic properties of realistic actions. This paper proposes the detection …
Second Life Complements The Internet For Reference Librarians, Florence Tang
Second Life Complements The Internet For Reference Librarians, Florence Tang
Georgia Library Quarterly
The article describes the Second Life culture as an interactive virtual environment from the perspective of a reference librarian. The cited advantages of using Second Life are synchronous interaction with other people, clearer appearance of chat text messages, incorporation of social norms, and accessibility. Among the noted obstacles to using Second Life are its price, intentional distress of other avatars, empty virtual public places, less hierarchy among participants, chatting via rapid typing, and quick change of landscapes and avatars.
Distribution-Based Concept Selection For Concept-Based Video Retrieval, Juan Cao, Hongfang Jing, Chong-Wah Ngo, Yongdong Zhang
Distribution-Based Concept Selection For Concept-Based Video Retrieval, Juan Cao, Hongfang Jing, Chong-Wah Ngo, Yongdong Zhang
Research Collection School Of Computing and Information Systems
Query-to-concept mapping plays one of the keys to concept-based video retrieval. Conventional approaches try to find concepts that are likely to co-occur in the relevant shots from the lexical or statistical aspects. However, the high probability of co-occurrence alone cannot ensure its effectiveness to distinguish the relevant shots from the irrelevant ones. In this paper, we propose distribution-based concept selection (DBCS) for query-to-concept mapping by analyzing concept score distributions of within and between relevant and irrelevant sets. In view of the imbalance between relevant and irrelevant examples, two variants of DBCS are proposed respectively by considering the two-sided and onesided …
Semantic Context Transfer Across Heterogeneous Sources For Domain Adaptive Video Search, Yu-Gang Jiang, Chong-Wah Ngo, Shih-Fu Chang
Semantic Context Transfer Across Heterogeneous Sources For Domain Adaptive Video Search, Yu-Gang Jiang, Chong-Wah Ngo, Shih-Fu Chang
Research Collection School Of Computing and Information Systems
Automatic video search based on semantic concept detectors has recently received significant attention. Since the number of available detectors is much smaller than the size of human vocabulary, one major challenge is to select appropriate detectors to response user queries. In this paper, we propose a novel approach that leverages heterogeneous knowledge sources for domain adaptive video search. First, instead of utilizing WordNet as most existing works, we exploit the context information associated with Flickr images to estimate query-detector similarity. The resulting measurement, named Flickr context similarity (FCS), reflects the co-occurrence statistics of words in image context rather than textual …
Scalable Detection Of Partial Near-Duplicate Videos By Visual-Temporal Consistency, Hung-Khoon Tan, Chong-Wah Ngo, Richang Hong, Tat-Seng Chua
Scalable Detection Of Partial Near-Duplicate Videos By Visual-Temporal Consistency, Hung-Khoon Tan, Chong-Wah Ngo, Richang Hong, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Following the exponential growth of social media, there now exist huge repositories of videos online. Among the huge volumes of videos, there exist large numbers of near-duplicate videos. Most existing techniques either focus on the fast retrieval of full copies or near-duplicates, or consider localization in a heuristic manner. This paper considers the scalable detection and localization of partial near-duplicate videos by jointly considering visual similarity and temporal consistency. Temporal constraints are embedded into a network structure as directed edges. Through the structure, partial alignment is novelly converted into a network flow problem where highly efficient solutions exist. To precisely …
Domain Adaptive Semantic Diffusion For Large Scale Context-Based Video Annotation, Yu-Gang Jiang, Jun Wang, Shih-Fu Chang, Chong-Wah Ngo
Domain Adaptive Semantic Diffusion For Large Scale Context-Based Video Annotation, Yu-Gang Jiang, Jun Wang, Shih-Fu Chang, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Learning to cope with domain change has been known as a challenging problem in many real-world applications. This paper proposes a novel and efficient approach, named domain adaptive semantic diffusion (DASD), to exploit semantic context while considering the domain-shift-of-context for large scale video concept annotation. Starting with a large set of concept detectors, the proposed DASD refines the initial annotation results using graph diffusion technique, which preserves the consistency and smoothness of the annotation over a semantic graph. Different from the existing graph learning methods which capture relations among data samples, the semantic graph treats concepts as nodes and the …
Localized Matching Using Earth Mover's Distance Towards Discovery Of Common Patterns From Small Image Samples, Hung-Khoon Tan, Chong-Wah Ngo
Localized Matching Using Earth Mover's Distance Towards Discovery Of Common Patterns From Small Image Samples, Hung-Khoon Tan, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
This paper proposes a new approach for the discovery of common patterns in a small set of images by region matching. The issues in feature robustness, matching robustness and noise artifact are addressed to delve into the potential of using regions as the basic matching unit. We novelly employ the many-to-many (M2M) matching strategy, specifically with the Earth Mover's Distance (EMD), to increase resilience towards the structural inconsistency from improper region segmentation. However, the matching pattern of M2M is dispersed and unregulated in nature, leading to the challenges of mining a common pattern while identifying the underlying transformation. To avoid …
A Latent Model For Visual Disambiguation Of Keyword-Based Image Search, Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia, Sujoy Roy
A Latent Model For Visual Disambiguation Of Keyword-Based Image Search, Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia, Sujoy Roy
Research Collection School Of Computing and Information Systems
The problem of polysemy in keyword-based image search arises mainly from the inherent ambiguity in user queries. We propose a latent model based approach that resolves user search ambiguity by allowing sense specific diversity in search results. Given a query keyword and the images retrieved by issuing the query to an image search engine, we first learn a latent visual sense model of these polysemous images. Next, we use Wikipedia to disambiguate the word sense of the original query, and issue these Wiki-senses as new queries to retrieve sense specific images. A sense-specific image classifier is then learnt by combining …
Wireless Networks: Spert: A Stateless Protocol For Energy-Sensitive Real-Time Routing For Wireless Sensor Network, Sohail Jabbar, Abid Ali Minhas, Raja Adeel Akhtar
Wireless Networks: Spert: A Stateless Protocol For Energy-Sensitive Real-Time Routing For Wireless Sensor Network, Sohail Jabbar, Abid Ali Minhas, Raja Adeel Akhtar
International Conference on Information and Communication Technologies
Putting constraints on performance of a system in the temporal domain, some times turns right into wrong and update into outdate. These are the scenarios where apposite value of time inveterate in the reality. But such timing precision not only requires tightly scheduled performance constraints but also requires optimal design and operation of all system components. Any malfunctioning at any relevant aspect may causes a serious disaster and even loss of human lives. Managing and interacting with such real-time system becomes much intricate when the resources are limited as in wireless sensor nodes. A wireless sensor node is typically comprises …
Networks - I: Collaborative 3d Digital Content Creation Exploiting A Grid Network, M. Gkion, M. Z. Patoli, A. Al-Barakati, W. Zhang, P. Newbury, M. White
Networks - I: Collaborative 3d Digital Content Creation Exploiting A Grid Network, M. Gkion, M. Z. Patoli, A. Al-Barakati, W. Zhang, P. Newbury, M. White
International Conference on Information and Communication Technologies
The increase in ease of the production of computer simulated graphics has opened new opportunities in the 3D industry. There are unlimited applications for the delivery of 3D Graphics especially concerning 3D multimedia presentation of digital content. Apart from aesthetic and entertaining reasons, experts apply computer simulations to visualize environments and to identify early errors or costs in order to limit the need of making real prototypes. Thus, 3D Graphics also minimize the time required for developing the final product. Existing 3D applications give partial support to users to engage in collaborative contribution for the production of a 3D model. …
Are Male And Female Avatars Perceived Equally In 3-D Virtual Worlds?, David Dewester, Fiona Fui-Hoon Nah, Sarah Gervais, Keng Siau
Are Male And Female Avatars Perceived Equally In 3-D Virtual Worlds?, David Dewester, Fiona Fui-Hoon Nah, Sarah Gervais, Keng Siau
Research Collection School Of Computing and Information Systems
Virtual worlds are three-dimensional, computer-generated worlds in which users take the form of avatars and use those avatars to interact with objects and other avatars in the virtual world. Virtual worlds are growing in importance in both educational institutions and businesses. Educational institutions have adopted virtual worlds as a medium for instructional delivery whereas businesses are using virtual worlds for recruitment, training, collaboration, and marketing. Given these emerging phenomena, a better understanding of behavioral and perceptual issues in virtual worlds is warranted. We propose a research model to study the interaction effects of gender stereotypicality of male and female avatars …
Large-Scale Near-Duplicate Web Video Search: Challenge And Opportunity, Wan-Lei Zhao, Song Tan, Chong-Wah Ngo
Large-Scale Near-Duplicate Web Video Search: Challenge And Opportunity, Wan-Lei Zhao, Song Tan, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
The massive amount of near-duplicate and duplicate web videos has presented both challenge and opportunity to multimedia computing. On one hand, browsing videos on Internet becomes highly inefficient for the need to repeatedly fast-forward videos of similar content. On the other hand, the tremendous amount of somewhat duplicate content also makes some traditionally difficult vision tasks become simple and easy. For example, annotating pictures can be as simple as recycling the tags of Internet images retrieved from image search engines. Such tasks, of either to eliminate or to recycle near-duplicates, can usually be achieved by the nearest neighbor search of …
A Bayesian Approach Integrating Regional And Global Features For Image Semantic Learning, Luong-Dong Nguyen, Ghim-Eng Yap, Ying Liu, Ah-Hwee Tan, Liang-Tien Chia, Joo-Hwee Lim
A Bayesian Approach Integrating Regional And Global Features For Image Semantic Learning, Luong-Dong Nguyen, Ghim-Eng Yap, Ying Liu, Ah-Hwee Tan, Liang-Tien Chia, Joo-Hwee Lim
Research Collection School Of Computing and Information Systems
In content-based image retrieval, the “semantic gap” between visual image features and user semantics makes it hard to predict abstract image categories from low-level features. We present a hybrid system that integrates global features (Gfeatures) and region features (R-features) for predicting image semantics. As an intermediary between image features and categories, we introduce the notion of mid-level concepts, which enables us to predict an image’s category in three steps. First, a G-prediction system uses G-features to predict the probability of each category for an image. Simultaneously, a R-prediction system analyzes R-features to identify the probabilities of mid-level concepts in that …
Exploring Inter-Concept Relationship With Context Space For Semantic Video Indexing, Xiao-Yong Wei, Yu-Gang Jiang, Chong-Wah Ngo
Exploring Inter-Concept Relationship With Context Space For Semantic Video Indexing, Xiao-Yong Wei, Yu-Gang Jiang, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Semantic concept detectors are often individually and independently developed. Using peripherally related concepts for leveraging the power of joint detection, which is referred to as context-based concept fusion (CBCF), has been one of the focus studies in recent years. This paper proposes the construction of a context space and the exploration of the space for CBCF. Context space considers the global consistency of concept relationship, addresses the problem of missing annotation, and is extensible for cross-domain contextual fusion. The space is linear and can be built by modeling the inter-concept relationship through annotation provided by either manual labeling or machine …
Energetic Path Finding Across Massive Terrain Data, Andrew N. Tsui
Energetic Path Finding Across Massive Terrain Data, Andrew N. Tsui
Master's Theses
Before there were airplanes, cars, trains, boats, or bicycles, the primary means of transportation was on foot. Unfortunately, many of the trails used by ancient travelers have long since been abandoned. We present a software tool which can help visualize and predict where these forgotten trails might lie through the use of a human-centered cost metric. By comparing the paths generated by our software with known historical trails, we demonstrate how the tool can indicate likely trails used by ancient travelers. In addition, this new tool provides novel visualizations to better help the user understand alternate paths, effect of terrain, …
Boundless Fluids Using The Lattice-Boltzmann Method, Kyle J. Haughey
Boundless Fluids Using The Lattice-Boltzmann Method, Kyle J. Haughey
Master's Theses
Computer-generated imagery is ubiquitous in today's society, appearing in advertisements, video games, and computer-animated movies among other places. Much of this imagery needs to be as realistic as possible, and animators have turned to techniques such as fluid simulation to create scenes involving substances like smoke, fire, and water. The Lattice-Boltzmann Method (LBM) is one fluid simulation technique that has gained recent popularity due to its relatively simple basic algorithm and the ease with which it can be distributed across multiple processors. Unfortunately, current LBM simulations also suffer from high memory usage and restrict free surface fluids to domains of …
A Revisit Of Generative Model For Automatic Image Annotation Using Markov Random Fields, Yu Xiang, Xiangdong Zhou, Tat-Seng Chua, Chong-Wah Ngo
A Revisit Of Generative Model For Automatic Image Annotation Using Markov Random Fields, Yu Xiang, Xiangdong Zhou, Tat-Seng Chua, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Much research effort on Automatic Image Annotation (AIA) has been focused on Generative Model, due to its well formed theory and competitive performance as compared with many well designed and sophisticated methods. However, when considering semantic context for annotation, the model suffers from the weak learning ability. This is mainly due to the lack of parameter setting and appropriate learning strategy for characterizing the semantic context in the traditional generative model. In this paper, we present a new approach based on Multiple Markov Random Fields (MRF) for semantic context modeling and learning. Differing from previous MRF related AIA approach, we …
Visualizing The Simulation Of 3-D Underwater Sensor Networks, Matthew T. Tran
Visualizing The Simulation Of 3-D Underwater Sensor Networks, Matthew T. Tran
Honors Scholar Theses
The majority of sensor network research deals with land-based networks, which are essentially two-dimensional, and thus the majority of simulation and animation tools also only handle such networks. Underwater sensor networks on the other hand, are essentially 3D networks because the depth at which a sensor node is located needs to be considered as well. Due to that additional dimension, specialized tools need to be used when conducting simulations for experimentation.
The School of Engineering’s Underwater Sensor Network (UWSN) lab is conducting research on underwater sensor networks and requires simulation tools for 3D networks. The lab has extended NS-2, a …
Alternative Task Bar: A Usability Study, Jordan Cote
Alternative Task Bar: A Usability Study, Jordan Cote
Honors Scholar Theses
An alternate approach to the task bar is proposed, one which makes use of modern computers' graphical abilities. This was accomplished with OpenGL, which is typically used in 3D scene generation. The application was integrated into the Desktop Environment in a novel way, in order to produce an arbitrary shaped window. The application was presented to survey participants, who were asked questions to reveal the feasibility of this type of task bar. Responses were positive, which encourages further development.
Real Time Driver Safety System, Gyuchoon Cho
Real Time Driver Safety System, Gyuchoon Cho
Masters Theses & Specialist Projects
The technology for driver safety has been developed in many fields such as airbag system, Anti-lock Braking System or ABS, ultrasonic warning system, and others. Recently, some of the automobile companies have introduced a new feature of driver safety systems. This new system is to make the car slower if it finds a driver’s drowsy eyes. For instance, Toyota Motor Corporation announced that it has given its pre-crash safety system the ability to determine whether a driver’s eyes are properly open with an eye monitor. This paper is focusing on finding a driver’s drowsy eyes by using face detection technology. …
Exploitation Of Geographic Information Systems For Vehicular Destination Prediction, Richard T. Muster
Exploitation Of Geographic Information Systems For Vehicular Destination Prediction, Richard T. Muster
Theses and Dissertations
Much of the recent successes in the Iraqi theater have been achieved with the aid of technology so advanced that celebrated journalist Bob Woodward recently compared it to the Manhattan Project of WWII. Intelligence, Surveillance, and Reconnaissance (ISR) platforms have emerged as the rising star of Air Force operational capabilities as they are enablers in the quest to track and disrupt terrorist and insurgent forces. This thesis argues that ISR systems have been severely under-exploited. The proposals herein seek to improve the machine-human interface of current ISR systems such that a predictive battle-space awareness may be achieved, leading to shorter …
A Framework For Analyzing Biometric Template Aging And Renewal Prediction, John W. Carls
A Framework For Analyzing Biometric Template Aging And Renewal Prediction, John W. Carls
Theses and Dissertations
Biometric technology and systems are modernizing identity capabilities. With maturing biometrics in full, rapid development, a higher accuracy of identity verification is required. An improvement to the security of biometric-based verification systems is provided through higher accuracy; ultimately reducing fraud, theft, and loss of resources from unauthorized personnel. With trivial biometric systems, a higher acceptance threshold to obtain higher accuracy rates increase false rejection rates and user unacceptability. However, maintaining the higher accuracy rate enhances the security of the system. An area of biometrics with a paucity of research is template aging and renewal prediction, specifically in regards to facial …
Architecting Human Operator Trust In Automation To Improve System Effectiveness In Multiple Unmanned Aerial Vehicles (Uav), Eric A. Cring, Adam G. Lenfestey
Architecting Human Operator Trust In Automation To Improve System Effectiveness In Multiple Unmanned Aerial Vehicles (Uav), Eric A. Cring, Adam G. Lenfestey
Theses and Dissertations
Current Unmanned Aerial System (UAS) designs require multiple operators for each vehicle, partly due to imperfect automation matched with the complex operational environment. This study examines the effectiveness of future UAS automation by explicitly addressing the human/machine trust relationship during system architecting. A pedigreed engineering model of trust between human and machine was developed and applied to a laboratory-developed micro-UAS for Special Operations. This unprecedented investigation answered three primary questions. Can previous research be used to create a useful trust model for systems engineering? How can trust be considered explicitly within the DoD Architecture Framework? Can the utility of architecting …
Image Processing For Multiple-Target Tracking On A Graphics Processing Unit, Michael A. Tanner
Image Processing For Multiple-Target Tracking On A Graphics Processing Unit, Michael A. Tanner
Theses and Dissertations
Multiple-target tracking (MTT) systems have been implemented on many different platforms, however these solutions are often expensive and have long development times. Such MTT implementations require custom hardware, yet offer very little flexibility with ever changing data sets and target tracking requirements. This research explores how to supplement and enhance MTT performance with an existing graphics processing unit (GPU) on a general computing platform. Typical computers are already equipped with powerful GPUs to support various games and multimedia applications. However, such GPUs are not currently being used in desktop MTT applications. This research explores if and how a GPU can …
Visual Word Proximity And Linguistics For Semantic Video Indexing And Near-Duplicate Retrieval, Yu-Gang Jiang, Chong-Wah Ngo
Visual Word Proximity And Linguistics For Semantic Video Indexing And Near-Duplicate Retrieval, Yu-Gang Jiang, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Bag-of-visual-words (BoW) has recently become a popular representation to describe video and image content. Most existing approaches, nevertheless, neglect inter-word relatedness and measure similarity by bin-to-bin comparison of visual words in histograms. In this paper, we explore the linguistic and ontological aspects of visual words for video analysis. Two approaches, soft-weighting and constraint-based earth mover’s distance (CEMD), are proposed to model different aspects of visual word linguistics and proximity. In soft-weighting, visual words are cleverly weighted such that the linguistic meaning of words is taken into account for bin-to-bin histogram comparison. In CEMD, a cross-bin matching algorithm is formulated such …
Scale-Rotation Invariant Pattern Entropy For Keypoint-Based Near-Duplicate Detection, Wan-Lei Zhao, Chong-Wah Ngo
Scale-Rotation Invariant Pattern Entropy For Keypoint-Based Near-Duplicate Detection, Wan-Lei Zhao, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Near-duplicate (ND) detection appears as a timely issue recently, being regarded as a powerful tool for various emerging applications. In the Web 2.0 environment particularly, the identification of near-duplicates enables the tasks such as copyright enforcement, news topic tracking, image and video search. In this paper, we describe an algorithm, namely Scale-Rotation invariant Pattern Entropy (SR-PE), for the detection of near-duplicates in large-scale video corpus. SR-PE is a novel pattern evaluation technique capable of measuring the spatial regularity of matching patterns formed by local keypoints. More importantly, the coherency of patterns and the perception of visual similarity, under the scenario …