Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (30)
- Software Engineering (27)
- Artificial Intelligence and Robotics (23)
- Social and Behavioral Sciences (18)
- Engineering (15)
-
- Theory and Algorithms (12)
- Other Computer Sciences (10)
- Programming Languages and Compilers (9)
- Arts and Humanities (8)
- Medicine and Health Sciences (8)
- Communication (7)
- Biomedical Engineering and Bioengineering (5)
- Education (5)
- Systems Architecture (5)
- Computer Engineering (4)
- Disability Studies (4)
- Mathematics (4)
- Numerical Analysis and Scientific Computing (4)
- Anatomy (3)
- Applied Mathematics (3)
- Data Science (3)
- Educational Technology (3)
- Electrical and Computer Engineering (3)
- Information Security (3)
- OS and Networks (3)
- Sense Organs (3)
- Aerospace Engineering (2)
- Institution
-
- Singapore Management University (57)
- City University of New York (CUNY) (9)
- University of Malaya (9)
- Old Dominion University (7)
- University of Arkansas, Fayetteville (4)
-
- Air Force Institute of Technology (3)
- Minnesota State University Moorhead (3)
- University of Central Florida (3)
- Washington University in St. Louis (3)
- California Polytechnic State University, San Luis Obispo (2)
- Montclair State University (2)
- Rochester Institute of Technology (2)
- San Jose State University (2)
- Southern Adventist University (2)
- The University of Akron (2)
- University of Dayton (2)
- Arcadia University (1)
- Central Washington University (1)
- Chapman University (1)
- Claremont Colleges (1)
- College of the Holy Cross (1)
- Dartmouth College (1)
- East Tennessee State University (1)
- Edith Cowan University (1)
- Embry-Riddle Aeronautical University (1)
- Harrisburg University of Science and Technology (1)
- James Madison University (1)
- Kennesaw State University (1)
- Kutztown University (1)
- Portland State University (1)
- Keyword
-
- Accessibility (7)
- Deep learning (4)
- HCI (4)
- Human Computer Interaction (4)
- Machine learning (4)
-
- Visualization (4)
- Eye tracking (3)
- Feature extraction (3)
- Screen reader (3)
- Virtual Reality (3)
- #antcenter (2)
- Android (2)
- Assistive technology (2)
- Augmented reality (2)
- Blink (2)
- Capacitive sensing (2)
- Cognitive performance (2)
- Computer Science (2)
- Computer graphics (2)
- Data visualization (2)
- Disabilities (2)
- Gallium nitride (2)
- Generative adversarial networks (2)
- Gesture recognition (2)
- Graphics (2)
- Human-computer interaction (2)
- Image processing (2)
- Knowledge Graph (2)
- LSTM (2)
- Programming (2)
- Publication
-
- Research Collection School Of Computing and Information Systems (57)
- Student Works (2020-2029) (9)
- Computer Science Faculty Publications (7)
- Open Educational Resources (6)
- Theses and Dissertations (4)
-
- McKelvey School of Engineering Graduate Student Theses & Dissertations (3)
- Student Academic Conference (3)
- Computer Science and Computer Engineering Undergraduate Honors Theses (2)
- Cybersecurity Undergraduate Research Showcase (2)
- Department of Computer Science Faculty Scholarship and Creative Works (2)
- Electronic Theses and Dissertations, 2020-2023 (2)
- Graduate Theses and Dissertations (2)
- MS in Computer Science Project Reports (2)
- Master's Theses (2)
- Williams Honors College, Honors Research Projects (2)
- All Master's Theses (1)
- Articles (1)
- Capstone Showcase (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Information Systems Faculty Publications (Archived) (1)
- Computer Science and Engineering Theses and Dissertations (1)
- Computer Science and Information Technology Faculty (1)
- Dartmouth College Ph.D Dissertations (1)
- Dissertations (1)
- Dissertations and Theses (1)
- Electronic Literature Organization Conference 2020 (1)
- English Honors Theses (1)
- Faculty Publications, Computer Science (1)
- Frameless (1)
- Harrisburg University Presidential Research Grants (1)
- Publication Type
- File Type
Articles 91 - 120 of 135
Full-Text Articles in Graphics and Human Computer Interfaces
Poster Abstract: Data Communication Using Switchable Privacy Glass, Changshuo Hu, Dong Ma, Mahbub Hassan, Wen Hu
Poster Abstract: Data Communication Using Switchable Privacy Glass, Changshuo Hu, Dong Ma, Mahbub Hassan, Wen Hu
Research Collection School Of Computing and Information Systems
Switchable privacy glass can electronically change its state between opaque and transparent. In this work, we propose to exploit the electronic configurability of switchable glass to modulate natural light, which can be demodulated by a nearby receiver with light sensing capability to realise data communication over natural light. A key advantage is that no energy is used to generate light, as it simply modulates the existing light in the nature. We demonstrate that the proposed data communication using switchable glass modulation can achieve 33.33 bits per second communication with a bit rate below 1% under a wide range of ambient …
Reinforced Negative Sampling Over Knowledge Graph For Recommendation, Xiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao, Meng Wang, Tat-Seng Chua
Reinforced Negative Sampling Over Knowledge Graph For Recommendation, Xiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao, Meng Wang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Properly handling missing data is a fundamental challenge in recommendation. Most present works perform negative sampling from unobserved data to supply the training of recommender models with negative signals. Nevertheless, existing negative sampling strategies, either static or adaptive ones, are insufficient to yield high-quality negative samples — both informative to model training and reflective of user real needs. In this work, we hypothesize that item knowledge graph (KG), which provides rich relations among items and KG entities, could be useful to infer informative and factual negative samples. Towards this end, we develop a new negative sampling model, Knowledge Graph Policy …
Techniques To Visualize Occluded Graph Elements For 2.5d Map Editing, Kazuyuki Fujita, Daigo Hayashi, Kotaro Hara, Kazuki Takashima, Yoshifumi Kitamura
Techniques To Visualize Occluded Graph Elements For 2.5d Map Editing, Kazuyuki Fujita, Daigo Hayashi, Kotaro Hara, Kazuki Takashima, Yoshifumi Kitamura
Research Collection School Of Computing and Information Systems
We propose an interface with two novel techniques to visualize occluded graph nodes and edges that help the user edit map data with a 2.5D geographical structure (e.g., multi-floor indoor maps). We first design a visualization technique —Repel Signification— that employs micro-animation to signify the graph elements that are overlapping with each other (and potentially erroneous). We also design a technique that enables the user to edit the occluded components with Expansion Interaction, which simultaneously visualizes both in-floor and across-floor occluded connections between the map elements. The combination of the two methods would enable the map editors (non-experts) to effectively …
The Impact Of Changing The Size Of Aircraft Radar Displays On Visual Search In The Cockpit, Justin R. Marsh
The Impact Of Changing The Size Of Aircraft Radar Displays On Visual Search In The Cockpit, Justin R. Marsh
Theses and Dissertations
Advances in sensor technology have enabled our fighter aircraft to find, fix, track, target, engage (F2T2E) at greater distances, providing the operator with more data within the battlefield. Modern aircraft are designed with larger displays while our legacy aircraft are being retrofitted with larger cockpit displays to enable display of the increased data. While this modification has been shown to enable improvements in human performance of many cockpit tasks, this effect is often not measured nor fully understood at a more generalizable level. This research outlines an approach to comparing human performance across two display sizes in future F-16 cockpits. …
Electronic Image Detectability Under Varying Illumination Conditions, Jeremy J. Miller
Electronic Image Detectability Under Varying Illumination Conditions, Jeremy J. Miller
Theses and Dissertations
Light in the built environment plays an essential role in the vision and the health of humans through non-visual receptors in the eyes. Unfortunately, image analysts and other Air Force personnel who engage in the detection of objects on softcopy displays are often required to work in very dimly-lit or dark environments as higher illumination reduces the contrast of displayed information. Literature has shown that increases in light exposure improves circadian rhythm entrainment and reduces the negative health consequences of insufficient lighting. This research examines the effects of indoor lighting to determine if increases in ambient illumination or changes to …
Image Enhanced Event Detection In News Articles, Meihan Tong, Shuai Wang, Yixin Cao, Bin Xu, Juaizi Li, Lei Hou, Tat-Seng Chua
Image Enhanced Event Detection In News Articles, Meihan Tong, Shuai Wang, Yixin Cao, Bin Xu, Juaizi Li, Lei Hou, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Event detection is a crucial and challenging sub-task of event extraction, which suffers from a severe ambiguity issue of trigger words. Existing works mainly focus on using textual context information, while there naturally exist many images accompanied by news articles that are yet to be explored. We believe that images not only reflect the core events of the text, but are also helpful for the disambiguation of trigger words. In this paper, we first contribute an image dataset supplement to ED benchmarks (i.e., ACE2005) for training and evaluation. We then propose a novel Dual Recurrent Multimodal Model, DRMM, to conduct …
Zero-Shot Ingredient Recognition By Multi-Relational Graph Convolutional Network, Jingjing Chen, Liangming Pan, Zhipeng Wei, Xiang Wang, Chong-Wah Ngo, Tat-Seng Chua
Zero-Shot Ingredient Recognition By Multi-Relational Graph Convolutional Network, Jingjing Chen, Liangming Pan, Zhipeng Wei, Xiang Wang, Chong-Wah Ngo, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Recognizing ingredients for a given dish image is at the core of automatic dietary assessment, attracting increasing attention from both industry and academia. Nevertheless, the task is challenging due to the difficulty of collecting and labeling sufficient training data. On one hand, there are hundred thousands of food ingredients in the world, ranging from the common to rare. Collecting training samples for all of the ingredient categories is difficult. On the other hand, as the ingredient appearances exhibit huge visual variance during the food preparation, it requires to collect the training samples under different cooking and cutting methods for robust …
Gdface: Gated Deformation For Multi-View Face Image Synthesis, Xuemiao Xu, Keke Li, Cheng Xu, Shengfeng He
Gdface: Gated Deformation For Multi-View Face Image Synthesis, Xuemiao Xu, Keke Li, Cheng Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Photorealistic multi-view face synthesis from a single image is an important but challenging problem. Existing methods mainly learn a texture mapping model from the source face to the target face. However, they fail to consider the internal deformation caused by the change of poses, leading to the unsatisfactory synthesized results for large pose variations. In this paper, we propose a Gated Deformable Face Synthesis Network to model the deformation of faces that aids the synthesis of the target face image. Specifically, we propose a dual network that consists of two modules. The first module estimates the deformation of two views …
Example-Based Colourization Via Dense Encoding Pyramids, Chufeng Xiao, Chu Han, Zhuming Zhang, Jing Qin, Tien-Tsin Wong, Guoqiang Han, Shengfeng He
Example-Based Colourization Via Dense Encoding Pyramids, Chufeng Xiao, Chu Han, Zhuming Zhang, Jing Qin, Tien-Tsin Wong, Guoqiang Han, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose a novel deep example-based image colourization method called dense encoding pyramid network. In our study, we define the colourization as a multinomial classification problem. Given a greyscale image and a reference image, the proposed network leverages large-scale data and then predicts colours by analysing the colour distribution of the reference image. We design the network as a pyramid structure in order to exploit the inherent multi-scale, pyramidal hierarchy of colour representations. Between two adjacent levels, we propose a hierarchical decoder–encoder filter to pass the colour distributions from the lower level to higher level in order to take both …
Accessibility Of Deepfakes, Andrew L. Collings
Accessibility Of Deepfakes, Andrew L. Collings
Cybersecurity Undergraduate Research Showcase
The danger posed by falsified media, commonly referred to as deepfakes, has been well researched and documented. The software Faceswap to was used to swap the faces of two politician (Joe Biden and Donald Trump). The testing was performed using an affordable consumer GPU (an AMD Radeon RX 570) over 100,000 iterations. The process and results for the two attempts with the best results (and largest differences) were recorded. The result was ultimately unconvincing, while the software was able to recreate the facial structure the lighting and skin tone did not blend at all.
Building Something With The Raspberry Pi, Richard Kordel
Building Something With The Raspberry Pi, Richard Kordel
Harrisburg University Presidential Research Grants
In 2017 Ryan Korn and I submitted a grant proposal in the annual Harrisburg University President’s Grant process. Our proposal was to partner with a local high school to install a classroom of 20 Raspberry Pi’s, along with the requisite peripherals. In that classroom students would be challenged to design something that combined programming with physical computing. In our presentation to the school we suggested that this project would give students the opportunity to be “amazing.”
As part of the grant, the top three students would be given scholarships to HU and the top five finalists would all be permitted …
Repurposing Visual Input Modalities For Blind Users: A Case Study Of Word Processors, Hae-Na Lee, Vikas Ashok, I.V. Ramakrishnan
Repurposing Visual Input Modalities For Blind Users: A Case Study Of Word Processors, Hae-Na Lee, Vikas Ashok, I.V. Ramakrishnan
Computer Science Faculty Publications
Visual 'point-and-click' interaction artifacts such as mouse and touchpad are tangible input modalities, which are essential for sighted users to conveniently interact with computer applications. In contrast, blind users are unable to leverage these visual input modalities and are thus limited while interacting with computers using a sequentially narrating screen-reader assistive technology that is coupled to keyboards. As a consequence, blind users generally require significantly more time and effort to do even simple application tasks (e.g., applying a style to text in a word processor) using only keyboard, compared to their sighted peers who can effortlessly accomplish the same tasks …
Rotate-And-Press: A Non-Visual Alternative To Point-And-Click, Hae-Na Lee, Vikas Ashok, I. V. Ramakrishnan
Rotate-And-Press: A Non-Visual Alternative To Point-And-Click, Hae-Na Lee, Vikas Ashok, I. V. Ramakrishnan
Computer Science Faculty Publications
Most computer applications manifest visually rich and dense graphical user interfaces (GUIs) that are primarily tailored for an easy-and-efficient sighted interaction using a combination of two default input modalities, namely the keyboard and the mouse/touchpad. However, blind screen-reader users predominantly rely only on keyboard, and therefore struggle to interact with these applications, since it is both arduous and tedious to perform the visual 'point-and-click' tasks such as accessing the various application commands/features using just keyboard shortcuts supported by screen readers.
In this paper, we investigate the suitability of a 'rotate-and-press' input modality as an effective non-visual substitute for the visual …
Svat4: A Computer Program For Visualization And Analysis Of Crystal Structures, Xingzhong Li
Svat4: A Computer Program For Visualization And Analysis Of Crystal Structures, Xingzhong Li
Nebraska Center for Materials and Nanoscience: Faculty Publications
SVAT4 is a computer program for interactive visualization of three-dimensional crystal structures, including chemical bonds and magnetic moments. A wide range of functions, e.g. revealing atomic layers and polyhedral clusters, are available for further structural analysis. Atomic sizes, colors, appearance, view directions and view modes (orthographic or perspective views) are adjustable. Customized work for the visualization and analysis can be saved and then reloaded. SVAT4 provides a template to simplify the process of preparation of a new data file. SVAT4 can generate high-quality images for publication and animations for presentations. The usability of SVAT4 is broadened by a software suite …
Nnv: The Neural Network Verification Tool For Deep Neural Networks And Learning-Enabled Cyber-Physical Systems, Hoang-Dung Tran, Xiaodong Yang, Diego Manzanas Lopez, Patrick Musau, Luan Viet Nguyen, Weiming Xiang, Stanley Bak, Taylor T. Johnson
Nnv: The Neural Network Verification Tool For Deep Neural Networks And Learning-Enabled Cyber-Physical Systems, Hoang-Dung Tran, Xiaodong Yang, Diego Manzanas Lopez, Patrick Musau, Luan Viet Nguyen, Weiming Xiang, Stanley Bak, Taylor T. Johnson
Computer Science Faculty Publications
This paper presents the Neural Network Verification (NNV) software tool, a set-based verification framework for deep neural networks (DNNs) and learning-enabled cyber-physical systems (CPS). The crux of NNV is a collection of reachability algorithms that make use of a variety of set representations, such as polyhedra, star sets, zonotopes, and abstract-domain representations. NNV supports both exact (sound and complete) and over-approximate (sound) reachability algorithms for verifying safety and robustness properties of feed-forward neural networks (FFNNs) with various activation functions. For learning-enabled CPS, such as closed-loop control systems incorporating neural networks, NNV provides exact and over-approximate reachability analysis schemes for linear …
Intelligent Cinematic Camera Control For Real-Time Graphics Applications, Ian Harris Meeder
Intelligent Cinematic Camera Control For Real-Time Graphics Applications, Ian Harris Meeder
Master's Theses
E-sports is currently estimated to be a billion dollar industry which is only growing in size from year to year. However the cinematography of spectated games leaves much to be desired. In most cases, the spectator either gets to control their own freely-moving camera or they get to see the view that a specific player sees. This thesis presents a system for the generation of cinematically-pleasing views for spectating real-time graphics applications. A custom real-time engine has been built to demonstrate the effect of this system on several different game modes with varying visual cinematic constraints, such as the rule …
Security Camera Using Raspberry Pi, Tejendra Khatri
Security Camera Using Raspberry Pi, Tejendra Khatri
Student Academic Conference
Making a security camera using raspberry pi utilizing OpenCV for facial recognition, upper body recognition or full-body recognition
Navigating Immersive And Interactive Vr Environments With Connected 360° Panoramas, Samuel Cosgrove
Navigating Immersive And Interactive Vr Environments With Connected 360° Panoramas, Samuel Cosgrove
Electronic Theses and Dissertations, 2020-2023
Emerging research is expanding the idea of using 360-degree spherical panoramas of real-world environments for use in "360 VR" experiences beyond video and image viewing. However, most of these experiences are strictly guided, with few opportunities for interaction or exploration. There is a desire to develop experiences with cohesive virtual environments created with 360 VR that allow for choice in navigation, versus scripted experiences with limited interaction. Unlike standard VR with the freedom of synthetic graphics, there are challenges in designing appropriate user interfaces (UIs) for 360 VR navigation within the limitations of fixed assets. To tackle this gap, we …
“Distance Learning” In The Ninth Century?: Micro-Cluster Analysis Of The Epistolary Network Of Alcuin After 796, William James Mattingly
“Distance Learning” In The Ninth Century?: Micro-Cluster Analysis Of The Epistolary Network Of Alcuin After 796, William James Mattingly
Theses and Dissertations--History
Scholars of eighth- and ninth-century education have assumed that intellectuals did not write works of Scriptural interpretation until that intellectual had a firm foundation in the seven liberal arts.This ensured that anyone who embarked on work of Scriptural interpretation would have the required knowledge and methods to read and interpret Scripture correctly. The potential for theological error and the transmission of those errors was too great unless the interpreter had the requisite training. This dissertation employs computistical methods, specifically the techniques of social network mapping and cluster analysis, to study closely the correspondence of Alcuin, a late-eighth- and early-ninth-century scholar …
Needfinding, Devorah Kletenik
Needfinding, Devorah Kletenik
Open Educational Resources
This activity guides students through the process needfinding to identify areas of need for their creation of a technology for the "public good." Students will conduct contextual inquiry to identify the needs of their target audience.
Personas, Scenarios And Storyboards, Devorah Kletenik
Personas, Scenarios And Storyboards, Devorah Kletenik
Open Educational Resources
This activity guides students towards the creation of personas, scenarios and storyboards for a product/website that they are creating.
Public Interest Technology: Coding For The Public Good, Devorah Kletenik
Public Interest Technology: Coding For The Public Good, Devorah Kletenik
Open Educational Resources
These slides are used to guide a discussion with students introducing them to the notion of public interest technology and coding for the public good. The lesson is intended to spark a discussion with students about different sorts of technology and their societal ramifications.
Accessibility: The Whys And The Hows, Devorah Kletenik
Accessibility: The Whys And The Hows, Devorah Kletenik
Open Educational Resources
This presentation introduces Computer Science students to the notion of accessibility: developing software for people with disabilities. This lesson provides a discussion of why accessibility is important (including the legal, societal and ethical benefits) as well as an overview of different types of impairments (visual, auditory, motor, neurological/cognitive) and how developers can make their software accessible to users with those disabilities. This lesson includes videos and links to readings and tutorials for students.
Coding For The Public Good: Front-End Website Design And Development, Devorah Kletenik
Coding For The Public Good: Front-End Website Design And Development, Devorah Kletenik
Open Educational Resources
This activity helps student design and develop a front-end of a website, from wireframes through HTML/CSS/Javascript. It includes design questions for students, including the invocation of Ben Schneiderman's eight golden rules for interface design.
Note: this activity assumes prior knowledge of web development. Since this activity is designed for an HCI course, with a focus on interface design, students are not expected to create a back-end for it. This activity can obviously be modified for a full-stack experience.
Accessibility Evaluation, Devorah Kletenik
Accessibility Evaluation, Devorah Kletenik
Open Educational Resources
This activity guides students through the evaluation of a website that they have created to see if it is accessible for users with disabilities. Students will simulate a number of different disabilities (e.g. visual impairments, color blindness, auditory impairments, motor impairments) to see if their website is accessible; they will also use automated W3 and WAVE tools to evaluate their sites. Students will consider the needs of users with disabilities by creating a persona and scenario of a user with disabilities interacting with their site. Finally, students will write up recommendations to change their site and implement the changes.
V-Slam And Sensor Fusion For Ground Robots, Ejup Hoxha
V-Slam And Sensor Fusion For Ground Robots, Ejup Hoxha
Dissertations and Theses
In underground, underwater and indoor environments, a robot has to rely solely on its on-board sensors to sense and understand its surroundings. This is the main reason why SLAM gained the popularity it has today. In recent years, we have seen excellent improvement on accuracy of localization using cameras and combinations of different sensors, especially camera-IMU (VIO) fusion. Incorporating more sensors leads to improvement of accuracy,but also robustness of SLAM. However, while testing SLAM in our ground robots, we have seen a decrease in performance quality when using the same algorithms on flying vehicles.We have an additional sensor for ground …
Automatic Gaze Classification For Aviators: Using Multi-Task Convolutional Networks As A Proxy For Flight Instructor Observation, Justin Wilson, Sandro Scielzo, Sukumaran Nair, Eric C. Larson
Automatic Gaze Classification For Aviators: Using Multi-Task Convolutional Networks As A Proxy For Flight Instructor Observation, Justin Wilson, Sandro Scielzo, Sukumaran Nair, Eric C. Larson
International Journal of Aviation, Aeronautics, and Aerospace
In this work, we investigate how flight instructors observe aviator scan patterns and assign quality to an aviator's gaze. We first establish the reliability of instructors to assign similar quality to an aviator's scan patterns, and then investigate methods to automate this quality using machine learning. In particular, we focus on the classification of gaze for aviators in a mixed-reality flight simulation. We create and evaluate two machine learning models for classifying gaze quality of aviators: a task-agnostic model and a multi-task model. Both models use deep convolutional neural networks to classify the quality of pilot gaze patterns for 40 …
Interactions Between Humans, Virtual Agent Characters And Virtual Avatars, Tamara Griffith
Interactions Between Humans, Virtual Agent Characters And Virtual Avatars, Tamara Griffith
Electronic Theses and Dissertations, 2020-2023
Simulations allow people to experience events as if they were happening in the real world in a way that is safer and less expensive than live training. Despite improvements in realism in simulated environments, one area that still presents a challenge is interpersonal interactions. The subtleties of what makes an interaction rich are difficult to define. We may never fully understand the complexity of human interchanges, however there is value in building on existing research into how individuals react to virtual characters to inform future investments. Virtual characters can either be automated through computational processes, referred to as agents, or …
Toward Efficient Automation Of Interpretable Machine Learning Boosting, Nathan Neuhaus
Toward Efficient Automation Of Interpretable Machine Learning Boosting, Nathan Neuhaus
All Master's Theses
Developing efficient automated methods for Interpretable Machine Learning (IML) is an important and long-term goal in the field of Artificial Intelligence. Currently the Machine Learning landscape is dominated by Neural Networks (NNs) and Support Vector Machines (SVMs), models which are often highly accurate. Despite high accuracy, such models are essentially “black boxes” and therefore are too risky for situations like healthcare where real lives are at stake. In such situations, so called “glass-box” models, such as Decision Trees (DTs), Bayesian Networks (BNs), and Logic Relational (LR) models are often preferred, however can succumb to accuracy limitations. Unfortunately, having to choose …
Neighbourhood Structure Preserving Cross-Modal Embedding For Video Hyperlinking, Yanbin Hao, Chong-Wah Ngo, Benoit Huet
Neighbourhood Structure Preserving Cross-Modal Embedding For Video Hyperlinking, Yanbin Hao, Chong-Wah Ngo, Benoit Huet
Research Collection School Of Computing and Information Systems
Video hyperlinking is a task aiming to enhance the accessibility of large archives, by establishing links between fragments of videos. The links model the aboutness between fragments for efficient traversal of video content. This paper addresses the problem of link construction from the perspective of cross-modal embedding. To this end, a generalized multi-modal auto-encoder is proposed.& x00A0;The encoder learns two embeddings from visual and speech modalities, respectively, whereas each of the embeddings performs self-modal and cross-modal translation of modalities. Furthermore, to preserve the neighbourhood structure of fragments, which is important for video hyperlinking, the auto-encoder is devised to model data …