Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (942)
- University of Dayton (114)
- Old Dominion University (99)
- Air Force Institute of Technology (98)
- California Polytechnic State University, San Luis Obispo (96)
-
- University of Arkansas, Fayetteville (89)
- University of Nebraska - Lincoln (51)
- City University of New York (CUNY) (48)
- Technological University Dublin (48)
- University of Malaya (43)
- San Jose State University (37)
- Dartmouth College (35)
- Embry-Riddle Aeronautical University (24)
- Clemson University (23)
- Purdue University (23)
- Rochester Institute of Technology (23)
- The University of Akron (22)
- Chapman University (20)
- Edith Cowan University (20)
- University of Kentucky (18)
- Michigan Technological University (16)
- University of Central Florida (15)
- Southern Adventist University (13)
- California State University, San Bernardino (12)
- Kennesaw State University (12)
- St. Mary's University (12)
- Nova Southeastern University (11)
- University of Minnesota Morris Digital Well (11)
- Louisiana State University (10)
- University of Nevada, Las Vegas (10)
- Keyword
-
- Virtual reality (62)
- Visualization (46)
- Computer graphics (38)
- Computer vision (37)
- Accessibility (36)
-
- Human-computer interaction (35)
- Augmented reality (33)
- Usability (31)
- Machine learning (29)
- Machine Learning (26)
- Artificial intelligence (25)
- Computer Science (25)
- Data visualization (25)
- Deep learning (24)
- Virtual Reality (23)
- HCI (22)
- Computer science (21)
- Eye tracking (20)
- Human computer interaction (20)
- User experience (20)
- Design (19)
- Deep Learning (16)
- Education (16)
- Feature extraction (15)
- Graph Neural Networks (15)
- Graphics (15)
- VR (15)
- Artificial Intelligence (14)
- Gamification (14)
- Image processing (14)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (916)
- Computer Science Faculty Publications (135)
- Theses and Dissertations (98)
- Master's Theses (50)
- Graduate Theses and Dissertations (43)
-
- Student Works (2000-2009) (33)
- Computer Science and Computer Engineering Undergraduate Honors Theses (31)
- 3-D Printed Model Structural Files (29)
- Publications and Research (28)
- Dartmouth College Master’s Theses (24)
- Williams Honors College, Honors Research Projects (22)
- Master's Projects (20)
- Conference papers (19)
- Frameless (19)
- Computer Science and Software Engineering (18)
- All Dissertations (17)
- Dissertations and Theses Collection (Open Access) (16)
- Dissertations, Master's Theses and Master's Reports (16)
- H-Workload 2017: Models and Applications (Works in Progress) (15)
- Computer Engineering (14)
- Electronic Theses and Dissertations (14)
- Theses : Honours (14)
- Honors Theses (13)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- AFIT Patents (11)
- CCAC Theses and Dissertations (11)
- Engineering Faculty Articles and Research (11)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (10)
- Student Works (2020-2029) (10)
- Inquiry: The University of Arkansas Undergraduate Research Journal (9)
- Publication Type
- File Type
Articles 421 - 450 of 2372
Full-Text Articles in Computer Sciences
The Impact Of Avatar Completeness On Embodiment And The Detectability Of Hand Redirection In Virtual Reality, Martin Feick, Andre Zenner, Simon Seibert, Anthony Tang, Antonio Krüger
The Impact Of Avatar Completeness On Embodiment And The Detectability Of Hand Redirection In Virtual Reality, Martin Feick, Andre Zenner, Simon Seibert, Anthony Tang, Antonio Krüger
Research Collection School Of Computing and Information Systems
To enhance interactions in VR, many techniques introduce offsets between the virtual and real-world position of users’ hands. Nevertheless, such hand redirection (HR) techniques are only effective as long as they go unnoticed by users—not disrupting the VR experience. While several studies consider how much unnoticeable redirection can be applied, these focus on mid-air floating hands that are disconnected from users’ bodies. Increasingly, VR avatars are embodied as being directly connected with the user’s body, which provide more visual cue anchoring, and may therefore reduce the unnoticeable redirection threshold. In this work, we studied more complete avatars and their effect …
Vaid: Indexing View Designs In Visual Analytics System, Lu Ying, Aoyu Wu, Haotian Li, Zikun Deng, Ji Lan, Jiang Wu, Yong Wang, Huamin Qu, Dazhen Deng, Yingcai Wu
Vaid: Indexing View Designs In Visual Analytics System, Lu Ying, Aoyu Wu, Haotian Li, Zikun Deng, Ji Lan, Jiang Wu, Yong Wang, Huamin Qu, Dazhen Deng, Yingcai Wu
Research Collection School Of Computing and Information Systems
Visual analytics (VA) systems have been widely used in various application domains. However, VA systems are complex in design, which imposes a serious problem: although the academic community constantly designs and implements new designs, the designs are difficult to query, understand, and refer to by subsequent designers. To mark a major step forward in tackling this problem, we index VA designs in an expressive and accessible way, transforming the designs into a structured format. We first conducted a workshop study with VA designers to learn user requirements for understanding and retrieving professional designs in VA systems. Thereafter, we came up …
Learning Nighttime Semantic Segmentation The Hard Way, Wenxi Liu, Jiaxin Cai, Qi Li, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu
Learning Nighttime Semantic Segmentation The Hard Way, Wenxi Liu, Jiaxin Cai, Qi Li, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu
Research Collection School Of Computing and Information Systems
Nighttime semantic segmentation is an important but challenging research problem for autonomous driving. The major challenges lie in the small objects or regions from the under-/over-exposed areas or suffer from motion blur caused by the camera deployed on moving vehicles. To resolve this, we propose a novel hard- class-aware module that bridges the main network for full-class segmentation and the hard-class network for segmenting aforementioned hard-class objects. In specific, it exploits the shared focus of hard-class objects from the dual-stream network, enabling the contextual information flow to guide the model to concentrate on the pixels that are hard to classify. …
Deep Reinforcement Learning Guided Improvement Heuristic For Job Shop Scheduling, Cong Zhang, Zhiguang Cao, Wen Song, Yaoxin Wu, Jie Zhang
Deep Reinforcement Learning Guided Improvement Heuristic For Job Shop Scheduling, Cong Zhang, Zhiguang Cao, Wen Song, Yaoxin Wu, Jie Zhang
Research Collection School Of Computing and Information Systems
Recent studies in using deep reinforcement learning (DRL) to solve Job-shop scheduling problems (JSSP) focus on construction heuristics. However, their performance is still far from optimality, mainly because the underlying graph representation scheme is unsuitable for modelling partial solutions at each construction step. This paper proposes a novel DRL-guided improvement heuristic for solving JSSP, where graph representation is employed to encode complete solutions. We design a Graph-Neural-Network-based representation scheme, consisting of two modules to effectively capture the information of dynamic topology and different types of nodes in graphs encountered during the improvement process. To speed up solution evaluation during improvement, …
Fashionregen: Llm‑Empowered Fashion Report Generation, Yujuan Ding, Yunshan Ma, Wenqi Fan, Yige Yao, Tat‑Seng Chua, Qing Li
Fashionregen: Llm‑Empowered Fashion Report Generation, Yujuan Ding, Yunshan Ma, Wenqi Fan, Yige Yao, Tat‑Seng Chua, Qing Li
Research Collection School Of Computing and Information Systems
Fashion analysis refers to the process of examining and evaluating trends, styles, and elements within the fashion industry to understand and interpret its current state, generating fashion reports. It is traditionally performed by fashion professionals based on their expertise and experience, which requires high labour cost and may also produce biased results for relying heavily on a small group of people. In this paper, to tackle the Fashion Report Generation (FashionReGen) task, we propose an intelligent Fashion Analyzing and Reporting system based the advanced Large Language Models (LLMs), debbed as GPT-FAR. Specifically, it tries to deliver FashionReGen based on effective …
An Empirical Study On The Efficacy Of Llm-Powered Chatbots In Basic Information Retrieval Tasks, Naja Faysal
An Empirical Study On The Efficacy Of Llm-Powered Chatbots In Basic Information Retrieval Tasks, Naja Faysal
Electronic Theses, Projects, and Dissertations
The rise of conversational user interfaces (CUIs) powered by large language models (LLMs) is transforming human-computer interaction. This study evaluates the efficacy of LLM-powered chatbots, trained on website data, compared to browsing websites for finding information about organizations across diverse sectors. A within-subjects experiment with 165 participants was conducted, involving similar information retrieval (IR) tasks using both websites (GUIs) and chatbots (CUIs). The research questions are: (Q1) Which interface helps users find information faster: LLM chatbots or websites? (Q2) Which interface helps users find more accurate information: LLM chatbots or websites?. The findings are: (Q1) Participants found information significantly faster …
Binder, Tyler A. Peaster, Lindsey M. Davenport, Madelyn Little, Alex Bales
Binder, Tyler A. Peaster, Lindsey M. Davenport, Madelyn Little, Alex Bales
ATU Scholars Symposium
Binder is a mobile application that aims to introduce readers to a book recommendation service that appeals to devoted and casual readers. The main goal of Binder is to enrich book selection and reading experience. This project was created in response to deficiencies in the mobile space for book suggestions, library management, and reading personalization. The tools we used to create the project include Visual Studio, .Net Maui Framework, C#, XAML, CSS, MongoDB, NoSQL, Git, GitHub, and Figma. The project’s selection of books were sourced from the Google Books repository. Binder aims to provide an intuitive interface that allows users …
Factors Influencing The Perceptions Of Human-Computer Interaction Curriculum Developers In Higher Education Institutions During Curriculum Design And Delivery, Cynthia Augustine, Salah Kabanda
Factors Influencing The Perceptions Of Human-Computer Interaction Curriculum Developers In Higher Education Institutions During Curriculum Design And Delivery, Cynthia Augustine, Salah Kabanda
The African Journal of Information Systems
Computer science (CS) and information systems students seeking to work as software developers upon graduating are often required to create software that has a sound user experience (UX) and meets the needs of its users. This includes addressing unique user, context, and infrastructural requirements. This study sought to identify the factors that influence the perceptions of human-computer interaction (HCI) curriculum developers in higher education institutions (HEIs) in developing economies of Africa when it comes to curriculum design and delivery. A qualitative enquiry was conducted and consisted of fourteen interviews with HCI curriculum developers and UX practitioners in four African countries. …
Immersive Japanese Language Learning Web Application Using Spaced Repetition, Active Recall, And An Artificial Intelligent Conversational Chat Agent Both In Voice And In Text, Marc Butler
MS in Computer Science Project Reports
In the last two decades various human language learning applications, spaced repetition software, online dictionaries, and artificial intelligent chat agents have been developed. However, there is no solution to cohesively combine these technologies into a comprehensive language learning application including skills such as speaking, typing, listening, and reading. Our contribution is to provide an immersive language learning web application to the end user which combines spaced repetition, a study technique used to review information at systematic intervals, and active recall, the process of purposely retrieving information from memory during a review session, with an artificial intelligent conversational chat agent both …
Image De‑Photobombing Benchmark, Vatsa S. Patel, Kunal Agrawal, Samah Baraheem, Amira Yousif, Tam Nguyen
Image De‑Photobombing Benchmark, Vatsa S. Patel, Kunal Agrawal, Samah Baraheem, Amira Yousif, Tam Nguyen
Computer Science Faculty Publications
Removing photobombing elements from images is a challenging task that requires sophisticated image inpainting techniques. Despite the availability of various methods, their effectiveness depends on the complexity of the image and the nature of the distracting element. To address this issue, we conducted a benchmark study to evaluate 10 state-of-the-art photobombing removal methods on a dataset of over 300 images. Our study focused on identifying the most effective image inpainting techniques for removing unwanted regions from images. We annotated the photobombed regions that require removal and evaluated the performance of each method using peak signal-to-noise ratio (PSNR), structural similarity index …
Improving Implicit Communication In Remote Collaboration Through Augmented Reality And Digital Twins, Nicholas Levergne
Improving Implicit Communication In Remote Collaboration Through Augmented Reality And Digital Twins, Nicholas Levergne
LSU Master's Theses
Large scale digital twinning projects are beginning to emerge across the tech industry. Within these projects is a desire to integrate augmented reality capabilities into industrial workflows. However, research on augmented reality technology for remote collaboration lacks ecologically valid studies of real world scenarios. Additionally, prior remote collaboration literature is focused on white-collar applications instead of blue-collar field work. Prior AR collaboration software is similarly limited, with most software allowing mixed camera views and annotation that requires participants to be stationary. This thesis introduces SpectAR, an augmented reality and desktop remote collaboration software suite developed in Unreal Engine 5.1.1. SpectAR …
Visualizing Routes With Ai-Discovered Street-View Patterns, Tsung Heng Wu, Md Amiruzzaman, Ye Zhao, Deepshikha Bhati, Jing Yang
Visualizing Routes With Ai-Discovered Street-View Patterns, Tsung Heng Wu, Md Amiruzzaman, Ye Zhao, Deepshikha Bhati, Jing Yang
Computer Science Faculty Publications
Street-level visual appearances play an important role in studying social systems, such as understanding the built environment, driving routes, and associated social and economic factors. It has not been integrated into a typical geographical visualization interface (e.g., map services) for planning driving routes. In this article, we study this new visualization task with several new contributions. First, we experiment with a set of AI techniques and propose a solution of using semantic latent vectors for quantifying visual appearance features. Second, we calculate image similarities among a large set of street-view images and then discover spatial imagery patterns. Third, we integrate …
Exploring Neural Networks For Developing A Chess Learning Platform With Integrated Ai Agent, Lauren Escobedo
Exploring Neural Networks For Developing A Chess Learning Platform With Integrated Ai Agent, Lauren Escobedo
Posters - 2024
Chess is a highly strategic, complex, and long-form game that has been popular for many centuries. Due to the aforementioned complexities of this game, new players often have a hard time learning how to effectively and successfully play. With the recent developments in machine learning algorithms, new opportunities arise to create artificially intelligent tutors - not only for chess, but for all subjects. This project aims to develop a product which investigates the integration of an artificially intelligent “coach”, named Chesster, to train the player, which is trained on a neural network machine learning algorithm.
Nookipedia, Shyann Francis
Nookipedia, Shyann Francis
Posters - 2024
Animal Crossing involves managing around 9000 items within a gameplay environment. This abundance of items makes it difficult for players to effectively track progress and navigate through their possessions. The primary motivation of Nookipedia is to enhance overall gameplay experience by streamlining achievement management, thereby reducing unnecessary interactions with non-playable characters (NPCs) and enabling players to focus more on core game activities. By improving inventory organization and reducing clutter, the aim is to create a smoother and more immersive gaming experience for players.
Lyraquist: Language Learning Via Music App, Vivian D'Souza, Siri Avula, Mahi Patel, Tanvi Singh, Ashley Bickham
Lyraquist: Language Learning Via Music App, Vivian D'Souza, Siri Avula, Mahi Patel, Tanvi Singh, Ashley Bickham
Senior Theses
Lyraquist is a new language learning mobile app that encourages the practice of a foreign language through music. Language learners can connect their Spotify Premium account to Lyraquist to listen to music in their target languages and utilize tools such as translation and vocabulary lists to facilitate language practice. Through integration with Spotify, users can import and create Spotify playlists and search the service’s entire catalog. By combining daily listening habits with several tasks associated with language learning in one place, Lyraquist hopes to be a useful language learning tool.
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research, Musa M. Azeem
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research, Musa M. Azeem
Senior Theses
The USC Faculty Dashboard is a web application designed to revolutionize how department heads, professors, and instructors monitor progress and make decisions, providing a centralized hub for efficient data storage and analysis. Currently, there’s a gap in tools tailored for department heads to concisely manage the performance of their department, which our platform aims to fill. The USC Faculty Dashboard offers easy access to upload and view student evaluation and research information, empowering department heads to evaluate the performance of faculty members and seamlessly track their research grants, publications, and expenditures. Furthermore, professors and instructors gain personalized performance analysis tools, …
Hop‑Based Heterogeneous Graph Transformer, Zixuan Yang, Xiao Wang, Yanhua Yu, Yuling Wang, Kangkang Lu, Zirui Guo, Xiting Qin, Yunshan Ma, Tat‑Seng Chua
Hop‑Based Heterogeneous Graph Transformer, Zixuan Yang, Xiao Wang, Yanhua Yu, Yuling Wang, Kangkang Lu, Zirui Guo, Xiting Qin, Yunshan Ma, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
The Graph Transformer (GT) has shown significant ability in processing graph-structured data, addressing limitations in graph neural networks, such as over-smoothing and over-squashing. However, the implementation of GT in real-world heterogeneous graphs (HGs) with complex topology continues to present numerous challenges. Firstly, a challenge arises in designing a tokenizer that is compatible with heterogeneity. Secondly, the complexity of the transformer hampers the acquisition of high-order neighbor information in HGs. In this paper, we propose a novel Hop-basedHeterogeneous Graph Transformer (H2Gormer) framework, paving a promising path for HGs to benefit from the capabilities of Transformers. We propose a Heterogeneous Hop-based Token …
Coca: Improving And Explaining Graph Neural Network-Based Vulnerability Detection Systems, Sicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo, Lili Bo, Bin Li, Wei Liu
Coca: Improving And Explaining Graph Neural Network-Based Vulnerability Detection Systems, Sicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo, Lili Bo, Bin Li, Wei Liu
Research Collection School Of Computing and Information Systems
Recently, Graph Neural Network (GNN)-based vulnerability detection systems have achieved remarkable success. However, the lack of explainability poses a critical challenge to deploy black-box models in security-related domains. For this reason, several approaches have been proposed to explain the decision logic of the detection model by providing a set of crucial statements positively contributing to its predictions. Unfortunately, due to the weakly-robust detection models and suboptimal explanation strategy, they have the danger of revealing spurious correlations and redundancy issue.In this paper, we propose Coca, a general framework aiming to 1) enhance the robustness of existing GNN-based vulnerability detection models to …
Terry Riley's "In C" For Mobile Ensemble, David B. Wetzel, Griffin Moe, George K. Thiruvathukal
Terry Riley's "In C" For Mobile Ensemble, David B. Wetzel, Griffin Moe, George K. Thiruvathukal
Computer Science: Faculty Publications and Other Works
This workshop presents a mobile-friendly Web Audio application for a “technology ensemble play-along” of Terry Riley’s 1964 composition In C. Attendees will join in a reading of In C using available web-enabled devices as musical instruments. We hope to demonstrate an accessible music-technology experience that relies on face-to-face interaction within a shared space. In this all-electronic implementation, no special musical or technical expertise is required.
Test-Time Augmentation For 3d Point Cloud Classification And Segmentation, Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung
Test-Time Augmentation For 3d Point Cloud Classification And Segmentation, Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung
Research Collection School Of Computing and Information Systems
Data augmentation is a powerful technique to enhance the performance of a deep learning task but has received less attention in 3D deep learning. It is well known that when 3D shapes are sparsely represented with low point density, the performance of the downstream tasks drops significantly. This work explores test-time augmentation (TTA) for 3D point clouds. We are inspired by the recent revolution of learning implicit representation and point cloud upsampling, which can produce high-quality 3D surface reconstruction and proximity-to-surface, respectively. Our idea is to leverage the implicit field reconstruction or point cloud upsampling techniques as a systematic way …
Iterative Graph Self-Distillation, Hanlin Zhang, Shuai Lin, Weiyang Liu, Pan Zhou, Jian Tang, Xiaodan Liang, Eric Xing
Iterative Graph Self-Distillation, Hanlin Zhang, Shuai Lin, Weiyang Liu, Pan Zhou, Jian Tang, Xiaodan Liang, Eric Xing
Research Collection School Of Computing and Information Systems
Recently, there has been increasing interest in the challenge of how to discriminatively vectorize graphs. To address this, we propose a method called Iterative Graph Self-Distillation (IGSD) which learns graph-level representation in an unsupervised manner through instance discrimination using a self-supervised contrastive learning approach. IGSD involves a teacher-student distillation process that uses graph diffusion augmentations and constructs the teacher model using an exponential moving average of the student model. The intuition behind IGSD is to predict the teacher network representation of the graph pairs under different augmented views. As a natural extension, we also apply IGSD to semi-supervised scenarios by …
Transiam: Aggregating Multi-Modal Visual Features With Locality For Medical Image Segmentation, Xuejian Li, Shiqiang Ma, Junhai Xu, Jijun Tang, Shengfeng He, Fei Guo
Transiam: Aggregating Multi-Modal Visual Features With Locality For Medical Image Segmentation, Xuejian Li, Shiqiang Ma, Junhai Xu, Jijun Tang, Shengfeng He, Fei Guo
Research Collection School Of Computing and Information Systems
Automatic segmentation of medical images plays an important role in the diagnosis of diseases. On single-modal data, convolutional neural networks have demonstrated satisfactory performance. However, multi-modal data encompasses a greater amount of information rather than single-modal data. Multi-modal data can be effectively used to improve the segmentation accuracy of regions of interest by analyzing both spatial and temporal information. In this study, we propose a dual-path segmentation model for multi-modal medical images, named TranSiam. Taking into account that there is a significant diversity between the different modalities, TranSiam employs two parallel CNNs to extract the features which are specific to …
Towards Understanding Convergence And Generalization Of Adamw, Pan Zhou, Xingyu Xie, Zhouchen Lin, Shuicheng Yan
Towards Understanding Convergence And Generalization Of Adamw, Pan Zhou, Xingyu Xie, Zhouchen Lin, Shuicheng Yan
Research Collection School Of Computing and Information Systems
AdamW modifies Adam by adding a decoupled weight decay to decay network weights per training iteration. For adaptive algorithms, this decoupled weight decay does not affect specific optimization steps, and differs from the widely used ℓ2-regularizer which changes optimization steps via changing the first- and second-order gradient moments. Despite its great practical success, for AdamW, its convergence behavior and generalization improvement over Adam and ℓ2-regularized Adam (ℓ2-Adam) remain absent yet. To solve this issue, we prove the convergence of AdamW and justify its generalization advantages over Adam and ℓ2-Adam. Specifically, AdamW provably converges but minimizes a dynamically regularized loss that …
Delving Into Multimodal Prompting For Fine-Grained Visual Classification, Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du, Shengfeng He, Zechao Li
Delving Into Multimodal Prompting For Fine-Grained Visual Classification, Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du, Shengfeng He, Zechao Li
Research Collection School Of Computing and Information Systems
Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts …
What Does One Billion Dollars Look Like?: Visualizing Extreme Wealth, William Mahoney Luckman
What Does One Billion Dollars Look Like?: Visualizing Extreme Wealth, William Mahoney Luckman
Dissertations, Theses, and Capstone Projects
The word “billion” is a mathematical abstraction related to “big,” but it is difficult to understand the vast difference in value between one million and one billion; even harder to understand the vast difference in purchasing power between one billion dollars, and the average U.S. yearly income. Perhaps most difficult to conceive of is what that purchasing power and huge mass of capital translates to in terms of power. This project blends design, text, facts, and figures into an interactive narrative website that helps the user better understand their position in relation to extreme wealth: https://whatdoesonebilliondollarslooklike.website/
The site incorporates …
From Canteen Food To Daily Meals: Generalizing Food Recognition To More Practical Scenarios, Guoshan Liu, Yang Jiao, Jingjing Chen, Bin Zhu, Yu-Gang Jiang
From Canteen Food To Daily Meals: Generalizing Food Recognition To More Practical Scenarios, Guoshan Liu, Yang Jiao, Jingjing Chen, Bin Zhu, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
The precise recognition of food categories plays a pivotal role for intelligent health management, attracting significant research attention in recent years. Prominent benchmarks, such as Food-101 and VIREO Food-172, provide abundant food image resources that catalyze the prosperity of research in this field. Nevertheless, these datasets are well-curated from canteen scenarios and thus deviate from food appearances in daily life. This discrepancy poses great challenges in effectively transferring classifiers trained on these canteen datasets to broader daily-life scenarios encountered by humans. Toward this end, we present two new benchmarks, namely DailyFood-172 and DailyFood-16, specifically designed to curate food images from …
M3sa: Multimodal Sentiment Analysis Based On Multi-Scale Feature Extraction And Multi-Task Learning, Changkai Lin, Hongju Cheng, Qiang Rao, Yang Yang
M3sa: Multimodal Sentiment Analysis Based On Multi-Scale Feature Extraction And Multi-Task Learning, Changkai Lin, Hongju Cheng, Qiang Rao, Yang Yang
Research Collection School Of Computing and Information Systems
Sentiment analysis plays an indispensable part in human-computer interaction. Multimodal sentiment analysis can overcome the shortcomings of unimodal sentiment analysis by fusing multimodal data. However, how to extracte improved feature representations and how to execute effective modality fusion are two crucial problems in multimodal sentiment analysis. Traditional work uses simple sub-models for feature extraction, and they ignore features of different scales and fuse different modalities of data equally, making it easier to incorporate extraneous information and affect analysis accuracy. In this paper, we propose a Multimodal Sentiment Analysis model based on Multi-scale feature extraction and Multi-task learning (M 3 SA). …
Efficient Unsupervised Video Hashing With Contextual Modeling And Structural Controlling, Jingru Duan, Yanbin Hao, Bin Zhu, Lechao Cheng, Pengyuan Zhou, Xiang Wang
Efficient Unsupervised Video Hashing With Contextual Modeling And Structural Controlling, Jingru Duan, Yanbin Hao, Bin Zhu, Lechao Cheng, Pengyuan Zhou, Xiang Wang
Research Collection School Of Computing and Information Systems
The most important effect of the video hashing technique is to support fast retrieval, which is benefiting from the high efficiency of binary calculation. Current video hash approaches are thus mainly targeted at learning compact binary codes to represent video content accurately. However, they may overlook the generation efficiency for hash codes, i.e., designing lightweight neural networks. This paper proposes an method, which is not only for computing compact hash codes but also for designing a lightweight deep model. Specifically, we present an MLP-based model, where the video tensor is split into several groups and multiple axial contexts are explored …
Learning From Machines: How Negative Feedback From Machines Improves Learning Between Humans, Tengjian Zou, Gokhan Ertug, Thomas Roulet
Learning From Machines: How Negative Feedback From Machines Improves Learning Between Humans, Tengjian Zou, Gokhan Ertug, Thomas Roulet
Research Collection Lee Kong Chian School Of Business
Prior studies on learning from failure primarily focus on how individuals learn from failure feedback given by other individuals. It is unclear whether and how the advent of machine feedback may influence individuals’ learning from failures. We suggest that failure feedback provided by machines facilitates learning in two ways. First, it focuses individuals’ attention on their failures, leading them to learn from these failures. Second, it serves as a catalyzer, motivating individuals to learn more from failure feedback given to them by other individuals as well. In addition, this catalyzing effect is stronger if the failure feedback from machines and …
Catnet: Cross-Modal Fusion For Audio-Visual Speech Recognition, Xingmei Wang, Jianchen Mi, Boquan Li, Yixu Zhao, Jiaxiang Meng
Catnet: Cross-Modal Fusion For Audio-Visual Speech Recognition, Xingmei Wang, Jianchen Mi, Boquan Li, Yixu Zhao, Jiaxiang Meng
Research Collection School Of Computing and Information Systems
Automatic speech recognition (ASR) is a typical pattern recognition technology that converts human speeches into texts. With the aid of advanced deep learning models, the performance of speech recognition is significantly improved. Especially, the emerging Audio–Visual Speech Recognition (AVSR) methods achieve satisfactory performance by combining audio-modal and visual-modal information. However, various complex environments, especially noises, limit the effectiveness of existing methods. In response to the noisy problem, in this paper, we propose a novel cross-modal audio–visual speech recognition model, named CATNet. First, we devise a cross-modal bidirectional fusion model to analyze the close relationship between audio and visual modalities. Second, …