Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (4482)
- Computer Sciences (4407)
- Social and Behavioral Sciences (2785)
- Business (1976)
- Databases and Information Systems (1609)
-
- Software Engineering (1285)
- Economics (978)
- International and Area Studies (864)
- Asian Studies (842)
- Artificial Intelligence and Robotics (817)
- Information Security (751)
- Law (735)
- Econometrics (504)
- Graphics and Human Computer Interfaces (498)
- Numerical Analysis and Scientific Computing (467)
- Finance and Financial Management (446)
- Psychology (429)
- Organizational Behavior and Theory (421)
- Engineering (383)
- Accounting (300)
- Programming Languages and Compilers (288)
- Sociology (250)
- Public Affairs, Public Policy and Public Administration (249)
- Computer Engineering (242)
- Corporate Finance (239)
- Arts and Humanities (232)
- Communication (222)
- Political Science (221)
- Theory and Algorithms (220)
- Education (207)
- Keyword
-
- Singapore (229)
- Machine learning (82)
- Deep learning (80)
- China (73)
- Artificial intelligence (53)
-
- Privacy (46)
- Social media (45)
- Empirical study (42)
- Trust (40)
- Innovation (39)
- Blockchain (37)
- Deep Learning (37)
- Training (37)
- Data mining (36)
- Security (36)
- Corporate governance (35)
- COVID-19 (34)
- Creativity (33)
- Software engineering (33)
- Culture (32)
- Natural language processing (32)
- Task analysis (32)
- Cloud computing (31)
- Classification (30)
- Neural networks (30)
- Reinforcement learning (30)
- Sustainability (30)
- Robustness (29)
- Large Language Models (28)
- Performance (28)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (4258)
- Research Collection Lee Kong Chian School Of Business (1089)
- Research Collection School Of Economics (776)
- Research Collection School of Social Sciences (713)
- Research Collection Yong Pung How School Of Law (670)
-
- Dissertations and Theses Collection (Open Access) (357)
- Research Collection School Of Accountancy (276)
- Research Collection College of Integrative Studies (136)
- Knowledge@SMU (55)
- Perspectives@SMU (47)
- Asian Management Insights (35)
- Research Collection Library (29)
- Dissertations and Theses Collection (18)
- Digital Narratives of Asia (17)
- Singapore Law Journal (Lexicon) (16)
- Sim Kee Boon Institute for Financial Economics (15)
- Social Space (14)
- CCX Research (11)
- Oral History Collection (11)
- Research Collection Lee Kong Chian School of Business (11)
- FORCE 2026 (8)
- Research Collection BNP Paribas Hedge Fund Centre (7)
- Centre for AI & Data Governance (2019-2025) (6)
- Research Collection School of Economics (6)
- Lien Centre for Social Innovation: Research (5)
- PhD Student’s Publications Collection (5)
- ROSA Research Briefs (5)
- SMU Corporate Reports (5)
- 2009 Yong Pung How Professorship of Law Lecture (4)
- Research Collection Centre for English Communication (4)
- Publication Type
Articles 331 - 360 of 8668
Full-Text Articles in Entire DC Network
Psychological Characteristics Of Board Secretaries And Information Disclosure Quality - An Empirical Analysis Based On The Questionnaire Survey And Interviews With Board Secretaries Of A-Share Listed Companies, Shuijun Dai
Dissertations and Theses Collection (Open Access)
The capital market functions fundamentally on the principles of information, and effective information disclosure acts as the cornerstone of its operation. At present, China's capital market is fully advancing the registration-based reform that is primarily focused on information disclosure. In this context, the secretary of the board of directors ("board secretary")—who serves as the legal officer tasked with overseeing information disclosure for A-share listed companies—plays a vital role in information release and transmission.
In the context of the comprehensive implementation of the registration system and establishment of an investor-oriented capital market, this study surveys over 400 board secretaries from A-share …
Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan
Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) demonstrate remarkable in-context learning capabilities but often struggle with complex, multi-step reasoning. Multi-Agent Debate (MAD) frameworks partially address these limitations by enabling iterative agent interactions. However, they neglect valuable historical insights by treating each new debate independently. In this paper, we propose Memory-Augmented MAD (MeMAD), a parameter-free memory-augmented MAD framework that systematically organizes and reuses past debate transcripts. MeMAD stores structured representations of successful and unsuccessful reasoning attempts enriched with self-reflections and peer feedback. It systematically retrieves them via semantic similarity at inference time to inform new reasoning tasks. Our experiments on challenging mathematical reasoning, scientific …
A System Framework To Symbolically Explore Intel Tdx Module Execution, Pansilu Pitigalaarachchillage, Xuhua Ding
A System Framework To Symbolically Explore Intel Tdx Module Execution, Pansilu Pitigalaarachchillage, Xuhua Ding
Research Collection School Of Computing and Information Systems
We present TDXplorer, the first dynamic symbolic analysis system for Intel's TDX Module, the software trusted computing base of TDX. Without using TDX hardware, an analyzer function on top of TDXplorer can not only apply dynamic analysis to control and instrument the TDX Module's execution, but also carry out symbolic execution for path exploration as well as security and functionality reasoning. The two types of analysis are seamlessly integrated in a way that symbolic execution is conducted directly upon the TDX Module's binary code and runtime states, which are shaped by using dynamic analysis techniques. We implement TDXplorer on Linux …
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Research Collection School Of Computing and Information Systems
The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …
The Excellent Legal Scholar, Seow Hon Tan
The Excellent Legal Scholar, Seow Hon Tan
Research Collection Yong Pung How School Of Law
The Excellent Legal Scholar: This article considers how virtues pan out in the life of the legal scholar, bearing in mind the purpose of legal scholarship and the identity of the legal scholar, who plays multifarious roles in today's research landscape. I consider how vision is important for the excellent legal scholar, bearing in mind that an aretaic account should be attentive to eudaimonia. I conclude with soul-searching questions for the legal scholar who endeavours to live an examined life that stands up to aretaic appraisal.
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
Dissertations and Theses Collection (Open Access)
This study develops a data-driven framework for optimal retail store location planning that integrates road network analysis, mobility data and optimization techniques. By addressing the limitations of traditional approaches that rely on outdated census data and manual site selection, this research offers a scalable and adaptable solution for retail expansion in diverse urban environments. Chapters 1 and 2 establish the foundational context and theoretical underpinnings of this research. Chapter 1 introduces the research problem and motivation, highlighting the limitations of existing approaches and defining three key research objectives: automating candidate site identification, improving footfall estimation, and developing a scalable multi-site …
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Dissertations and Theses Collection (Open Access)
The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …
Design-Led Leadership: A Mixed Methods Study On Understanding How Design Thinking Capabilities Impact Leadership Effectiveness, Tareq Waleed Muhmood
Design-Led Leadership: A Mixed Methods Study On Understanding How Design Thinking Capabilities Impact Leadership Effectiveness, Tareq Waleed Muhmood
Dissertations and Theses Collection (Open Access)
“Widespread design thinking among organisation leaders is desirable for the creation of a humanly and sustainable future” (R. J. Boland et al., 2008a)
This study investigates the impact of design thinking capabilities on leadership effectiveness through a mixed-methods research design.
While design thinking is widely used in product and service innovation, its application as a leadership development framework remains underexplored. To address this gap, the study employed both quantitative and qualitative methods, including pre- and post-intervention survey of leaders, their supervisors, and subordinates within an education business in Vietnam (EBV), as well as semi-structured interviews with senior leaders at a …
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang
Research Collection School Of Computing and Information Systems
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and extensive world knowledge. However, whether these MLLMs possess human-like compositional reasoning abilities remains an open problem. To unveil their reasoning behaviors, we first curate a Multimodal Assumptive Reasoning Benchmark (MARS-Bench) in this paper. Interestingly, we find that most prevalent MLLMs can be easily fooled by the introduction of a presupposition into the question, whereas such presuppositions appear naive to human reasoning. Besides, we also propose a simple yet effective method, Active Deduction (AD), a novel reinforcement learning paradigm to encourage …
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Research Collection School Of Computing and Information Systems
Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status — the condition or transformation of ingredients and tools — offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time …
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
3Dvisual grounding aims to identify and localize objects in a 3Dspacebasedontextualdescriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsisten cies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a frame work that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Rela tion Decoupling (SRD) module restructures complex multi anchor queries into a set of targeted single-anchor state ments, generating a structured set of perspective-aware de scriptions that clarify positional relationships. These de composed representations serve as the foundation for the Multi-view …
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu
Research Collection School Of Computing and Information Systems
We present EmoShortcuts1, a novel social Mixed Reality (MR) framework that enhances emotional expression by dynamically augmenting avatar body gestures to reflect users’ emotional states. While social MR enables immersive remote interactions through avatars, conveying emotions remains challenging due to limitations in head-mounted display (HMD) tracking (e.g., missing lower-body movements, such as stomping or defensive postures), and users’ tendency to deprioritize nonverbal expressions during multitasking. EmoShortcuts addresses these challenges by introducing an augmentation framework that generates expressive body gestures even when users’ physical movements are restricted. We conducted a formative study with 12 participants to identify key challenges in emotional …
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Research Collection School Of Computing and Information Systems
In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Research Collection School Of Computing and Information Systems
Website owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable …
Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang
Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Chain-of-thought (CoT) reasoning greatly improves the interpretability and problem-solving abilities of multimodal large language models (MLLMs). However, existing ap proaches focus on text CoT, limiting their ability to lever age visual cues. Visual CoT remains underexplored, and the only work [35] is based on supervised fine-tuning that relies on extensive labeled bounding-box data and is hard to generalize to unseen cases. In this paper, we introduce Unsupervised Visual CoT (UV-CoT), a novel framework for image-level CoT reasoning via preference optimization. UV-CoTperforms preference comparisons between model generated bounding boxes (one is preferred and the other is dis-preferred), eliminating the need for …
Morphology-Aware Hrv Estimation From Wrist Ppg In Sedentary Scenarios, Changshuo Hu, Hung Manh Pham, Dong Ma
Morphology-Aware Hrv Estimation From Wrist Ppg In Sedentary Scenarios, Changshuo Hu, Hung Manh Pham, Dong Ma
Research Collection School Of Computing and Information Systems
Photoplethysmography (PPG) is widely used in wearable devices for non-invasive heart rate variability (HRV) monitoring. While most prior work focuses on mitigating motion artifacts, recent studies highlight that even subtle contact pressure variations can distort waveform morphology and lead to inaccurate HRV estimates. In this work, we propose a morphology-aware deep learning framework that conditions HRV estimation on beat-level waveform types. Our model jointly encodes the raw PPG waveform and a sequence of pressure-induced morphology labels using parallel encoders, integrates them via cross-attention, and predicts normal-to-normal (NN) intervals and beat count to support downstream HRV computation. Evaluated on the public …
Unpacking Singapore's Leasehold Relativity Table: An Empirical And Legal Analysis, Koon Shing Kwong, Jing Rong Goh, Seng Wei, Edward Ti
Unpacking Singapore's Leasehold Relativity Table: An Empirical And Legal Analysis, Koon Shing Kwong, Jing Rong Goh, Seng Wei, Edward Ti
Research Collection School Of Economics
In Singapore, most land is state-owned, with the state generally issuing leasehold estates via state leases of not more than 99 years1, depending on the intended land use. Naturally, the value of a leasehold estate, which erodes over time as the lease approaches the end of its term, is a key component of the premium charged for lease renewals, or the tax imposed for permission given in relation to a development that would increase the value of the land. By law, the state valuation of leasehold land is prescribed by a leasehold relativity table colloquially known as ‘Bala’s Curve’ or …
Weak Identification Of Long Memory With Implications For Volatility Modeling;, Jia Li, Peter C. B. Phillips, Shuping Shi, Jun Yu
Weak Identification Of Long Memory With Implications For Volatility Modeling;, Jia Li, Peter C. B. Phillips, Shuping Shi, Jun Yu
Research Collection School Of Economics
This paper explores implications of weak identification in common ‘long memory’ and recent ‘rough’ approaches to modeling volatility dynamics of financial assets. We unveil an asymptotic near-observational equivalence between a long memory model with weak autoregressive dynamics and a rough model with a near-unit autoregressive root. Standard methods struggle to distinguish them, and conventional asymptotics are invalid. We propose an identification-robust approach to construct confidence sets that reveal the uncertainty and aid inference. Empirical studies based on realized volatility and trading volume often fail to statistically reject either model, thereby providing evidence of their potential coexistence.
Rethinking Teaching Evaluation Reports: Designing Ai-Transformed Student Feedback For Instructor Engagement, Ruoxi Shang, Keri Mallari, Au Wei Bin Yeong, Ken Yasuhara, Anthony Tang, Gary Hsieh
Rethinking Teaching Evaluation Reports: Designing Ai-Transformed Student Feedback For Instructor Engagement, Ruoxi Shang, Keri Mallari, Au Wei Bin Yeong, Ken Yasuhara, Anthony Tang, Gary Hsieh
Research Collection School Of Computing and Information Systems
Student feedback is critical for improving teaching, yet instructors often avoid reading evaluations due to emotional burden and information overload. We present a systematic exploration of how language models can distill and transform student evaluations into adaptive, actionable insights. Through a systematic design space exploration combining 4 feedback strategies (removing harmful content, paraphrasing criticism, sandwiching negatives, adding constructive suggestions) with 4 presentation formats (themes, cards, letters, chatbots), we created six AI-augmented prototypes of teaching evaluations. Interviews with 16 post-secondary instructors revealed that effective use of AI in feedback processing should: (1) support action formation through focused views and divergent thinking, …
Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu
Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu
Research Collection School Of Computing and Information Systems
Prewriting is the process of generating and organising ideas before a first draft. It consists of a combination of informal, iterative, and semi-structured strategies such as visual diagramming, which poses a challenge for collaborating with large language models (LLMs) in a turn-taking conversational manner. We present Polymind, a visual diagramming tool that leverages multiple LLM-powered agents to support prewriting. The system features a parallel collaboration workflow in place of the turn-taking conversational interactions. It defines multiple ''microtasks'' to simulate group collaboration scenarios such as collaborative writing and group brainstorming. Instead of repetitively prompting a chatbot for various purposes, Polymind enables …
Conditional Attribute-Based Pre: Definition And Construction From Lwe, Lisha Yao, Jian Weng, Pengfei Wu, Guofeng Tang, Guomin Yang, Haiyang Xue, Robert H. Deng
Conditional Attribute-Based Pre: Definition And Construction From Lwe, Lisha Yao, Jian Weng, Pengfei Wu, Guofeng Tang, Guomin Yang, Haiyang Xue, Robert H. Deng
Research Collection School Of Computing and Information Systems
Attribute-based proxy re-encryption (AB-PRE) is a crucial variant of proxy re-encryption. It allows a proxy with a re-encryption key to transform a delegator’s ciphertext associated with an access policy into another ciphertext associated with a new access policy, enabling delegatees with matching attributes to decrypt the transformed ciphertext. However, a key limitation of AB-PRE is that the delegator cannot control which ciphertexts are transformed. As a result, the proxy, once given the re-encryption key, indiscriminately transforms all ciphertexts, effectively switching their underlying policies—an issue known as the all-or-nothing problem. It limits the system’s flexibility and practicality in real-world use cases.In …
Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He
Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He
Research Collection School Of Computing and Information Systems
Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieve cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's …
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He
Research Collection School Of Computing and Information Systems
Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces cross-image stroke attention, a mechanism embedded within self-attention layers to establish fine-grained semantic correspondences and enable accurate stroke attribute transfer. This allows our method to adaptively integrate reference stroke characteristics into content images while maintaining structural integrity. Additionally, we develop adaptive contrast enhancement and semanticfocused attention to reinforce content preservation and foreground emphasis. Stroke2Sketch effectively synthesizes stylistically faithful sketches that closely …
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Research Collection School Of Computing and Information Systems
Tactile graphics on a refreshable display have proven effective in enabling visually impaired people to comprehend pictorial content. To further evaluate the effectiveness of refreshable tactile displays in blind education, we designed tactile data comics, a method that combines step-by-step presentation of tactile graphics with verbal narration. We conducted a user study with sixteen visually impaired students to compare tactile data comics against verbal-only and static tactile graphics. Our findings show that tactile data comics significantly improve participants’ comprehension and engagement during the learning experience. These empirical results suggest that the integration of refreshable tactile displays and tactile data comics …
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Research Collection School Of Computing and Information Systems
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …
Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang
Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
Traditional ship detection methods primarily rely on single-modal approaches, such as visible or infrared images, which limit their application in complex scenarios involving varying lighting conditions and heavy fog. To address this issue, we explore the advantages of short-wave infrared (SWIR) and long-wave infrared (LWIR) in ship detection and propose a novel single-stage image fusion detection algorithm called LSFDNet. This algorithm leverages feature interaction between the image fusion and object detection subtask networks, achieving remarkable detection performance and generating visually impressive fused images. To further improve the saliency of objects in the fused images and improve the performance of the …
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Research Collection School Of Computing and Information Systems
In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …
Information-Bottleneck Driven Binary Neural Network For Change Detection, Kaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao, Cheng-Zhong Xu, Hui Kong
Information-Bottleneck Driven Binary Neural Network For Change Detection, Kaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao, Cheng-Zhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
In this paper, we propose Binarized Change Detection (BiCD), the first binary neural network (BNN) designed specifically for change detection. Conventional network binarization approaches, which directly quantize both weights and activations in change detection models, severely limit the network's ability to represent input data and distinguish between changed and unchanged regions. This results in significantly lower detection accuracy compared to real-valued networks. To overcome these challenges, BiCD enhances both the representational power and feature separability of BNNs, improving detection performance. Specifically, we introduce an auxiliary objective based on the Information Bottleneck (IB) principle, guiding the encoder to retain essential input …
Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He
Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
The power of visual language models is showcased in visual understanding tasks, where language-guided models achieve impressive flexibility and precision. In this paper, we ex tend this capability to the challenging domain of image matting by framing it as a soft grounding problem, enabling a single diffusion model to handle diverse objects, textures, and transparencies, all directed by descriptive text prompts. Our method teaches the diffusion model to ground alpha mattes by guiding it through a process of instance-level localization and transparency estimation. First, we introduce an intermediate objective that trains the model to accurately localize semantic components of the …