Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 4446

Full-Text Articles in Entire DC Network

Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo Aug 2026

Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …


A Fuzzy–Neutrosophic Suitability Index For Selecting An Appropriate Reasoning Model Under Vagueness, Incompleteness, And Conflict, Nada A. Nabeeh, Ahmed Samy Jul 2026

A Fuzzy–Neutrosophic Suitability Index For Selecting An Appropriate Reasoning Model Under Vagueness, Incompleteness, And Conflict, Nada A. Nabeeh, Ahmed Samy

Neutrosophic Systems with Applications

Fuzzy reasoning and neutrosophic reasoning are both used to handle uncertainty, but they are not intended for the same uncertainty structure. Fuzzy reasoning is suitable when uncertainty appears mainly as gradual vagueness, where a value may belong to a concept such as ``high risk'' or ``good performance'' to a certain degree. In this case, a membership value is often sufficient. Neutrosophic reasoning is more suitable when the problem also contains incomplete information, undecided evidence, or conflict between sources. In such cases, one membership degree may be too limited because it cannot represent support, rejection, and indeterminacy separately. This study introduces …


Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy Jul 2026

Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy

Neutrosophic Systems with Applications

User Interface (UI) design can be seen as an essential aspect of human-computer interaction (HCI) and makes communication easier between people and technology. In today's digital economy, interface quality has become one of the most important business concerns, since it has a direct impact on customer satisfaction and retention while affecting revenue. Although creating user-centered and accessible interfaces is crucial, doing so is a difficult and time-consuming process, which leads to burnout for many usability professionals. Although conventional artificial intelligence (AI) was utilized for design assessment and automation, the arrival of generative AI technology has created new possibilities for automated …


Ai-Powered Resume Screening, Sang Suh, Numery Zaber Jul 2026

Ai-Powered Resume Screening, Sang Suh, Numery Zaber

Faculty Publications

Traditional resume screening is manual, slow, and susceptible to bias, and it struggles to keep pace with today’s application volumes. This paper presents a dual-engine, AI-powered resume screening system designed for transparency and reproducibility. The primary (classical) pipeline encodes resumes and job descriptions using Sentence-BERT (SBERT), computes a resume–job match score via cosine similarity, classifies candidates into 25 job categories using XGBoost, and provides model interpretability through SHAP. In parallel, a prompted large language model (LLM) baseline (GPT-4o/4o-mini) outputs a match score and predicted category for comparative analysis. A Streamlit-based interface integrates both engines to support recruiter workflows and human-in-the-loop …


Surveyception: An Exploration Of Deceptive Survey Forms, Muhammad Danish Jul 2026

Surveyception: An Exploration Of Deceptive Survey Forms, Muhammad Danish

Computer Science ETDs

Survey platforms such as Google Forms and Microsoft Forms are widely used for feedback, data collection, and engagement, but scammers increasingly exploit them to distribute phishing and deceptive attacks. This thesis presents a large-scale study of survey-form abuse across ten major providers. We collected 140,000 forms from three sources: public posts on X, search-engine results, and web pages from the top 10 million DomCop-ranked domains. Using automated filtering and manual qualitative review, we identified 2,645 forms requesting sensitive information and classified 566 as scams. These forms used techniques including phishing, private-secret theft, account and personal-data harvesting, financial deception, and psychological …


Coordinating Meaning With Ai System Cards: A Thematic Analysis, Jennifer Rene French Cyrek Jul 2026

Coordinating Meaning With Ai System Cards: A Thematic Analysis, Jennifer Rene French Cyrek

Doctoral Dissertations and Projects

As the meaning of AI risk remains unsettled across sociotechnical and public discourse, AI system cards are an emergent, yet understudied, genre of technical documentation through which AI technology developers publicly frame new AI system capabilities including risks. This thematic content analysis study examines how AI technology developers coordinate meaning regarding risk and responsible development in stewardship of AI. Guided by a constitutive view of communication and systems theory, second-order cybernetics, and the cybernetic tradition, this study analyzes a purposive corpus of AI system cards collected from 2023-2025 using thematic content analysis and the hierarchy of meaning heuristic from coordinated …


Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau Jul 2026

Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau

Presidential Fellows Articles and Research

This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached from the dual perspective of a researcher trained in electrical engineering, pattern recognition, machine learning, and neural networks, who has also spent most of his life engaged in the arts as actor, stage and film director, writer, composer, and visual artist, and in philosophy. Drawing on Margaret Boden’s foundational framework — both her three properties of creativity (novelty, surprise, and value) and her three types of creative processes (combinatorial, exploratory, and transformational) — the paper argues that AI systems are structurally incapable of creativity in …


Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee Jul 2026

Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee

Computer Science Faculty Research & Creative Works

Reproducibility must validate architectural robustness, not just numerical accuracy. We evaluate ColBERT-v2 and ConstBERT across five dimensions, finding that while ConstBERT reproduces within 0.05% MRR@10 on MS-MARCO, both models show a drop of 86-97% on long, narrative queries (TREC ToT 2025). Ablations prove this failure is architectural: performance plateaus at 20 words because the MaxSim operator's uniform token weighting cannot distinguish signal from filler noise. Furthermore, undocumented backend parameters create an 8-point gap due to ConstBERT's sparse centroid coverage, and fine-tuning with 3x more data actually degrades performance by up to 29%. We conclude that architectural constraints in multi-vector retrieval …


Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman Jul 2026

Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman

Computer Science Faculty Research & Creative Works

Debugging consumes a substantial portion of the software development lifecycle, yet researchers do not yet understand well the effectiveness of Large Language Models (LLMs) in this task. Competitive programming offers a rich benchmark for such evaluation, given its diverse problem domains and strict efficiency requirements. We present an empirical study of LLM-based debugging on competitive programming problems and introduce DePro, a test-case-driven approach that assists programmers by correcting existing code rather than generating new solutions. DePro combines brute-force reference generation, stress testing, and iterative LLM-guided refinement to efficiently identify and resolve errors. Experiments on 13 faulty user submissions from Codeforces …


Cnkg: Harnessing Large Language Models For Cognitive Neuroscience Knowledge Graph Construction, Ali Sarabadani, Kheirollah Rahsepar Fard, Hamid Dalvand Jul 2026

Cnkg: Harnessing Large Language Models For Cognitive Neuroscience Knowledge Graph Construction, Ali Sarabadani, Kheirollah Rahsepar Fard, Hamid Dalvand

Turkish Journal of Electrical Engineering and Computer Sciences

Textual resources are among the most valuable sources of information in cognitive neuroscience (CN) for understanding and investigating brain activity and cognitive processes. Extracting and constructing knowledge graphs (KGs) from these texts can facilitate medical research by providing deeper insights into neurological diseases and brain function. In recent years, the use of large language models (LLMs) in natural language processing (NLP) has become increasingly widespread, significantly enhancing the extraction of meaningful information from large volumes of text. This study proposes a novel approach for constructing and evaluating a specialized knowledge graph, termed the cognitive neuroscience knowledge graph (CNKG), from scientific …


Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi Jul 2026

Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi

College of Population Health Faculty Papers

BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.

METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence …


Genedit: A Context-Aware Equity, Diversity And Inclusion Principles Integration Tool For Software Engineering Education, Chetan Arora, Ajanie Kodagoda Bammanna Arachchige, Jinchun Du, Muhammad Aamir Cheema, Aster Cosmos, Antonette Shibani, Vasudha Malhotra, Naeem Janjua, Afaq Shah Jul 2026

Genedit: A Context-Aware Equity, Diversity And Inclusion Principles Integration Tool For Software Engineering Education, Chetan Arora, Ajanie Kodagoda Bammanna Arachchige, Jinchun Du, Muhammad Aamir Cheema, Aster Cosmos, Antonette Shibani, Vasudha Malhotra, Naeem Janjua, Afaq Shah

Research outputs 2022 to 2026

Equity, diversity and inclusion (EDI) is widely acknowledged as essential in software engineering (SE), yet day-to-day integration into teaching remains uncommon due to time pressures, low instructor confidence, and fragmented resources. We present GenEDIt, a purpose-built, LLM-backed chatbot that helps educators weave EDI into SE education artefacts without altering intended learning outcomes. Unlike generic chat interfaces, GenEDIt provides a tailored UI and workflow, and embeds retrieval over a vetted EDI knowledge base. The tool supports five educator-centred modes - (1) methods for integrating EDI into current activities; (2) examples/datasets; (3) EDI-integrated assessments and rubrics; (4) reflective prompts; and (5) rapid …


How Can Accessibility In Computing Education Be Improved Through Hci Research?, Dr David Santandreu Calonge, Linda Smail, Firuz Kamalov, Dima Yousef, Melody Sylvain Jul 2026

How Can Accessibility In Computing Education Be Improved Through Hci Research?, Dr David Santandreu Calonge, Linda Smail, Firuz Kamalov, Dima Yousef, Melody Sylvain

All Works

Accessibility - accommodation of diverse sensory, motor, cognitive, and linguistic needs - remains critically underrepresented in computing education despite broad societal acknowledgment of its importance. This position paper argues that systemic inequities and persistent barriers to integrating accessibility and inclusion - ranging from curricular inertia to limited faculty capacity - require new methodological approaches drawn from Human-Computer Interaction (HCI) research. HCI provides tested frameworks that pair participatory design with inclusive pedagogy, guided by empirical evaluation to drive systemic change. We identify three key areas where HCI can advance accessibility education: (1) embedding accessibility within core curricula through design-based pedagogies, (2) …


Heterogeneous Graph-Augmented Contrastive Learning For Extreme Multi-Class Fiqh Classification, Ali A. Jalil Jul 2026

Heterogeneous Graph-Augmented Contrastive Learning For Extreme Multi-Class Fiqh Classification, Ali A. Jalil

Al-Bahir

  • Background/Introduction: Fine-grained text classification in the field of Islamic Jurisprudence (Fiqh) is difficult because of the structural interdependence of the legal concepts and the extremely multi-class long-tail data distribution (667 classes with 5,979 samples, 52.2% of which contain less than 5 samples). The main problem with traditional flat classifiers is that they assume that target classes are independent and orthogonal output neurons which discards very important relational semantics.
  • Objectives: This paper seeks to remediate this extreme imbalance and maintain structural taxonomy by modeling the structural space of classification label space itself as an object to be learned, while giving a …


Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen Jul 2026

Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen

Research Collection School Of Computing and Information Systems

Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …


Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong Jul 2026

Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong

Research Collection School Of Computing and Information Systems

As AI coding agents become embedded in software development workflows, developers are beginning to operationalize ethical principles by encoding behavioral rules into repository-level context files for AI agents, such as AGENTS.md files. Rather than examining the ethics of AI agents in the abstract, this vision paper investigates how ethics and values are already being translated for AI agents into actionable instructions that shape agent behavior. Through a preliminary investigation, we find that developers are already embedding guidance related to fairness, accessibility, sustainability, tone, and privacy. These artifacts function as a developer-authored governance layer, translating abstract principles into situated, natural-language directives …


A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes Jul 2026

A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes

Research Collection School Of Computing and Information Systems

Agentic AI coding tools such as Claude Code and OpenAI Codex execute multi-step coding tasks with limited human oversight. To steer these tools, developers create repository-level configuration artifacts (e.g., Markdown files) for configuration mechanisms such as Context Files, Skills, Rules, and Hooks. There is no curated dataset yet that captures these configurations at scale. This dataset, collected from open-source GitHub repositories, fills that gap. We selected 40,585 actively maintained repositories through metadata filtering, classified them using GPT-5.2 to identify 36,710 as belonging to engineered software projects, and systematically detected configuration artifacts in these repositories. The dataset covers 4,738 repositories across …


Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang Jul 2026

Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …


Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen Jul 2026

Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

Research Collection School Of Computing and Information Systems

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …


Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo Jul 2026

Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo

Research Collection School Of Computing and Information Systems

Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …


Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun Jul 2026

Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun

Research Collection School Of Computing and Information Systems

Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …


Physics-Informed Machine Learning For Predictive Digital Twins In Greenhouses, Hoang Kim Tran Jul 2026

Physics-Informed Machine Learning For Predictive Digital Twins In Greenhouses, Hoang Kim Tran

Graduate Theses and Dissertations

Greenhouses are widely used to create controlled environments for crop production, where indoor climate variables such as air temperature, relative humidity, CO2 concentration, radiation, and soil or substrate conditions directly affect plant growth, yield, energy consumption, and resource use. However, greenhouse climates are highly dynamic because they are influenced by complex interactions among outdoor weather, greenhouse structure, crop transpiration, heating systems, ventilation, screens, fans, irrigation, and other actuators. Building digital twins for greenhouses provides a powerful way to represent, monitor, and predict these climate dynamics. A predictive digital twin can simulate future indoor climate states based on current greenhouse conditions, …


Localizing And Repairing Backdoor-Sensitive Layers In Pre-Trained Language Models For Secure Fine-Tuning, Sunanda Das Jul 2026

Localizing And Repairing Backdoor-Sensitive Layers In Pre-Trained Language Models For Secure Fine-Tuning, Sunanda Das

Graduate Theses and Dissertations

Task-agnostic backdoor attacks can contaminate pre-trained language models (PLMs) in a way that survives downstream adaptation, even under full fine-tuning, making it difficult for practitioners to trust third-party checkpoints. Existing defenses often rely on privileged assumptions (e.g., access to poisoned data or trigger/target knowledge), thereby limiting their applicability in realistic settings. We present DiSec (Disentanglement of potentially adversarial weights for Secure fine-tuning), a robust and label-efficient purification framework that uses only clean auxiliary text and does not rely on downstream supervision or attack signatures. DiSec elicits model-internal signals from this clean data to separate suspicious parameter components that are inconsistent …


There Is No Free Benchmark: An Institutional View Of Legal Ai Benchmarking, Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho Jul 2026

There Is No Free Benchmark: An Institutional View Of Legal Ai Benchmarking, Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho

Faculty Scholarship

Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations” — wherein models make up facts, law, and precedent — leading Chief Justice Roberts to spotlight this risk in his annual report on the judiciary. We argue that there is a need for public AI benchmarking in law. First, relative to other AI application domains, the legal AI ecosystem lacks legibility — there is little information about the design and performance of many commercial …


Advancing Social Media Analytics And Personalized Generation Via Transfer Learning, Discourse-Aware Modeling, And Collaborative Modeling, Gibson Nkhata Jul 2026

Advancing Social Media Analytics And Personalized Generation Via Transfer Learning, Discourse-Aware Modeling, And Collaborative Modeling, Gibson Nkhata

Graduate Theses and Dissertations

Social media platforms have become central to information exchange, shaping public opinion across social, political, and economic domains. However, the massive volume of user-generated content, combined with its informal, nuanced, and often noisy nature, presents significant challenges for automated analysis and generation. Tasks such as stance detection, rumor verification, and personalized content generation are further complicated by sarcasm, evolving discourse structures, and diverse user preferences. Addressing these challenges requires models that can effectively leverage linguistic nuance, conversational dynamics, and collaborative user signals. Transfer learning has emerged as a powerful paradigm for improving performance in low-resource and complex language understanding tasks. …


Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto Jul 2026

Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto

School of Accountancy Faculty Publications

This study examines whether large language models exhibit systematic country-contingent differential treatment in financial fraud detection. Analyzing 30,000 synthetic transactions with identical statistical properties across three country attributions (United States, Great Britain, and China), we find LLMs assign significantly higher fraud probabilities to Chinese-attributed transactions (36.2%) compared to Western countries (≈30–31%), resulting in accuracy disparities of 67% versus 74%. The gap remains stable across five independent experimental replications and persists when using Chinese language prompts, ruling out linguistic effects. Bias mitigation strategies, such as requiring explanations or explicit country neutrality instructions, reduce but fail to eliminate these disparities. Testing across …


Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo Jul 2026

Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo

Research Collection School Of Computing and Information Systems

This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona …


Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma Jul 2026

Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma

Research Collection School Of Computing and Information Systems

Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …


Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma Jul 2026

Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma

Research Collection School Of Computing and Information Systems

Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …


Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang Jul 2026

Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang

Research Collection School Of Computing and Information Systems

Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …