Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (1105)
- TÜBİTAK (199)
- Wright State University (191)
- Technological University Dublin (182)
- Old Dominion University (126)
-
- Missouri University of Science and Technology (112)
- San Jose State University (111)
- City University of New York (CUNY) (92)
- Neutrosophic Systems with Applications (89)
- MBZUAI (70)
- Dartmouth College (67)
- Brigham Young University (65)
- University of Texas at El Paso (63)
- Chulalongkorn University (62)
- University of Nebraska - Lincoln (54)
- Zayed University (54)
- Boise State University (48)
- Edith Cowan University (48)
- Kennesaw State University (47)
- New Jersey Institute of Technology (44)
- California Polytechnic State University, San Luis Obispo (43)
- Embry-Riddle Aeronautical University (42)
- Purdue University (41)
- Chapman University (40)
- University of Nebraska at Omaha (40)
- University of Arkansas, Fayetteville (38)
- Montclair State University (36)
- University of South Florida (36)
- Air Force Institute of Technology (35)
- Portland State University (34)
- Keyword
-
- Machine learning (184)
- Natural language processing (181)
- Artificial intelligence (139)
- Natural Language Processing (126)
- Machine Learning (121)
-
- Deep learning (117)
- Artificial Intelligence (91)
- Large language models (84)
- Social media (80)
- Computer Science (78)
- Twitter (63)
- NLP (60)
- Deep Learning (59)
- Sentiment analysis (58)
- Large Language Models (55)
- Computational linguistics (53)
- Department of Computer Science and Engineering (53)
- Text mining (51)
- Neural networks (48)
- Data mining (44)
- Generative AI (44)
- AI (42)
- Computer science (39)
- Classification (37)
- Information retrieval (36)
- Social Media (36)
- Fuzzy logic (35)
- Semantics (34)
- LLM (32)
- BERT (31)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (1030)
- Turkish Journal of Electrical Engineering and Computer Sciences (199)
- Theses and Dissertations (160)
- Master's Projects (99)
- Conference papers (93)
-
- Neutrosophic Systems with Applications (89)
- Kno.e.sis Publications (84)
- Dissertations (70)
- Computer Science Faculty Research & Creative Works (66)
- Chulalongkorn University Theses and Dissertations (Chula ETD) (62)
- Browse all Theses and Dissertations (61)
- Computer Science Faculty Publications (58)
- All Works (54)
- Faculty Publications (49)
- Computer Science Faculty Publications and Presentations (45)
- Publications and Research (44)
- Articles (40)
- Electronic Theses and Dissertations (39)
- Computer Science and Engineering Faculty Publications (38)
- Natural Language Processing Faculty Publications (38)
- Departmental Technical Reports (CS) (36)
- Dissertations and Theses Collection (Open Access) (34)
- USF Tampa Graduate Theses and Dissertations (34)
- Computer Science: Faculty Publications (33)
- Dissertations, Theses, and Capstone Projects (32)
- CCAC Theses and Dissertations (30)
- Theses (30)
- Faculty Scholarship (27)
- Masters Theses (26)
- Department of Computer Science Technical Reports (25)
- Publication Type
- File Type
Articles 31 - 60 of 4446
Full-Text Articles in Entire DC Network
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …
A Fuzzy–Neutrosophic Suitability Index For Selecting An Appropriate Reasoning Model Under Vagueness, Incompleteness, And Conflict, Nada A. Nabeeh, Ahmed Samy
A Fuzzy–Neutrosophic Suitability Index For Selecting An Appropriate Reasoning Model Under Vagueness, Incompleteness, And Conflict, Nada A. Nabeeh, Ahmed Samy
Neutrosophic Systems with Applications
Fuzzy reasoning and neutrosophic reasoning are both used to handle uncertainty, but they are not intended for the same uncertainty structure. Fuzzy reasoning is suitable when uncertainty appears mainly as gradual vagueness, where a value may belong to a concept such as ``high risk'' or ``good performance'' to a certain degree. In this case, a membership value is often sufficient. Neutrosophic reasoning is more suitable when the problem also contains incomplete information, undecided evidence, or conflict between sources. In such cases, one membership degree may be too limited because it cannot represent support, rejection, and indeterminacy separately. This study introduces …
Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy
Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy
Neutrosophic Systems with Applications
User Interface (UI) design can be seen as an essential aspect of human-computer interaction (HCI) and makes communication easier between people and technology. In today's digital economy, interface quality has become one of the most important business concerns, since it has a direct impact on customer satisfaction and retention while affecting revenue. Although creating user-centered and accessible interfaces is crucial, doing so is a difficult and time-consuming process, which leads to burnout for many usability professionals. Although conventional artificial intelligence (AI) was utilized for design assessment and automation, the arrival of generative AI technology has created new possibilities for automated …
Ai-Powered Resume Screening, Sang Suh, Numery Zaber
Ai-Powered Resume Screening, Sang Suh, Numery Zaber
Faculty Publications
Traditional resume screening is manual, slow, and susceptible to bias, and it struggles to keep pace with today’s application volumes. This paper presents a dual-engine, AI-powered resume screening system designed for transparency and reproducibility. The primary (classical) pipeline encodes resumes and job descriptions using Sentence-BERT (SBERT), computes a resume–job match score via cosine similarity, classifies candidates into 25 job categories using XGBoost, and provides model interpretability through SHAP. In parallel, a prompted large language model (LLM) baseline (GPT-4o/4o-mini) outputs a match score and predicted category for comparative analysis. A Streamlit-based interface integrates both engines to support recruiter workflows and human-in-the-loop …
Surveyception: An Exploration Of Deceptive Survey Forms, Muhammad Danish
Surveyception: An Exploration Of Deceptive Survey Forms, Muhammad Danish
Computer Science ETDs
Survey platforms such as Google Forms and Microsoft Forms are widely used for feedback, data collection, and engagement, but scammers increasingly exploit them to distribute phishing and deceptive attacks. This thesis presents a large-scale study of survey-form abuse across ten major providers. We collected 140,000 forms from three sources: public posts on X, search-engine results, and web pages from the top 10 million DomCop-ranked domains. Using automated filtering and manual qualitative review, we identified 2,645 forms requesting sensitive information and classified 566 as scams. These forms used techniques including phishing, private-secret theft, account and personal-data harvesting, financial deception, and psychological …
Coordinating Meaning With Ai System Cards: A Thematic Analysis, Jennifer Rene French Cyrek
Coordinating Meaning With Ai System Cards: A Thematic Analysis, Jennifer Rene French Cyrek
Doctoral Dissertations and Projects
As the meaning of AI risk remains unsettled across sociotechnical and public discourse, AI system cards are an emergent, yet understudied, genre of technical documentation through which AI technology developers publicly frame new AI system capabilities including risks. This thematic content analysis study examines how AI technology developers coordinate meaning regarding risk and responsible development in stewardship of AI. Guided by a constitutive view of communication and systems theory, second-order cybernetics, and the cybernetic tradition, this study analyzes a purposive corpus of AI system cards collected from 2023-2025 using thematic content analysis and the hierarchy of meaning heuristic from coordinated …
Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau
Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau
Presidential Fellows Articles and Research
This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached from the dual perspective of a researcher trained in electrical engineering, pattern recognition, machine learning, and neural networks, who has also spent most of his life engaged in the arts as actor, stage and film director, writer, composer, and visual artist, and in philosophy. Drawing on Margaret Boden’s foundational framework — both her three properties of creativity (novelty, surprise, and value) and her three types of creative processes (combinatorial, exploratory, and transformational) — the paper argues that AI systems are structurally incapable of creativity in …
Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
Computer Science Faculty Research & Creative Works
Reproducibility must validate architectural robustness, not just numerical accuracy. We evaluate ColBERT-v2 and ConstBERT across five dimensions, finding that while ConstBERT reproduces within 0.05% MRR@10 on MS-MARCO, both models show a drop of 86-97% on long, narrative queries (TREC ToT 2025). Ablations prove this failure is architectural: performance plateaus at 20 words because the MaxSim operator's uniform token weighting cannot distinguish signal from filler noise. Furthermore, undocumented backend parameters create an 8-point gap due to ConstBERT's sparse centroid coverage, and fine-tuning with 3x more data actually degrades performance by up to 29%. We conclude that architectural constraints in multi-vector retrieval …
Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman
Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman
Computer Science Faculty Research & Creative Works
Debugging consumes a substantial portion of the software development lifecycle, yet researchers do not yet understand well the effectiveness of Large Language Models (LLMs) in this task. Competitive programming offers a rich benchmark for such evaluation, given its diverse problem domains and strict efficiency requirements. We present an empirical study of LLM-based debugging on competitive programming problems and introduce DePro, a test-case-driven approach that assists programmers by correcting existing code rather than generating new solutions. DePro combines brute-force reference generation, stress testing, and iterative LLM-guided refinement to efficiently identify and resolve errors. Experiments on 13 faulty user submissions from Codeforces …
Cnkg: Harnessing Large Language Models For Cognitive Neuroscience Knowledge Graph Construction, Ali Sarabadani, Kheirollah Rahsepar Fard, Hamid Dalvand
Cnkg: Harnessing Large Language Models For Cognitive Neuroscience Knowledge Graph Construction, Ali Sarabadani, Kheirollah Rahsepar Fard, Hamid Dalvand
Turkish Journal of Electrical Engineering and Computer Sciences
Textual resources are among the most valuable sources of information in cognitive neuroscience (CN) for understanding and investigating brain activity and cognitive processes. Extracting and constructing knowledge graphs (KGs) from these texts can facilitate medical research by providing deeper insights into neurological diseases and brain function. In recent years, the use of large language models (LLMs) in natural language processing (NLP) has become increasingly widespread, significantly enhancing the extraction of meaningful information from large volumes of text. This study proposes a novel approach for constructing and evaluating a specialized knowledge graph, termed the cognitive neuroscience knowledge graph (CNKG), from scientific …
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
College of Population Health Faculty Papers
BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.
METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence …
Genedit: A Context-Aware Equity, Diversity And Inclusion Principles Integration Tool For Software Engineering Education, Chetan Arora, Ajanie Kodagoda Bammanna Arachchige, Jinchun Du, Muhammad Aamir Cheema, Aster Cosmos, Antonette Shibani, Vasudha Malhotra, Naeem Janjua, Afaq Shah
Genedit: A Context-Aware Equity, Diversity And Inclusion Principles Integration Tool For Software Engineering Education, Chetan Arora, Ajanie Kodagoda Bammanna Arachchige, Jinchun Du, Muhammad Aamir Cheema, Aster Cosmos, Antonette Shibani, Vasudha Malhotra, Naeem Janjua, Afaq Shah
Research outputs 2022 to 2026
Equity, diversity and inclusion (EDI) is widely acknowledged as essential in software engineering (SE), yet day-to-day integration into teaching remains uncommon due to time pressures, low instructor confidence, and fragmented resources. We present GenEDIt, a purpose-built, LLM-backed chatbot that helps educators weave EDI into SE education artefacts without altering intended learning outcomes. Unlike generic chat interfaces, GenEDIt provides a tailored UI and workflow, and embeds retrieval over a vetted EDI knowledge base. The tool supports five educator-centred modes - (1) methods for integrating EDI into current activities; (2) examples/datasets; (3) EDI-integrated assessments and rubrics; (4) reflective prompts; and (5) rapid …
How Can Accessibility In Computing Education Be Improved Through Hci Research?, Dr David Santandreu Calonge, Linda Smail, Firuz Kamalov, Dima Yousef, Melody Sylvain
How Can Accessibility In Computing Education Be Improved Through Hci Research?, Dr David Santandreu Calonge, Linda Smail, Firuz Kamalov, Dima Yousef, Melody Sylvain
All Works
Accessibility - accommodation of diverse sensory, motor, cognitive, and linguistic needs - remains critically underrepresented in computing education despite broad societal acknowledgment of its importance. This position paper argues that systemic inequities and persistent barriers to integrating accessibility and inclusion - ranging from curricular inertia to limited faculty capacity - require new methodological approaches drawn from Human-Computer Interaction (HCI) research. HCI provides tested frameworks that pair participatory design with inclusive pedagogy, guided by empirical evaluation to drive systemic change. We identify three key areas where HCI can advance accessibility education: (1) embedding accessibility within core curricula through design-based pedagogies, (2) …
Heterogeneous Graph-Augmented Contrastive Learning For Extreme Multi-Class Fiqh Classification, Ali A. Jalil
Heterogeneous Graph-Augmented Contrastive Learning For Extreme Multi-Class Fiqh Classification, Ali A. Jalil
Al-Bahir
- Background/Introduction: Fine-grained text classification in the field of Islamic Jurisprudence (Fiqh) is difficult because of the structural interdependence of the legal concepts and the extremely multi-class long-tail data distribution (667 classes with 5,979 samples, 52.2% of which contain less than 5 samples). The main problem with traditional flat classifiers is that they assume that target classes are independent and orthogonal output neurons which discards very important relational semantics.
- Objectives: This paper seeks to remediate this extreme imbalance and maintain structural taxonomy by modeling the structural space of classification label space itself as an object to be learned, while giving a …
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Research Collection School Of Computing and Information Systems
Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …
Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong
Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong
Research Collection School Of Computing and Information Systems
As AI coding agents become embedded in software development workflows, developers are beginning to operationalize ethical principles by encoding behavioral rules into repository-level context files for AI agents, such as AGENTS.md files. Rather than examining the ethics of AI agents in the abstract, this vision paper investigates how ethics and values are already being translated for AI agents into actionable instructions that shape agent behavior. Through a preliminary investigation, we find that developers are already embedding guidance related to fairness, accessibility, sustainability, tone, and privacy. These artifacts function as a developer-authored governance layer, translating abstract principles into situated, natural-language directives …
A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes
A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes
Research Collection School Of Computing and Information Systems
Agentic AI coding tools such as Claude Code and OpenAI Codex execute multi-step coding tasks with limited human oversight. To steer these tools, developers create repository-level configuration artifacts (e.g., Markdown files) for configuration mechanisms such as Context Files, Skills, Rules, and Hooks. There is no curated dataset yet that captures these configurations at scale. This dataset, collected from open-source GitHub repositories, fills that gap. We selected 40,585 actively maintained repositories through metadata filtering, classified them using GPT-5.2 to identify 36,710 as belonging to engineered software projects, and systematically detected configuration artifacts in these repositories. The dataset covers 4,738 repositories across …
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Research Collection School Of Computing and Information Systems
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Research Collection School Of Computing and Information Systems
Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …
Physics-Informed Machine Learning For Predictive Digital Twins In Greenhouses, Hoang Kim Tran
Physics-Informed Machine Learning For Predictive Digital Twins In Greenhouses, Hoang Kim Tran
Graduate Theses and Dissertations
Greenhouses are widely used to create controlled environments for crop production, where indoor climate variables such as air temperature, relative humidity, CO2 concentration, radiation, and soil or substrate conditions directly affect plant growth, yield, energy consumption, and resource use. However, greenhouse climates are highly dynamic because they are influenced by complex interactions among outdoor weather, greenhouse structure, crop transpiration, heating systems, ventilation, screens, fans, irrigation, and other actuators. Building digital twins for greenhouses provides a powerful way to represent, monitor, and predict these climate dynamics. A predictive digital twin can simulate future indoor climate states based on current greenhouse conditions, …
Localizing And Repairing Backdoor-Sensitive Layers In Pre-Trained Language Models For Secure Fine-Tuning, Sunanda Das
Localizing And Repairing Backdoor-Sensitive Layers In Pre-Trained Language Models For Secure Fine-Tuning, Sunanda Das
Graduate Theses and Dissertations
Task-agnostic backdoor attacks can contaminate pre-trained language models (PLMs) in a way that survives downstream adaptation, even under full fine-tuning, making it difficult for practitioners to trust third-party checkpoints. Existing defenses often rely on privileged assumptions (e.g., access to poisoned data or trigger/target knowledge), thereby limiting their applicability in realistic settings. We present DiSec (Disentanglement of potentially adversarial weights for Secure fine-tuning), a robust and label-efficient purification framework that uses only clean auxiliary text and does not rely on downstream supervision or attack signatures. DiSec elicits model-internal signals from this clean data to separate suspicious parameter components that are inconsistent …
There Is No Free Benchmark: An Institutional View Of Legal Ai Benchmarking, Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho
There Is No Free Benchmark: An Institutional View Of Legal Ai Benchmarking, Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho
Faculty Scholarship
Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations” — wherein models make up facts, law, and precedent — leading Chief Justice Roberts to spotlight this risk in his annual report on the judiciary. We argue that there is a need for public AI benchmarking in law. First, relative to other AI application domains, the legal AI ecosystem lacks legibility — there is little information about the design and performance of many commercial …
Advancing Social Media Analytics And Personalized Generation Via Transfer Learning, Discourse-Aware Modeling, And Collaborative Modeling, Gibson Nkhata
Graduate Theses and Dissertations
Social media platforms have become central to information exchange, shaping public opinion across social, political, and economic domains. However, the massive volume of user-generated content, combined with its informal, nuanced, and often noisy nature, presents significant challenges for automated analysis and generation. Tasks such as stance detection, rumor verification, and personalized content generation are further complicated by sarcasm, evolving discourse structures, and diverse user preferences. Addressing these challenges requires models that can effectively leverage linguistic nuance, conversational dynamics, and collaborative user signals. Transfer learning has emerged as a powerful paradigm for improving performance in low-resource and complex language understanding tasks. …
Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto
Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto
School of Accountancy Faculty Publications
This study examines whether large language models exhibit systematic country-contingent differential treatment in financial fraud detection. Analyzing 30,000 synthetic transactions with identical statistical properties across three country attributions (United States, Great Britain, and China), we find LLMs assign significantly higher fraud probabilities to Chinese-attributed transactions (36.2%) compared to Western countries (≈30–31%), resulting in accuracy disparities of 67% versus 74%. The gap remains stable across five independent experimental replications and persists when using Chinese language prompts, ruling out linguistic effects. Bias mitigation strategies, such as requiring explanations or explicit country neutrality instructions, reduce but fail to eliminate these disparities. Testing across …
Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo
Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo
Research Collection School Of Computing and Information Systems
This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona …
Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma
Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma
Research Collection School Of Computing and Information Systems
Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …
Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma
Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma
Research Collection School Of Computing and Information Systems
Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …
Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang
Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang
Research Collection School Of Computing and Information Systems
Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …