Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 3481 - 3495 of 3495

Full-Text Articles in Computer Sciences

Adapting Online Customer Reviews For Blind Users: A Case Study Of Restaurant Reviews, Mohan Sunkara, Akshay Kolgar Nayak, Sandeep Kalari, Yash Prakash, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok Jan 2025

Adapting Online Customer Reviews For Blind Users: A Case Study Of Restaurant Reviews, Mohan Sunkara, Akshay Kolgar Nayak, Sandeep Kalari, Yash Prakash, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok

Computer Science Faculty Publications

Online reviews have become an integral aspect of consumer decision-making on e-commerce websites, especially in the restaurant industry. Unlike sighted users who can visually skim through the reviews, perusing reviews remains challenging for blind users, who rely on screen reader assistive technology that supports predominantly one-dimensional narration of content via keyboard shortcuts. In an interview study, we uncovered numerous pain points of blind screen reader users with online restaurant reviews, notably, the listening fatigue and frustration after going through only the first few reviews. To address these issues, we developed QuickCue assistive tool that performs aspect-focused sentiment-driven summarization to reorganize …


Adversarially Attacking Graph Properties And Sparsification In Graph Learning, Chunjiang Zhu, Blake Gaines, Jing Deng, Jinbo Bi Jan 2025

Adversarially Attacking Graph Properties And Sparsification In Graph Learning, Chunjiang Zhu, Blake Gaines, Jing Deng, Jinbo Bi

Computer Science Faculty Publications

Graph neural networks and graph transformers explicitly or implicitly rely on fundamental properties of the underlying graph, such as spectral properties and shortest-path distances. However, it is still not clear how these graph properties are vulnerable to adversarial attacks and what impacts this has on the downstream graph learning. Moreover, while graph sparsification has been used to improve computational cost of learning over graphs, its susceptibility to adversarial attacks has not been studied. In this paper, we study adversarial attacks on graph properties and graph sparsification and their impacts on downstream graph learning, paving the way for how to protect …


Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput Jan 2025

Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput

Computer Science Faculty Publications

In this study, we address the mounting challenge of monitoring high throughput computing clusters running computationally intensive jobs, which increasingly strains system administrators. We develop autoencoders that analyze traces of Linux kernel CPU metrics to capture salient system features by producing robust compressed embeddings for various downstream tasks. In addition, we employ graph neural networks to incorporate contextual information from surrounding CPUs and assess their performance. We also demonstrate the enhanced job differentiation achieved by increasing the sampling rate of these traces. Our models are evaluated based on their ability to generate meaningful latent representations, detect anomalies, and distinguish between …


From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer Jan 2025

From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer

Computer Science Faculty Publications

Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …


S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala Jan 2025

S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala

Computer Science Faculty Publications

Feature Distillation (FD) strategies are proven to be effective in mitigating Catastrophic Forgetting (CF) seen in Class Incremental Learning (CIL). However, current FD approaches enforce strict alignment of feature magnitudes and directions across incremental steps, limiting the model’s ability to adapt to new knowledge. In this paper, we propose Structurally Stable Incremental Learning (S²IL), a FD method for CIL that mitigates forgetting by focusing on preserving the overall spatial patterns of features which promote flexible (plasticity) yet stable representations that preserve old knowledge (stability). We also demonstrate that our proposed method S²IL achieves strong incremental accuracy and outperforms other FD …


Benchmarking And Improving Foundation Model Dietary Estimates From Meal Images, Yongcheng Mu, Jiangwen Sun, Jing He Jan 2025

Benchmarking And Improving Foundation Model Dietary Estimates From Meal Images, Yongcheng Mu, Jiangwen Sun, Jing He

Computer Science Faculty Publications

Accurate quantifying dietary contents, such as calories, proteins, carbohydrates, and fats, from an image of a meal plate is vital for managing diabetes. Recently, Large Multimodal Models (LMMs) have excelled in complex vision-language tasks due to their use of very large, highly diverse data. This study benchmarked the use of seven LMMs that include full and lightweight models of GPT, Gemini, and Llama for nutrition estimation based on Google's Nutrition5k dataset and our own phone-collected DonateAndLearn dataset. We analyzed the performance of LMMs and the RGB-D fusion model, in which the RGB-D model was specifically trained using Nutrition5k data. On …


Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He Jan 2025

Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He

Computer Science Faculty Publications

DeepSSETracer is a method for segmenting protein secondary structure from medium-resolution (5-10Å) cryogenic electron microscopy (cryo-EM) density maps. We conducted experiments and ablation studies to examine the effects of normalization methods, max-pooling, activation functions, and loss calculation region on DeepSSETracer. By combining multiple technical improvements, the performance of the new version, DeepSSETracer 2.0, was significantly enhanced compared to DeepSSETracer 1.1. On a set of 77 test cases, the weighted average per-voxel F1 score increased from 62.1% to 70.3% for helix detection, and from 47.8% to 62.5% for β-sheet detection. While each of the five modifications in the network enhanced the …


Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li Jan 2025

Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li

Computer Science Faculty Publications

Large Language Models (LLMs) exhibit strong natural language processing capabilities but also pose significant privacy risks, particularly regarding the leakage of Personally Identifiable Information (PII) embedded in their training data. Existing PII extraction methods suffer from the limitations of low success rates or impracticality for large-scale PII extraction. In this study, we propose a novel PII extraction approach based on enhanced few-shot learning techniques, which achieves efficient and cost-effective PII retrieval without relying on fine-tuning or jailbreaking. We evaluated our approach on both open-source and closed-source LLMs. The experimental results demonstrate that, for non-targeted PII extraction, the attack success rate …


An Optimized Generalized Multi-Color Point Implicit Solver For Intel Gpus Using Oneapi Esimd, Joseph Wassell, Mohammad Zubair, Aaron Walden, Gabriel Nastac, Eric Nielsen, Timothée Ewart Jan 2025

An Optimized Generalized Multi-Color Point Implicit Solver For Intel Gpus Using Oneapi Esimd, Joseph Wassell, Mohammad Zubair, Aaron Walden, Gabriel Nastac, Eric Nielsen, Timothée Ewart

Computer Science Faculty Publications

This paper presents an efficient implementation of a linear-solver kernel relevant to FUN3D, a suite of computational fluid dynamics software developed at NASA’s Langley Research Center. The linear solver is optimized for a range of block sizes commonly used in FUN3D. The implementation targets Aurora, the Argonne Leadership Computing Facility’s (ALCF) exascale machine featuring Intel Data Center Max 1550 GPUs. The linear solver’s performance is memory bandwidth-bound due to its low arithmetic intensity. The primary performance challenges stem from variable matrix row lengths and indirect memory access patterns inherent in unstructured-grid applications. Variable block sizes introduce additional complexity through differing …


Understanding Pii Leakage In Large Language Models: A Systematic Survey, Shuai Cheng, Zhao Li, Shu Meng, Mengxia Ren, Haitao Xu, Shuai Hao, Chuan Yue, Fang Zhang Jan 2025

Understanding Pii Leakage In Large Language Models: A Systematic Survey, Shuai Cheng, Zhao Li, Shu Meng, Mengxia Ren, Haitao Xu, Shuai Hao, Chuan Yue, Fang Zhang

Computer Science Faculty Publications

Large Language Models (LLMs) have demonstrated exceptional success across a variety of tasks, particularly in natural language processing, leading to their growing integration into numerous facets of daily life. However, this widespread deployment has raised substantial privacy concerns, especially regarding personally identifiable information (PII), which can be directly associated with specific individuals. The leakage of such information presents significant real-world privacy threats. In this paper, we conduct a systematic investigation into existing research on PII leakage in LLMs, encompassing commonly utilized PII datasets, evaluation metrics, and current studies on both PII leakage attacks and defensive strategies. Finally, we identify unresolved …


Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li Jan 2025

Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li

Computer Science Faculty Publications

Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a …


Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh Jan 2025

Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh

Computer Science Faculty Publications

Predicting the length of stay (LoS) is important for hospital administration, as it helps allocate proper resources, such as bed management and hospital staffing. Patients' Electronic Health Records (EHRs) contain highly relevant data for LoS prediction; however, their integration and effective use in predictive modeling for accurately estimating LoS remain challenging. To address this, we propose a homogeneous Graph Neural Network (GNN)-based framework for predicting LoS. This method employs a comprehensive data fusion strategy based on the hospital Visit-based Similarity Graph (VSG), which integrates diverse multi-modal clinical features into a coherent, homogeneous graph representation. Next, this VSG is fed into …


Humans Vs. Llms On Open Domain Scientific Claim Verification: A Baseline Study, Benjamin Curtis, Stefania Dzhaman, Matthew Maisonave, Jian Wu Jan 2025

Humans Vs. Llms On Open Domain Scientific Claim Verification: A Baseline Study, Benjamin Curtis, Stefania Dzhaman, Matthew Maisonave, Jian Wu

Computer Science Faculty Publications

Verifying scientific claims is challenging for the general public because most people lack domain knowledge. Manual verification by subject domain experts is accurate, but it is obviously not scalable to meet the rising number of scientific claims on the Web. Whether the emerging large language models and large reasoning models can be used for scientific claim verification, and how their performances compare to humans, are still research questions. To this end, we developed a new benchmark MSVEC2 that consists of 138 claims from credible fact verification websites and science news outlets. Two tasks were given to both human and LLM …


Securing Secrets: Exploring The Aes Encryption And Key Security Capabilities Of Chatgpt, Kayla Taylor Jan 2025

Securing Secrets: Exploring The Aes Encryption And Key Security Capabilities Of Chatgpt, Kayla Taylor

Student Works

The development and increasing accessibility of generative artificial intelligence (AI) tools and large language models (LLMs) have allowed cryptographers to explore a variety of cryptanalysis problems in dynamic and interactive ways. Prompt engineering, the process by which input text is tested and refined to elicit a desired response from LLMs, is a nascent area of research that remains largely unexplored in many contexts, including cryptography. This study will explore the potential applications and limitations of prompt engineering in the context of Advanced Encryption Standard (AES) encryption and key security with OpenAI’s ChatGPT (GPT-4o) through two main objectives: First, given a …


The Great Scrape: The Clash Between Scraping And Privacy, Daniel J. Solove, Woodrow Hartzog Jan 2025

The Great Scrape: The Clash Between Scraping And Privacy, Daniel J. Solove, Woodrow Hartzog

Faculty Scholarship

Artificial intelligence (AI) systems depend on massive quantities of data, often gathered by “scraping”—the automated extraction of large amounts of data from the internet. A great deal of scraped data contains people’s personal information. This personal data provides the grist for AI tools such as facial recognition, deep fakes, and generative AI. Although scraping enables web searching, archiving of records, and meaningful scientific research, scraping for AI can also be objectionable and even harmful to individuals and society.

Organizations are scraping at an escalating pace and scale, even though many privacy laws are seemingly incongruous with the practice. In this …