Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer Science Faculty Publications and Presentations

Discipline
Institution
Keyword
Publication Year

Articles 1 - 30 of 404

Full-Text Articles in Computer Sciences

Vision Transformers And Convolutional Neural Networks For Land Use Scene Classification, Arun D. Kulkarni Jul 2026

Vision Transformers And Convolutional Neural Networks For Land Use Scene Classification, Arun D. Kulkarni

Computer Science Faculty Publications and Presentations

Land use scene classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable resource management. In recent years, deep learning methods have significantly advanced the state-of-the-art, with Convolutional Neural Networks (CNNs) dominating the field because of their strong ability to capture local spatial features. However, the emergence of Vision Transformers (ViTs) has introduced a new paradigm that models long-range dependencies through self attention mechanisms, potentially enabling improved global context understanding. This study presents a comparative assessment of Vision Transformers and CNN-based architectures for remote sensing land use scene classification. Representative CNN models, such …


An Integrated Framework For Memory-Centric Analysis: From Trace Collection To Co-Design, Dhruv Gajaria, Prajwal Challa, Yasodha Suriyakumar, Joseph Manzano, Nathan Tallent, Andrés Márquez May 2026

An Integrated Framework For Memory-Centric Analysis: From Trace Collection To Co-Design, Dhruv Gajaria, Prajwal Challa, Yasodha Suriyakumar, Joseph Manzano, Nathan Tallent, Andrés Márquez

Computer Science Faculty Publications and Presentations

IntroductionThe memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems-poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing contention effects, bandwidth saturation, and interference patterns that emerge at larger scales. These limitations reflect a processor-centric design philosophy—in both performance analysis tools and system co-design methodologies—that is increasingly misaligned with …


Agnostic Tomography Of Stabilizer Product States, Sabee Grewal, Vishnu Iyer, William Kretschmer, Daniel Liang Mar 2026

Agnostic Tomography Of Stabilizer Product States, Sabee Grewal, Vishnu Iyer, William Kretschmer, Daniel Liang

Computer Science Faculty Publications and Presentations

We define a quantum learning task called agnostic tomography, where given copies of an arbitrary state ρ and a class of quantum states C, the goal is to output a succinct description of a state that approximates ρ at least as well as any state in C (up to some small error ε). This task generalizes ordinary quantum tomography of states in C and is more challenging because the learning algorithm must be robust to perturbations of ρ. We give an efficient agnostic tomography algorithm for the class C of n-qubit stabilizer product states. Assuming ρ has fidelity at least …


A Cryptographic Perspective On The Verifiability Of Quantum Advantage, Nai-Hui Chia, Honghao Fu, Fang Song, Penghui Yao Mar 2026

A Cryptographic Perspective On The Verifiability Of Quantum Advantage, Nai-Hui Chia, Honghao Fu, Fang Song, Penghui Yao

Computer Science Faculty Publications and Presentations

In recent years, achieving verifiable quantum advantage on a NISQ device has emerged as an important open problem in quantum information. The sampling-based quantum advantages are not known to have efficient verification methods. This article investigates the verification of quantum advantage from a cryptographic perspective. We establish a strong connection between the verifiability of quantum advantage and cryptographic and complexity primitives, including efficiently samplable, statistically far but computationally indistinguishable pairs of (mixed) quantum states (EFI), pseudorandom states (PRS), and variants of minimum circuit size problems (MCSP). Specifically, we prove that a) a sampling-based quantum advantage is either verifiable or can …


Clustering Of Temporal And Visual Data: Recent Advancements, Priyanka Mudgal Jan 2026

Clustering Of Temporal And Visual Data: Recent Advancements, Priyanka Mudgal

Computer Science Faculty Publications and Presentations

Clustering plays a central role in uncovering latent structure within both temporal and visual data. It enables critical insights in various domains including healthcare, finance, surveillance, autonomous systems, and many more. With the growing volume and complexity of time-series and image-based datasets, there is an increasing demand for robust, flexible, and scalable clustering algorithms. Although these modalities differ—time-series being inherently sequential and vision data being spatial—they exhibit common challenges such as high dimensionality, noise, variability in alignment and scale, and the need for interpretable groupings. This survey presents a comprehensive review of recent advancements in clustering methods that are adaptable …


Wild@Fire2025: Overview Of Word-Level Code-Mixed Language Identification In Dravidian Languages, Ameeta Agrawal, Asha Hegde, Sharal Coelho, Sabur Butt, Fazlourrahman Balouchzahi, Sudha V, Shashirekha Hosahalli Lakshmaiah Jan 2026

Wild@Fire2025: Overview Of Word-Level Code-Mixed Language Identification In Dravidian Languages, Ameeta Agrawal, Asha Hegde, Sharal Coelho, Sabur Butt, Fazlourrahman Balouchzahi, Sudha V, Shashirekha Hosahalli Lakshmaiah

Computer Science Faculty Publications and Presentations

Code-mixing is considered as a linguistic phenomenon that combines several languages into one text. It has now become very common in multilingual societies, especially in digital communication. Word-Level Identification of Languages in Dravidian Languages (WILD) - a Code-mixed Language Identification (CoLI) in Dravidian languages shared task, organized as a part of Forum for Information Retrieval and Evaluation (FIRE) 2025, put forward these challenges to the researchers by asking them to develop models capable of classifying words in code-mixed texts involving Dravidian languages - Tamil, Telugu, Malayalam, Kannada, and Tulu, which are interwoven with English. It poses significant challenges due to …


Adaptive Image Acquisition Algorithms For Resource-Constrained Single-Photon Cameras, Yeganeh Jalalpour, Wu-Chi Feng Dec 2025

Adaptive Image Acquisition Algorithms For Resource-Constrained Single-Photon Cameras, Yeganeh Jalalpour, Wu-Chi Feng

Computer Science Faculty Publications and Presentations

Emerging single-photon camera (SPC) technologies have unique challenges in data acquisition and processing. Unlike conventional sensors that produce a single 8- to 16-bit brightness value per pixel, SPCs record photon arrivals with many more samples per pixel, using high floating-point precision for each photon collected. This means that they must handle potentially millions of timestamps, especially at higher spatial resolutions and in the presence of ambient light, creating bottlenecks within the pixel circuitry. To address these challenges associated with SPCs, this paper proposes adaptive algorithms designed to efficiently distribute hardware resources among groups of pixels. By selectively subsampling the data …


Hint-Guided Video Frame Interpolation For Video Compression, Pan Tan, Wu-Chi Feng Dec 2025

Hint-Guided Video Frame Interpolation For Video Compression, Pan Tan, Wu-Chi Feng

Computer Science Faculty Publications and Presentations

Traditional video compression continues to advance, but the gainsin efficiency are diminishing and come at the cost of higher compu-tational complexity. Despite achieving competitive rate-distortionresults, current neural video codecs (NVCs) generally lack sup-port for a wide range of quality levels, often requiring multiplemodels to achieve flexible rate control, which increases both train-ing cost and deployment complexity. To address the limitations ofboth traditional codecs and current NVCs, we propose a hybridvideo compression framework that integrates traditional codecswith hint-guided video frame interpolation (VFI), a learning-basedtechnique for synthesizing intermediate frames. By using decodedreference frames and leveraging compressed-domain hints to guideinterpolation, our method improves …


Understanding The Role Of Sentiment And Emotion For Predicting Forced Displacement, Helge Marahrens, Ameeta Agrawal, Ali Arab, Katharine Donato, Yaguang Liu, Nathan Wycoff, Mohamed Ahmed, Colin Hwang, Lina Laghzaoui, Kate Liggio, Multiple Additional Authors Oct 2025

Understanding The Role Of Sentiment And Emotion For Predicting Forced Displacement, Helge Marahrens, Ameeta Agrawal, Ali Arab, Katharine Donato, Yaguang Liu, Nathan Wycoff, Mohamed Ahmed, Colin Hwang, Lina Laghzaoui, Kate Liggio, Multiple Additional Authors

Computer Science Faculty Publications and Presentations

Digital trace data play an important role determining where and when people will move during migration crises because of their detailed temporal and spatial granularity. Yet, identifying variables that reliably serve as early indicators of movement remains a challenging task. Within this context, we conduct an in-depth analysis of two types of variables that can be constructed from social media data – sentiment and emotion. Sentiment is conceptually broad and easier to detect from social media posts, while emotion is conceptually nuanced and more difficult to determine. We investigate the potential of both sentiment and emotion of Twitter/X posts as …


A Case Study On The Effectiveness Of Llms In Verification With Proof Assistants, Barış Bayazıt, Yao Li, Xujie Si Oct 2025

A Case Study On The Effectiveness Of Llms In Verification With Proof Assistants, Barış Bayazıt, Yao Li, Xujie Si

Computer Science Faculty Publications and Presentations

Large language models (LLMs) can potentially help with verification using proof assistants by automating proofs. However, it is unclear how effective LLMs are in this task. In this paper, we perform a case study based on two mature Rocq projects: the hs-to-coq tool and Verdi. We evaluate the effectiveness of LLMs in generating proofs by both quantitative and qualitative analysis. Our study finds that: (1) external dependencies and context in the same source file can significantly help proof generation; (2) LLMs perform great on small proofs but can also generate large proofs; (3) LLMs perform differently on different verification projects; …


Improved Fpt Approximation For Sum Of Radii Clustering With Mergeable Constraints, Sayan Bandyapadhyay, Tainzhi Chen Sep 2025

Improved Fpt Approximation For Sum Of Radii Clustering With Mergeable Constraints, Sayan Bandyapadhyay, Tainzhi Chen

Computer Science Faculty Publications and Presentations

In this work, we study k-min-sum-of-radii (k-MSR) clustering under mergeable constraints. k-MSR seeks to group data points using a set of up to k balls, such that the sum of the radii of the balls is minimized. A clustering constraint is called mergeable if merging two clusters satisfying the constraint, results in a cluster that also satisfies the constraint. Many popularly studied constraints are mergeable, including fairness constraints and lower bound constraints. In our work, we design a (4 + ϵ)-approximation for k-MSR under any given mergeable constraint with runtime 2 O( k ϵ ·log2 k ϵ )n 4 , …


Approximation And Parameterized Algorithms For Covering With Disks Of Two Types Of Radii, Sayan Bandyapadhyay, Eli Mitchell Aug 2025

Approximation And Parameterized Algorithms For Covering With Disks Of Two Types Of Radii, Sayan Bandyapadhyay, Eli Mitchell

Computer Science Faculty Publications and Presentations

We study the Discrete Covering with Two Types of Radii problem motivated by its application in wireless networks. In this problem, the goal is to assign either small-range high frequency or large-range low frequency to each access point, maximizing the number of users in high-frequency regions while ensuring that each user is in the range of an access point. Unlike other weighted covering problems, our problem requires satisfying two simultaneous objectives, which calls for novel approaches that leverage the underlying geometry of the problem. In our work, we present two new algorithms: the first is a polynomial-time (2.5 + ϵ)-approximation, …


Towards Scalable Schema Mapping Using Large Language Models, Christopher Buss, Mahdis Safari, Arash Termehchy, David Maier, Stefan Lee Aug 2025

Towards Scalable Schema Mapping Using Large Language Models, Christopher Buss, Mahdis Safari, Arash Termehchy, David Maier, Stefan Lee

Computer Science Faculty Publications and Presentations

The growing need to integrate information from many diverse sources poses significant scalability challenges for data integration systems. These systems often rely on manually written schema mappings, which are complex and costly to maintain. While recent advances suggest that large language models (LLMs) can assist in automating schema mapping, key challenges remain. We motivate future research in schema mapping generation by highlighting key challenges, presenting a competitive bidirectional schema matching pipeline, and exploring the limitations of current methods for generating more complex mappings.


Imputation Via Domain Adaptation: Rethinking Variable Subset Forecasting From Knowledge Transfer, Runchang Liang, Qi Hao, Yue Gao, Kunpeng Liu, Lu Jiang, Pengyang Wang, Minghao Yin Aug 2025

Imputation Via Domain Adaptation: Rethinking Variable Subset Forecasting From Knowledge Transfer, Runchang Liang, Qi Hao, Yue Gao, Kunpeng Liu, Lu Jiang, Pengyang Wang, Minghao Yin

Computer Science Faculty Publications and Presentations

Multivariate time series forecasting in practical deployment faces a critical challenge termed Variable Subset Forecasting (VSF), where certain variables accessible during training are entirely missing during inference. This creates a stark discrepancy between the training (source domain with full variables) and inference (target domain with partial variables) environments, disrupting cross-variable dependencies and fragmenting global temporal patterns. Existing imputation methods, limited to transferring local knowledge (e.g., temporal neighbors or pairwise correlations), fail to capture essential global dynamics, leading to severe performance degradation under distribution shifts. To address these challenges, we redefine VSF as a cross-domain knowledge transfer problem and propose VIDA, …


Coli@Fire2024: Findings Of Word-Level Code-Mixed Language Identification In Dravidian Languages, Asha Hegde, Fazlourrahman Balouchzahi, Sabur Butt, Sharal Coelho, Kavya G, Harshitha S. Kumar, Sonith D, Shashirekha H. L., Ameeta Agrawal Jul 2025

Coli@Fire2024: Findings Of Word-Level Code-Mixed Language Identification In Dravidian Languages, Asha Hegde, Fazlourrahman Balouchzahi, Sabur Butt, Sharal Coelho, Kavya G, Harshitha S. Kumar, Sonith D, Shashirekha H. L., Ameeta Agrawal

Computer Science Faculty Publications and Presentations

Code-mixing, a linguistic phenomenon where multiple languages are blended within a single text, has become increasingly prevalent in multilingual societies, particularly in digital communication. The CoLI-Dravidian shared task, organized as part of Forum for Information Retrieval and Evaluation (FIRE) 2024, aimed to address these challenges by inviting researchers to develop models capable of classifying words in code-mixed texts involving Dravidian languages — Tamil, Kannada, Malayalam, and Tulu - interwoven with English. The task presents significant challenges due to the complexity of linguistic structures, mixed-language tokens, and dialectal variations, especially in low-resource languages like those in the Dravidian family. The participating …


Robust Contraction Decomposition For Minor-Free Graphs And Its Applications, Bandyapadhyay Sayan, William Lochet, Daniel Lokshtanov, Dániel Marx, Pranabendu Misra, Multiple Additional Authors Jun 2025

Robust Contraction Decomposition For Minor-Free Graphs And Its Applications, Bandyapadhyay Sayan, William Lochet, Daniel Lokshtanov, Dániel Marx, Pranabendu Misra, Multiple Additional Authors

Computer Science Faculty Publications and Presentations

We prove a robust contraction decomposition theorem for H-minor-free graphs, which states that given an H-minor-free graph G and an integer p, one can partition in polynomial time the vertices of G into p sets Z₁,… ,Z_p such that tw(G/(Z_i ⧵ Z')) = O(p + |Z'|) for all i ∈ [p] and Z' ⊆ Z_i. Here, tw(⋅) denotes the treewidth of a graph and G/(Z_i ⧵ Z') denotes the graph obtained from G by contracting all edges with both endpoints in Z_i ⧵ Z'. Our result generalizes earlier results by Klein [SICOMP 2008] and Demaine et al. [STOC 2011] based …


Advanced Machine Learning Techniques For Social Support Detection On Social Media, Olga Kolesnikova, Moein Shahiki Tash, Zahra Ahani, Ameeta Agrawal, Raúl Monroy, Grigori Sidorov May 2025

Advanced Machine Learning Techniques For Social Support Detection On Social Media, Olga Kolesnikova, Moein Shahiki Tash, Zahra Ahani, Ameeta Agrawal, Raúl Monroy, Grigori Sidorov

Computer Science Faculty Publications and Presentations

The widespread use of social media highlights the need to understand its impact, particularly the role of online social support. In this study, we present a dataset of YouTube comments, initially comprising 66,272 entries, which was refined to 42,695, with a subset of 10,000 comments selected for detailed analysis without additional filtering. The dataset is annotated for three classification tasks: (1) distinguishing supportive from non-supportive comments, (2) determining whether the support is directed at an individual or a group, and (3) further categorizing group support into six subtypes (Nation, LGBTQ, Black Community, Women, Religion, and Other). To address data imbalances …


Deep Learning Classification Of Drainage Crossings Based On High-Resolution Dem-Derived Geomorphological Information, Michael Edidem, Bill Xu, Ruopu Li, Di Wu, Banafsheh Rekabdar, Guangxing Wang May 2025

Deep Learning Classification Of Drainage Crossings Based On High-Resolution Dem-Derived Geomorphological Information, Michael Edidem, Bill Xu, Ruopu Li, Di Wu, Banafsheh Rekabdar, Guangxing Wang

Computer Science Faculty Publications and Presentations

High-resolution digital elevation models (HRDEMs) from LiDAR and InSAR technologies have significantly improved the accuracies of mapping hydrographic features such as river boundaries, streamlines, and waterbodies over large areas. However, drainage crossings that facilitate the passage of drainage flows beneath roads are not often represented in HRDEMs, resulting in erratic or distorted hydrographic features. At present, drainage crossing datasets are largely missing or available with variable quality. While previous studies have investigated basic convolutional neural network (CNN) models for drainage crossing characterization, it remains unclear if advanced deep learning models will improve the accuracy of drainage crossing classification. Although HRDEM-derived …


Unremarkable To Remarkable Ai Agent: Exploring Boundaries Of Agent Intervention For Adults With And Without Cognitive Impairment, Mai Lee Chang, Samantha Reig, Alicia (Hyun Jin) Lee, Anna Huang, Hugo Simão, Nara Han, Neeta M. Khanuja, Abdullah Ubed Mohammad Ali, Rebekah Martinez, John Zimmerman, Jodi Forlizzi, Aaron Steinfeld May 2025

Unremarkable To Remarkable Ai Agent: Exploring Boundaries Of Agent Intervention For Adults With And Without Cognitive Impairment, Mai Lee Chang, Samantha Reig, Alicia (Hyun Jin) Lee, Anna Huang, Hugo Simão, Nara Han, Neeta M. Khanuja, Abdullah Ubed Mohammad Ali, Rebekah Martinez, John Zimmerman, Jodi Forlizzi, Aaron Steinfeld

Computer Science Faculty Publications and Presentations

As the population of older adults increases, there is a growing need for support for them to age in place. This is exacerbated by the growing number of individuals struggling with cognitive decline and shrinking number of youth who provide care for them. Artificially intelligent agents could provide cognitive support to older adults experiencing memory problems, and they could help informal caregivers with coordination tasks. To better understand this possible future, we conducted a speed dating with storyboards study to reveal invisible social boundaries that might keep older adults and their caregivers from accepting and using agents. We found that …


Next Arrival And Destination Prediction Via Spatiotemporal Embedding With Urban Geography And Human Mobility Data, Pengjiang Li, Zaitian Wang, Xinhao Zhang, Pengfei Wang, Kunpeng Liu Mar 2025

Next Arrival And Destination Prediction Via Spatiotemporal Embedding With Urban Geography And Human Mobility Data, Pengjiang Li, Zaitian Wang, Xinhao Zhang, Pengfei Wang, Kunpeng Liu

Computer Science Faculty Publications and Presentations

With the development of transportation networks, countless trajectory data are accumulated, and understanding human mobility from traffic data could be helpful for smart cities, urban computing, and urban planning. Extracting valuable insights from traffic data, such as taxi trajectories, can significantly improve residents’ daily lives. There are many studies on spatiotemporal data mining. As we know, arrival prediction or regional function detection encompasses important tasks for traffic management and urban planning. However, trajectory data are often mutilated because of personal privacy and hardware limitations, i.e., we usually can only obtain partial trajectory information. In this paper, we develop an embedding …


Foundation Models Boost Low-Level Perceptual Similarity Metrics, Abhijay Ghildyal, Nabajeet Barman, Saman Zadtootaghaj Jan 2025

Foundation Models Boost Low-Level Perceptual Similarity Metrics, Abhijay Ghildyal, Nabajeet Barman, Saman Zadtootaghaj

Computer Science Faculty Publications and Presentations

For full-reference image quality assessment (FR-IQA) using deep-learning approaches, the perceptual similarity score between a distorted image and a reference image is typically computed as a distance measure between features extracted from a pretrained CNN or more recently, a Transformer network. Often, these intermediate features require further fine-tuning or processing with additional neural network layers to align the final similarity scores with human judgments. So far, most IQA models based on foundation models have primarily relied on the final layer or the embedding for the quality score estimation. In contrast, this work explores the potential of utilizing the intermediate features …


Drta: Dynamic Reward Scaling For Reinforcement Learning In Time Series Anomaly Detection, Bahareh Golchin, Banafsheh Rekabdar, Kunpeng Liu Jan 2025

Drta: Dynamic Reward Scaling For Reinforcement Learning In Time Series Anomaly Detection, Bahareh Golchin, Banafsheh Rekabdar, Kunpeng Liu

Computer Science Faculty Publications and Presentations

Anomaly detection in time series data is important for applications in finance, healthcare, sensor networks, and industrial monitoring. Traditional methods usually struggle with limited labeled data, high false-positive rates, and difficulty generalizing to novel anomaly types. To overcome these challenges, we propose a reinforcement learning-based framework that integrates dynamic reward shaping, Variational Autoencoder (VAE), and active learning, called DRTA. Our method uses an adaptive reward mechanism that balances exploration and exploitation by dynamically scaling the effect of VAE-based reconstruction error and classification rewards. This approach enables the agent to detect anomalies effectively in low-label systems while maintaining high precision and …


Scitopic: Enhancing Topic Discovery In Scientific Literature Through Advanced Llm, Pengjiang Li, Zaitian Wang, Xinhao Zhang, Ran Zhang, Lu Jiang, Pengfei Wang, Yuanchun Zhou Jan 2025

Scitopic: Enhancing Topic Discovery In Scientific Literature Through Advanced Llm, Pengjiang Li, Zaitian Wang, Xinhao Zhang, Ran Zhang, Lu Jiang, Pengfei Wang, Yuanchun Zhou

Computer Science Faculty Publications and Presentations

Topic discovery in scientific literature provides valuable insights for researchers to identify emerging trends and explore new avenues for investigation, facilitating easier scientific information retrieval. Many machine learning methods, particularly deep embedding techniques, have been applied to discover research topics. However, most existing topic discovery methods rely on word embedding to capture the semantics and lack a comprehensive understanding of scientific publications, struggling with complex, high-dimensional text relationships. Inspired by the exceptional comprehension of textual information by large language models (LLMs), we propose an advanced topic discovery method enhanced by LLMs to improve scientific topic identification, namely SciTopic. Specifically, we …


Polynomial-Time Constant-Approximation For Fair Sum-Of-Radii Clustering, Sina Nezhad, Sayan Bandyapadhyay, Tianzhi Chen Jan 2025

Polynomial-Time Constant-Approximation For Fair Sum-Of-Radii Clustering, Sina Nezhad, Sayan Bandyapadhyay, Tianzhi Chen

Computer Science Faculty Publications and Presentations

In a seminal work, Chierichetti et al. [20] introduced the (t,k)-fair clustering problem: Given a set of red points and a set of blue points in a metric space, a clustering is called fair if the number of red points in each cluster is at most t times and at least 1/t times the number of blue points in that cluster. The goal is to compute a fair clustering with at most k clusters that optimizes certain objective function. Considering this problem, they designed a polynomial-time O(1)- and O(t)-approximation for the k-center and the k-median objective, respectively. Recently, Carta et …


“Lost-In-The-Later”: Framework For Quantifying Contextual Grounding In Large Language Models, Yufei Tao, Adam Hiatt, Rahul Seetharaman, Ameeta Agrawal Jan 2025

“Lost-In-The-Later”: Framework For Quantifying Contextual Grounding In Large Language Models, Yufei Tao, Adam Hiatt, Rahul Seetharaman, Ameeta Agrawal

Computer Science Faculty Publications and Presentations

Large language models (LLMs) are capable of leveraging both contextual and parametric knowledge but how they prioritize and integrate these sources remains underexplored. We introduce CoPE, a novel framework for systematically quantifying contextual grounding in LLMs. CoPE distinguishes between contextual knowledge (CK) and parametric knowledge (PK), enabling fine-grained attribution across languages and tasks. Using our newly created MultiWikiAtomic dataset in English, Spanish, and Danish, we analyze how LLMs integrate context, prioritize information, and incorporate PK in open-ended question answering. We find that across models and languages, only around 50 to 76 percent of outputs are grounded in the given context, …


Retrieval-Augmented Feature Generation For Domain-Specific Classification, Xinhao Zhang, Jinghan Zhang, Fengran Mo, Dakshak Keerthi Chandra, Yu-Zhong Chen, Fei Xie, Kunpeng Liu Jan 2025

Retrieval-Augmented Feature Generation For Domain-Specific Classification, Xinhao Zhang, Jinghan Zhang, Fengran Mo, Dakshak Keerthi Chandra, Yu-Zhong Chen, Fei Xie, Kunpeng Liu

Computer Science Faculty Publications and Presentations

Feature generation can significantly enhance learning outcomes, particularly for tasks with limited data. An effective way to improve feature generation is to expand the current feature space using existing features and enriching the informational content. However, generating new, interpretable features usually requires domain-specific knowledge on top of the existing features. In this paper, we introduce a Retrieval-Augmented Feature Generation method, RAFG, to generate useful and explainable features specific to domain classification tasks. To increase the interpretability of the generated features, we conduct knowledge retrieval among the existing features in the domain to identify potential feature associations. These associations are expected …


Osa-Diff: An Origin Sampling Based Adversarial Attack Using Diffusion Models, Shayan Jalalipour, Banafsheh Rekabdar Jan 2025

Osa-Diff: An Origin Sampling Based Adversarial Attack Using Diffusion Models, Shayan Jalalipour, Banafsheh Rekabdar

Computer Science Faculty Publications and Presentations

Diffusion models are becoming an increasingly popular emerging technology, however their use in adversarial attacks remains a scarcely explored topic. We show that diffusion models can be used to create end-to-end hidden adversarial perturbations with high rate of success, and propose a novel diffusion based adversarial attack that allows for substantially faster training time (through improved convergence on high quality images) and with substantially less computational overhead than typical diffusion model training


Tabular Data-Centric Ai: Challenges, Techniques And Future Perspectives, Yanjie Fu, Dongjie Wang, Hui Xiong, Kunpeng Liu Oct 2024

Tabular Data-Centric Ai: Challenges, Techniques And Future Perspectives, Yanjie Fu, Dongjie Wang, Hui Xiong, Kunpeng Liu

Computer Science Faculty Publications and Presentations

Tabular data are the most widely used data formats in almost every application domain, such as, biology, ecology, and material science. The purpose of tabular data-centric AI is to use AI to augment the predictive power of tabular data to get better AI. Tabular data-centric AI is essential because it can reconstruct distance measures, reshape discriminative patterns, and improve data AI readiness (structural, predictive, interaction, and expression levels), which is significant in industries and real-world deployments. Therefore, our tutorial is designed to capture the interest of professionals with expertise in artificial intelligence, machine learning, and data mining, as well as …


Story Of Your Lazy Function's Life: A Bidirectional Demand Semantics For Mechanized Cost Analysis Of Lazy Programs, Li-Yao Xia, Laura Israel, Maite Kramarz, Nicholas Coltharp, Koen Claessen, Stephanie Weirich, Yao Li Aug 2024

Story Of Your Lazy Function's Life: A Bidirectional Demand Semantics For Mechanized Cost Analysis Of Lazy Programs, Li-Yao Xia, Laura Israel, Maite Kramarz, Nicholas Coltharp, Koen Claessen, Stephanie Weirich, Yao Li

Computer Science Faculty Publications and Presentations

Lazy evaluation is a powerful tool that enables better compositionality and potentially better performance in functional programming, but it is challenging to analyze its computation cost. Existing works either require manually annotating sharing, or rely on separation logic to reason about heaps of mutable cells. In this paper, we propose a bidirectional demand semantics that allows for extrinsic reasoning about the computation cost of lazy programs without relying on special program logics. To show the effectiveness of our approach, we apply the demand semantics to a variety of case studies including insertion sort, selection sort, Okasaki's banker's queue, and the …


True Contraction Decomposition And Almost Eth-Tight Bipartization For Unit-Disk Graphs, Sayan Bandyapadhyay, William Lochet, Daniel Lokshtanov, Saket Saurabh, Jie Xue Jul 2024

True Contraction Decomposition And Almost Eth-Tight Bipartization For Unit-Disk Graphs, Sayan Bandyapadhyay, William Lochet, Daniel Lokshtanov, Saket Saurabh, Jie Xue

Computer Science Faculty Publications and Presentations

We prove a structural theorem for unit-disk graphs, which (roughly) states that given a set D of n unit disks inducing a unit-disk graph ...