Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (108)
- Engineering (68)
- Computer Engineering (46)
- Databases and Information Systems (41)
- Social and Behavioral Sciences (37)
-
- Electrical and Computer Engineering (28)
- Data Science (24)
- Numerical Analysis and Scientific Computing (23)
- Other Computer Sciences (22)
- Software Engineering (18)
- Medicine and Health Sciences (16)
- Business (14)
- Communication (14)
- Information Security (14)
- Arts and Humanities (12)
- Linguistics (9)
- Social Media (9)
- Theory and Algorithms (9)
- Law (8)
- Library and Information Science (8)
- Computational Linguistics (6)
- Education (6)
- Operations Research, Systems Engineering and Industrial Engineering (6)
- Bioinformatics (5)
- Graphics and Human Computer Interfaces (5)
- Life Sciences (5)
- Public Affairs, Public Policy and Public Administration (5)
- Scholarly Publishing (5)
- Institution
-
- Singapore Management University (47)
- Old Dominion University (30)
- TÜBİTAK (17)
- Technological University Dublin (13)
- Wright State University (9)
-
- Dartmouth College (7)
- New Jersey Institute of Technology (7)
- Boise State University (6)
- City University of New York (CUNY) (6)
- United Arab Emirates University (6)
- Virginia Commonwealth University (6)
- Zayed University (6)
- Brigham Young University (5)
- Loyola University Chicago (5)
- California Polytechnic State University, San Luis Obispo (4)
- San Jose State University (4)
- University of Arkansas Little Rock (4)
- University of Arkansas, Fayetteville (4)
- University of Denver (4)
- Western Michigan University (4)
- Air Force Institute of Technology (3)
- Edith Cowan University (3)
- Mississippi State University (3)
- Missouri University of Science and Technology (3)
- Montclair State University (3)
- Purdue University (3)
- Southern Methodist University (3)
- The Texas Medical Center Library (3)
- University of Kentucky (3)
- University of Nebraska at Omaha (3)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (37)
- Theses and Dissertations (18)
- Turkish Journal of Electrical Engineering and Computer Sciences (17)
- Dissertations (15)
- Electronic Theses and Dissertations (10)
-
- Computer Science Faculty Publications (8)
- All Works (6)
- Browse all Theses and Dissertations (6)
- Conference papers (6)
- Master's Theses (6)
- VMASC Publications (6)
- Dissertations and Theses Collection (Open Access) (5)
- Boise State University Theses and Dissertations (4)
- Computer Science Faculty Proceedings & Presentations (3)
- Computer Science Senior Theses (3)
- Computer Science Theses & Dissertations (3)
- Computer Science: Faculty Publications and Other Works (3)
- Dissertations, Theses, and Capstone Projects (3)
- Doctoral Dissertations (3)
- Engineering Management & Systems Engineering Faculty Publications (3)
- Master's Projects (3)
- SMU Data Science Review (3)
- Theses and Dissertations--Computer Science (3)
- Thesis/ Dissertation Defenses (3)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (3)
- Articles (2)
- CGU Faculty Publications and Research (2)
- Computer Science Faculty Research & Creative Works (2)
- Computer Science and Engineering Faculty Publications (2)
- Computer Science and Engineering Theses - Archive (2)
- Publication Type
Articles 1 - 30 of 293
Full-Text Articles in Computer Sciences
Large Language Models In Nlp: Evolution, Architectural Trends, And Open Challenges, Haseeb Javed, Babar Shah, Farman Ali, Daehan Kwak
Large Language Models In Nlp: Evolution, Architectural Trends, And Open Challenges, Haseeb Javed, Babar Shah, Farman Ali, Daehan Kwak
All Works
The rise of Large Language Models (LLMs) has transformed how Natural Language Processing (NLP) and its subdomains are approached. Recent technological advancements have driven this transformation. This study offers researchers a detailed overview of LLMs, comparing them with traditional rule-based systems, statistical techniques, machine learning, neural networks, and the rise of transformer-based architectures. From a wider perspective, language models such as GPT, BERT, T5, PaLM, and LLaMA have facilitated the transformation of entire sectors, including healthcare and business, due to their highly scalable nature. Despite their wide range of applications, LLMs face numerous challenges, such as output biases, limited interpretability, …
Semantic Shields: Automating Critical Infrastructure Defense Via Nlp-Driven Ransomware Profiling, Henry Trowbridge, Ian Zalcberg, Ryan Schley, Carter Yagemann, Natasha Phan, Srikar Maduposu, Vimal Buck
Semantic Shields: Automating Critical Infrastructure Defense Via Nlp-Driven Ransomware Profiling, Henry Trowbridge, Ian Zalcberg, Ryan Schley, Carter Yagemann, Natasha Phan, Srikar Maduposu, Vimal Buck
Military Cyber Affairs
Ransomware poses a growing threat to critical infrastructure, where successful attacks can disrupt operational technology (OT) and industrial control systems (ICS) with significant public safety consequences. However, attributing ransomware incidents to specific threat actors remains challenging due to ransomware-as-a-service ecosystems, actor rebranding, and the obfuscation of traditional indicators of compromise. This paper presents Semantic Shields, an NLP-driven attribution framework that leverages BERT-generated semantic embeddings and DBSCAN clustering to profile ransomware actors through the linguistic characteristics of ransom notes. Using a dataset of 295 ransom notes from 189 distinct threat groups, the framework achieved an 87.2% true positive clustering rate and …
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Master's Theses
Legislators frequently discuss the same policy issues across multiple hearings and legislative sessions, sometimes maintaining consistent positions and other times modifying or reframing their stance over time. Understanding how these positions evolve is important for analyzing political discourse and democratic accountability, yet identifying such shifts at scale remains difficult.
We introduce TRACE (Temporal Rhetorical Analysis and Consistency Evaluation), a system built on the Digital Democracy Database (DDDB) for detecting rhetorical inconsistency in California legislative hearing testimony. TRACE organizes utterances into speaker-anchored timelines indexed by bill and session, then applies hybrid semantic retrieval — combining dense BGE embeddings with BM25 lexical …
Optimizing Gated Rnns, Joshua Paul Fechete
Optimizing Gated Rnns, Joshua Paul Fechete
Honors Projects
Gated recurrent neural networks such as Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU) help fix instability present in normal recurrent neural networks. This allows them to be used for various real-world tasks, and due to their architecture, they are uniquely qualified to handle variable sized input such as text. However, even before training can begin on a machine learning model, various hyperparameters must be chosen to decide how the model will be architectured. Choosing good hyperparameters is vital for creating a model that performs well but is not larger and more computationally expensive to run than it needs …
Ai Powered Student Support And Development Platform, Wadha Alyammahi
Ai Powered Student Support And Development Platform, Wadha Alyammahi
Thesis/ Dissertation Defenses
Over the past few years, academic stress and the issue of mental well-being of students has become a pressing issue in the context of educational settings, influencing academic achievements and the general quality of life to a considerable extent. There is a growing use of digital platforms by students as a source of academic support, but most of the solutions available do not support the emotional state of the users or offer any other personalized and context-sensitive support. To address this difficulty, this paper introduces the design and development of a stress-aware conversational support system of students using advanced natural …
You Can’T Spell Audit Without Ai: The Current Uses Of Artificial Intelligence In Audit, Jena Perkins
You Can’T Spell Audit Without Ai: The Current Uses Of Artificial Intelligence In Audit, Jena Perkins
Senior Honors Theses
The accounting profession continuously adapts to the innovations provided by the broader context in which it exists. Artificial intelligence (AI) is a forerunner among tools used to enhance and optimize auditing services within the accounting profession. The realm of AI offers advancements to procedures used within an audit to detect misstatements. Based on the proprietary platforms developed by Big 4 accounting firms, AI is a key component in maintaining an advanced approach towards auditing.
The Application Of Natural Language Processing Towards Auditing Of Unstructured Data: A Design Science Approach, Dennis K. Amoatey
The Application Of Natural Language Processing Towards Auditing Of Unstructured Data: A Design Science Approach, Dennis K. Amoatey
Electronic Theses and Dissertations
Financial auditors must manually review large volumes of unstructured text that may include contracts, internal policies, footnotes, and journal entry descriptions. This time-intensive process introduces risk of human error and inconsistency. Despite advances in automation, no systematic approach exists for applying Natural Language Processing (NLP) to this problem at scale. Using a design science approach, this study develops a framework that demonstrates how NLP techniques can be incorporated across key phases in the audit process, including planning, internal controls evaluation, evidence gathering, and reporting. Initial evaluation through expert feedback had a mix of responses. While some argued difficulty with data …
Advancing Context-Aware Detection Of Socially Harmful Discourse Using Transformer-Based Models, Santosh Chapagain
Advancing Context-Aware Detection Of Socially Harmful Discourse Using Transformer-Based Models, Santosh Chapagain
All Graduate Theses and Dissertations, Fall 2023 to Present
Social media platforms are a central part of modern communication, shaping how people share ideas, build communities, and discuss social issues. While these spaces can support connection and self expression, they also enable the spread of harmful language such as hate speech. At the same time, social media is an important place where members of marginalized communities, including sexual and gender minorities, express stress, discrimination, and emotional challenges in ways that are often indirect and context dependent.
This research examines whether modern artificial intelligence systems can better identify harmful language and expressions of minority stress in online posts. The study …
Dementia Detection In Low-Resource Languages: Evaluating Translation-Assisted Transfer Learning For Multilingual Clinical Assessment, Kylar A. Deloach
Dementia Detection In Low-Resource Languages: Evaluating Translation-Assisted Transfer Learning For Multilingual Clinical Assessment, Kylar A. Deloach
Honors Theses
Alzheimer's disease (AD) is a growing global health concern, with millions of people affected worldwide and cases expected to rise significantly in the coming decades. Early detection is critical for patient treatment and care, and recent advances in natural language processing (NLP) have shown promise in identifying linguistic markers associated with AD. However, most existing work has focused on English, leaving speakers of other languages with limited access to such tools. This study investigates how effective AD detection models trained on English data are at transferring to Greek, a low-resource language with limited dementia-related speech data available. We propose a …
Analogy2kg: An Automatic Pipeline For Deriving Knowledge Graphs From Long-Text Analogies, Kara Combs, Lance E. Champagne, Bruce A. Cox, Christine M. Schubert Kabban, Trevor Bihl, Grace Lemming
Analogy2kg: An Automatic Pipeline For Deriving Knowledge Graphs From Long-Text Analogies, Kara Combs, Lance E. Champagne, Bruce A. Cox, Christine M. Schubert Kabban, Trevor Bihl, Grace Lemming
Faculty Publications
Analogical reasoning is an increasingly popular, lightweight solution to enable large language model (LLM)-level reasoning without computational complexity. Still, it has yet to be adopted due to its reliance on strictly hand-formatted data. Therefore, we propose Analogy2KG (“Analogy to Knowledge Graph”), as an automatic pipeline that transforms text into a KG format via a fine-tuned version of information extraction (IE) algorithms for long-text analogies. The need to verify that the complex underlying analogical structure of the data is maintained was done via paired samples tests in the creation and validation of this pipeline. Graph density was used to evaluate the …
Automated Software Size Measurement Using Multilingual Domain-Adapted Language Models, Samet Tenekeci̇, Hüseyi̇n Ünlü, Burak Keçeci̇, Muhammed Efe İnci̇r, Onur Demi̇rörs
Automated Software Size Measurement Using Multilingual Domain-Adapted Language Models, Samet Tenekeci̇, Hüseyi̇n Ünlü, Burak Keçeci̇, Muhammed Efe İnci̇r, Onur Demi̇rörs
Turkish Journal of Electrical Engineering and Computer Sciences
Software Size Measurement (SSM) is crucial for estimating required project effort as well as budget and schedule. However, many small and medium-sized companies struggle to apply objective SSM due to limited resources and lack of expertise. This often leads to inaccurate estimates and project overruns. There is a need for practical, low-resource solutions that support these tasks without requiring expert involvement. Motivated by this challenge, this study proposes an automated software size measurement approach that formulates the measurement task as supervised regression over natural language requirements, using domain-adapted transformer models. We construct large-scale Turkish and English software engineering corpora to …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Effective Wordle Heuristics, Ronald I. Greenberg
Effective Wordle Heuristics, Ronald I. Greenberg
Computer Science: Faculty Publications and Other Works
While previous researchers have performed an exhaustive search to determine an optimal Wordle strategy, that computation is very time consuming and produced a strategy using words that are unfamiliar to most people. With Wordle solutions being gradually eliminated (with a new puzzle each day and no reuse), an improved strategy could be generated each day, but the computation time makes a daily exhaustive search impractical. This paper shows that simple heuristics allow for fast generation of effective strategies and that little is lost by guessing only words that are possible solution words rather than more obscure words.
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
All Graduate Theses, Dissertations, and Other Capstone Projects
With healthcare systems under growing pressure from rising patient volumes and shrinking consultation windows, improving how patients communicate with physicians has become essential to delivering quality care. Yet patients routinely arrive at appointments unable to clearly describe their symptoms, recall their medical history, or articulate concerns, contributing to miscommunication, diagnostic inefficiency, and pre-visit anxiety. This study introduces PreVisit AI, a conversational system designed to address this gap through structured, knowledge-based patient preparation. The system is built on a Retrieval-Augmented Generation (RAG) architecture combining HuggingFace sentence embeddings (all-MiniLM-L6-v2), a Chroma vector store, and Google’s Gemini language model over a curated seven-document …
Mapping Multiclass-Targeted Hate Speech In Online Discourse: An Open Dataset, Sanaa Kaddoura, Sumaia Al-Kohlani
Mapping Multiclass-Targeted Hate Speech In Online Discourse: An Open Dataset, Sanaa Kaddoura, Sumaia Al-Kohlani
All Works
Online social networks have become central spaces for public discourse, where hostile and discriminatory language toward social groups can cause psychological and social consequences for marginalized communities. Although multiple public hate speech datasets are available, many rely on binary categorization practices that obscure linguistic, cultural, and contextual variation across targeted groups. As a result, minority and less visible forms of hate speech remain insufficiently documented and analyzed. This discussion paper examines methodological limitations in existing hate speech annotation schemes and presents a re-annotation framework applied to the HatEval2019 dataset. The proposed framework introduces target-specific multiclass labels that distinguish subcategories of …
Advancing Hate Speech Detection: Binary And Multiclass Approaches From Traditional Methods To Parameter-Efficient And Ontology-Guided Language Models, Mahmoud Abusaqer
Advancing Hate Speech Detection: Binary And Multiclass Approaches From Traditional Methods To Parameter-Efficient And Ontology-Guided Language Models, Mahmoud Abusaqer
Graduate Theses/Dissertations
The widespread proliferation of hate speech on social media platforms poses significant challenges for content moderation and user safety, requiring automated systems that are simultaneously accurate, efficient, and capable of fine-grained distinctions. This thesis investigates hate speech detection through five published manuscripts organized into two complementary threads: binary detection (hateful vs. non-hateful) and multiclass detection across demographic targeting categories. The binary thread progresses from a broad 38-model baseline spanning traditional machine learning, deep learning, and transformer architectures (where RoBERTa reaches 91.48% accuracy and CatBoost remains competitive at 88.60%) to parameter-efficient adaptation, in which Low-Rank Adaptation (LoRA) of large language models …
A Comprehensive Survey Of Prompt Engineering Techniques In Large Language Models, Tonmoy Debnath, Md Nurul Absar Siddiky, Muhammad Enayetur Rahman, Prosenjit Das, Antu Kumar Guha, Muhammad Rezaur Rahman, H. M. Dipu Kabir
A Comprehensive Survey Of Prompt Engineering Techniques In Large Language Models, Tonmoy Debnath, Md Nurul Absar Siddiky, Muhammad Enayetur Rahman, Prosenjit Das, Antu Kumar Guha, Muhammad Rezaur Rahman, H. M. Dipu Kabir
Electrical & Computer Engineering Faculty Publications
Prompt engineering has arisen as a pivotal discipline in optimizing the performance of Large Language Models (LLMs) by structuring inputs to enhance coherence, accuracy, and task alignment. This paper comprehensively surveys various prompting techniques, systematically categorizing them according to their application domains and methodological foundations. Fundamental approaches like zero-shot and few-shot prompting are examined along with advanced strategies, including chain-of-thought reasoning, retrieval-augmented generation, and self-consistency mechanisms. A rigorous qualitative analysis is conducted to evaluate each technique's strengths, limitations, and optimal use cases, offering a structured framework for selecting the most effective prompting strategies. Theoretical insights and empirical findings are consolidated …
Large Language Models (Llms) For Clinical Note Generation: International Classification Of Disease (Icd) Code, Knowledge Graph (Kg) And Prompt Evaluation, Ivan P. Makohon
Large Language Models (Llms) For Clinical Note Generation: International Classification Of Disease (Icd) Code, Knowledge Graph (Kg) And Prompt Evaluation, Ivan P. Makohon
Computer Science Theses & Dissertations
In the past decade, a surge in the amount of electronic health record (EHR) data in the United States occurred, driven by a favorable policy environment created by the Health Information Technology for Economic and Clinical Health (HITECH) Act of 2009 and the 21st Century Cures Act of 2016. Clinical notes for patients’ assessments, diagnoses, and treatments are captured in these EHRs in free-form text by physicians, who spend a considerable amount of time entering them. Manually writing these notes is time-consuming, increasing patient waiting times and potentially delaying diagnoses. Large language models (LLMs), such as GPT-4o, possess the ability …
Emotion Analysis And Neural Language Models For Classification, Andrew Mackey
Emotion Analysis And Neural Language Models For Classification, Andrew Mackey
Graduate Theses and Dissertations
Emotion analysis is a branch of artificial intelligence and natural language processing focused on recognizing emotions hidden throughout various forms of digital data, including text, images, and multi-modal representations. In this dissertation, we present four published and planned works that investigate different methodologies for natural language analysis tasks using deep learning techniques. The first published work we present considers the task of identifying fake news using various text and emotion representations. We demonstrate that emotion representations combined with word embedding techniques can improve the accuracy of fake news classification. Our second published work further investigates the fake news classification task …
Research On Cgf-Oriented Natural Language Interaction Framework, Xinmeng Li, Kai Xu, Yue Hu, Hesong Huang, Quanjun Yin
Research On Cgf-Oriented Natural Language Interaction Framework, Xinmeng Li, Kai Xu, Yue Hu, Hesong Huang, Quanjun Yin
Journal of System Simulation
Abstract: To address the mismatch between existing natural language interaction frameworks and training tasks in simulation-based military training, which limits smooth interaction between trainees and Computer Generated Forces (CGF), this paper proposes a Natural Language Interaction framework for Computer Generated Forces (NLI4CGF). The framework analyzes the logic and functional requirements of natural language interaction between trainees and CGF, and establishes an interaction architecture tailored for military simulation training scenarios. It supports semantic parsing and knowledge query tasks within a prototype system developed for infantry squad simulation training. Experimental results demonstrate that the proposed model performs effectively, meets the requirements of …
Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh
Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh
Doctoral
Image captioning models enable us to automatically generate natural language image descriptions for previously unseen images. It combines the two fields of computer vision and natural language generation, allowing models to interpret the con tent of an image and communicate that knowledge through natural language text.
Research into image captioning has the potential benefit of reducing the gap in digital information availability between fully sighted individuals and those who are visually impaired. However, automatically generated captions often fail to provide the required level of detail and specificity to achieve this goal. Furthermore, current standard evaluation methods are insufficient at measuring …
The Challenge Of Achieving Attributability In Multilingual Table-To-Text Generation With Question-Answer Blueprints, Aden Haussmann
The Challenge Of Achieving Attributability In Multilingual Table-To-Text Generation With Question-Answer Blueprints, Aden Haussmann
International Journal of Undergraduate Research and Creative Activities
Generating faithful text descriptions from data tables is a significant challenge in Natural Language Processing (NLP), especially for the world’s many low-resource languages. This paper investigates whether Question-Answer (QA) blueprints—an intermediate planning step where a model first asks and answers questions about the data—can improve the factual accuracy of multilingual table-to-text generation. This novel approach is tested on the TaTA dataset, which includes several African languages, by finetuning models with and without these blueprints.
The results show a key distinction: while the QA blueprint method improves performance for English-only models, these gains disappear in the multilingual setting. This paper’s analysis …
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Research Collection School Of Computing and Information Systems
Automated hate speech detection is an important tool in combating the spread of hate speech, particularly in social media. Numerous methods have been developed for the task, including a recent proliferation of deep-learning based approaches. A variety of datasets have also been developed, exemplifying various manifestations of the hate-speech detection problem. We present here a largescale empirical comparison of deep and shallow hate-speech detection methods, mediated through the three most commonly used datasets. Our goal is to illuminate progress in the area, and identify strengths and weaknesses in the current state-of-the-art. We particularly focus our analysis on measures of practical …
Computational Analogies In The Era Of Large Language Models, Amarakoon Mudiyanselage Thilini Wijesiriwardene
Computational Analogies In The Era Of Large Language Models, Amarakoon Mudiyanselage Thilini Wijesiriwardene
Theses and Dissertations
Analogical reasoning is an important part of human cognition requiring the integration of abstract reasoning, pattern recognition, and background knowledge. Despite significant advances in language modeling, the capacity of current methods to accurately identify, model, and evaluate analogies remains fundamentally underexplored.
Analogies enable individuals to perceive deep similarities between superficially different situations. Effective analogy-making requires integrating knowledge about the external world with abstract reasoning and pattern recognition capabilities. While current language models (LMs), trained on massive textual corpora using autoregressive or masked objectives, achieve impressive performance across Natural Language Processing (NLP) tasks such as text generation, summarization, and classification, their …
Drawing On Uncertainty Methodologies Of Neutrosophic Hypersoft Sets In Cognitive Computing-Driven Healthcare Systems, Mona Mohamed, Nurhan Alaa
Drawing On Uncertainty Methodologies Of Neutrosophic Hypersoft Sets In Cognitive Computing-Driven Healthcare Systems, Mona Mohamed, Nurhan Alaa
Neutrosophic Systems with Applications
A new paradigm called cognitive computing simulates human reasoning and decision-making through integrating advanced techniques such as artificial intelligence (AI) and natural language processing (NLP). Cognitive computing systems, in contrast to traditional systems, can handle both structured and unstructured data, adjust to new information, and offer context-sensitive insights. This study examines how cognitive computing improves decision-making, personalization, and human-machine collaboration in various fields. Cognitive computing in the healthcare sector processes clinical notes, imaging data, and electronic health records to help physicians with diagnosis, treatment planning, and patient engagement. This study examines key applications, including their role in diagnostic support, where …
Sugar: A Sequence Unfolding Based Transformer Model For Group Activity Recognition, Yash U. Gondkar
Sugar: A Sequence Unfolding Based Transformer Model For Group Activity Recognition, Yash U. Gondkar
Graduate Masters Theses
Large Language Models have improved significantly in the past couple of years due to the adoption of transformers. However, transformers still find it challenging to process videos due to limited context size caused by their quadratic computing cost. Therefore, we studied a booming field in machine learning which powers applications like social scene analysis and video surveillance systems called Group Activity Recognition (GAR). We found that recent models were able to achieve more than 90% accuracy on popular datasets like the Volleyball dataset, however, it turned out that even they relied on transformers.
Therefore, in this work, we developed a …
Domain Obedient Deep Learning, Soumadeep Saha
Domain Obedient Deep Learning, Soumadeep Saha
Doctoral Theses
Deep learning, a family of data-driven artificial intelligence techniques, has shown immense promise in a plethora of applications, and it has even outpaced experts in several domains. However, unlike symbolic approaches to learning, these methods fall short when it comes to abiding by and learning from pre-existing established principles. This is a significant deficit for deployment in critical applications such as robotics, medicine, industrial automation, etc. For a decision system to be considered for adoption in such fields, it must demonstrate the ability to adhere to specified constraints, an ability missing in deep learning-based approaches. Exploring this problem serves as …
Development And Validation Of Venous Thromboembolism-Bidirectional Encoder Representations From Transformers (Vte-Bert) Natural Language Processing Model, Omid Jafari, Shengling Ma, Barbara D Lam, Jun Y Jiang, Emily Zhou, Mrinal Ranjan, Justine Ryu, Raka Bandyo, Arash Maghsoudi, Bo Peng, Christopher I Amos, Abiodun Oluyomi, Nathanael R Fillmore, Jennifer La, Ang Li
Development And Validation Of Venous Thromboembolism-Bidirectional Encoder Representations From Transformers (Vte-Bert) Natural Language Processing Model, Omid Jafari, Shengling Ma, Barbara D Lam, Jun Y Jiang, Emily Zhou, Mrinal Ranjan, Justine Ryu, Raka Bandyo, Arash Maghsoudi, Bo Peng, Christopher I Amos, Abiodun Oluyomi, Nathanael R Fillmore, Jennifer La, Ang Li
Faculty, Staff and Students Publications
Background: Accurate and rapid phenotyping of venous thromboembolism (VTE) in longitudinal studies is important. A natural language processing (NLP) tool externally validated in representative patients is lacking.
Objectives: To train and validate an efficient NLP model to detect incident VTE event.
Methods: We designed a novel NLP platform, NLPMed, to assist thrombosis researchers with data preprocessing, phenotype annotation, language model finetuning, and NLP application. Using clinical notes, discharge summaries, and radiology reports from patients with cancer at 2 healthcare institutions, we finetuned Bio_Clinical Bidirectional Encoder Representations from Transformers (BERT) to develop VTE-BERT. The new model was trained to detect acute …
Enhancing Multi-Step Stock Price Forecasting With Social Media Sentiment And Engagement Metrics, Damilare Olaniyan
Enhancing Multi-Step Stock Price Forecasting With Social Media Sentiment And Engagement Metrics, Damilare Olaniyan
Electronic Theses and Dissertations
This thesis investigates whether social media sentiment can improve the accuracy of stock price prediction beyond traditional historical data. While financial markets have long relied on structured numerical indicators, the growing influence of public discourse on platforms like Twitter has introduced new opportunities for extracting market-relevant signals from unstructured text. The study focuses on four major technology firms and combines sentiment features derived from Twitter with historical stock prices in a hybrid machine learning framework. Engagement-weighted sentiment, linguistic complexity, and polarity intensity were extracted using natural language processing techniques and incorporated into classification and regression models. Results show that including …
Advancing Fake News Detection With Graph Neural Network And Deep Learning, Haji Gul, Feras Al-Obeidat, Muhammad Wasim, Adnan Amin, Fernando Moreira
Advancing Fake News Detection With Graph Neural Network And Deep Learning, Haji Gul, Feras Al-Obeidat, Muhammad Wasim, Adnan Amin, Fernando Moreira
All Works
In the modern era of digital technology, the rapid distribution of news via social media platforms substantially contributes to the propagation of false information, presenting challenges in upholding the accuracy and reliability of information. This study presents an updated approach that utilizes graph neural networks (GNNs) alongside with advanced deep learning techniques to improve the identification of false information. In contrast to traditional approaches that primarily rely on analyzing text and assessing the credibility of sources, our methodology utilizes the structural information of news propagation networks. This allows for a detailed comprehension of the interconnections and patterns that are indicative …