Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Engineering

Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 23 of 23

Full-Text Articles in Computational Linguistics

Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming Yu, Bin Deng, Zhengang Zhang Jul 2025

Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming Yu, Bin Deng, Zhengang Zhang

Journal of Scientific Information Research

[Purpose/significance] This study addresses the issue of inadequate perception of entity boundaries in traditional character-level modeling-based named entity recognition models by integrating syntax information containing entity boundary features into the task using a multi-head graph attention network with dense connections. This integration enhances the effectiveness of named entity recognition.

[Method/process] This study proposes a Syntax-enhanced Boundary-aware Named Entity Recognition Model (SynBNER), which utilizes BERT for text semantic representation and integrates syntax information using a dense-connected graph attention network. This integration incorporates implicit entity boundary information from syntax information into word representations, thereby enhancing the model's entity boundary perception capability.

[Result/conclusion] …


From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie Apr 2025

From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie

Undergraduate Theses

Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …


On The Provenance Of Software Systems: Automating Software Traceability With Knowledge Graph And Large Language Model Synergy, Tyler Procko Apr 2025

On The Provenance Of Software Systems: Automating Software Traceability With Knowledge Graph And Large Language Model Synergy, Tyler Procko

Doctoral Dissertations and Master's Theses

The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the study of being, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life …


Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri Jan 2024

Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri

Civil Engineering Dissertations - Archive

Infrastructure projects impact a broad range of stakeholders, particularly local communities, whose engagement is critical for successful outcomes. Despite the importance of public engagement in these projects, traditional methods of capturing and analyzing public opinion often fail to fully represent the diverse, genuine perspectives involved. This has led to conflicts between community members and project sponsors. On the other hand, despite advancements in text analytics, including natural language processing (NLP) and its subfields such as topic modeling, sentiment analysis, and neural networks, their functionalities and effectiveness in analyzing public opinion in the domain of infrastructure projects have not been fully …


Executive Order On The Safe, Secure, And Trustworthy Development And Use Of Artificial Intelligence, Joseph R. Biden Oct 2023

Executive Order On The Safe, Secure, And Trustworthy Development And Use Of Artificial Intelligence, Joseph R. Biden

Copyright, Fair Use, Scholarly Communication, etc.

Section 1. Purpose. Artificial intelligence (AI) holds extraordinary potential for both promise and peril. Responsible AI use has the potential to help solve urgent challenges while making our world more prosperous, productive, innovative, and secure. At the same time, irresponsible use could exacerbate societal harms such as fraud, discrimination, bias, and disinformation; displace and disempower workers; stifle competition; and pose risks to national security. Harnessing AI for good and realizing its myriad benefits requires mitigating its substantial risks. This endeavor demands a society-wide effort that includes government, the private sector, academia, and civil society.

My Administration places the highest urgency …


Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian) Mar 2023

Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)

Library Philosophy and Practice (e-journal)

Abstract

Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …


Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat Jan 2023

Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat

Graduate Theses/Dissertations

Social media has become a domain that involves a lot of hate speech. Some users feel entitled to engage in abusive conversations by sending abusive messages, tweets, or photos to other users. It is critical to detect hate speech and prevent innocent users from becoming victims. In this study, I explore the effectiveness and performance of various machine learning methods employing text processing techniques to create a robust system for hate speech identification. I assess the performance of Naïve Bayes, Support Vector Machines, Decision Trees, Random Forests, Logistic Regression, and K Nearest Neighbors using three distinct datasets sourced from social …


Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin Jun 2022

Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin

Data and Test Instruments

How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.


Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian Apr 2022

Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian

Northeast Journal of Complex Systems (NEJCS)

In this study, we introduce a new network feature for detecting suicidal ideation from clinical texts and conduct various additional experiments to enrich the state of knowledge. We evaluate statistical features with and without stopwords, use lexical networks for feature extraction and classification, and compare the results with standard machine learning methods using a logistic classifier, a neural network, and a deep learning method. We utilize three text collections. The first two contain transcriptions of interviews conducted by experts with suicidal (n=161 patients that experienced severe ideation) and control subjects (n=153). The third collection consists of interviews conducted by experts …


Chaprates, Brinly Xavier, Micole Amanda Marietta, Nidhi Vedantam May 2020

Chaprates, Brinly Xavier, Micole Amanda Marietta, Nidhi Vedantam

Student Scholar Symposium Abstracts and Posters

On the Chapman campus, through taking and choosing various classes, there is a significant need for communication and feedback between students and peers, professors, tutors, and study groups. With this, we wanted to create an application that enables users from various majors to not only easily and effectively communicate with various people in their field, but one that also enables them to give and receive feedback on various classes through a rating system. We believe that the application will aid students in a myriad of specific ways, including being involved in study groups and getting tutoring help, determining which classes …


Size Matters: The Impact Of Training Size In Taxonomically-Enriched Word Embeddings, Alfredo Maldonado, Filip Klubicka, John D. Kelleher Oct 2019

Size Matters: The Impact Of Training Size In Taxonomically-Enriched Word Embeddings, Alfredo Maldonado, Filip Klubicka, John D. Kelleher

Articles

Word embeddings trained on natural corpora (e.g., newspaper collections, Wikipedia or the Web) excel in capturing thematic similarity (“topical relatedness”) on word pairs such as ‘coffee’ and ‘cup’ or ’bus’ and ‘road’. However, they are less successful on pairs showing taxonomic similarity, like ‘cup’ and ‘mug’ (near synonyms) or ‘bus’ and ‘train’ (types of public transport). Moreover, purely taxonomy-based embeddings (e.g. those trained on a random-walk of WordNet’s structure) outperform natural-corpus embeddings in taxonomic similarity but underperform them in thematic similarity. Previous work suggests that performance gains in both types of similarity can be achieved by enriching natural-corpus embeddings with …


The Design And Implementation Of Aida: Ancient Inscription Database And Analytics System, M. Parvez Rashid Jul 2019

The Design And Implementation Of Aida: Ancient Inscription Database And Analytics System, M. Parvez Rashid

School of Computing: Dissertations, Theses, and Student Research

AIDA, the Ancient Inscription Database and Analytic system can be used to translate and analyze ancient Minoan language. The AIDA system currently stores three types of ancient Minoan inscriptions: Linear A, Cretan Hieroglyph and Phaistos Disk inscriptions. In addition, AIDA provides candidate syllabic values and translations of Minoan words and inscriptions into English. The AIDA system allows the users to change these candidate phonetic assignments to the Linear A, Cretan Hieroglyph and Phaistos symbols. Hence the AIDA system provides for various scholars not only a convenient online resource to browse Minoan inscriptions but also provides an analysis tool to explore …


Advanced Recurrent Network-Based Hybrid Acoustic Models For Low Resource Speech Recognition, Jian Kang, Wei-Qiang Zhang, Wei-Wei Liu, Jia Liu, Michael T. Johnson Jul 2018

Advanced Recurrent Network-Based Hybrid Acoustic Models For Low Resource Speech Recognition, Jian Kang, Wei-Qiang Zhang, Wei-Wei Liu, Jia Liu, Michael T. Johnson

Electrical and Computer Engineering Faculty Publications

Recurrent neural networks (RNNs) have shown an ability to model temporal dependencies. However, the problem of exploding or vanishing gradients has limited their application. In recent years, long short-term memory RNNs (LSTM RNNs) have been proposed to solve this problem and have achieved excellent results. Bidirectional LSTM (BLSTM), which uses both preceding and following context, has shown particularly good performance. However, the computational requirements of BLSTM approaches are quite heavy, even when implemented efficiently with GPU-based high performance computers. In addition, because the output of LSTM units is bounded, there is often still a vanishing gradient issue over multiple layers. …


Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin Jan 2018

Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin

Data and Test Instruments

Automatic clickbait detection is a relatively novel task in natural language processing (NLP) and machine learning (ML). “Clickbait” is a hyperlink created primarily to attract attention to its target content. This article introduces a binary classifier, the Language and Information Technology Research Lab (LiT.RL, pronounced “literal”) Clickbait Detector, which automatically distinguishes clickbait from nonclickbait. We used NLP and ML for 38 textual features, contrasting clickbait with “headlinese.” When tested on 11,000 hyperlinks, it achieves 94 per cent accuracy using a support vector machine. Integrated with the LiT.RL News Verification Browser, a downloadable stand-alone research tool, the Clickbait Detector user interface …


A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim Jan 2016

A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim

All Graduate Theses, Dissertations, and Other Capstone Projects

This work explores the trade-offs between time and frequency information during the feature extraction process of an automatic speech recognition (ASR) system using wavelet transform (WT) features instead of Mel-frequency cepstral coefficients (MFCCs) and the benefits of combining the WTs and the MFCCs as inputs to an ASR system. A virtual machine from the Speech Recognition Virtual Kitchen resource (www.speechkitchen.org) is used as the context for implementing a wavelet signal processing module in a speech recognition system. Contributions include a comparison of MFCCs and WT features on small and large vocabulary tasks, application of combined MFCC and WT features on …


A Computational Translation Of The Phaistos Disk, Peter Revesz Oct 2015

A Computational Translation Of The Phaistos Disk, Peter Revesz

School of Computing: Conference and Workshop Papers

For over a century the text of the Phaistos Disk remained an enigma without a convincing translation. This paper presents a novel semi-automatic translation method that uses for the first time a recently discovered connection between the Phaistos Disk symbols and other ancient scripts, including the Old Hungarian alphabet. The connection between the Phaistos Disk script and the Old Hungarian alphabet suggested the possibility that the Phaistos Disk language may be related to Proto-Finno-Ugric, Proto-Ugric, or Proto-Hungarian. Using words and suffixes from those languages, it is possible to translate the Phaistos Disk text as an ancient sun hymn, possibly connected …


A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz Oct 2015

A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz

School of Computing: Conference and Workshop Papers

Crete was the birthplace of several ancient writings, including the Cretan Hieroglyphs, the Linear A and the Linear B scripts. Out of these three only Linear B is deciphered. The sound values of the Cretan Hieroglyph and the Linear A symbols are unknown and attempts to reconstruct them based on Linear B have not been fruitful. In this paper, we compare the ancient Cretan scripts with four other Mediterranean and Black Sea scripts, namely Phoenician, South Arabic, Greek and Old Hungarian. We provide a computational study of the evolution of the three Cretan and four other scripts. This study encompasses …


An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler Dec 2014

An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler

LSU New Orleans Theses and Dissertations

This thesis performs an empirical analysis of Word2Vec by comparing its output to WordNet, a well-known, human-curated lexical database. It finds that Word2Vec tends to uncover more of certain types of semantic relations than others -- with Word2Vec returning more hypernyms, synonomyns and hyponyms than hyponyms or holonyms. It also shows the probability that neighbors separated by a given cosine distance in Word2Vec are semantically related in WordNet. This result both adds to our understanding of the still-unknown Word2Vec and helps to benchmark new semantic tools built from word vectors.


Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee Jan 2014

Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee

Articles

In a situated dialogue, misunderstandings may arise if the participants perceive or interpret the environment in different ways. In human-computer dialogue this may be due the sensor errors. We present an experiment system and a series of experiments in which we investigate this problem.


Using Textual Features To Predict Popular Content On Digg, Paul H. Miller Apr 2011

Using Textual Features To Predict Popular Content On Digg, Paul H. Miller

Department of English: Dissertations, Theses, and Student Research

Over the past few years, collaborative rating sites, such as Netflix, Digg and Stumble, have become increasingly prevalent sites for users to find trending content.  I used various data mining techniques to study Digg, a social news site, to examine the influence of content on popularity.  What influence does content have on popularity, and what influence does content have on users’ decisions?  Overwhelmingly, prior studies have consistently shown that predicting popularity based on content is difficult and maybe even inherently impossible.  The same submission can have multiple outcomes and content neither determines popularity, nor individual user decisions.  My results show …


Study Of Stemming Algorithms, Savitha Kodimala Dec 2010

Study Of Stemming Algorithms, Savitha Kodimala

UNLV Theses, Dissertations, Professional Papers, and Capstones

Automated stemming is the process of reducing words to their roots. The stemmed words are typically used to overcome the mismatch problems associated with text searching.


In this thesis, we report on the various methods developed for stemming. In particular, we show the effectiveness of n-gram stemming methods on a collection of documents.


Minimum Mean Square Error Spectral Peak Envelope Estimation For Automatic Vowel Classification, Jaishree Venugopal Jul 2001

Minimum Mean Square Error Spectral Peak Envelope Estimation For Automatic Vowel Classification, Jaishree Venugopal

Electrical & Computer Engineering Theses & Dissertations

Spectral feature computations continue to be a very difficult problem for accurate machine recognition of speech. In this work, which focuses on vowels, a new spectral peak envelope method for vowel classification is developed, based on a missing frequency components model of speech recognition. According to the missing frequency components model, vowel recognition depends only on the spectral (harmonic) peaks. Smoothing and interpolation of the spectra, performed in the standard cepstral analysis method commonly used in automatic speech recognition, actually loses valuable information and results in reduced recognition accuracy. The new method for feature extraction presented in this thesis is …


Visual Speech Training Aid For The Deaf, Subhashri Venkat Jul 1990

Visual Speech Training Aid For The Deaf, Subhashri Venkat

Electrical & Computer Engineering Theses & Dissertations

A computer-based vowel articulation training aid has been developed. A "continuous" acoustic-phonetic transformation is performed to map speech parameters to a lower dimensionality display space. There are two possible approaches to this transformation problem. The transformation could be either linear or a combination nonlinear/linear. The nonlinear transformation is performed using a multi-layered feedforward neural network with linear output layers. Speech parameters are extracted either from an analog filter bank arrangement (band energies) or by a digital signal processing procedure (Discrete Cosine Transform Coefficients). The speech parameters obtained from both methods correspond to the spectral envelope of the speech signals. The …