Using Ai For Qualitative Labeling: Consistency And Comparisons,
2024
Rollins College
Using Ai For Qualitative Labeling: Consistency And Comparisons, James Temple
Honors Program Theses
This paper continues research that evaluates the capacity of artificial intelligence (AI) to perform qualitative coding tasks. The previous study found that AI models lacked consistency with themselves and did not agree with human coded data. Since that study, AI’s general level of intelligence has increased. Hence, this study re-evaluates how well the newest set of AI models (Claude 3 and Gemini) can perform qualitative coding tasks. When tested, the new AI models perform about the same or better than previous models depending on the metric tested. While Gemini and Claude 3 do not agree with human output any more …
Linguistic Inquiry And Word Count (Liwc) For Text Analysis,
2024
Virginia Commonwealth University
Linguistic Inquiry And Word Count (Liwc) For Text Analysis, Isabella Lenzini, Destiny Fore Msw, Anna W. Wright Phd
IRBEH/Spit for Science Publications and Presentations
No abstract provided.
Streamlining Public Engagement In Transportation Projects Using Text Analytics,
2024
University of Texas at Arlington
Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri
Civil Engineering Dissertations - Archive
Infrastructure projects impact a broad range of stakeholders, particularly local communities, whose engagement is critical for successful outcomes. Despite the importance of public engagement in these projects, traditional methods of capturing and analyzing public opinion often fail to fully represent the diverse, genuine perspectives involved. This has led to conflicts between community members and project sponsors. On the other hand, despite advancements in text analytics, including natural language processing (NLP) and its subfields such as topic modeling, sentiment analysis, and neural networks, their functionalities and effectiveness in analyzing public opinion in the domain of infrastructure projects have not been fully …
A Computational Investigation Of English Spelling,
2024
University of Kentucky
A Computational Investigation Of English Spelling, John Winstead
Theses and Dissertations--Linguistics
This thesis examines the predictability and regularity of English orthography through computational methods. The primary objective is to use n-gram models to predict missing letters in English words by exploiting contextual information from adjacent letters. The study evaluates the impact of dataset size, word length, letter position, and vowel presence on the predictive accuracy of these models, uncovering patterns and structures inherent to English spelling.
The research utilizes a range of datasets, including the Carnegie Mellon University Pronouncing Dictionary, the Brown Corpus, the Corpus of Late Modern English Texts, the Lampeter Corpus of Early Modern English Tracts, and the Open …
A Computer-Assisted Approach To Lexical Borrowing In Northeast Caucasian Languages,
2024
University of Kentucky
A Computer-Assisted Approach To Lexical Borrowing In Northeast Caucasian Languages, Bonnie Eleanor Wren-Hardin
Theses and Dissertations--Linguistics
The disambiguation of loanwords and cognates can be a challenge, especially in areas where there has been intense language contact over an extended period of time, when the contact is between genetically related languages, and when the number of languages involved is large Over the past several decades, more and more computational approaches to automatic cognate and borrowing detection have been created in an attempt to ease the load of examining hundreds to thousands of individual lexemes, as well as determine language family relationships with allegedly greater accuracy. While these methods are not perfect and cannot replace the knowledge or …
Guilty Machines: On Ab-Sens In The Age Of Ai,
2023
Virginia Commonwealth University
Guilty Machines: On Ab-Sens In The Age Of Ai, Dylan Lackey, Katherine Weinschenk
Critical Humanities
For Lacan, guilt arises in the sublimation of ab-sens (non-sense) into the symbolic comprehension of sen-absexe (sense without sex, sense in the deficiency of sexual relation), or in the maturation of language to sensibility through the effacement of sex. Though, as Slavoj Žižek himself points out in a recent article regarding ChatGPT, the split subject always misapprehends the true reason for guilt’s manifestation, such guilt at best provides a sort of evidence for the inclusion of the subject in the order of language, acting as a necessary, even enjoyable mark of the subject’s coherence (or, more importantly, the subject’s separation …
Executive Order On The Safe, Secure, And Trustworthy Development And Use Of Artificial Intelligence,
2023
United States Office of the President
Executive Order On The Safe, Secure, And Trustworthy Development And Use Of Artificial Intelligence, Joseph R. Biden
Copyright, Fair Use, Scholarly Communication, etc.
Section 1. Purpose. Artificial intelligence (AI) holds extraordinary potential for both promise and peril. Responsible AI use has the potential to help solve urgent challenges while making our world more prosperous, productive, innovative, and secure. At the same time, irresponsible use could exacerbate societal harms such as fraud, discrimination, bias, and disinformation; displace and disempower workers; stifle competition; and pose risks to national security. Harnessing AI for good and realizing its myriad benefits requires mitigating its substantial risks. This endeavor demands a society-wide effort that includes government, the private sector, academia, and civil society.
My Administration places the highest urgency …
Towards Interpretable Machine Reading Comprehension With Mixed Effects Regression And Exploratory Prompt Analysis,
2023
CUNY Graduate Center
Towards Interpretable Machine Reading Comprehension With Mixed Effects Regression And Exploratory Prompt Analysis, Luca Del Signore
Dissertations, Theses, and Capstone Projects
We investigate the properties of natural language prompts that determine their difficulty in machine reading comprehension tasks. While much work has been done benchmarking language model performance at the task level, there is considerably less literature focused on how individual task items can contribute to interpretable evaluations of natural language understanding. Such work is essential to deepening our understanding of language models and ensuring their responsible use as a key tool in human machine communication. We perform an in depth mixed effects analysis on the behavior of three major generative language models, comparing their performance on a large reading comprehension …
A Computational Analysis Of Volodymyr Zelenskyy's Public Diplomacy Discourse In Times Of Crisis,
2023
Pepperdine University
A Computational Analysis Of Volodymyr Zelenskyy's Public Diplomacy Discourse In Times Of Crisis, Amber Brittain-Hale, Amber Brittain-Hale
Education Division Scholarship
In this study, we delve into the public diplomacy discourse of Ukrainian President Volodymyr Zelenskyy during the ongoing crisis of the Russo-Ukrainian War. We aim to conduct a computational analysis of Zelenskyy's English, Russian, and Ukrainian speeches, exploring the linguistic patterns and code-switching employed in his discourse. The study period encompasses Russia’s build-up to and full-scale invasion of Ukraine from May 2019 to May 30, 2023. This time frame is crucial as it captures the dynamic development of the crisis and the expansion of Zelenskyy's presidency, providing a unique context for analyzing his public diplomacy efforts. By utilizing Linguistic Inquiry …
Ideology Prediction From Scarce And Biased Supervision: Learn To Disregard The “What” And Focus On The “How”!,
2023
CUHK Shenzhen
Ideology Prediction From Scarce And Biased Supervision: Learn To Disregard The “What” And Focus On The “How”!, Chen Chen, Dylan Walker, Venkatesh Saligrama
Business Faculty Articles and Research
We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selection bias. We propose a novel statistical model that decomposes the document embeddings into a linear superposition of two vectors; a latent neutral context vector independent of ideology, and a latent position vector aligned with ideology. We train an end-to-end model that has intermediate contextual and positional vectors as outputs. At deployment time, our model predicts labels for input documents …
Destined Failure,
2023
Rhode Island School of Design
Destined Failure, Chengjun Pan
Masters Theses
I attempt to examine the complex structure of human communication, explaining why it is bound to fail. By reproducing experienceable phenomena, I demonstrate how they can expose communication structure and reveal the limitations of our perception and symbolization.I divide the process of communication into six stages: input, detection, symbolization, dictionary, interpretation, and output. In this thesis, I examine the flaws and challenges that arise in the first five stages. I argue that reception acts as a filter and that understanding relies on a symbolic system that is full of redundancies. Therefore, every interpretation is destined to be a deviation.
The Sociolinguistics Of Code-Switching In Hong Kong’S Digital Landscape: A Mixed-Methods Exploration Of Cantonese-English Alternation Patterns On Whatsapp,
2023
The Chinese University of Hong Kong
The Sociolinguistics Of Code-Switching In Hong Kong’S Digital Landscape: A Mixed-Methods Exploration Of Cantonese-English Alternation Patterns On Whatsapp, Wilkinson Daniel Wong Gonzales, Yuen Man Tsang
Journal of English and Applied Linguistics
This paper examines the prevalence of Cantonese-English code mixing in Hong Kong through an under-researched digital medium. Prior research on this code-alternation practice has often been limited to exploring either the social or linguistic constraints of code-switching in spoken or written communication. Our study takes a holistic approach to analyzing code-switching in a hybrid medium that exhibits features of both spoken and written discourse. We specifically analyze the code-switching patterns of 24 undergraduates from a Hong Kong university on WhatsApp and examine how both social and linguistic factors potentially constrain these patterns. Utilizing a self-compiled sociolinguistic corpus as well as …
Neural Network Vs. Rule-Based G2p: A Hybrid Approach To Stress Prediction And Related Vowel Reduction In Bulgarian,
2023
CUNY Graduate Center
Neural Network Vs. Rule-Based G2p: A Hybrid Approach To Stress Prediction And Related Vowel Reduction In Bulgarian, Maria Karamihaylova
Dissertations, Theses, and Capstone Projects
An effective grapheme-to-phoneme (G2P) conversion system is a critical element of speech synthesis. Rule-based systems were an early method for G2P conversion. In recent years, machine learning tools have been shown to outperform rule-based approaches in G2P tasks. We investigate neural network sequence-to-sequence modeling for the prediction of syllable stress and resulting vowel reductions in the Bulgarian language. We then develop a hybrid G2P approach which combines manually written grapheme-to-phoneme mapping rules with neural network-enabled syllable stress predictions by inserting stress markers in the predicted stress position of the transcription produced by the rule-based finite-state transducer. Finally, we apply vowel …
Topics For He But Not For She: Quantifying And Classifying Gender Bias In The Media,
2023
CUNY Graduate Center
Topics For He But Not For She: Quantifying And Classifying Gender Bias In The Media, Tyler J. Lanni
Dissertations, Theses, and Capstone Projects
In this study, we used computational techniques to analyze the language used in news articles to describe female and male politicians. Our corpus included 370 subtexts for male candidates and 374 subtexts for female candidates, gathered through the New York Times API. We conducted two experiments: an LDA topic analysis to explore the data, and a logistic regression to classify the subtexts as either male or female. Our analysis revealed some noteworthy findings that suggest the possibility of developing a gender bias classifier in the future. However, to create a more robust understanding of bias, additional research and data are …
Evaluating Neural Networks As Cognitive Models For Learning Quasi-Regularities In Language,
2023
CUNY Graduate Center
Evaluating Neural Networks As Cognitive Models For Learning Quasi-Regularities In Language, Xiaomeng Ma
Dissertations, Theses, and Capstone Projects
Many aspects of language can be categorized as quasi-regular: the relationship between the inputs and outputs is systematic but allows many exceptions. Common domains that contain quasi-regularity include morphological inflection and grapheme-phoneme mapping. How humans process quasi-regularity has been debated for decades. This thesis implemented modern neural network models, transformer models, on two tasks: English past tense inflection and Chinese character naming, to investigate how transformer models perform quasi-regularity tasks. This thesis focuses on investigating to what extent the models' performances can represent human behavior. The results show that the transformers' performance is very similar to human behavior in many …
Ai Approaches To Understand Human Deceptions, Perceptions, And Perspectives In Social Media,
2023
New Jersey Institute of Technology
Ai Approaches To Understand Human Deceptions, Perceptions, And Perspectives In Social Media, Chih-Yuan Li
Dissertations
Social media platforms have created virtual space for sharing user generated information, connecting, and interacting among users. However, there are research and societal challenges: 1) The users are generating and sharing the disinformation 2) It is difficult to understand citizens' perceptions or opinions expressed on wide variety of topics; and 3) There are overloaded information and echo chamber problems without overall understanding of the different perspectives taken by different people or groups.
This dissertation addresses these three research challenges with advanced AI and Machine Learning approaches. To address the fake news, as deceptions on the facts, this dissertation presents Machine …
Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers,
2023
East Tennessee State University
Predicting High-Cap Tech Stock Polarity: A Combined Approach Using Support Vector Machines And Bidirectional Encoders From Transformers, Ian L. Grisham
Electronic Theses and Dissertations
The abundance, accessibility, and scale of data have engendered an era where machine learning can quickly and accurately solve complex problems, identify complicated patterns, and uncover intricate trends. One research area where many have applied these techniques is the stock market. Yet, financial domains are influenced by many factors and are notoriously difficult to predict due to their volatile and multivariate behavior. However, the literature indicates that public sentiment data may exhibit significant predictive qualities and improve a model’s ability to predict intricate trends. In this study, momentum SVM classification accuracy was compared between datasets that did and did not …
Improving Sign Recognition With Phonology,
2023
University of Southern California
Improving Sign Recognition With Phonology, Lee Kezar, Jesse Thomason, Zed Sevcikova Sehyr
Communication Sciences and Disorders Faculty Articles and Research
We use insights from research on American Sign Language (ASL) phonology to train models for isolated sign language recognition (ISLR), a step towards automatic sign language understanding. Our key insight is to explicitly recognize the role of phonology in sign production to achieve more accurate ISLR than existing work which does not consider sign language phonology. We train ISLR models that take in pose estimations of a signer producing a single sign to predict not only the sign but additionally its phonological characteristics, such as the handshape. These auxiliary predictions lead to a nearly 9% absolute gain in sign recognition …
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War,
2023
Southern Methodist University
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War, Yucheol Shin, Yvan Sojdehei, Limin Zheng, Brad Blanchard
SMU Data Science Review
The Ukrainian-Russian war has garnered significant attention worldwide, with fake news obstructing the formation of public opinion and disseminating false information. This scholarly paper explores the use of unsupervised learning methods and the Bidirectional Encoder Representations from Transformers (BERT) to detect fake news in news articles from various sources. BERT topic modeling is applied to cluster news articles by their respective topics, followed by summarization to measure the similarity scores. The hypothesis posits that topics with larger variances are more likely to contain fake news. The proposed method was evaluated using a dataset of approximately 1000 labeled news articles related …
Research On Topic Discovery And Evolution Trend Based On Temporal Keyword Characteristics Analysis,
2023
College of Information Engineering, Nanjing University of Finance and Economics, Nanjing 210023
Research On Topic Discovery And Evolution Trend Based On Temporal Keyword Characteristics Analysis, Shuqing Li, Juntao Zhu, Wan Wang
Journal of Scientific Information Research
[Purpose/significance]Excavating the research topics in a large number of articles, sorting out the evolution context and correlation of the research topics, predicting the frontier hot spots of the topics can be helpful to enhance the scientificity and vividness of the evolution results.[Method/precess]This paper puts forward the concept of time series influence factor as an important feature in keyword extraction, uses the method of time window to mine and identify topics by using topic model, and makes visual analysis. By applying time series model in the field of deep learning, the purpose of predicting topic popularity is achieved.[Result/concluson]It is verified that …
