Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (98)
- Computer Sciences (92)
- Arts and Humanities (47)
- Artificial Intelligence and Robotics (43)
- Discourse and Text Linguistics (26)
-
- Communication (25)
- Engineering (23)
- Semantics and Pragmatics (23)
- Psychology (22)
- Library and Information Science (20)
- Computer Engineering (19)
- Applied Linguistics (18)
- Psycholinguistics and Neurolinguistics (17)
- Phonetics and Phonology (15)
- First and Second Language Acquisition (13)
- Language Description and Documentation (13)
- Communication Technology and New Media (12)
- Anthropological Linguistics and Sociolinguistics (11)
- Cognition and Perception (11)
- Data Science (11)
- Medicine and Health Sciences (11)
- Databases and Information Systems (10)
- Other Computer Sciences (10)
- Syntax (10)
- Digital Humanities (9)
- Electrical and Computer Engineering (9)
- Social Media (9)
- Institution
-
- City University of New York (CUNY) (67)
- Technological University Dublin (30)
- Montclair State University (18)
- University of Kentucky (16)
- Brigham Young University (10)
-
- Chapman University (7)
- Dartmouth College (5)
- University of Nebraska - Lincoln (5)
- Western University (5)
- East Tennessee State University (4)
- University of Malaya (4)
- Air Force Institute of Technology (3)
- California Polytechnic State University, San Luis Obispo (3)
- Hunan Provincial Institute of Scientific and Technology Information (3)
- Minnesota State University, Mankato (3)
- Old Dominion University (3)
- Portland State University (3)
- San Jose State University (3)
- University of Central Florida (3)
- Binghamton University (2)
- Boise State University (2)
- Claremont Colleges (2)
- New Jersey Institute of Technology (2)
- Purdue University (2)
- The University of Akron (2)
- University of Louisville (2)
- Ursinus College (2)
- Valparaiso University (2)
- Bellarmine University (1)
- Bowling Green State University (1)
- Keyword
-
- Natural Language Processing (19)
- Computational linguistics (18)
- Natural language processing (16)
- Machine learning (12)
- NLP (10)
-
- Computational Linguistics (9)
- Machine Learning (7)
- Corpus linguistics (6)
- Deep learning (6)
- Dialogue (6)
- WordNet (6)
- BERT (5)
- Linguistics (5)
- Prepositions (5)
- Social media (5)
- Text classification (5)
- AI (4)
- Artificial intelligence (4)
- Digital humanities (4)
- Information retrieval (4)
- Natural language processing (Computer science) (4)
- Prosody (4)
- Situated Dialog (4)
- Sociolinguistics (4)
- Spatial Language (4)
- Spatial Templates (4)
- Twitter (4)
- Artificial Intelligence (3)
- Authorship attribution (3)
- Automatic Speech Recognition (3)
- Publication Year
- Publication
-
- Dissertations, Theses, and Capstone Projects (57)
- Conference papers (18)
- Department of Computer Science Faculty Scholarship and Creative Works (13)
- Faculty Publications (11)
- Theses and Dissertations--Linguistics (10)
-
- Publications and Research (9)
- Articles (6)
- Communication Sciences and Disorders Faculty Articles and Research (4)
- Conference Papers (4)
- Department of Linguistics Faculty Scholarship and Creative Works (4)
- Theses and Dissertations (4)
- Data and Test Instruments (3)
- Dissertations (3)
- Electronic Theses and Dissertations (3)
- Journal of Scientific Information Research (3)
- Student Works (2020-2029) (3)
- CGU Faculty Publications and Research (2)
- Commonwealth Computational Summit (2)
- Computer Science Summer Fellows (2)
- Electrical & Computer Engineering Theses & Dissertations (2)
- Electronic Literature Organization Conference 2020 (2)
- Journal of Tolkien Research (2)
- LING 590/Internet Language (2)
- Library Philosophy and Practice (e-journal) (2)
- Master's Theses (2)
- Masters Theses (2)
- Northeast Journal of Complex Systems (NEJCS) (2)
- Other Resources (2)
- Proceedings from the Document Academy (2)
- School of Computing: Conference and Workshop Papers (2)
- Publication Type
- File Type
Articles 151 - 180 of 250
Full-Text Articles in Computational Linguistics
Speech Perception In “Bubble” Noise: Korean Fricatives And Affricates By Native And Non-Native Korean Listeners, Jiyoung Choi
Speech Perception In “Bubble” Noise: Korean Fricatives And Affricates By Native And Non-Native Korean Listeners, Jiyoung Choi
Dissertations, Theses, and Capstone Projects
The current study examines acoustic cues used by second language learners of Korean to discriminate between Korean fricatives and affricates in noise and how these cues relate to those used by native Korean listeners. Stimuli consist of naturally-spoken consonant-vowel-consonant-vowel (CVCV) syllables: /sɑdɑ/, /s*ɑdɑ/, /tʃɑdɑ/, /tʃhɑdɑ/, and /tʃ*ɑdɑ/. In this experiment, the “bubble noise” methodology of Mandel at al. (2016) was used to identify the time-frequency locations of important cues in each utterance, i.e., where audibility of the location is significantly correlated with correct identification of the utterance in noise. Results show that non-native Korean listeners can discriminate between …
Describing Doggo-Speak: Features Of Doggo Meme Language, Jennifer Bivens
Describing Doggo-Speak: Features Of Doggo Meme Language, Jennifer Bivens
Dissertations, Theses, and Capstone Projects
Doggo-speak is a specialized way of writing most commonly associated with captions on Doggo memes, humorous images of dogs shared in online communities. This paper will explore linguistic features of Doggo-speak through analysis of social media posts by Doggo fan pages. It will use the discussed features as inputs to five machine learning classifiers and will show, through this classification task, that the discussed features are sufficient for distinguishing between Doggo-speak and more general English text.
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Dissertations, Theses, and Capstone Projects
Automatic Identification of personality in conversational speech has many applications in natural language processing such as leader identification in a meeting, adaptive dialogue systems, and dating websites. However, the widespread acceptance of automatic personality recognition through lexical and vocal characteristics is limited by the variability of error rate in a general purpose model among speakers from different demographic groups. While other work reports accuracy, we explored error rates of automatic personality recognition task using classification models for different genders and native language groups (L1). We also present a statistical experiment showing the influence of gender and L1 on the relation …
Automatic Analysis Of Musical Lyrics, Joanna Gormley
Automatic Analysis Of Musical Lyrics, Joanna Gormley
Honors Senior Capstone Projects
Is music getting less sophisticated over time? That is the question which this study aims to answer, with the goal of improving upon previous analysis done on the topic. The blog posts which inspired this project lacked accuracy and dimensionality. Realizing that a larger data set of songs would make a significant difference in the precision of our analysis, we set out to design a piece of software constructed with the capability to analyze several thousand songs. Mimicking previous works which analyzed sophistication of music, the software focuses on the lyrics of songs. Three metrics were used in order to …
Role Of Information Technology In Development Of Eritrean Language - ኣበርክቶ ቴክኖሎጂ ሓበሬታ ኣብ ምምዕባል ቋንቋታት ኤርትራ, Filmon Gebreyesus Ph.D
Role Of Information Technology In Development Of Eritrean Language - ኣበርክቶ ቴክኖሎጂ ሓበሬታ ኣብ ምምዕባል ቋንቋታት ኤርትራ, Filmon Gebreyesus Ph.D
Symposium on Eritrean Literature
Information technology has been affecting us in every day of our lives, especially social media has been the main means of communication in our society. But, all the access to this current and ever-growing technology has always been limited to using it in English, Arab or other languages because our language didn’t come up to speed with the current technology.
Though there has been lots of efforts to develop Tigrigna or other languages application programs to help us use our language, there are still lots of gaps that could be filled to achieve the competence of our languages. In light …
Does The Test Work? Evaluating A Web-Based Language Placement Test, Avizia Long, Sun-Young Shin, Kimberly Geeslin, Erik Willis
Does The Test Work? Evaluating A Web-Based Language Placement Test, Avizia Long, Sun-Young Shin, Kimberly Geeslin, Erik Willis
Faculty Publications
In response to the need for examples of test validation from which everyday language programs can benefit, this paper reports on a study that used Bachman’s (2005) assessment use argument (AUA) framework to examine evidence to support claims made about the intended interpretations and uses of scores based on a new web-based Spanish language placement test. The test, which consisted of 100 items distributed across five item types (sound discrimination, grammar, listening comprehension, reading comprehension, and vocabulary), was tested with 2,201 incoming first-year and transfer students at a large, Midwestern public university. Analyses of internal consistency and validity revealed the …
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Data and Test Instruments
Automatic clickbait detection is a relatively novel task in natural language processing (NLP) and machine learning (ML). “Clickbait” is a hyperlink created primarily to attract attention to its target content. This article introduces a binary classifier, the Language and Information Technology Research Lab (LiT.RL, pronounced “literal”) Clickbait Detector, which automatically distinguishes clickbait from nonclickbait. We used NLP and ML for 38 textual features, contrasting clickbait with “headlinese.” When tested on 11,000 hyperlinks, it achieves 94 per cent accuracy using a support vector machine. Integrated with the LiT.RL News Verification Browser, a downloadable stand-alone research tool, the Clickbait Detector user interface …
A Markedly Different Approach: Investigating Pie Stops Using Modern Empirical Methods, Phillip Barnett
A Markedly Different Approach: Investigating Pie Stops Using Modern Empirical Methods, Phillip Barnett
Theses and Dissertations--Linguistics
In this thesis, I investigate a decades-old problem found in the stop system of Proto-Indo-European (PIE). More specifically, I will be investigating the paucity of */b/ in the forms reconstructed for the ancient, hypothetical language. As cross-linguistic evidence and phonological theory alone have fallen short of providing a satisfactory answer, herein will I employ modern empirical methods of linguistic investigation, namely laboratory phonology experiments and computational database analysis. Following Byrd 2015, I advocate for an examination of synchronic phenomena and behavior as a method for investigating diachronic change.
In Chapter 1, I present an overview of the various proposed phonological …
Exploring The Functional And Geometric Bias Of Spatial Relations Using Neural Language Models, Simon Dobnik, Mehdi Ghanimifard, John D. Kelleher
Exploring The Functional And Geometric Bias Of Spatial Relations Using Neural Language Models, Simon Dobnik, Mehdi Ghanimifard, John D. Kelleher
Conference papers
The challenge for computational models of spatial descriptions for situated dialogue systems is the integration of information from different modalities. The semantics of spatial descriptions are grounded in at least two sources of information: (i) a geometric representation of space and (ii) the functional interaction of related objects that. We train several neural language models on descriptions of scenes from a dataset of image captions and examine whether the functional or geometric bias of spatial descriptions reported in the literature is reflected in the estimated perplexity of these models. The results of these experiments have implications for the creation of …
#Hashtags: A Look At The Evaluative Roles Of Hashtags On Twitter, Leah Rose Schaede
#Hashtags: A Look At The Evaluative Roles Of Hashtags On Twitter, Leah Rose Schaede
Theses and Dissertations--Linguistics
Social media has become a large part of today’s pop culture and keeping up with what is going on not only in our social circles, but around the world. It has given many a platform to unite their causes, build fandoms, and share their commentary with the world. A tool in helping group posts together or give commentary on a thought is the hashtag. In this paper I explore the evaluative roles of hashtags in social media discourse, specifically on Twitter. I use a sample of randomly selected tweets from the Twitter API stream I collected and compiled myself. I …
Cloud‐Based Text Analytics Harvesting, Cleaning And Analyzing Corporate Earnings Conference Calls, Michael Chuancai Zhang, Vikram Gazula, Dan Stone, Hong Xie
Cloud‐Based Text Analytics Harvesting, Cleaning And Analyzing Corporate Earnings Conference Calls, Michael Chuancai Zhang, Vikram Gazula, Dan Stone, Hong Xie
Commonwealth Computational Summit
No abstract provided.
Cloud-Based Text Analytics: Harvesting, Cleaning And Analyzing Corporate Earnings Conference Calls, Michael Chuancai Zhang, Vikram Gazula, Dan Stone, Hong Xie
Cloud-Based Text Analytics: Harvesting, Cleaning And Analyzing Corporate Earnings Conference Calls, Michael Chuancai Zhang, Vikram Gazula, Dan Stone, Hong Xie
Commonwealth Computational Summit
Does management language cohesion in earnings conference calls matter to the capital market? As a part of the research on the above question, and taking advantage of the modern IT technologies, this project:
- harvested 115,882 earnings conference call transcripts from SeekingAlpha.com
- parsed and structured 89,988 transcripts using regular expressions in Stata
- analyzed 179,976 text files using Amazon Elastic Compute Cloud (Amazon EC2), which
- saved almost 2 years (675 days) of the project time
A Sentiment Analysis Of Language & Gender Using Word Embedding Models, Ellyn Rolleston Keith
A Sentiment Analysis Of Language & Gender Using Word Embedding Models, Ellyn Rolleston Keith
Dissertations, Theses, and Capstone Projects
Since Robin Lakoff started the conversation around language and gender with her 1975 essay “Language and Woman’s Place,” extensive work has been done on analyzing sociolinguistics associated with gender. While much work has been done on the differences between how men and women use language, there is less research to be found on language about women as opposed to language about men. In this work, I build a word embedding model from a corpus of Wikipedia film summaries and use this model to create lists of words associated with men and words associated with women. I then use sentiment analysis …
Back To The Future: Logic And Machine Learning, Simon Dobnik, John D. Kelleher
Back To The Future: Logic And Machine Learning, Simon Dobnik, John D. Kelleher
Conference papers
In this paper we argue that since the beginning of the natural language processing or computational linguistics there has been a strong connection between logic and machine learning. First of all, there is something logical about language or linguistic about logic. Secondly, we argue that rather than distinguishing between logic and machine learning, a more useful distinction is between top-down approaches and data-driven approaches. Examining some recent approaches in deep learning we argue that they incorporate both properties and this is the reason for their very successful adoption to solve several problems within language technology.
Quantitative Criticism Of Literary Relationships, Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James Brofos, Jorge A. Bonilla Lopez, Lea Schroeder
Quantitative Criticism Of Literary Relationships, Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James Brofos, Jorge A. Bonilla Lopez, Lea Schroeder
Dartmouth Scholarship
Authors often convey meaning by referring to or imitating prior works of literature, a process that creates complex networks of literary relationships (“intertextuality”) and contributes to cultural evolution. In this paper, we use techniques from stylometry and machine learning to address subjective literary critical questions about Latin literature, a corpus marked by an extraordinary concentration of intertextuality. Our work, which we term “quantitative criticism,” focuses on case studies involving two influential Roman authors, the playwright Seneca and the historian Livy. We find that four plays related to but distinct from Seneca’s main writings are differentiated from the rest of the …
Es-Esa: An Information Retrieval Prototype Using Explicit Semantic Analysis And Elasticsearch, Brian D. Sloan
Es-Esa: An Information Retrieval Prototype Using Explicit Semantic Analysis And Elasticsearch, Brian D. Sloan
Dissertations, Theses, and Capstone Projects
Many modern information retrieval systems work by using keyword search to locate documents in an inverted index by matching those documents based on terms in a user’s query. While highly effective for many use-cases, one notable drawback to simple keyword-based searching is that the contextual knowledge surrounding the user’s underlying information need may be lost, particularly if the user’s query terms are ambiguous or have multiple meanings. Research in the field of semantic search aims to make progress towards resolving this. One methodology in particular, explicit semantic analysis, works by modeling a document not only as a set of …
Robot Perception Errors And Human Resolution Strategies In Situated Human-Robot Dialogue, Niels Schütte, Brian Mac Namee, John D. Kelleher
Robot Perception Errors And Human Resolution Strategies In Situated Human-Robot Dialogue, Niels Schütte, Brian Mac Namee, John D. Kelleher
Articles
Errors in visual perception may cause problems in situated dialogues. We investigated this problem through an experiment in which human participants interacted through a natural language dialogue interface with a simulated robot.We introduced errors into the robot’s perception, and observed the resulting problems in the dialogues and their resolutions.We then introduced different methods for the user to request information about the robot’s understanding of the environment. We quantify the impact of perception errors on the dialogues, and investigate resolution attempts by users at a structural level and at the level of referring expressions.
Automatic Idiom Recognition With Word Embeddings, Jing Peng, Anna Feldman
Automatic Idiom Recognition With Word Embeddings, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms and their literal counterparts do not appear in the same contexts. We propose two approaches: (1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …
Generating Amharic Present Tense Verbs: A Network Morphology & Datr Account, T. Michael W. Halcomb
Generating Amharic Present Tense Verbs: A Network Morphology & Datr Account, T. Michael W. Halcomb
Theses and Dissertations--Linguistics
In this thesis I attempt to model, that is, computationally reproduce, the natural transmission (i.e. inflectional regularities) of twenty present tense Amharic verbs (i.e. triradicals beginning with consonants) as used by the language’s speakers. I root my approach in the linguistic theory of network morphology (NM) and model it using the DATR evaluator. In Chapter 1, I provide an overview of Amharic and discuss the fidel as an abugida, the verb system’s root-and-pattern morphology, and how radicals of each lexeme interacts with prefixes and suffixes. I offer an overview of NM in Chapter 2 and DATR in Chapter 3. In …
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
We present a new methodological approach which combines both naturally-occurring speech harvested on the web and speech data elicited in the laboratory. This proof-of-concept study examines the phenomenon of focus sensitivity in English, in which the interpretation of particular grammatical constructions (e.g., the comparative) is sensitive to the location of prosodic prominence. Machine learning algorithms (support vector machines and linear discriminant analysis) and human perception experiments are used to cross-validate the web-harvested and lab-elicited speech. Results con rm the theoretical predictions for location of prominence in comparative clauses and the advantages using both web-harvested and lab-elicited speech. The most robust …
Towards Multipurpose Readability Assessment, Ion Madrazo
Towards Multipurpose Readability Assessment, Ion Madrazo
Boise State University Theses and Dissertations
Readability refers to the ease with which a reader can understand a text. Automatic readability assessment has been widely studied over the past 50 years. However, most of the studies focus on the development of tools that apply either to a single language, domain, or document type. This supposes duplicate efforts for both developers, who need to integrate multiple tools in their systems, and final users, who have to deal with incompatibilities among the readability scales of different tools. In this manuscript, we present MultiRead, a multipurpose readability assessment tool capable of predicting the reading difficulty of texts of varied …
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Conference papers
In this paper we examine how people negotiate, interpret and repair the frame of reference (FoR) in online text based dialogues discussing spatial scenes in Swedish. We describe work-in-progress in which participants are given different perspectives of the same scene and asked to locate several objects that are only shown on one of their pictures. This task requires participants to coordinate on FoR in order to identify the missing objects. This study has implications for situated dialogue systems.
Lexical Variation, Lexical Innovation, And Speaker Motivations: A Historical Psycholinguistic Approach, Jason Timm Dr.
Lexical Variation, Lexical Innovation, And Speaker Motivations: A Historical Psycholinguistic Approach, Jason Timm Dr.
Linguistics ETDs
Speakers commonly re-purpose existing forms in the mental lexicon to create novel form-meaning. Contemporary evidence that such innovation processes have occurred historically is attested in varying degrees of polysemy in the mental lexicon. This dissertation considers speaker motivations underlying these innnovation processes historically. Strong synchronic relationships between frequency and degree of polysemy, on one hand, and frequency and lexical access, on the other hand, have traditionally been interpreted as evidence for the primacy of economic motivations in processes of lexical innovation. In contrast, the cognitive processes that most commonly facilitate innovation, metaphor and metonymy, have largely been described as processes …
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
Dissertations, Theses, and Capstone Projects
In recent years, Twitter has become a popular testing ground for techniques in authorship attribution. This is due to both the ease of building large corpora as well as the challenges associated with the character limit imposed by the service and the writing styles that have developed as a result. As both false and genuine claims of hacked Twitter accounts have made international news, there is an increasing need for this type of work. For newer Twitter accounts, however, there is little training data. Thus, this study looks to lay the groundwork for cross-domain authorship attribution: training on one source …
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
Dissertations, Theses, and Capstone Projects
This project evaluates four mainstream taggers on a representative collection of child-adult’s dialogues from Child Language Data Exchange System. The nine children’s files from Valian corpora and part of Eve corpora have been manually labeled, and rewrote with LARC tagset. They served as gold standard corpora in the training and testing process. Four taggers: CLAN MOR tagger, ACOPOST trigram tagger, Stanford parser, and Ver. 1.14 of Brill tagger have been tested by 10-fold cross validation. By analyzing what kinds of assumptions the tagger made about category assignment lead to failing, we identify several problematic cases of tagging. By comparing the …
Latent Semantic Indexing In The Discovery Of Cyber-Bullying In Online Text, Jacob L. Bigelow
Latent Semantic Indexing In The Discovery Of Cyber-Bullying In Online Text, Jacob L. Bigelow
Computer Science Summer Fellows
The rise in the use of social media and particularly the rise of adolescent use has led to a new means of bullying. Cyber-bullying has proven consequential to youth internet users causing a need for a response. In order to effectively stop this problem we need a verified method of detecting cyber-bullying in online text; we aim to find that method. For this project we look at thirteen thousand labeled posts from Formspring and create a bank of words used in the posts. First the posts are cleaned up by taking out punctuation, normalizing emoticons, and removing high and low …
Detection Of Cyberbullying In Sms Messaging, Bryan W. Bradley
Detection Of Cyberbullying In Sms Messaging, Bryan W. Bradley
Computer Science Summer Fellows
Cyberbullying is a type of bullying that uses technology such as cell phones to harass or malign another person. To detect acts of cyberbullying, we are developing an algorithm that will detect cyberbullying in SMS (text) messages. Over 80,000 text messages have been collected by software installed on cell phones carried by participants in our study. This paper describes the development of the algorithm to detect cyberbullying messages, using the cell phone data collected previously. The algorithm works by first separating the messages into conversations in an automated way. The algorithm then analyzes the conversations and scores the severity and …
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Dissertations, Theses, and Capstone Projects
While event extraction and automatic summarization have taken great strides in the realm of news stories, fictional narratives like fairy tales have not been so fortunate. A number of challenges arise from the literary elements present in fairy tales that are not found in more straightforward corpora of natural language, such as archaic expressions and sentence structures. To aid in summarization of fictional texts, I created an class - a template for a digital object, in this case a semantic and story event - that captures elements predicted to help classify events as important for inclusion. I wrote a processor …
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Dissertations, Theses, and Capstone Projects
A person’s writing style is uniquely quantifiable and can serve reliably as a biometric. A writer who wishes to remain anonymous can use a number of privacy technologies but can still be identified simply by the words they choose to use — how frequently they use common words like “of,” for instance. Nondescript is a web tool designed first to identify the user’s writing style in terms of word frequency from a given writing sample and document, then to suggest how the author can change their document to lessen its probability of being attributed to them. While Nondescript does not …
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Dissertations, Theses, and Capstone Projects
The process of producing written text is complex and constrained by pressures that range from physical to psychological. In a series of three sets of experiments, this thesis demonstrates the effects of linguistic context on the timing patterns of the production of keystrokes. We elucidate the effect of linguistic context at three different levels of granularity: The first set of experiments illustrate how the nontraditional syntax of a single linguistic construct, the multi-word expression, can create significant changes in keystroke production patterns. This set of experiments is followed by a set of experiments that test the hypothesis on the entire …