Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (98)
- Computer Sciences (92)
- Arts and Humanities (47)
- Artificial Intelligence and Robotics (43)
- Discourse and Text Linguistics (26)
-
- Communication (25)
- Engineering (23)
- Semantics and Pragmatics (23)
- Psychology (22)
- Library and Information Science (20)
- Computer Engineering (19)
- Applied Linguistics (18)
- Psycholinguistics and Neurolinguistics (17)
- Phonetics and Phonology (15)
- First and Second Language Acquisition (13)
- Language Description and Documentation (13)
- Communication Technology and New Media (12)
- Anthropological Linguistics and Sociolinguistics (11)
- Cognition and Perception (11)
- Data Science (11)
- Medicine and Health Sciences (11)
- Databases and Information Systems (10)
- Other Computer Sciences (10)
- Syntax (10)
- Digital Humanities (9)
- Electrical and Computer Engineering (9)
- Social Media (9)
- Institution
-
- City University of New York (CUNY) (67)
- Technological University Dublin (30)
- Montclair State University (18)
- University of Kentucky (16)
- Brigham Young University (10)
-
- Chapman University (7)
- Dartmouth College (5)
- University of Nebraska - Lincoln (5)
- Western University (5)
- East Tennessee State University (4)
- University of Malaya (4)
- Air Force Institute of Technology (3)
- California Polytechnic State University, San Luis Obispo (3)
- Hunan Provincial Institute of Scientific and Technology Information (3)
- Minnesota State University, Mankato (3)
- Old Dominion University (3)
- Portland State University (3)
- San Jose State University (3)
- University of Central Florida (3)
- Binghamton University (2)
- Boise State University (2)
- Claremont Colleges (2)
- New Jersey Institute of Technology (2)
- Purdue University (2)
- The University of Akron (2)
- University of Louisville (2)
- Ursinus College (2)
- Valparaiso University (2)
- Bellarmine University (1)
- Bowling Green State University (1)
- Keyword
-
- Natural Language Processing (19)
- Computational linguistics (18)
- Natural language processing (16)
- Machine learning (12)
- NLP (10)
-
- Computational Linguistics (9)
- Machine Learning (7)
- Corpus linguistics (6)
- Deep learning (6)
- Dialogue (6)
- WordNet (6)
- BERT (5)
- Linguistics (5)
- Prepositions (5)
- Social media (5)
- Text classification (5)
- AI (4)
- Artificial intelligence (4)
- Digital humanities (4)
- Information retrieval (4)
- Natural language processing (Computer science) (4)
- Prosody (4)
- Situated Dialog (4)
- Sociolinguistics (4)
- Spatial Language (4)
- Spatial Templates (4)
- Twitter (4)
- Artificial Intelligence (3)
- Authorship attribution (3)
- Automatic Speech Recognition (3)
- Publication Year
- Publication
-
- Dissertations, Theses, and Capstone Projects (57)
- Conference papers (18)
- Department of Computer Science Faculty Scholarship and Creative Works (13)
- Faculty Publications (11)
- Theses and Dissertations--Linguistics (10)
-
- Publications and Research (9)
- Articles (6)
- Communication Sciences and Disorders Faculty Articles and Research (4)
- Conference Papers (4)
- Department of Linguistics Faculty Scholarship and Creative Works (4)
- Theses and Dissertations (4)
- Data and Test Instruments (3)
- Dissertations (3)
- Electronic Theses and Dissertations (3)
- Journal of Scientific Information Research (3)
- Student Works (2020-2029) (3)
- CGU Faculty Publications and Research (2)
- Commonwealth Computational Summit (2)
- Computer Science Summer Fellows (2)
- Electrical & Computer Engineering Theses & Dissertations (2)
- Electronic Literature Organization Conference 2020 (2)
- Journal of Tolkien Research (2)
- LING 590/Internet Language (2)
- Library Philosophy and Practice (e-journal) (2)
- Master's Theses (2)
- Masters Theses (2)
- Northeast Journal of Complex Systems (NEJCS) (2)
- Other Resources (2)
- Proceedings from the Document Academy (2)
- School of Computing: Conference and Workshop Papers (2)
- Publication Type
- File Type
Articles 91 - 120 of 250
Full-Text Articles in Computational Linguistics
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nonetheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, …
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Theses and Dissertations
An underexplored aspect of the abundant linguistic data available for computational analysis is the potential for conducting large-scale comparative analyses between different discourses as a more rigorous complement to existing methods. A significant obstacle in comparative studies is the requirement that a researcher distinguish between discursive distinctions that are superficial and those that reflect deeper, structural relationships between the discourses being compared. In this dissertation, I describe two methodologies for comparing discourses in terms of their underlying structures, drawing on distributional semantic models and information theory. These methodologies are explored within three case studies in which the discourses of several …
Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross
Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross
Conference papers
Fluid interaction between virtual agents and humans requires the understanding of many issues of conversational pragmatics. One such issue is the interaction between communication strategy and personality. As a step towards developing models of personality driven pragmatics policies, in this paper, we present our initial experiment to explore differences in user interaction with two contrasting avatar personalities. Each user saw a single personality in a video-call setting and gave feedback on the interaction. Our expectations, that a more extroverted outgoing positive personality would be a more successful tutor, were only partially confirmed. While this personality did induce longer conversations in …
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
Dissertations, Theses, and Capstone Projects
Nearly thirty years ago, the United States Supreme Court revaluated the criteria for accepting forensic science and expert testimony, challenging Forensic Linguistics to assert itself as a reputable science. Much work has been produced in the interim to that end, but much still needs to be accomplished to satisfy the judicial standards. Computational linguistics has the potential to provide that necessary analytical framework. This paper’s intent is two-fold. First, there are two competing theories on the proper features necessary to identify an unknown author. Four features were drawn from the syntactic computational linguistics tradition and four from computational stylometry to …
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Dissertations, Theses, and Capstone Projects
This dissertation presents the first implementation of label imputation for the task of homograph disambiguation using 1) transcribed audio, and 2) parallel, or translated, corpora. For label imputation from parallel corpora, a hypothesis of interlingual alignment between homograph pronunciations and text word forms is developed and formalized. Both audio and parallel corpora label imputation techniques are tested empirically in experiments that compare homograph disambiguation model performance using: 1) hand-labeled training data, and 2) hand-labeled training data augmented with label-imputed data. Regularized, multinomial logistic regression and pre-trained ALBERT, BERT, and XLNet language models fine-tuned as token classifiers are developed for homograph …
Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor
Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor
Dissertations, Theses, and Capstone Projects
This paper investigates recent English loanwords in Russian and explores ways in which computational methods can help further theoretical research. The goal of the study is two-fold: to find new, previously unattested loanwords borrowed over the last decade and to examine the rate of adaptation of the new borrowings, attested by the degree to which they conform to the constraints of the Russian language. First, we train a finite-state pipeline that combines character n-gram language models, which encode phonotactic and lexical properties of loanwords, with a binary classifier to detect loanwords. The model achieves state-of-the-art performance results during evaluation, surpassing …
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Dissertations, Theses, and Capstone Projects
In a quick search online, one can find many tools which use information from news headlines to make predictions concerning the trajectory of a given stock. But what if we went further, looking instead into the text of the article, to extract this and other information? Here, the goal is to extract the sentence in which a stock ticker symbol is mentioned from a news article, then determine sentiment and subjectivity values from that sentence, and finally make a prediction on whether or not the value of that stock will go up or not in a 24-hour timespan. Bloomberg News …
The Public Innovations Explorer: A Geo-Spatial & Linked-Data Visualization Platform For Publicly Funded Innovation Research In The United States, Seth Schimmel
Dissertations, Theses, and Capstone Projects
The Public Innovations Explorer (https://sethsch.github.io/innovations-explorer/app/index.html) is a web-based tool created using Node.js, D3.js and Leaflet.js that can be used for investigating awards made by Federal agencies and departments participating in the Small Business Innovation Research (SBIR) and Small Business Technology Transfer (STTR) grant-making programs between 2008 and 2018. By geocoding the publicly available grants data from SBIR.gov, the Public Innovations Explorer allows users to identify companies performing publicly-funded innovative research in each congressional district and obtain dynamic district-level summaries of funding activity by agency and year. Applying spatial clustering techniques on districts' employment levels across major economic sectors provides users …
Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos
Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos
Electronic Theses and Dissertations
This study investigates one of the Polish language’s most arbitrary cases: the genitive masculine inanimate singular. It collects and ranks several guidelines to help language learners discern its proper usage and also introduces a framework to provide detailed feedback regarding arbitrary cases. The study tests this framework by implementing and evaluating a hybrid grammar checker called PLPrepare. PLPrepare performs similarly to other grammar checkers and is able to detect genitive case usages and provide feedback based on a number of error classifications.
Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos
Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos
World Languages & Cultures Department Publications
SEÑAL is a modular program that automatizes and facilitates the assessment of Spanish L2 compositions. The tool can extract syntactic and lexical information, while also assessing grammar, and Spanish L2 proficiency. A computational tool that can process written learners’ corpora and extract measures of language development has enormous practical value in Spanish and modern language departments alike.
An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli
An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli
Communication Sciences and Disorders Faculty Articles and Research
"We are four researchers who study psycholinguistics, linguistics, neuroscience and deaf education. Our team of deaf and hearing scientists worked with a group of software engineers to create the ASL-LEX database that anyone can use for free. We cataloged information on nearly 3,000 signs and built a visual, searchable and interactive database that allows scientists and linguists to work with ASL in entirely new ways."
Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci
Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci
The Journal of Electronic Theses and Dissertations
Although the Arabic language is spoken by more than 300 million people and is one of the six official languages of the United Nations (UN), there has been less research done on Arabic text data (compared to English) in the realm of machine learning, especially in text classification. In the past decade, Arabic data such as news, tweets, etc. have begun to receive some attention. Although automatic text classification plays an important role in improving the browsability and accessibility of data, Electronic Theses and Dissertations (ETDs) have not received their fair share of attention, in spite of the huge number …
An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals
An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals
Faculty Publications
Many computerized tools for comparing phonetic transcriptions have been proposed and shared in the past; however, previous tools are relatively difficult to access and incorporate into clinical and research practice, or require users to learn additional phonetic symbol systems. The purpose of this project was to develop and test a readily available web-based application for quantitatively comparing phonetic transcriptions that are input using International Phonetic Alphabet (IPA) symbols. A web-based computer application was developed to allow for IPA phonetic transcription comparison. A point-and-click keyboard was developed to provide support for character input of the full IPA, as well as most …
The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory
The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory
Communication Sciences and Disorders Faculty Articles and Research
ASL-LEX is a publicly available, large-scale lexical database for American Sign Language (ASL). We report on the expanded database (ASL-LEX 2.0) that contains 2,723 ASL signs. For each sign, ASL-LEX now includes a more detailed phonological description, phonological density and complexity measures, frequency ratings (from deaf signers), iconicity ratings (from hearing non-signers and deaf signers), transparency (“guessability”) ratings (from non-signers), sign and videoclip durations, lexical class, and more. We document the steps used to create ASL-LEX 2.0 and describe the distributional characteristics for sign properties across the lexicon and examine the relationships among lexical and phonological properties of signs. Correlation …
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
Dissertations, Theses, and Capstone Projects
Code-switching is the linguistic phenomenon where a multilingual person alternates between two or more languages in a conversation, whether that be spoken or written. This thesis studies the automatic detection of code-switching occurring specifically between English and Spanish in two corpora.
Twitter and other social media sites have provided an abundance of linguistic data that is available to researchers to perform countless experiments. Collecting the data is fairly easy if a study is on monolingual text, but if a study requires code-switched data, this becomes a complication as APIs only accept one language as a parameter. This thesis focuses on …
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
Dissertations, Theses, and Capstone Projects
As a subdomain of author profiling, gender prediction (sometimes called gender inference) has received a substantial amount of attention—both as a task in itself, and for other downstream analyses. Throughout the existing literature various statistical and machine learning methods have been applied to extract features in order to either characterize and differentiate female and male writing styles, or simply to achieve maximum accuracy on gender prediction as a binary classification task. However, researchers often do not disclose how they conceptualize gender nor do they consider the implications that gender prediction has for non-binary and trans individuals. Along with an overview …
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Library Faculty Publications
This year’s iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems’ predictions. Transformer-based …
What You Do Or What You Say? An Examination Of Analyst Reactions To Prototypical And Non-Prototypical Ceos Linguistic And Competitive Behaviors, Courtney Hart
Theses and Dissertations--Management
Non-prototypical CEOs are those that process different demographic characteristics from a target reference group. In the US, a non-prototypical CEO is both white and male. While the negative responses to non-prototypical leaders based on race and gender have been well documented, we know less on what these leaders do that may influence biased evaluations. In this dissertation I took an impression management view to examine analysts’ evaluative bias (AEB) on prototypical and non-prototypical CEOs hiding linguistic behaviors and competitive aggressiveness. Specifically, I examined hiding linguistic behaviors on quarterly conference calls and two attributes of competitive repertoire will be researched. Drawing …
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
We present the results and the main findings of the NLP4IF-2021 shared tasks. Task 1 focused on fighting the COVID-19 infodemic in social media, and it was offered in Arabic, Bulgarian, and English. Given a tweet, it asked to predict whether that tweet contains a verifiable claim, and if so, whether it is likely to be false, is of general interest, is likely to be harmful, and is worthy of manual fact-checking; also, whether it is harmful to society, and whether it requires the attention of policy makers. Task 2 focused on censorship detection, and was offered in Chinese. A …
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
Student Works (2020-2029)
While it has been established that discursive practices involving writing are highly situated and context-bound, few local studies have examined the processes involved in the construction of texts and the shaping of the intended text. Recent perspectives on writing as ways of working and acting, and texts as social action suggest that studying the processes behind the construction of a text could provide a clearer understanding of the discursive practices involved in the shaping of the text. This is seen as relevant to professional writing context in the present study as texts are often shaped to facilitate the social action …
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Dissertations, Theses, and Capstone Projects
Recent advances in deep learning have greatly improved the ability of researchers to develop effective machine translation systems. In particular, the application of modern neural architectures, such as the Transformer, has achieved state-of-the-art BLEU scores in many translation tasks. However, it has been found that even state-of-the-art neural machine translation models can suffer from certain implicit biases, such as gender bias (Lu et al., 2019). In response to this issue, researchers have proposed various potential solutions: some have proposed approaches that inject missing gender information into models, while others have attempted modifying the training data itself. We focus on mitigating …
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Dissertations, Theses, and Capstone Projects
This thesis presents experiments with using representation learning to explore how neural networks learn. Neural networks which take text as input create internal representations of the text during their training. Recent work has found that these representations can be used to perform other downstream linguistic tasks, such as part-of-speech (POS) tagging. This demonstrates that the neural networks are learning linguistic information and storing this information in the representations. We focus on the representations created by neural machine translation (NMT) models and whether they can be used in POS tagging. We train 5 NMT models including an auto-encoder. We extract the …
A Semantic Prosody Analysis Of Swear Words In A Corpus Of English Songs, Hashim Hazri Shahreen
A Semantic Prosody Analysis Of Swear Words In A Corpus Of English Songs, Hashim Hazri Shahreen
Student Works (2020-2029)
This study looks at the semantic prosody of swear words found in a corpus of English songs by looking at their collocations using a corpus software; AntConc. A total of 545 songs were chosen based on the Billboard year-end chart from 2011 to 2016 with 243,689 number of tokens and 8,139 number of word types. Word lists and concordance lines were generated from the corpus for the data analysis. The analysis on the concordance lines showed that not all swear words with negative-based meaning possessed negative semantic prosody. Despite possessing negative semantic prosody, 10 out of 15 swear words possessed …
Lulling Waters: A Poetry Reading For Real-Time Music Generation Through Emotion Mapping, Ashley Muniz, Toshihisa Tsuruoka
Lulling Waters: A Poetry Reading For Real-Time Music Generation Through Emotion Mapping, Ashley Muniz, Toshihisa Tsuruoka
Electronic Literature Organization Conference 2020
Through a poetic narrative, “Lulling Waters” tells the story of a whale overcoming the loss of his mother, who passed away from ingesting plastic, as he attempts to escape from the polluted oceanic world. The live performance of this poem utilizes a software system called Soundwriter, which was developed with the goal of enriching the oral storytelling experience through music. This video demonstrates how Soundwriter’s real-time hybrid system was able to analyze “Lulling Waters” through its lexical and auditory features. Emotionally salient words were given ratings based on arousal, valence, and dominance while the emotionally charged prosodic features of the …
Poetry For Seers Or The Peruvian Visual Poetic Tradition In Front Of New Media, Michael Hurtado, Pamela Medina, Enrique García, Michael Prado
Poetry For Seers Or The Peruvian Visual Poetic Tradition In Front Of New Media, Michael Hurtado, Pamela Medina, Enrique García, Michael Prado
Electronic Literature Organization Conference 2020
Since the first decades of the twentieth century, Peruvian poetic tradition has been characterized by experimental uses of language. Among these possibilities, some records tensioned this medium from the link with the plastic arts, as in the case of the poetry of José María Eguren, while others opted for the playing with the spatiality and visuality of the blank sheet, such as in the case of the work of Carlos Oquendo de Amat. However, it is not until the appearance of the poetry of César Vallejo, specifically with a poems like Trilce in 1922, that these breakages force us to …
Identifying Facets Of Reader-Generated Online Reviews Of Children’S Books Based On A Textual Analysis Approach, Yunseon Choi, Soohyung Joo
Identifying Facets Of Reader-Generated Online Reviews Of Children’S Books Based On A Textual Analysis Approach, Yunseon Choi, Soohyung Joo
Information Science Faculty Publications
With the increasing popularity of social media, online reviews have become one of the primary information sources for book selection. Prior studies have analyzed online reviews, mostly in the domain of business. However, little research has examined the content of online book reviews of children’s books. Book reviews generated by book readers contain different aspects of information, such as opinions, feedback, or emotional responses, from the perspectives of readers. This study explores what aspects of the books are addressed in readers’ reviews, and then it intends to identify categorical features or facets of online book reviews of children’s books. We …
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
University Honors Theses
CBOW and Skip Gram are two NLP techniques to produce word embedding models that are accurate and performant. They were invented in the seminal paper by T. Mikolov et al. and have since observed optimizations such as negative sampling and subsampling. This paper implements a fully-optimized version of these models using Py-Torch and runs them through a toy sentiment/subject analysis. It is weakly observed that different corpus types affect the skew of word embeddings such that fictional corpus are better suited for sentiment analysis and non-fictional for subject analysis.
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
University Honors Theses
The automatic extraction of keyphrases from scholarly papers is a necessary step for many Natural Language Processing (NLP) tasks, including text retrieval, machine translation, and text summarization. However, due to the different grammatical and semantic intricacies of languages, this is a highly language-dependent task. Many free and open source implementations of state-of-the-art keyphrase extraction techniques exist, but they are not adapted for processing Russian text. Furthermore, the multi-linguistic character of scholarly papers in the field of Russian computational linguistics and NLP introduces additional complexity to keyphrase extraction. This paper describes a free and open source program as a proof of …
Enhancing The Performance Of Ir-Based Traceability Recovery Of Requirement Artifacts Using Noun Phrases, Dafaalla Abdelrahman Mashahi Khalafalla
Enhancing The Performance Of Ir-Based Traceability Recovery Of Requirement Artifacts Using Noun Phrases, Dafaalla Abdelrahman Mashahi Khalafalla
Student Works (2020-2029)
Requirement traceability can be considered as a measure of software quality to help achieve validation, verification, and reusability. Neglecting traceability leads to less maintainable software. Creating traceability links after-the-fact, known as traceability recovery, is a tedious and time-consuming process when it is done manually. Therefore, information retrieval (IR) methods have been used to automatically identify traceability links between the artifacts. However, as a result of limitations of the software engineer and the IR techniques, the performance of the IR methods is negatively affected. There is no IR method that is able to recover traceability links between artifacts with high precision …
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Dissertations, Theses, and Capstone Projects
The UMETRICS database (Universities: Measuring the Effects of Research on Innovation, Competitiveness, and Science) contains rich information on grants from sponsored federal and non-federal research for 32 universities over a 15-year period. It is hosted at IRIS (Institute for Research on Innovation and Science, University of Michigan) and serves as a rich source of university administrative data; however, it does not contain information on research fields. Categorizing grants data by research field can help to measure results of investment in research and science and provide evidence for the data-driven policy-making; yet administrative data often lacks this type of categorization. In …