Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (23)
- Computer Sciences (21)
- Artificial Intelligence and Robotics (13)
- Scholarly Communication (10)
- Education (8)
-
- Numerical Analysis and Scientific Computing (6)
- Databases and Information Systems (5)
- Medicine and Health Sciences (5)
- Public Affairs, Public Policy and Public Administration (4)
- Business (3)
- Cataloging and Metadata (3)
- Engineering (3)
- Higher Education (3)
- Science and Technology Policy (3)
- Statistics and Probability (3)
- Arts and Humanities (2)
- Information Literacy (2)
- International and Comparative Education (2)
- Sports Sciences (2)
- Accessibility (1)
- Algebraic Geometry (1)
- Business Administration, Management, and Operations (1)
- Communication Sciences and Disorders (1)
- Curriculum and Instruction (1)
- Design of Experiments and Sample Surveys (1)
- Disability and Equity in Education (1)
- Education Economics (1)
- Keyword
-
- Artificial intelligence (5)
- Bibliometric analysis (4)
- Natural language processing (4)
- Computational linguistics (3)
- Large language models (3)
-
- Natural language understanding (3)
- Academic libraries (2)
- Affordable course content (2)
- Computer science (2)
- Deep learning (2)
- Electronic theses and dissertations (2)
- Hypothesis extraction (2)
- Hypothesis generation (2)
- Machine learning (2)
- Methodology (2)
- Methods (2)
- Natural language inference (2)
- OER (2)
- Publishing (2)
- Qualitative (2)
- Research (2)
- Scholarly publishing (2)
- Scientific claim verification (2)
- Sentiment analysis (2)
- Sports (2)
- Summarization (2)
- Text mining (2)
- Topic modeling (2)
- ACC (1)
- Academic capitalism (1)
- Publication Year
- Publication
-
- Computer Science Faculty Publications (16)
- Open Access Week (3)
- Educational Leadership & Workforce Development Faculty Publications (2)
- Libraries Faculty & Staff Publications (2)
- Management Faculty Publications (2)
-
- Rehabilitation Sciences Faculty Publications (2)
- Electrical & Computer Engineering Faculty Publications (1)
- Engineering Management & Systems Engineering Faculty Publications (1)
- Human Movement Studies & Special Education Faculty Publications (1)
- Information Technology & Decision Sciences Faculty Publications (1)
- Libraries Faculty & Staff Presentations (1)
- Philosophy Faculty Publications (1)
- Physics Faculty Publications (1)
- Psychology Faculty Publications (1)
- STEMPS Faculty Publications (1)
- Speech-Language Pathology Faculty Publications (1)
- VMASC Publications (1)
- Publication Type
Articles 1 - 30 of 38
Full-Text Articles in Scholarly Publishing
A Review Of Replications Over 88 Years Of The American Speech-Language Hearing Association Journal Publications, Andrea L.B. Ford, Jessica S. Riccardi, Austin Thompson, Helen L. Long, James C. Borders, Scott R. Schroeder, Alisa Baron, Mariam El Amin, Danika L. Pfeiffer
A Review Of Replications Over 88 Years Of The American Speech-Language Hearing Association Journal Publications, Andrea L.B. Ford, Jessica S. Riccardi, Austin Thompson, Helen L. Long, James C. Borders, Scott R. Schroeder, Alisa Baron, Mariam El Amin, Danika L. Pfeiffer
Speech-Language Pathology Faculty Publications
Purpose: This review aimed to determine the prevalence of replication studies in the field of Communication Sciences and Disorders and examine longitudinal changes in this prevalence across American Speech-Language-Hearing Association (ASHA) journals from 1936 to 2024. It is a conceptual replication of Cook and colleagues' review of replications in special education journals.
Method: Our data source comprised 17,843 journal articles published in seven ASHA peer-reviewed journals between 1936 and 2024. We identified 3,903 articles, excluding duplicates, that had at least one instance of replicat*. Two research team members screened articles independently, and then one research team member coded each article, …
Modeling Rank Distribution And The Relative Importance Factor Index In Discrete Power-Law Models: Application To Social Resilience Using The Scopus Database, Brian Llinas, Jose Padilla, Humberto Llinas, Erika Frydenlund, Katherine Palacio
Modeling Rank Distribution And The Relative Importance Factor Index In Discrete Power-Law Models: Application To Social Resilience Using The Scopus Database, Brian Llinas, Jose Padilla, Humberto Llinas, Erika Frydenlund, Katherine Palacio
VMASC Publications
Prior research on power-law distributions has primarily focused on modeling frequency patterns, with less attention given to rank distributions and how ranked positions reflect relative importance among elements. In discrete power-law distributions, frequency-based metrics often provide limited discrimination in the tail, where elements may exhibit similar counts but differ in relative dominance. These patterns are especially evident, for instance, in academic publishing, where keywords, affiliations, and citations commonly exhibit power-law behavior. To address this limitation, we introduce the Relative Importance Factor (RIF) Index, a statistical measure derived from the estimated discrete power-law rank distribution rather than an additional independent parameter. …
Can An Experienced Qualitative Researcher Distinguish Ai From Human Qualitative Content Analysis?, Alexandra T. Lucas, Jianna Ramos, Maria Bajwa, Aaron Calhoun, Mark W. Scerbo, Janice C. Palaganas
Can An Experienced Qualitative Researcher Distinguish Ai From Human Qualitative Content Analysis?, Alexandra T. Lucas, Jianna Ramos, Maria Bajwa, Aaron Calhoun, Mark W. Scerbo, Janice C. Palaganas
Psychology Faculty Publications
Background
Artificial intelligence (AI) has become increasingly embedded in research workflows. Large language models (LLMs) are being used to code segments of text, organise codes into themes and interpret patterns within contexts. Recent comparisons between human and AI analyses demonstrate up to 80% thematic overlap, yet humans consistently exhibit deeper interpretive integration and contextual understanding. This study assesses whether experienced researchers can distinguish between entirely human-generated and AI-generated qualitative content analyses of a simulation debriefing.
Methods
We conducted a qualitative descriptive study comparing human-generated qualitative content analysis (QCA) with ChatGPT-4o-generated QCA using a single focus group transcript on emotion management …
Ensuring Rich Rigor Of Qualitative Methodologies In Behavior Analytic Research, Daria K. Lorio-Barsten, Selena J. Layden
Ensuring Rich Rigor Of Qualitative Methodologies In Behavior Analytic Research, Daria K. Lorio-Barsten, Selena J. Layden
Human Movement Studies & Special Education Faculty Publications
Quantitative methods remain the hallmark of research in applied behavior analysis. Yet, such methods frequently fail to capture the nuances of context where behavior analysis is practiced. Therefore, qualitative methods can provide complementary means to gain deeper insight into changes in socially significant behavior. We believe that researchers within the field of behavior analysis have much to gain from embracing qualitative methodologies. We propose that more researchers can and should consider conducting rigorous qualitative research to elevate the voices of the participants and relate the depth and complexities of their nuanced experiences. This article discusses Tracy’s “big tent” quality criteria …
Affordable Course Content And Open Education Resources For Undergraduate Courses Teaching Fundamentals Of Wireless Communications And Networking, Dimitrie C. Popescu, Otilia Popescu
Affordable Course Content And Open Education Resources For Undergraduate Courses Teaching Fundamentals Of Wireless Communications And Networking, Dimitrie C. Popescu, Otilia Popescu
Electrical & Computer Engineering Faculty Publications
Wireless communication systems and networks along with the services they provide have become an essential component of the modern 21st century society, fueling job growth in the wireless industry and increasing the need for engineers specialized in wireless communication systems. As a consequence, over the past two decades, undergraduate courses teaching fundamentals of wireless communication systems and networks have become common in electrical and computer engineering and technology programs. At the same time, the number of textbooks dedicated to wireless systems and networks published by mainstream publishers has also grown, with availability in various formats and offerings and a significant …
Exploring Marshall–Olkin Models Through Bibliometric And Topic Modeling Approaches Uses Latent Dirichlet Allocation (1981-2025): A Study Based On Scopus Data, Humberto Llinás, Brian Llinás, Carlos López, Daniela Nuñez
Exploring Marshall–Olkin Models Through Bibliometric And Topic Modeling Approaches Uses Latent Dirichlet Allocation (1981-2025): A Study Based On Scopus Data, Humberto Llinás, Brian Llinás, Carlos López, Daniela Nuñez
Computer Science Faculty Publications
The Marshall–Olkin family of distributions has gained increasing attention in fields such as reliability engineering, survival analysis, financial risk modeling, and actuarial science because of its flexibility in modeling dependence among events and its wide range of extensions. Despite its growing relevance, a systematic understanding of how research on Marshall–Olkin models has evolved over time is still limited. This study addresses this gap by combining bibliometric techniques with topic modeling to analyze the structure and evolution of the scientific literature on Marshall–Olkin models. The analysis includes all 266 peer-reviewed publications on Marshall–Olkin models indexed in Scopus between 1981 and 2025. …
Open Scholarly Information Systems: Status Quo, Challenges, Opportunities, Hannah Bast, Guillaume Cabanac, Paolo Manghi, Jian Wu, Marcel R. Ackermann
Open Scholarly Information Systems: Status Quo, Challenges, Opportunities, Hannah Bast, Guillaume Cabanac, Paolo Manghi, Jian Wu, Marcel R. Ackermann
Computer Science Faculty Publications
Over the past 30 years, a rich ecosystem of scholarly information systems has developed that openly provide their services to the scientific community. These systems include aggregators of bibliographic metadata (e.g., DBLP, OpenCitations, OpenAIRE Graph, OpenAlex, ORKG, Semantic Scholar, CiteSeerX, and CORE); publication, data, and software repositories (e.g., Arxiv.org, Figshare, Zenodo, Software Heritage, and Dataverse); and PID authorities (e.g., ORCID, ROR, Crossref, and DataCite). This interdisciplinary Dagstuhl Seminar "Open Scholarly Information Systems: Status Quo, Challenges, Opportunities" (25381) was the first of its kind to bring together practitioners from this ecosystem, as well as researchers investigating related questions or relying on …
Women's And Men's Authorship Experiences: A Prospective Meta Analysis, George C. Banks, Lisa M. Rasmussen, Scott Tonidandel, Jeffrey M. Pollack, Mary M. Hausfeld, Courtney Williams, Betsy H. Albritton, Joseph A. Allen, Nicolas Bastardoz, John H. Batchelor, Andrew A. Bennett, Roman Briker, Christopher M. Castille, Bart A. De Jong, Elise Demeter, Justin A. Desimone, James G. Field, Maria Figueroa-Armijos, M. Fernanda Garcia, William L. Gardner, J. Jeffrey Gish, Laura M. Giurge, Claudia N. Gonzalez-Brambila, M. Gloria González-Morales, Lorenz Graf-Vlachy, Roopak Kumar Gupta, Amanda S. Hinojosa, Zion Howard, Sven Kepes, Tine Köhler, Dejun Tony Kong, Markus Langer, Teng Lat Loi, Liam P. Maher, Chao Miao, Murad A. Mithani, Lakshmi Balachandran Nair, William G. Obenauer, Ernest H. O'Boyle, Jason R. Pierce, Deborah M. Powell, Roni Reiter-Palmon, Deborah E. Rupp, Srinivasan Tatachari, Jane S. Thomas, Tiia Vissak, Jako Volschenk, Chen Wang, Christopher E. Whelpley, Hans-Georg Wolff, Haley M. Woznyj, Tao Yang
Women's And Men's Authorship Experiences: A Prospective Meta Analysis, George C. Banks, Lisa M. Rasmussen, Scott Tonidandel, Jeffrey M. Pollack, Mary M. Hausfeld, Courtney Williams, Betsy H. Albritton, Joseph A. Allen, Nicolas Bastardoz, John H. Batchelor, Andrew A. Bennett, Roman Briker, Christopher M. Castille, Bart A. De Jong, Elise Demeter, Justin A. Desimone, James G. Field, Maria Figueroa-Armijos, M. Fernanda Garcia, William L. Gardner, J. Jeffrey Gish, Laura M. Giurge, Claudia N. Gonzalez-Brambila, M. Gloria González-Morales, Lorenz Graf-Vlachy, Roopak Kumar Gupta, Amanda S. Hinojosa, Zion Howard, Sven Kepes, Tine Köhler, Dejun Tony Kong, Markus Langer, Teng Lat Loi, Liam P. Maher, Chao Miao, Murad A. Mithani, Lakshmi Balachandran Nair, William G. Obenauer, Ernest H. O'Boyle, Jason R. Pierce, Deborah M. Powell, Roni Reiter-Palmon, Deborah E. Rupp, Srinivasan Tatachari, Jane S. Thomas, Tiia Vissak, Jako Volschenk, Chen Wang, Christopher E. Whelpley, Hans-Georg Wolff, Haley M. Woznyj, Tao Yang
Management Faculty Publications
The opaqueness of author naming and ordering, when coupled with power dynamics, can lead to a number of disadvantages in academic careers. In this commentary, we investigate gender differences in authorship experiences in a large prospective meta-analytic study (k = 46; n = 3,565; 12 countries). We find that women’s and men’s authorship experiences differ significantly with women reporting greater prevalence of problematic behaviors. We present seven actionable recommendations for improving the receipt and reporting of intellectual credit. Such actions are needed to ensure fairness in authorship, which is one of the most powerful factors in academics’ career outcomes.
Dominance Of Leading Business Schools In Top Journals: Insights For Increasing Institutional Representation, Rodrigo Romero-Silva, Erika Marsillac, Sander De Leeuw
Dominance Of Leading Business Schools In Top Journals: Insights For Increasing Institutional Representation, Rodrigo Romero-Silva, Erika Marsillac, Sander De Leeuw
Information Technology & Decision Sciences Faculty Publications
The competitive push for business schools to publish in prestigious journals has resulted in a disproportionate number of papers in prestigious Management and Operations Research/Management Science (OR/MS) journals coming from a select group of institutions. Our analysis shows the Matthew effect of prestigious journals favors established schools with 51.2% of papers in 18 Management ABS 4* journals and 61.3% of papers in 3 OR/MS ABS 4* journals involving authors from the 100 top business schools identified by the University of Texas at Dallas (UTD). Citation patterns are similarly concentrated among papers authored by scholars from UTD-listed business schools, with nearly …
Techno-Pedagogic Discourse And The Online Learning Assetization Regime, David F. Ayers
Techno-Pedagogic Discourse And The Online Learning Assetization Regime, David F. Ayers
Educational Leadership & Workforce Development Faculty Publications
Research on academic capitalism has critiqued the commodification of knowledge, but it has not critiqued assetization. The difference between commodification and assetization is more than a technical distinction. Extracting value from assets requires a different regime of coordination than what is required for extracting value from commodities. Through an analysis of 132 texts related to the online course review process at 16 research universities in the USA, I propose and problematize an online learning assetization regime which transforms discipline-based knowledge into digital content that can be owned, controlled, and managed as a university asset. This process potentially alienates faculty from …
An 11-Year (2012-2022) Review Of Journal Of Athletic Training Publication Study Designs And Sample Sizes, Zachary K. Winkelmann, Samantha E. Scarneo-Miller, Emily C. Smith, Ryan M. Argetsinger, Lindsey E. Eberman
An 11-Year (2012-2022) Review Of Journal Of Athletic Training Publication Study Designs And Sample Sizes, Zachary K. Winkelmann, Samantha E. Scarneo-Miller, Emily C. Smith, Ryan M. Argetsinger, Lindsey E. Eberman
Rehabilitation Sciences Faculty Publications
Background
Research findings must be representative by creating a sample of individuals, ensuring the results can be generalized and applicable to a larger population, which has historically been guided by a power analysis. However, the varied research design methods require a unique approach to sampling and a formula for recruitment and size. Therefore, the purpose of this study was to analyze historical data from published manuscripts in the Journal of Athletic Training (JAT) relative to study design and sample sizes. A secondary purpose was to further explore metrics for survey-based research.
Methods
This descriptive analysis explored 1267 publications in each …
A Bibliometric Analysis Of Ai-Driven Healthcare Literature Containing Kos Keywords: Trends, Themes, And Gaps, Julaine Clunis, Eric Asare
A Bibliometric Analysis Of Ai-Driven Healthcare Literature Containing Kos Keywords: Trends, Themes, And Gaps, Julaine Clunis, Eric Asare
STEMPS Faculty Publications
As artificial intelligence (AI) becomes increasingly embedded in healthcare applications, concerns have emerged around the trustworthiness, interpretability, and context-awareness of these systems. Knowledge Organization Systems (KOS) hold considerable potential to address these challenges by supporting semantic standardization, explainability, and domain alignment. This study presents a bibliometric analysis of scholarly publications referencing both AI and healthcare concepts to examine how KOS are positioned within this evolving discourse. The findings indicate that while early literature frequently and explicitly referenced KOS—such as ontologies, controlled vocabularies, and classification systems—their visibility has declined relative to newer paradigms such as machine learning and large language models. …
Leveraging Transformer-Based Ocr Model With Generative Data Augmentation For Engineering Document Recognition, Wael Khallouli, Mohammad Shahab Uddin, Andres Sousa-Poza, Jiang Li, Samuel Kovacic
Leveraging Transformer-Based Ocr Model With Generative Data Augmentation For Engineering Document Recognition, Wael Khallouli, Mohammad Shahab Uddin, Andres Sousa-Poza, Jiang Li, Samuel Kovacic
Engineering Management & Systems Engineering Faculty Publications
The long-standing practice of document-based engineering has resulted in the accumulation of a large number of engineering documents across various industries. Engineering documents, such as 2D drawings, continue to play a significant role in exchanging information and sharing knowledge across multiple engineering processes. However, these documents are often stored in non-digitized formats, such as paper and portable document format (PDF) files, making automation difficult. As digital engineering transforms processes in many industries, digitizing engineering documents presents a crucial challenge that requires advanced methods. This research addresses the problem of automatically extracting textual content from non-digitized legacy engineering documents. We introduced …
An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor
An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor
Computer Science Faculty Publications
Research in Bengali Natural Language Processing (BNLP) is rapidly expanding. Despite being one of the most widely spoken languages in the world, BNLP research remains insufficient, particularly in Bengali speech recognition. The languages rich morphology, agglutinative structure, and diverse dialects make text and speech processing especially challenging. However, these challenges can be addressed with effective preprocessing techniques. Various organizations in Bangladesh and West Bengal are integrating Natural Language Processing (NLP) into their services, but without a thorough understanding of preprocessing, these implementations remain incomplete. Applying proper preprocessing techniques to the Bengali language will serve as a foundation for developing robust …
A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás
A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás
Computer Science Faculty Publications
P-adic analysis, introduced by Kurt Hensel in the early 20th century, has developed into a fundamental area of mathematical research with broad applications in number theory, algebraic geometry, and mathematical physics. This study aims to examine the thematic evolution and scholarly impact of p-adic research through a comprehensive topic modeling and bibliometric analysis. Using classical bibliometric techniques (e.g., performance analysis, co-authorship, and co-citation networks) combined with Latent Dirichlet Allocation (LDA), we analyzed 7388 peer-reviewed documents published between 1965 and 2024. The computational workflow was conducted using R (version 4.4.1) and VOSviewer (version 1.6.20), which enabled the identification of 20 distinct …
From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer
From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer
Computer Science Faculty Publications
Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …
Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu
Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu
Computer Science Faculty Publications
Science news is increasingly important in connecting scientists and the public by sharing discoveries and innovations. With the rise of large language models (LLMs), there is potential to automate science news creation, but concerns exist about the quality of LLM-generated news versus human-written news. This paper explores whether LLMs can outperform humans in distinguishing between human-written and LLM-generated news. Inspired by the Chain-of-Thought prompting method, we designed a simple yet effective variant called Guided Few-shot (GFS), which encodes the characteristics of news of two types with examples. Our experiments indicated that GFS with just a single example effectively boosted the …
From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer
From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer
Computer Science Faculty Publications
Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …
Jufon Vaikutus Yliopistoissa Työskentelevien Tutkijoiden Julkaisukanavia Koskeviin Päätöksiin, Melina Aarnikoivu, Charles Mathies, Nelli Piattoeva
Jufon Vaikutus Yliopistoissa Työskentelevien Tutkijoiden Julkaisukanavia Koskeviin Päätöksiin, Melina Aarnikoivu, Charles Mathies, Nelli Piattoeva
Educational Leadership & Workforce Development Faculty Publications
Yliopistojen rahoitusmallissa tärkeää roolia näyttelevä Julkaisufoorumi (JUFO) on virallisesti tarkoitettu laajempien julkaisumäärien tarkasteluun. On kuitenkin viitteitä siitä, että JUFOa käytetään laajasti myös yksittäisten tutkijoiden arviointiin. Tässä monimenetelmätutkimuksessa tarkastelemme sitä, miten JUFO-tasot vaikuttavat yksilöiden julkaisukanavia koskeviin päätöksiin, sekä sitä, miten tutkijat kokevat JUFOn roolin omassa työssään sekä tiedejulkaisemisessa laajemmin. Aineistonamme on keväällä 2023 toteutettu verkkokysely (N = 276), johon vastasi eri uravaiheiden ja alojen tutkijoita ympäri Suomea. Analysoimme vastaukset sekä määrällisiä (T- ja ANOVA-testit) sekä laadullisia (induktiivinen sisällönanalyysi) menetelmiä käyttäen. Tulosten perusteella JUFO-tasot vaikuttavat merkittävästi tutkijoiden enemmistön (54,9 %) päätöksiin siitä, millä perusteella julkaisukanava valitaan. Tasoilla oli suurempi merkitys tekniikan ja …
Can Large Language Models Discern Evidence For Scientific Hypotheses? Case Studies In The Social Sciences, Sai Koneru, Jian Wu, Sarah Rajtmajer
Can Large Language Models Discern Evidence For Scientific Hypotheses? Case Studies In The Social Sciences, Sai Koneru, Jian Wu, Sarah Rajtmajer
Computer Science Faculty Publications
Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies …
Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu
Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu
Computer Science Faculty Publications
The iterative character of work in machine learning (ML) and artificial intelligence (AI) and reliance on comparisons against benchmark datasets emphasize the importance of reproducibility in that literature. Yet, resource constraints and inadequate documentation can make running replications particularly challenging. Our work explores the potential of using downstream citation contexts as a signal of reproducibility. We introduce a sentiment analysis framework applied to citation contexts from papers involved in Machine Learning Reproducibility Challenges in order to interpret the positive or negative outcomes of reproduction attempts. Our contributions include training classifiers for reproducibility-related contexts and sentiment analysis, and exploring correlations between …
Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle
Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
Changes made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 …
Building Datasets To Support Information Extraction And Structure Parsing From Electronic Theses And Dissertations, William A. Ingram, Jian Wu, Sampanna Yashwant Kahu, Javaid Akbar Manzoor, Bipasha Banerjee, Aman Ahuja, Muntabir Hasan Choudhury, Lamia Salsabil, Winston Shields, Edward A. Fox
Building Datasets To Support Information Extraction And Structure Parsing From Electronic Theses And Dissertations, William A. Ingram, Jian Wu, Sampanna Yashwant Kahu, Javaid Akbar Manzoor, Bipasha Banerjee, Aman Ahuja, Muntabir Hasan Choudhury, Lamia Salsabil, Winston Shields, Edward A. Fox
Computer Science Faculty Publications
Despite the millions of electronic theses and dissertations (ETDs) publicly available online, digital library services for ETDs have not evolved past simple search and browse at the metadata level. We need better digital library services that allow users to discover and explore the content buried in these long documents. Recent advances in machine learning have shown promising results for decomposing documents into their constituent parts, but these models and techniques require data for training and evaluation. In this article, we present high-quality datasets to train, evaluate, and compare machine learning methods in tasks that are specifically suited to identify and …
The 50 Most Cited Papers On Rugby Since 2000 Reveal A Focus Primarily On Strength And Conditioning In Elite Male Players, Katherine J. Hunzinger, Eric Schussler
The 50 Most Cited Papers On Rugby Since 2000 Reveal A Focus Primarily On Strength And Conditioning In Elite Male Players, Katherine J. Hunzinger, Eric Schussler
Rehabilitation Sciences Faculty Publications
We sought to conduct a bibliometric analysis and review of the most cited publications relating to rugby since 2000 in order to identify topics of interest and those that warrant further investigations. Clarivate Web of Science database was used to perform a literature search using the search term "rugby." The top 200 papers by citation count were extracted and reviewed for the inclusion criteria: all subjects were rugby players. The top 50 manuscripts were included for analysis of author, publication year, country of lead authors, institution, journal name and impact factor, topic, participant sex, and level of rugby. The total …
A Self-Regulating System For Assessing Scientific Predictive Power, Ted C. Rogers
A Self-Regulating System For Assessing Scientific Predictive Power, Ted C. Rogers
Physics Faculty Publications
I propose a method for tracking and assessing scientific progress using a prediction consensus algorithm designed for the purpose. The protocol obviates the need for centralized referees to generate scientific questions, gather predictions, and assess the accuracy or success of those predictions. It relies instead on crowd wisdom and a system of checks and balances for all tasks. It is intended to take the form of a web-based, searchable database. I describe a prototype implementation that I call Ex Quaerum. The main purpose of the present document is to motivate the project, to explain it's underlying philosophy, to explain the …
D-Lib Magazine Pioneered Web-Based Scholarly Communication, Michael L. Nelson, Herbert Van De Sompel
D-Lib Magazine Pioneered Web-Based Scholarly Communication, Michael L. Nelson, Herbert Van De Sompel
Computer Science Faculty Publications
The web began with a vision of, as stated by Tim Berners-Lee in 1991, “that much academic information should be freely available to anyone”. For many years, the development of the web and the development of digital libraries and other scholarly communications infrastructure proceeded in tandem. A milestone occurred in July, 1995, when the first issue of D-Lib Magazine was published as an online, HTML-only, open access magazine, serving as the focal point for the then emerging digital library research community. In 2017 it ceased publication, in part due to the maturity of the community it served as well as …
Extractive Research Slide Generation Using Windowed Labeling Ranking, Athar Sefid, Prasenjit Mitra, Jian Wu, C. Lee Giles
Extractive Research Slide Generation Using Windowed Labeling Ranking, Athar Sefid, Prasenjit Mitra, Jian Wu, C. Lee Giles
Computer Science Faculty Publications
Presentation slides generated from original research papers provide an efficient form to present research innovations. Manually generating presentation slides is labor-intensive. We propose a method to automatically generates slides for scientific articles based on a corpus of 5000 paper-slide pairs compiled from conference proceedings websites. The sentence labeling module of our method is based on SummaRuNNer, a neural sequence model for extractive summarization. Instead of ranking sentences based on semantic similarities in the whole document, our algorithm measures the importance and novelty of sentences by combining semantic and lexical features within a sentence window. Our method outperforms several baseline methods …
Undergraduate Student Perspectives On Textbook Costs And Implications For Academic Success, Lucinda Rush Wittkower, Leo S. Lo
Undergraduate Student Perspectives On Textbook Costs And Implications For Academic Success, Lucinda Rush Wittkower, Leo S. Lo
Libraries Faculty & Staff Publications
To provide more affordable course content to our students and faculty, local data on how students perceive textbook expenses and how the costs impact student success would be necessary in order to advocate to faculty and other stakeholders. This survey, conducted at a mid-sized research public institution, aims to explore student perceptions of textbooks and how these perceptions influence academic success. The results reveal that students feel that the cost of required textbooks is unreasonable and that students are more likely to purchase required textbooks for in-major classes than for elective or general education courses. The most common means of …
Acknowledgement Entity Recognition In Cord-19 Papers, Jian Wu, Pei Wang, Xin Wei, Sarah Rajtmajer, C. Lee Giles, Christopher Griffin
Acknowledgement Entity Recognition In Cord-19 Papers, Jian Wu, Pei Wang, Xin Wei, Sarah Rajtmajer, C. Lee Giles, Christopher Griffin
Computer Science Faculty Publications
Acknowledgements are ubiquitous in scholarly papers. Existing acknowledgement entity recognition methods assume all named entities are acknowledged. Here, we examine the nuances between acknowledged and named entities by analyzing sentence structure. We develop an acknowledgement extraction system, AckExtract based on open-source text mining software and evaluate our method using manually labeled data. AckExtract uses the PDF of a scholarly paper as input and outputs acknowledgement entities. Results show an overall performance of F1=0.92. We built a supplementary database by linking CORD-19 papers with acknowledgement entities extracted by AckExtract including persons and organizations and find that only up to …
Opening Books And The National Corpus Of Graduate Research, William A. Ingram, Edward A. Fox, Jian Wu
Opening Books And The National Corpus Of Graduate Research, William A. Ingram, Edward A. Fox, Jian Wu
Computer Science Faculty Publications
Virginia Tech University Libraries, in collaboration with Virginia Tech Department of Computer Science and Old Dominion University Department of Computer Science, request $505,214 in grant funding for a 3-year project, the goal of which is to bring computational access to book-length documents, demonstrating that with Electronic Theses and Dissertations (ETDs). The project is motivated by the following library and community needs. (1) Despite huge volumes of book-length documents in digital libraries, there is a lack of models offering effective and efficient computational access to these long documents. (2) Nationwide open access services for ETDs generally function at the metadata level. …