Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (40)
- Physical Sciences and Mathematics (40)
- Databases and Information Systems (14)
- Arts and Humanities (9)
- Engineering (4)
-
- Information Security (4)
- American Studies (3)
- Cataloging and Metadata (3)
- Computer Engineering (3)
- Feminist, Gender, and Sexuality Studies (3)
- American Literature (2)
- Communication (2)
- Digital Humanities (2)
- Lesbian, Gay, Bisexual, and Transgender Studies (2)
- American Popular Culture (1)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (1)
- Applied Ethics (1)
- Artificial Intelligence and Robotics (1)
- Data Science (1)
- Digital Communications and Networking (1)
- Education (1)
- Educational Assessment, Evaluation, and Research (1)
- English Language and Literature (1)
- Geography (1)
- Human Geography (1)
- Indigenous Studies (1)
- Internet Law (1)
- Keyword
-
- Web archives (36)
- Digital preservation (17)
- Web archiving (16)
- Internet archives (11)
- Archive-It (7)
-
- Digital libraries (7)
- Internet Archive (7)
- Memento (7)
- Mementos (7)
- URls (7)
- Search engines (6)
- Metadata (5)
- Archives (4)
- JavaScript (4)
- Javascript (4)
- Social media (4)
- TimeMaps (4)
- Wayback Machine (4)
- Web archive (4)
- Heritrix (3)
- Storytelling (3)
- Analysis (2)
- Analysis tools (2)
- Archival records (2)
- Archives & records (2)
- Big data (2)
- Collections (2)
- Computer science (2)
- Curated web collections (2)
- Data sets (2)
- Publication Year
- Publication
- Publication Type
- File Type
Articles 1 - 30 of 71
Full-Text Articles in Archival Science
Cluster Conversation: (Re)Writing Our Histories, (Re)Building Feminist Worlds: Working Toward Hope In The Archives, Ruth Osorio, Lamaya Williams, Megan Mcintyre
Cluster Conversation: (Re)Writing Our Histories, (Re)Building Feminist Worlds: Working Toward Hope In The Archives, Ruth Osorio, Lamaya Williams, Megan Mcintyre
English Faculty Publications
[Introduction] "Hope is not like a lottery ticket you can sit on the sofa and clutch, feeling lucky. [...] Hope is an ax you break down doors with in an emergency." - Rebecca Solnit
In 2018, Cheryl Glenn wrote, "The work of feminist rhetorical historiography is far from done; in fact, it has just begun-and it is anchored in hope." Following Glenn, we explore hope in this cluster as a methodological imperative in the archives. Informed by theorists Paulo Freire, bell hooks, Rebecca Solnit, and Cornel West, the writers in this Cluster Conversation envision hope as a radical orientation toward …
Not Here, Go There: Analyzing Redirection Patterns On The Web, Kritika Garg, Sawood Alam, Dietrich Ayala, Michele C. Weigle, Michael L. Nelson
Not Here, Go There: Analyzing Redirection Patterns On The Web, Kritika Garg, Sawood Alam, Dietrich Ayala, Michele C. Weigle, Michael L. Nelson
Computer Science Faculty Publications
URI redirections are integral to web management, supporting structural changes, SEO optimization, and security. However, their complexities affect usability, SEO performance, and digital preservation. This study analyzed 11 million unique redirecting URIs, following redirections up to 10 hops per URI, to uncover patterns and implications of redirection practices. Our findings revealed that 50% of the URIs terminated successfully, while 50% resulted in errors, including 0.06% exceeding 10 hops. Canonical redirects, such as HTTP to HTTPS transitions, were prevalent, reflecting adherence to SEO best practices. Non-canonical redirects, often involving domain or path changes, highlighted significant web migrations, rebranding, and security risks. …
Github Repository Complexity Leads To Diminished Web Archive Availability, David Calano, Michael Nelson, Michele Weigle
Github Repository Complexity Leads To Diminished Web Archive Availability, David Calano, Michael Nelson, Michele Weigle
Computer Science Faculty Publications
Software is often developed using versioned controlled software, such as Git, and hosted on centralized Web hosts, such as GitHub and GitLab. These Web hosted software repositories are made available to users in the form of traditional HTML Web pages for each source file and directory, as well as a presentational home page and various descriptive pages. We examined more than 12,000 Web hosted Git repository project home pages, primarily from GitHub, to measure how well their presentational components are preserved in the Internet Archive, as well as the source trees of the collected GitHub repositories to assess the extent …
Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle
Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle
Computer Science Faculty Publications
In this case study, we use web archives to analyze 8,824 webpages that were taken offline and subsequently put back online, thus experiencing a “near death experience.” We enumerate the stages of a webpage’s near death experience, including the change from a successful HTTP status code to non-successful and back, the intermediate stage with markers such as an under construction banner, and an analysis of how the pages came back differently.
Surfacing Text Changes In Archived Webpages, Lesley Frew
Surfacing Text Changes In Archived Webpages, Lesley Frew
Computer Science Theses & Dissertations
Webpages change over time, and web archives hold copies of historical versions of webpages. Users of web archives, such as journalists, want to find and view changes on webpages over time. However, the current search interfaces for web archives do not adequately support this task. For the web archives that include a full-text search feature, multiple versions of the same webpage that match the search query are shown individually without enumerating changes, or are grouped together in a way that hides changes. We present a change text search engine that allows users to find changes in webpages. We describe the …
Learner Assessment, Elizabeth Burns
Learner Assessment, Elizabeth Burns
STEMPS Faculty Publications
The article discusses the importance of learner assessment in school libraries, highlighting the need for school librarians to incorporate assessment practices aligned with the National School Library Standards. Assessment in school libraries focuses on measuring competencies aligned with real-world information-seeking behaviors rather than traditional grades and testing. The article emphasizes the role of diagnostic, formative, and summative assessments in tracking learner progress and supporting the overall library program. It also underscores the significance of using learner data to establish library goals and showcase the impact of school libraries on academic achievement.
Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle
Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
Changes made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 …
Robots Still Outnumber Humans In Web Archives In 2019, But Less Than In 2015 And 2012, Himarsha R. Jayanetti, Kritika Garg, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Robots Still Outnumber Humans In Web Archives In 2019, But Less Than In 2015 And 2012, Himarsha R. Jayanetti, Kritika Garg, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
The significance of the web and the crucial role of web archives in its preservation highlight the necessity of understanding how users, both human and robot, access web archive content, and how best to satisfy this disparate needs of both types of users. To identify robots and humans in web archives and analyze their respective access patterns, we used the Internet Archive’s (IA) Wayback Machine access logs from 2012, 2015, and 2019, as well as Arquivo.pt’s (Portuguese Web Archive) access logs from 2019. We identified user sessions in the access logs and classified those sessions as human or robot based …
Assessing The Prevalence And Archival Rate Of Uris To Git Hosting Platforms In Scholarly Publications, Emily Escamilla
Assessing The Prevalence And Archival Rate Of Uris To Git Hosting Platforms In Scholarly Publications, Emily Escamilla
Computer Science Theses & Dissertations
The definition of scholarly content has expanded to include the data and source code that contribute to a publication. While major archiving efforts to preserve conventional scholarly content, typically in PDFs (e.g., LOCKSS, CLOCKSS, Portico), are underway, no analogous effort has yet emerged to preserve the data and code referenced in those PDFs, particularly the scholarly code hosted online on Git Hosting Platforms (GHPs). Similarly, Software Heritage is working to archive public source code, but there is value in archiving the surrounding ephemera that provide important context to the code while maintaining their original URIs. In current implementations, source code …
Supporting Account-Based Queries For Archived Instagram Posts, Himarsha R. Jayanetti
Supporting Account-Based Queries For Archived Instagram Posts, Himarsha R. Jayanetti
Computer Science Theses & Dissertations
Social media has become one of the primary modes of communication in recent times, with popular platforms such as Facebook, Twitter, and Instagram leading the way. Despite its popularity, Instagram has not received as much attention in academic research compared to Facebook and Twitter, and its significant role in contemporary society is often overlooked. Web archives are making efforts to preserve social media content despite the challenges posed by the dynamic nature of these sites. The goal of our research is to facilitate the easy discovery of archived copies, or mementos, of all posts belonging to a specific Instagram account …
Hashes Are Not Suitable To Verify Fixity Of The Public Archived Web, Mohamed Aturban, Martin Klein, Herbert Van De Sompel, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Hashes Are Not Suitable To Verify Fixity Of The Public Archived Web, Mohamed Aturban, Martin Klein, Herbert Van De Sompel, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
Web archives, such as the Internet Archive, preserve the web and allow access to prior states of web pages. We implicitly trust their versions of archived pages, but as their role moves from preserving curios of the past to facilitating present day adjudication, we are concerned with verifying the fixity of archived web pages, or mementos, to ensure they have always remained unaltered. A widely used technique in digital preservation to verify the fixity of an archived resource is to periodically compute a cryptographic hash value on a resource and then compare it with a previous hash value. If the …
The Dsa Toolkit Shines Light Into Dark And Stormy Archives, Shawn Morgan Jones, Himarsha R. Jayanetti, Alex Osborne, Paul Koerbin, Klein Martin, Michele C. Weigle, Michael L. Nelson
The Dsa Toolkit Shines Light Into Dark And Stormy Archives, Shawn Morgan Jones, Himarsha R. Jayanetti, Alex Osborne, Paul Koerbin, Klein Martin, Michele C. Weigle, Michael L. Nelson
Computer Science Faculty Publications
Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web archive collections are often configured to revisit the same original resource multiple times. This is incredibly useful for understanding an unfolding news story or the evolution of an organization. Unfortunately, over time, some of these original resources can go off-topic and no longer suit the purpose for which the collection was originally created. They can go off-topic due to web site …
This Old Vase: Ancient Art And Primary Source Instruction In The Archives, Laraann Canner
This Old Vase: Ancient Art And Primary Source Instruction In The Archives, Laraann Canner
Libraries Faculty & Staff Publications
No abstract provided.
Information Activism: A Queer History Of Lesbian Media Technologies, Dawn Betts-Green
Information Activism: A Queer History Of Lesbian Media Technologies, Dawn Betts-Green
STEMPS Faculty Publications
No abstract provided.
A Crisis Of Erasure: Transgender And Gender-Nonconforming Populations Navigating Breast Cancer Health Information, Curtis Shane Tenney, Karl J. Surkan, Lynette Hammond Gerido, Dawn Betts-Green
A Crisis Of Erasure: Transgender And Gender-Nonconforming Populations Navigating Breast Cancer Health Information, Curtis Shane Tenney, Karl J. Surkan, Lynette Hammond Gerido, Dawn Betts-Green
STEMPS Faculty Publications
In this paper, we use the topic of breast cancer as an example of health crisis erasure in both informational and institutional contexts, particularly within the transgender and gender-nonconforming population. Breast cancer health information conforms and defaults to conventional cultural associations with femininity, as is the case with pregnancy and other “single-sex” conditions (Surkan, 2015). Many health information and research practices normalize sexualities, pathologize non-normative gender (Drescher et al., 2012; Fish, 2008; Müller, 2018), and fail to recognize gender-nonconforming categories (Frohard‐Dourlent et al., 2017). Because breast cancer health information is sexually normalized, an information boundary exists for the LGBTQ+ community, …
Automatic Metadata Extraction Incorporating Visual Features From Scanned Electronic Theses And Dissertations, Muntabir Hasan Choudhury, Himarsha R. Jayanetti, Jian Wu, William A. Ingram, Edward A. Fox
Automatic Metadata Extraction Incorporating Visual Features From Scanned Electronic Theses And Dissertations, Muntabir Hasan Choudhury, Himarsha R. Jayanetti, Jian Wu, William A. Ingram, Edward A. Fox
Computer Science Faculty Publications
Electronic Theses and Dissertations (ETDs) contain domain knowledge that can be used for many digital library tasks, such as analyzing citation networks and predicting research trends. Automatic metadata extraction is important to build scalable digital library search engines. Most existing methods are designed for born-digital documents, so they often fail to extract metadata from scanned documents such as ETDs. Traditional sequence tagging methods mainly rely on text-based features. In this paper, we propose a conditional random field (CRF) model that combines text-based and visual features. To verify the robustness of our model, we extended an existing corpus and created a …
Afterlives Of Indigenous Archives: Essays In Honor Of "The Occom Circle" [Book Review], Drew Lopenzina
Afterlives Of Indigenous Archives: Essays In Honor Of "The Occom Circle" [Book Review], Drew Lopenzina
English Faculty Publications
(First paragraph) Afterlives of Indigenous Archives takes its title from Anishinaabe author Gerald Vizenor who is, in turn, repurposing a quote from French theorist Jacques Derrida who, in his 1995 work, Archive Fever, referred to the archive as that which gestures toward “an excess of life,” something that “resists annihilation” (183). This excess, or “afterlife,” of the archive remains, for Vizenor at least, an unexpected location of Indigenous survivance—a site from which, despite every violent attempt to colonially contain and collapse Native presence, it is still possible to carry something forward from the ruins of representation. With this in mind, …
Legal And Technical Issues For Text And Data Mining In Greece, Maria Kanellopoulou - Botti, Marinos Papadopoulos, Christos Zampakolas, Paraskevi Ganatsiou
Legal And Technical Issues For Text And Data Mining In Greece, Maria Kanellopoulou - Botti, Marinos Papadopoulos, Christos Zampakolas, Paraskevi Ganatsiou
Computer Ethics - Philosophical Enquiry (CEPE) Proceedings
Web harvesting and archiving pertains to the processes of collecting from the web and archiving of works that reside on the Web. Web harvesting and archiving is one of the most attractive applications for libraries which plan ahead for their future operation. When works retrieved from the Web are turned into archived and documented material to be found in a library, the amount of works that can be found in said library can be far greater than the number of works harvested from the Web. The proposed participation in the 2019 CEPE Conference aims at presenting certain issues related to …
Shakespeare's Globe Archive: Theatres, Players & Performance, Rob Tench
Shakespeare's Globe Archive: Theatres, Players & Performance, Rob Tench
Libraries Faculty & Staff Publications
No abstract provided.
Web Archives At The Nexus Of Good Fakes And Flawed Originals, Michael L. Nelson
Web Archives At The Nexus Of Good Fakes And Flawed Originals, Michael L. Nelson
Computer Science Faculty Publications
[Summary] The authenticity, integrity, and provenance of resources we encounter on the web are increasingly in question. While many people are inured to the possibility of altered images, the easy accessibility of powerful software tools that synthesize audio and video will unleash a torrent of convincing “deepfakes” into our social discourse. Archives will no longer be monopolized by a countable number of institutions such as governments and publishers, but will become a competitive space filled with social engineers, propagandists, conspiracy theorists, and aspiring Hollywood directors. While the historical record has never been singular nor unmalleable, current technologies empower an unprecedented …
Subjectivity And Methodology In The Arch'i'Ve, Elizabeth J. Vincelette
Subjectivity And Methodology In The Arch'i'Ve, Elizabeth J. Vincelette
English Faculty Publications
This article explores methodologies from the fields of library archival science, human geography, composition and rhetoric, and established editorial practices in English studies. By elaborating on the role of a researcher’s subjectivity in archival creation, this work expands the conversation regarding methodology and archives, especially how archives present us with new ways of seeing and making narratives during the editorial decision-making involved in their creation. Writing about my own experience, I privilege the researcher’s point of view with a narrative about my construction of a digital archive. With archival research, we should promote the revelation of methods and methodology to …
Off Topic Memento Toolkit, Shawn M. Jones, Michele C. Weigle, Michael L. Nelson
Off Topic Memento Toolkit, Shawn M. Jones, Michele C. Weigle, Michael L. Nelson
Computer Science Faculty Publications
Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web archive collections are often configured to revisit the same original resource multiple times. This is incredibly useful for understanding an unfolding news story or the evolution of an organization. Unfortunately, over time, some of these original resources can go off-topic and no longer suit the purpose for which the collection was originally created. They can go off-topic due to web site …
205.3 The Many Shapes Of Archive-It, Shawn Jones, Michael L. Nelson, Alexander Nwala, Michele C. Weigle
205.3 The Many Shapes Of Archive-It, Shawn Jones, Michael L. Nelson, Alexander Nwala, Michele C. Weigle
Computer Science Faculty Publications
Web archives, a key area of digital preservation, meet the needs of journalists, social scientists, historians, and government organizations. The use cases for these groups often require that they guide the archiving process themselves, selecting their own original resources, or seeds, and creating their own web archive collections. We focus on the collections within Archive-It, a subscription service started by the Internet Archive in 2005 for the purpose of allowing organizations to create their own collections of archived web pages, or mementos. Understanding these collections could be done via their user-supplied metadata or via text analysis, but the metadata is …
Swimming In A Sea Of Javascript Or: How I Learned To Stop Worrying And Love High-Fidelity Replay, John A. Berlin, Michael L. Nelson, Michele C. Weigle
Swimming In A Sea Of Javascript Or: How I Learned To Stop Worrying And Love High-Fidelity Replay, John A. Berlin, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
[First paragraph] Preserving and replaying modern web pages in high-fidelity has become an increasingly difficult task due to the increased usage of JavaScript. Reliance on server-side rewriting alone results in live-leakage and or the inability to replay a page due to the preserved JavaScript performing an action not permissible from the archive. The current state-of-the-art high fidelity archival preservation and replay solutions rely on handcrafted client-side URL rewriting libraries specifically tailored for the archive, namely Webrecoder's and Pywb's wombat.js [12]. Web archives not utilizing client-side rewriting rely on server-side rewriting that misses URLs used in a manner not accounted for …
Client-Assisted Memento Aggregation Using The Prefer Header, Mat Kelly, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Client-Assisted Memento Aggregation Using The Prefer Header, Mat Kelly, Sawood Alam, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
[First paragraph] Preservation of the Web ensures that future generations have a picture of how the web was. Web archives like Internet Archive's Wayback Machine, WebCite, and archive.is allow individuals to submit URIs to be archived, but the captures they preserve then reside at the archives. Traversing these captures in time as preserved by multiple archive sources (using Memento [8]) provides a more comprehensive picture of the past Web than relying on a single archive. Some content on the Web, such as content behind authentication, may be unsuitable or inaccessible for preservation by these organizations. Furthermore, this content may be …
It Is Hard To Compute Fixity On Archived Web Pages, Mohamed Aturban, Michael L. Nelson, Michele C. Weigle
It Is Hard To Compute Fixity On Archived Web Pages, Mohamed Aturban, Michael L. Nelson, Michele C. Weigle
Computer Science Faculty Publications
[Introduction] Checking fixity in web archives is performed to ensure archived resources, or mementos (denoted by URI-M) have remained unaltered since when they were captured. The final report of the PREMIS Working Group [2] defines information used for fixity as "information used to verify whether an object has been altered in an undocumented or unauthorized way." The common technique for checking fixity is to generate a current hash value (i.e., a message digest or a checksum) for a file using a cryptographic hash function (e.g., SHA-256) and compare it to the hash value generated originally. If they have different hash …
Avoiding Zombies In Archival Replay Using Serviceworker, Sawood Alam, Mat Kelly, Michele C. Weigle, Michael L. Nelson
Avoiding Zombies In Archival Replay Using Serviceworker, Sawood Alam, Mat Kelly, Michele C. Weigle, Michael L. Nelson
Computer Science Faculty Publications
[First paragraph] A Composite Memento is an archived representation of a web page with all the page requisites such as images and stylesheets. All embedded resources have their own URIs, hence, they are archived independently. For a meaningful archival replay, it is important to load all the page requisites from the archive within the temporal neighborhood of the base HTML page. To achieve this goal, archival replay systems try to rewrite all the resource references to appropriate archived versions before serving HTML, CSS, or JS. However, an effective server-side URL rewriting is difficult when URLs are generated dynamically using JavaScript. …
Using Web Archives To Enrich The Live Web Experience Through Storytelling, Yasmin Alnoamany
Using Web Archives To Enrich The Live Web Experience Through Storytelling, Yasmin Alnoamany
Computer Science Theses & Dissertations
Much of our cultural discourse occurs primarily on the Web. Thus, Web preservation is a fundamental precondition for multiple disciplines. Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources. Understanding the contents and boundaries of these archived collections is a challenge for most people, resulting in the paradox of the larger the collection, the harder it is to understand. Meanwhile, as the sheer volume of data grows on the Web, "storytelling" is becoming a popular …
Combining Heritrix And Phantomjs For Better Crawling Of Pages With Javascript, Justin F. Brunelle, Michele C. Weigle, Michael L. Nelson
Combining Heritrix And Phantomjs For Better Crawling Of Pages With Javascript, Justin F. Brunelle, Michele C. Weigle, Michael L. Nelson
Computer Science Presentations
PDF of a powerpoint presentation from the International Internet Preservation Consortium (IIPC) 2016 Conference in Reykjavik, Iceland, April 11, 2016. Also available on Slideshare.
Storytelling For Summarizing Collections In Web Archives, Yasmin Alnoamany, Michele C. Weigle, Michael L. Nelson
Storytelling For Summarizing Collections In Web Archives, Yasmin Alnoamany, Michele C. Weigle, Michael L. Nelson
Computer Science Presentations
PDF of a powerpoint presentation from the Coalition for Networked Information (CNI) Spring 2016 Membership Meeting in San Antonio, Texas, April 5, 2016. Also available on Slideshare.