Skip to main content
SHOW DETAILS
eye
Title
Date Archived
Creator
by Internet Archive Web Group
collection

eye 6,644

This collection contains web crawl data for a random selection of 500k (0.5 million) Crossref DOI redirects, including the doi.org redirect requests. The intent of this crawl is to gather loose statistics on the number of failing redirects, number of host websites that block automated crawling, and a corpus of HTML landing pages for metadata extraction (eg, "signposting" HTTP headers, linked data HTML metadata, semantic markup). Total size of (uncompressed) WARC data is 50 GB,...
Internet Archive Research Publication Crawls
Internet Archive Research Publication Crawls
collection
21,177
ITEMS
112.5M
VIEWS
by Internet Archive Web Group
collection

eye 112.5M

A series of open web crawls targeting journal articles, technical memos, essays, datasets, and other research publications. This collection contains WARC and CDX files that end up in Wayback ( https://web.archive.org ). See also bibliographic metadata corpuses at  https://archive.org/details/ia_biblio_metadata
UNPAYWALL-PDF-CRAWL-2020-11
UNPAYWALL-PDF-CRAWL-2020-11
collection
199
ITEMS
1.8M
VIEWS
by Internet Archive Web Group
collection

eye 1.8M

Web PDF GROBID Corpus (July 2019)
Web PDF GROBID Corpus (July 2019)
collection
10
ITEMS
17
VIEWS
by Internet Archive Web Group
collection

eye 17

DOAJ-CRAWL-2020-11
DOAJ-CRAWL-2020-11
collection
102
ITEMS
959,382
VIEWS
by Internet Archive Web Group
collection

eye 959,382

DOI-LANDING-CRAWL-2018-06
DOI-LANDING-CRAWL-2018-06
collection
279
ITEMS
3.4M
VIEWS
by Internet Archive Web Group
collection

eye 3.4M

arXiv Content Crawl (2019-10)
arXiv Content Crawl (2019-10)
collection
37
ITEMS
79,989
VIEWS
by Internet Archive Web Group
collection

eye 79,989

CORE-UPSTREAM-CRAWL-2018-11
CORE-UPSTREAM-CRAWL-2018-11
collection
741
ITEMS
1.9M
VIEWS
by Internet Archive Web Group
collection

eye 1.9M

Crawl of "upstream" URLs from CORE (core.ac.uk) metadata dump. Only a partial seedlist of files crawled.
SEMSCHOLAR-DIRECT-PDF-CRAWL-2020-02
SEMSCHOLAR-DIRECT-PDF-CRAWL-2020-02
collection
1,011
ITEMS
1.6M
VIEWS
by Internet Archive Web Group
collection

eye 1.6M

UNPAYWALL-PDF-CRAWL-2019-04
UNPAYWALL-PDF-CRAWL-2019-04
collection
641
ITEMS
5.9M
VIEWS
by Internet Archive Web Group
collection

eye 5.9M

UNPAYWALL-PDF-CRAWL-2020-03
UNPAYWALL-PDF-CRAWL-2020-03
collection
344
ITEMS
2M
VIEWS
by Internet Archive Web Group
collection

eye 2M

MAG-PDF-CRAWL-2020-03
MAG-PDF-CRAWL-2020-03
collection
489
ITEMS
4.2M
VIEWS
by Internet Archive Web Group
collection

eye 4.2M

OA-DOI-CRAWL-2020-12
OA-DOI-CRAWL-2020-12
collection
191
ITEMS
1.6M
VIEWS
by Internet Archive Web Group
collection

eye 1.6M

UNPAYWALL-PDF-CRAWL-2020-05
UNPAYWALL-PDF-CRAWL-2020-05
collection
282
ITEMS
1.8M
VIEWS
by Internet Archive Web Group
collection

eye 1.8M

UNPAYWALL-PDF-CRAWL-2021-05
UNPAYWALL-PDF-CRAWL-2021-05
collection
123
ITEMS
973,161
VIEWS
by Internet Archive Web Group
collection

eye 973,161

Wide Web Targeted PDF Crawling (2017)
Wide Web Targeted PDF Crawling (2017)
collection
922
ITEMS
3.2M
VIEWS
by Internet Archive Web Group
collection

eye 3.2M

ARXIV-PUBMEDCENTRAL-CRAWL-2020-04
ARXIV-PUBMEDCENTRAL-CRAWL-2020-04
collection
60
ITEMS
114,355
VIEWS
by Internet Archive Web Group
collection

eye 114,355

OA-DOI-CRAWL-2020-02
OA-DOI-CRAWL-2020-02
collection
278
ITEMS
3.5M
VIEWS
by Internet Archive Web Group
collection

eye 3.5M

OAI-PMH-CRAWL-2020-06
OAI-PMH-CRAWL-2020-06
collection
2,946
ITEMS
5.9M
VIEWS
by Internet Archive Web Group
collection

eye 5.9M

Web PDF GROBID Corpus (June 2019)
Web PDF GROBID Corpus (June 2019)
collection
10
ITEMS
50
VIEWS
by Internet Archive Web Group
collection

eye 50

OA-JOURNAL-CRAWL-2019-08
OA-JOURNAL-CRAWL-2019-08
collection
201
ITEMS
2.9M
VIEWS
by Internet Archive Web Group
collection

eye 2.9M

PubMed Central Crawl (2019-10)
PubMed Central Crawl (2019-10)
collection
216
ITEMS
456,372
VIEWS
by Internet Archive Web Group
collection

eye 456,372

UNPAYWALL-PDF-CRAWL-2018-07
UNPAYWALL-PDF-CRAWL-2018-07
collection
1,241
ITEMS
15.8M
VIEWS
by Internet Archive Web Group
collection

eye 15.8M

Web archive data from a crawl of open access PDF URLs provided by Unpaywall.
DATACITE-DOI-CRAWL-2020-01
DATACITE-DOI-CRAWL-2020-01
collection
1,417
ITEMS
4M
VIEWS
by Internet Archive Web Group
collection

eye 4M

PUBMEDCENTRAL-CRAWL-2020-02
PUBMEDCENTRAL-CRAWL-2020-02
collection
108
ITEMS
262,794
VIEWS
by Internet Archive Web Group
collection

eye 262,794

by Internet Archive Web Group
collection

eye 1,409

This collection holds database snapshots (SQL) and bulk metadata exports (JSON and TSV) from https:///fatcat.wiki (an Internet Archive service)
DIRECT-OA-CRAWL-2019
DIRECT-OA-CRAWL-2019
collection
2,566
ITEMS
5.6M
VIEWS
by Internet Archive Web Group
collection

eye 5.6M

Web PDF Training Sets
Web PDF Training Sets
collection
6
ITEMS
179
VIEWS
by Internet Archive Web Group
collection

eye 179

MAG-PDF-CRAWL-2020-07
MAG-PDF-CRAWL-2020-07
collection
196
ITEMS
1.8M
VIEWS
by Internet Archive Web Group
collection

eye 1.8M

Open Access Journal Test Crawl (2018)
Open Access Journal Test Crawl (2018)
collection
794
ITEMS
11.5M
VIEWS
by Internet Archive Web Group
collection

eye 11.5M

OMICS-DOI-LANDING-CRAWL-2019-04
OMICS-DOI-LANDING-CRAWL-2019-04
collection
4
ITEMS
14,146
VIEWS
by Internet Archive Web Group
collection

eye 14,146

This crawl started in April 2019, as an informal collaboration with Crossref. Crawling a smallish number (100k) DOI redirects and landing pages (plus PDF outlinks, and maybe a couple other hops) for a single large publisher (OMICS, which has multiple subsidiaries). Intent is to get reasonably good capture that can be used as canonical preservation copies of the landing pages. Secondary goal is to get decent fulltext capture coverage.
MSAG-PDF-CRAWL-2017
collection
1,855
ITEMS
13M
VIEWS
by Internet Archive Web Group
collection

eye 13M

Microsoft Academic Graph public corpus (Feb 2016) PDF URLs, filtered to remove large sites (pubmed, citeseerx, arxiv) and already-crawled URLs.
Topics: papers, journals
SCIELO-CRAWL-2020-07
SCIELO-CRAWL-2020-07
collection
41
ITEMS
204,694
VIEWS
by Internet Archive Web Group
collection

eye 204,694

Bulk Bibliographic Metadata
Bulk Bibliographic Metadata
collection
234
ITEMS
17,643
VIEWS
by Internet Archive Web Group
collection

eye 17,643

This collection contains both external ("upstream") metadata dumps and Internet Archive generated databases and reports on our holdings of papers, books, and other documents.
Scholarly TDM Corpora
Scholarly TDM Corpora
collection
44
ITEMS
24
VIEWS
by Internet Archive Web Group
collection

eye 24

Access-restricted text and data-mining corpora. If you are interested in getting access to work with this content, contact info@archive.org
PLATFORM-CRAWL-2020
PLATFORM-CRAWL-2020
collection
649
ITEMS
522,698
VIEWS
by Internet Archive Web Group
collection

eye 522,698

OA-JOURNAL-CRAWL-2020-07
OA-JOURNAL-CRAWL-2020-07
collection
1,923
ITEMS
10.7M
VIEWS
by Internet Archive Web Group
collection

eye 10.7M