Skip to main content

37
UPLOADS


More right-solid

Show sorted alphabetically

Show sorted alphabetically

More right-solid
SHOW DETAILS
eye
Title
Date Archived
Creator
Scholarly TDM Corpora
Scholarly TDM Corpora
collection
44
ITEMS
22
VIEWS
by Internet Archive Web Group
collection

eye 22

Access-restricted text and data-mining corpora. If you are interested in getting access to work with this content, contact info@archive.org
PLATFORM-CRAWL-2020
PLATFORM-CRAWL-2020
collection
649
ITEMS
386,390
VIEWS
by Internet Archive Web Group
collection

eye 386,390

OA-JOURNAL-CRAWL-2020-07
OA-JOURNAL-CRAWL-2020-07
collection
1,923
ITEMS
9.5M
VIEWS
by Internet Archive Web Group
collection

eye 9.5M

UNPAYWALL-PDF-CRAWL-2020-11
UNPAYWALL-PDF-CRAWL-2020-11
collection
199
ITEMS
1.6M
VIEWS
by Internet Archive Web Group
collection

eye 1.6M

by Internet Archive Web Group
collection

eye 6,476

This collection contains web crawl data for a random selection of 500k (0.5 million) Crossref DOI redirects, including the doi.org redirect requests. The intent of this crawl is to gather loose statistics on the number of failing redirects, number of host websites that block automated crawling, and a corpus of HTML landing pages for metadata extraction (eg, "signposting" HTTP headers, linked data HTML metadata, semantic markup). Total size of (uncompressed) WARC data is 50 GB,...
Internet Archive Research Publication Crawls
Internet Archive Research Publication Crawls
collection
21,054
ITEMS
100.1M
VIEWS
by Internet Archive Web Group
collection

eye 100.1M

A series of open web crawls targeting journal articles, technical memos, essays, datasets, and other research publications. This collection contains WARC and CDX files that end up in Wayback ( https://web.archive.org ). See also bibliographic metadata corpuses at  https://archive.org/details/ia_biblio_metadata
Wide Web Targeted PDF Crawling (2017)
Wide Web Targeted PDF Crawling (2017)
collection
922
ITEMS
3M
VIEWS
by Internet Archive Web Group
collection

eye 3M

OA-DOI-CRAWL-2020-12
OA-DOI-CRAWL-2020-12
collection
191
ITEMS
1.4M
VIEWS
by Internet Archive Web Group
collection

eye 1.4M

MAG-PDF-CRAWL-2020-03
MAG-PDF-CRAWL-2020-03
collection
489
ITEMS
3.6M
VIEWS
by Internet Archive Web Group
collection

eye 3.6M

UNPAYWALL-PDF-CRAWL-2020-05
UNPAYWALL-PDF-CRAWL-2020-05
collection
282
ITEMS
1.6M
VIEWS
by Internet Archive Web Group
collection

eye 1.6M

UNPAYWALL-PDF-CRAWL-2021-05
UNPAYWALL-PDF-CRAWL-2021-05
collection
123
ITEMS
847,305
VIEWS
by Internet Archive Web Group
collection

eye 847,305

MSAG-PDF-CRAWL-2017
collection
1,855
ITEMS
11.6M
VIEWS
by Internet Archive Web Group
collection

eye 11.6M

Microsoft Academic Graph public corpus (Feb 2016) PDF URLs, filtered to remove large sites (pubmed, citeseerx, arxiv) and already-crawled URLs.
Topics: papers, journals
MAG-PDF-CRAWL-2020-07
MAG-PDF-CRAWL-2020-07
collection
196
ITEMS
1.6M
VIEWS
by Internet Archive Web Group
collection

eye 1.6M

SCIELO-CRAWL-2020-07
SCIELO-CRAWL-2020-07
collection
41
ITEMS
188,303
VIEWS
by Internet Archive Web Group
collection

eye 188,303

OMICS-DOI-LANDING-CRAWL-2019-04
OMICS-DOI-LANDING-CRAWL-2019-04
collection
4
ITEMS
13,827
VIEWS
by Internet Archive Web Group
collection

eye 13,827

This crawl started in April 2019, as an informal collaboration with Crossref. Crawling a smallish number (100k) DOI redirects and landing pages (plus PDF outlinks, and maybe a couple other hops) for a single large publisher (OMICS, which has multiple subsidiaries). Intent is to get reasonably good capture that can be used as canonical preservation copies of the landing pages. Secondary goal is to get decent fulltext capture coverage.
Open Access Journal Test Crawl (2018)
Open Access Journal Test Crawl (2018)
collection
794
ITEMS
10.6M
VIEWS
by Internet Archive Web Group
collection

eye 10.6M

Bulk Bibliographic Metadata
Bulk Bibliographic Metadata
collection
223
ITEMS
17,205
VIEWS
by Internet Archive Web Group
collection

eye 17,205

This collection contains both external ("upstream") metadata dumps and Internet Archive generated databases and reports on our holdings of papers, books, and other documents.
Web PDF GROBID Corpus (July 2019)
Web PDF GROBID Corpus (July 2019)
collection
10
ITEMS
17
VIEWS
by Internet Archive Web Group
collection

eye 17

arXiv Content Crawl (2019-10)
arXiv Content Crawl (2019-10)
collection
37
ITEMS
64,880
VIEWS
by Internet Archive Web Group
collection

eye 64,880

DOI-LANDING-CRAWL-2018-06
DOI-LANDING-CRAWL-2018-06
collection
279
ITEMS
3.2M
VIEWS
by Internet Archive Web Group
collection

eye 3.2M

CORE-UPSTREAM-CRAWL-2018-11
CORE-UPSTREAM-CRAWL-2018-11
collection
741
ITEMS
1.5M
VIEWS
by Internet Archive Web Group
collection

eye 1.5M

Crawl of "upstream" URLs from CORE (core.ac.uk) metadata dump. Only a partial seedlist of files crawled.
DOAJ-CRAWL-2020-11
DOAJ-CRAWL-2020-11
collection
102
ITEMS
858,478
VIEWS
by Internet Archive Web Group
collection

eye 858,478

SEMSCHOLAR-DIRECT-PDF-CRAWL-2020-02
SEMSCHOLAR-DIRECT-PDF-CRAWL-2020-02
collection
1,011
ITEMS
1.4M
VIEWS
by Internet Archive Web Group
collection

eye 1.4M

UNPAYWALL-PDF-CRAWL-2019-04
UNPAYWALL-PDF-CRAWL-2019-04
collection
641
ITEMS
5.2M
VIEWS
by Internet Archive Web Group
collection

eye 5.2M

UNPAYWALL-PDF-CRAWL-2020-03
UNPAYWALL-PDF-CRAWL-2020-03
collection
344
ITEMS
1.8M
VIEWS
by Internet Archive Web Group
collection

eye 1.8M

Web PDF GROBID Corpus (June 2019)
Web PDF GROBID Corpus (June 2019)
collection
10
ITEMS
49
VIEWS
by Internet Archive Web Group
collection

eye 49

OAI-PMH-CRAWL-2020-06
OAI-PMH-CRAWL-2020-06
collection
2,946
ITEMS
4.8M
VIEWS
by Internet Archive Web Group
collection

eye 4.8M

OA-DOI-CRAWL-2020-02
OA-DOI-CRAWL-2020-02
collection
278
ITEMS
3.2M
VIEWS
by Internet Archive Web Group
collection

eye 3.2M

ARXIV-PUBMEDCENTRAL-CRAWL-2020-04
ARXIV-PUBMEDCENTRAL-CRAWL-2020-04
collection
60
ITEMS
101,993
VIEWS
by Internet Archive Web Group
collection

eye 101,993

by Internet Archive Web Group
collection

eye 1,358

This collection holds database snapshots (SQL) and bulk metadata exports (JSON and TSV) from https:///fatcat.wiki (an Internet Archive service)
Web PDF Training Sets
Web PDF Training Sets
collection
6
ITEMS
178
VIEWS
by Internet Archive Web Group
collection

eye 178

DIRECT-OA-CRAWL-2019
DIRECT-OA-CRAWL-2019
collection
2,566
ITEMS
5.1M
VIEWS
by Internet Archive Web Group
collection

eye 5.1M

PUBMEDCENTRAL-CRAWL-2020-02
PUBMEDCENTRAL-CRAWL-2020-02
collection
108
ITEMS
234,015
VIEWS
by Internet Archive Web Group
collection

eye 234,015

UNPAYWALL-PDF-CRAWL-2018-07
UNPAYWALL-PDF-CRAWL-2018-07
collection
1,241
ITEMS
14.4M
VIEWS
by Internet Archive Web Group
collection

eye 14.4M

Web archive data from a crawl of open access PDF URLs provided by Unpaywall.
PubMed Central Crawl (2019-10)
PubMed Central Crawl (2019-10)
collection
216
ITEMS
403,156
VIEWS
by Internet Archive Web Group
collection

eye 403,156

OA-JOURNAL-CRAWL-2019-08
OA-JOURNAL-CRAWL-2019-08
collection
201
ITEMS
2.7M
VIEWS
by Internet Archive Web Group
collection

eye 2.7M

DATACITE-DOI-CRAWL-2020-01
DATACITE-DOI-CRAWL-2020-01
collection
1,417
ITEMS
3.6M
VIEWS
by Internet Archive Web Group
collection

eye 3.6M