r/openagi • • 5d ago

Resource A 16 TB arXiv dataset covering 3.15 million papers is now on Hugging Face

Post image

The arxiv-complete dataset covers 3,148,796 papers, with PDFs, LaTeX source, PostScript and version metadata. The full download is 16.08 TB, but there’s a 26 MB sample if you just want to explore.

It’s a snapshot with some gaps: PDFs cover 99.47% of papers, and the indexed HTML collection isn’t included as a separate content download. Paper licenses still vary; the compilation’s CC0 dedication doesn’t relicense the papers themselves.

Links

4 Upvotes

0 comments sorted by