arXiv fields in every export
The paper record covers arXiv identifier, version number, title, abstract, primary category, cross listed categories, submission date per version, latest version date, comments field, journal reference where the author has supplied one, and DOI where a published version exists.
Authors come as rows with surname, given name, position in the list and affiliation as supplied. Affiliation on preprints is inconsistently populated and inconsistently written, so it is normalised to institutions with the raw string preserved, exactly as on our literature work elsewhere.
Version history is a first class part of this dataset rather than an afterthought. Each version carries its own date and, where the author supplied one, its own comment describing what changed. A paper that has been revised four times over two years is telling a different story from one posted once, and abstracts do change between versions in ways that matter to anyone doing text analysis.
Cross listing is preserved as a list rather than reduced to the primary category. Interdisciplinary work is precisely what most research trend analysis is looking for, and it lives in the cross listings.
Where a published journal version is later linked, that link is captured, along with the retrieval timestamp on every row.