PubMed fields in every export
The citation record covers PubMed identifier, DOI where present, article title, abstract, journal title and abbreviation, ISSN, volume, issue, pagination, publication date and the electronic publication date where they differ, publication type, and language.
Authors come as rows rather than a string. Each carries surname, given name, initials, position in the author list, corresponding author flag where marked, and the affiliation as published. Author position matters: first and last author carry very different meaning in biomedical publishing, and a flattened author field throws that away.
Subject indexing is delivered structured. MeSH descriptors and qualifiers are kept as separate terms with their major topic flags intact, which is what makes precise selection possible. Filtering a corpus by a controlled vocabulary term is a categorically better operation than matching words in a title, and this source is one of the few where that vocabulary exists.
Identifiers that connect outward are extracted deliberately: trial registration numbers, grant and funding identifiers, and full text links including PubMed Central identifiers where the article is open. Those are the fields that let a literature dataset join to a trial registry, to a funder database or to the text itself.
Affiliations are normalised to institutions with the raw string preserved, and every row carries the retrieval timestamp.