NPR fields in every export
The story record covers canonical URL, headline, teaser, section and programme where the item came from a show, byline, publication timestamp, last updated timestamp, item type and the story identifier.
Item type separates written story, broadcast segment with audio, segment with transcript, and the combinations, because those are genuinely different objects. A three minute segment with a fifty word web summary is not a thin article, and only typing prevents it being counted as one.
Audio fields are captured where published: duration, the programme it aired on and the air date, which can differ from the web publication timestamp and frequently does.
Transcripts are collected where NPR publishes them, and they are the substantive text for broadcast items. Word count for those rows comes from the transcript rather than the summary, which is the only way length based measures mean anything on this source.
Source is recorded per row: national or which member station, so regional coverage can be separated from national without recollecting anything.