BBC News fields in every export
The article record is the spine: canonical URL, headline, the summary line where one is published, section path, publisher edition, language and the article identifier the BBC uses in its own URLs. That identifier is what lets you follow one story across a rewrite rather than treating each version as a new item.
Timestamps come in a pair and both matter. Published is when the story first appeared, updated is when it was last touched. A story with six hours between them is not the same event as a story with six seconds, and any measure of how fast the BBC moved on something needs both. We keep them separately and never overwrite one with the other.
Attribution is thinner on the BBC than on a newspaper: many stories carry a desk rather than a named reporter, and correspondents are often credited in the body instead of a byline field. We record what is published, mark the rest as unattributed and do not guess.
The rest is context. Section and topic tags, the position a story held on its section front and when that position was observed, the lead image URL and its caption, word count of the publicly rendered body, outbound links, and whether the page is an article, a live page or a video item. Live pages also carry the entry count and the timestamp of the last entry, because a live page that stopped updating three hours ago is a different thing from one that is still running.
Every row carries the collection timestamp. On a site that rewrites in place, a field without a time attached is not a fact about anything.