Guardian fields in every export
The article record covers canonical URL, headline, standfirst, section and its path, the full list of keyword tags, byline and contributor, publication timestamp, last modified timestamp, edition, and the article identifier used across Guardian URLs.
Tags are the field that repays the most attention here. Guardian articles carry several, mixing subject, place, organisation and format, and they are published rather than guessed at. That makes filtering precise: everything tagged to a company, everything tagged to a competition, everything tagged both to a country and to a policy area. We keep them as a list rather than flattening them into a string.
Front observations sit alongside the articles. Each pass over a section front records which stories were on it, in what order, under which headline at that moment, and when the pass ran. That is where promotion and rewriting become visible, and it is the part no API returns.
Live blogs are collected as entries: entry text, entry timestamp, author where credited, and any embedded quote, image or link, all keyed to the parent blog. The parent keeps an entry count and the timestamp of its most recent entry.
Then the usual context: language, word count of the publicly rendered body, lead image URL and caption, outbound links, item type for video and audio pieces, and the collection timestamp on every row.