Der Spiegel fields in every export
The article record covers canonical URL, headline, standfirst, section and subsection, byline, publication timestamp, last modified timestamp, language, edition, item type and the article identifier used in Spiegel URLs.
The paywall flag is a first class field. Every row records whether the article was behind SPIEGEL Plus at the time of collection and how much text rendered publicly, so nobody downstream mistakes a two paragraph opening for a full article. Silent truncation is the single most common way a German monitoring dataset misleads people.
Language and edition are recorded separately because the German and English sides are different publications rather than translations. A story existing in one and not the other is itself a finding, and the schema has to make that visible instead of collapsing both into one row.
Front observations run alongside the articles: which stories sat on the front or a section front, in what order and when. Spiegel moves things around through the day, and prominence is a much better proxy for reach than a mention count.
Then the context fields: topic tags where published, lead image URL and caption, word count of the publicly rendered portion, outbound links, and the collection timestamp on every row.