Onet fields in every export
The article record covers canonical URL, headline, standfirst, vertical and section, byline where published, publication timestamp, last updated timestamp, item type and the article identifier.
Vertical is a separate field from section because on a portal they are different things. A story in the business vertical and a story in the business section of the news vertical are not the same object, and a client comparing coverage across Polish media needs to know which they are holding.
Text is stored in Polish with its diacritics intact, never transliterated in place. Polish diacritics are routinely stripped in URLs and slugs but present in body text, and normalising them away at collection time destroys the ability to match accurately later.
Brand and product matches are delivered with the surface form that was found, not just a flag. When a name appears in an inflected form, the row records which form, which is what lets a client audit the matching rather than trust it.
Then the usual context: topic tags where published, word count, lead image URL, front position from repeated observation, outbound links, and the collection timestamp on every row.