CNN fields in every export
The item record covers canonical URL, headline, summary line, section and subsection, edition, byline where one is published, publication timestamp, last updated timestamp and the item identifier used in CNN URLs.
Item type is a first class field and it is the one that saves the analysis. Article, video package, gallery, live page and interactive are recorded explicitly, so a video item with forty words of caption is never mistaken for a thin article. Where a video is the item, its duration is captured instead of a word count.
Front observations run alongside: which items sat on the front or a section front, in what order and at what time. On a broadcaster this matters more than usual, because promotion turns over faster than on a newspaper.
Live pages are collected as a parent record plus their stream of timestamped entries, each with its own text and author where credited. Collapsing a live page into one article throws away the reporting inside it.
Then the context fields: topic tags where published, lead image URL and caption, outbound links, publicly rendered body text and its word count for text items, and the collection timestamp on every row.