Ars Technica fields in every export
The article record covers canonical URL, headline, standfirst, section, author, publication timestamp, last updated timestamp, item type and the article identifier.
Body text is collected in full where publicly rendered, with word count, because on a long-form source the length is a meaningful attribute rather than a side effect. Multi page articles are reassembled into one record instead of arriving as fragments.
Comment threads are collected with structure intact: comment identifier, parent, author handle, timestamp, depth and text. Flattening the thread destroys the argument, and on this source the argument is a substantial part of what a client is buying.
Entity extraction runs over both the article and the thread. Products, vendors and technologies mentioned are normalised to identifiers, because the mention that matters is frequently in a reply comparing your product with an alternative rather than in the headline.
Commenter handles are treated as identifiers rather than people: no attempt to link them to real identities and no cross site profiling, and we say so before anybody asks.