Al Jazeera fields in every export
The article record covers canonical URL, headline, summary, section, service, language, text direction, byline, publication timestamp, last modified timestamp, item type and the article identifier used in the URL.
Service and language are separate fields for a reason. Service tells you which newsroom produced the item; language tells you the script the row is in. Keeping them apart makes it possible to ask whether a story was covered by the Arabic service at all, which is a different question from whether an Arabic version of a known story exists.
Entity matching across scripts is delivered rather than left to you. A named organisation, place or person is linked across the two services using a normalised identifier, so a company that appears under three transliterations arrives as one entity with the surface forms recorded. Doing this downstream, after the data has landed in two scripts, is far harder than doing it during collection.
Both timestamps are kept separately, front observations record which stories were promoted on which index and when, and the usual context fields follow: topic tags where published, lead image URL and caption, word count, outbound links, and the collection timestamp on every row.
Text is stored in the original script with direction recorded, never transliterated in place. Transliteration and translation, where wanted, are separate labelled fields.