Why source context is a column, not a caveat
Regional media analysis goes wrong most often at the aggregation step, and the cause is always the same: outlets treated as interchangeable observers.
They are not. Ownership, country of publication and editorial orientation shape what gets covered and how, and an aggregate sentiment or coverage figure computed across a mixed set of outlets without those fields is really a measurement of which outlets happened to be in the sample. Recording the context makes the aggregate decomposable - you can see whether a finding holds across outlet types or is driven by one.
The second reason is the English-language question, the same one that applies to any English outlet in a non-English market. This is a selection for an international and regional English readership, not a sample of Arabic-language media, and a complete regional picture needs Arabic sources that are a separate and harder project.
The third is multi-country coverage. A Gulf-based outlet reporting across several countries produces articles whose subject country differs from the publication country, and keeping those apart is what stops regional coverage being attributed to the wrong market.
The fourth is transliteration. Arabic names appear in many romanised forms, and matching across outlets - let alone to Arabic-language sources - requires that variation to be preserved rather than normalised away at collection.