Deutsche Welle articles, editions and the ID behind every URL
A Deutsche Welle scraper has to start from what DW is: Germany's international broadcaster, a public broadcaster funded by federal taxes under the Deutsche Welle Act, on air since 1953 and based in Bonn and Berlin. It writes for audiences outside Germany, so dw.com is not a German site with an English corner. It is 31 language editions, written by separate language desks, each under its own path code. /en/, /de/, /es/, /ru/, /uk/ and /ar/ explain themselves, while /sw/ is Kiswahili, /ha/ Hausa, /am/ Amharic, /fa-ir/ Persian, /fa-af/ Dari, /ps/ Pashto, and /pt-br/ and /pt-002/ split Portuguese between Brazil and Africa. There is no paywall and no subscriber tier, so every story renders in full.
Every address ends in a numeric ID with a type prefix: a- for articles, live- for live blogs, video-, audio-, g- for galleries, t- for topic pages, s- for sections and program- for shows. The slug in front of the ID is decoration, since an address with the ID alone lands on the story, but the language path is part of the key: an English ID under /de/ leads nowhere. IDs are unique across the editions with one exception. /zh-hant/ is the Traditional Chinese twin of /zh/ and reuses its IDs, so Chinese stories are deduplicated, not counted twice.
Each edition also has a headlines page with its last seven days of output, an A to Z topic index and a sitemap set of its own, and most editions publish RSS feeds. Together they turn the site into a Deutsche Welle database keyed on one number.
Get a Quote