Deutsche Welle Scraper for Articles, Video and Audio in 31 Editions

Deutsche Welle publishes 31 language editions written by separate desks, and it rewrites headlines after publication. We collect every edition and keep every version of the story.

Deutsche Welle Scraper
Solutions

How we scrape Deutsche Welle and run your own Deutsche Welle API

We scrape Deutsche Welle as a managed service. Discovery runs on the cheap surfaces - Deutsche Welle RSS feeds, the news sitemap and each edition's headlines page - and the fields come from the story pages, which carry the full record. Backfill walks the topic pages. Rows are deduplicated on the numeric ID, feed links lose their maca tracking parameter, and every headline version is stored with its timestamp. You get CSV, JSON, XLSX or an endpoint we host: a Deutsche Welle API of your own, with field names agreed up front. Anti-bot handling, proxy rotation and CAPTCHA solving are part of the service; paths the robots file closes, such as search and the site's graph endpoint, stay closed, and nothing needs a login. The price of a Deutsche Welle feed depends on editions, cadence and backfill depth.

What a Deutsche Welle dataset holds: four headlines, three dates

One DW story is several records in one, and to extract Deutsche Welle data cleanly the export keeps them apart. These are the fields we collect for each article:

  • Identity - numeric ID, content type, edition path, language, canonical URL and the p.dw.com short link.
  • Four headline fields - the display headline, the original name that stays in the slug and the RSS item, a separate SEO title for the browser tab and the social card title, which usually repeats the original name; plus teaser and short teaser. They drift apart: a straight news headline in the slug can sit under a question-style headline written hours later.
  • Three dates - first publication, last publication and last modification. The date DW shows on the page and in the feed follows the last republish, so first publication is kept as a field of its own.
  • Taxonomy - category IDs shared across editions (Politics carries the same ID in the German, Russian, Arabic and Kiswahili editions), the localized thematic label, internal topic and region labels written in German whatever the edition, ISO country codes and links to topic pages.
  • Keywords - editorial tags that appear in the news sitemap only while a story sits in its window of about two days; after that, topic links are the lasting tags.
  • Byline - the author credit as printed and the agency line next to it, such as "with dpa, Reuters" or "with AFP, dpa". We keep the credit on the story and skip author profile pages, portraits and social handles.
  • Flags - genre (news or dossier), opinion marker and an editorial priority of high or neutral.
  • Body and media - full text with its inline links to topics and earlier stories, word count, lead image ID with caption and agency photo credit, attached video and audio, related stories.

Deutsche Welle live blogs add a list of posts, each with its own ID, title, timestamp and text. Stories from the 2000s carry a free-text credit line instead of a structured byline, no category or country code and a single date; those fields stay empty rather than guessed.

What a Deutsche Welle dataset holds: four headlines, three dates
Deutsche Welle video, podcasts, TV schedule and DW Learn German

Deutsche Welle video, podcasts, TV schedule and DW Learn German

The same ID system covers DW's media, and each type brings fields of its own.

  • Deutsche Welle video - title, teaser, duration, programme (DW News, Focus on Europe, Tomorrow Today, The 77 Percent and dozens more), dates, thumbnail and stream address. We record whether a video has subtitles and in which language; the subtitle files sit on a path the robots file closes, so transcripts stay out.
  • Audio - episode title, host, programme, duration and MP3 link. Deutsche Welle podcast programmes also have feeds of their own with keywords per episode, and the long-running ones reach back nearly a decade.
  • Galleries - image titles, captions, alt text and photo credits.
  • TV schedule - the English, Spanish and Arabic live channel pages list the coming slots: start, end, programme, episode title and teaser.
  • DW Learn German - run in 19 interface languages, it grades news by level: Kurz und leicht video news at A2, Top-Thema mit Vokabeln at B1 and Langsam gesprochene Nachrichten at B2. The Deutsche Welle B1 articles carry the full manuscript, an MP3 and a glossary of headwords with gender, plural and a plain German definition, which makes these Deutsche Welle German articles a leveled corpus rather than news for native readers. Exercise and test pages are closed to crawlers and stay out.

Deutsche Welle articles, editions and the ID behind every URL

A Deutsche Welle scraper has to start from what DW is: Germany's international broadcaster, a public broadcaster funded by federal taxes under the Deutsche Welle Act, on air since 1953 and based in Bonn and Berlin. It writes for audiences outside Germany, so dw.com is not a German site with an English corner. It is 31 language editions, written by separate language desks, each under its own path code. /en/, /de/, /es/, /ru/, /uk/ and /ar/ explain themselves, while /sw/ is Kiswahili, /ha/ Hausa, /am/ Amharic, /fa-ir/ Persian, /fa-af/ Dari, /ps/ Pashto, and /pt-br/ and /pt-002/ split Portuguese between Brazil and Africa. There is no paywall and no subscriber tier, so every story renders in full.

Every address ends in a numeric ID with a type prefix: a- for articles, live- for live blogs, video-, audio-, g- for galleries, t- for topic pages, s- for sections and program- for shows. The slug in front of the ID is decoration, since an address with the ID alone lands on the story, but the language path is part of the key: an English ID under /de/ leads nowhere. IDs are unique across the editions with one exception. /zh-hant/ is the Traditional Chinese twin of /zh/ and reuses its IDs, so Chinese stories are deduplicated, not counted twice.

Each edition also has a headlines page with its last seven days of output, an A to Z topic index and a sitemap set of its own, and most editions publish RSS feeds. Together they turn the site into a Deutsche Welle database keyed on one number.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why Deutsche Welle headlines are worth tracking edition by edition

Deutsche Welle data scraping pays off because DW is Germany talking to the world, and each desk decides what its own audience hears: one broadcaster, one legal mandate, some thirty agendas. The German edition is far from the largest. Over the three months to late September 2026 the article sitemaps listed about 2,950 articles and live blogs in Spanish, 2,770 in Russian, 2,230 in Ukrainian, 1,880 in English and 1,040 in German.

  • Regional coverage in regional languages - the Persian, Dari and Pashto desks on Iran and Afghanistan, DW Hausa and DW Kiswahili on West and East Africa, the Russian and Ukrainian desks on the war. For media monitoring in those markets, what the German public broadcaster tells local audiences in their own language is the signal, not a translation of an English wire.
  • Agenda comparison - shared category IDs and ISO country codes let you count what each desk covers without translating a word.
  • Headline changes - DW rewrites and republishes, and stored versions show which stories were reframed, when, and in which editions.
  • Research corpora - full, unpaywalled news text in Amharic, Hausa, Kiswahili, Pashto, Dari, Bengali and Urdu is scarce, and DW's own rules allow text and data mining for non-commercial research.

Nothing is cut short by a paywall, so word counts and models compare fairly across editions.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Buying Deutsche Welle data from ScrapeIt

ScrapeIt is a managed web scraping agency: we build, run and repair the collectors, and you receive the data. A Deutsche Welle project usually opens with a sample - one week of three editions you choose, with headline versions, category IDs and bylines - so you can test the schema against your use before a schedule is agreed. Files land in S3, GCS, SFTP, a webhook or your warehouse, and we keep the pipeline working as DW changes its pages.

FAQ

Does Deutsche Welle have an API?

Not a public one. DW does not publish a developer API: there is no documentation, key sign-up or developer terms. The website loads its data from an internal graph endpoint that the robots file closes to crawlers, and the JSON interface behind the DW apps is undocumented. What DW does publish for machines is RSS and Atom feeds for most editions and their main sections, podcast feeds with MP3 files, and news, video and topic sitemaps. Its partner channels are about rebroadcasting, not data: AudioDepot gives radio stations free MP3 programmes, DW Transtel sells TV programmes, and reuse beyond the free-use rules goes through a licence request.

How far back does the Deutsche Welle archive go?

Well beyond the sitemaps, which cover about three months. Topic pages go back fifty stories per page through each topic's full history, and the English Bundesliga topic reaches November 2001. Site search is closed to crawlers, so a backfill is built from topic pages, the A to Z index and the links between stories, and addresses built on the numeric ID stay stable. Older records are thinner: stories from the 2000s carry a free-text credit instead of a structured byline, no category or country code and a single date, and we deliver those fields empty rather than inferred.

Is the Deutsche Welle RSS feed enough on its own?

No, but it is the cheapest trigger. DW publishes its feeds in RDF, RSS 2.0 and Atom; the English list runs from all stories and top stories to Germany, World, Europe, Business, Sports and more. An item carries the ID, headline, teaser, category and a date rounded to the minute that follows the last republish. There is no byline, no tags and no body text, and the all-stories feed is a selection: in September 2026 it carried roughly one English article in five. We use the feeds to spot new IDs and take the fields from the story pages.

Can you scrape Deutsche Welle video and audio, or only text?

Both, as metadata and links. For video we collect title, teaser, duration, programme, dates, thumbnail and stream address, and note whether subtitles exist; for audio the episode, host, programme, duration and MP3 link, plus the podcast feeds. We do not download or re-host the media: these are DW's copyrighted broadcasts, and the subtitle files sit on a path closed to crawlers. If you need the footage itself, DW Transtel and AudioDepot are the licensed routes, and we say so at scoping.

Is it legal to scrape Deutsche Welle, and what may we do with the data?

DW's own rules draw the line. Its free-use page allows text and data mining for academic or research purposes when the use is non-commercial, only the necessary data is copied and nothing is republished or altered. The robots file adds a notice that automated collection needs DW's prior written permission, and commercial reuse goes through DW's licensing request form. We work only on public pages, keep to the paths the robots file leaves open and pace requests. Whether your use fits the free-use rules or needs a licence is a question for your counsel before the project starts.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582