Der Spiegel Scraper for German News Monitoring

An English-only monitoring setup is blind in the largest media market in Europe. Spiegel is where that gap is most expensive.

Der Spiegel Scraper
Solutions

Managed Der Spiegel scraping, run end to end by us

ScrapeIt runs the Spiegel collector as a managed service. You name the sections, subjects and editions; we build the crawl, run it on your cadence and hand back CSV, JSON, Excel or a push into your warehouse. German language matching, including compound and umlaut variants of your brand and product names, is part of the setup rather than something you configure afterwards.

Cadence is per source. The headline overview justifies frequent passes; individual sections and archive work run slower. Given the daily volume, a considered schedule is also what keeps the cost sensible.

We collect what the site renders publicly to an ordinary visitor and flag SPIEGEL Plus rather than working around it. We do not use subscriber credentials. Spiegel journalism is copyrighted and their terms restrict automated collection and reuse, so the dataset is for analysis rather than republication. Bring the use case to your own counsel before the project starts.

Der Spiegel fields in every export

The article record covers canonical URL, headline, standfirst, section and subsection, byline, publication timestamp, last modified timestamp, language, edition, item type and the article identifier used in Spiegel URLs.

The paywall flag is a first class field. Every row records whether the article was behind SPIEGEL Plus at the time of collection and how much text rendered publicly, so nobody downstream mistakes a two paragraph opening for a full article. Silent truncation is the single most common way a German monitoring dataset misleads people.

Language and edition are recorded separately because the German and English sides are different publications rather than translations. A story existing in one and not the other is itself a finding, and the schema has to make that visible instead of collapsing both into one row.

Front observations run alongside the articles: which stories sat on the front or a section front, in what order and when. Spiegel moves things around through the day, and prominence is a much better proxy for reach than a mention count.

Then the context fields: topic tags where published, lead image URL and caption, word count of the publicly rendered portion, outbound links, and the collection timestamp on every row.

Der Spiegel fields in every export
The English edition, headline overview and archive

The English edition, headline overview and archive

The English international edition is worth collecting alongside the German side, not instead of it. It tells you what Spiegel considered significant enough to put in front of an international audience, which is an editorial judgement in its own right. Rows carry the edition, so the two can be compared or analysed separately without recollecting anything.

The headline overview page is the efficient discovery surface. It lists recent output in publication order across sections, which means a project can stay current with a light schedule instead of crawling the whole site. We read it frequently, take the news sitemap as a cross check, and reserve deeper section passes for the desks the client actually cares about.

Historical backfill is practical for recent years. Article URLs are stable and section and topic pages paginate, so a retrospective on a company or a subject is a normal one time job separate from the ongoing feed. The paywall flag applies to historical rows exactly as it does to current ones, so an archive pull is honest about what it contains.

Translation, where a client wants it, is a separate labelled field. We never overwrite the German source text with a translation, because the original is what any later verification has to be done against.

How Spiegel is structured, and where the paywall sits

Spiegel publishes an enormous amount in German and a smaller, curated English edition under its international section. The two are not translations of each other: the English edition is a selection, so a story that matters in Germany may have no English counterpart at all. Any monitoring built only on the English side will simply not see most of the coverage.

Access is mixed. A large part of the site is free, and a subscriber tier marked SPIEGEL Plus sits over the rest. On a paid article the headline, the summary and an opening portion render publicly and the remainder does not. That boundary defines what a collector can honestly deliver, and it is a boundary we do not cross.

Structurally the site is arranged into sections with a dedicated headline overview page that lists recent output in publication order, which is a far cheaper discovery surface than crawling every section. Alongside it a German news sitemap covers recently published items, and the site documents its own RSS feeds.

Volume is the practical constraint. The front alone renders tens of thousands of words, and the daily output is large enough that a naive crawl becomes expensive quickly. Sensible projects here run off the headline overview and the sitemap, and reserve full section passes for the desks that actually matter to the client.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why German coverage is where English-only monitoring fails

The failure is quiet, which is what makes it expensive. A dashboard configured with English sources reports a clean number, and the number is wrong because most of the coverage in the largest media market in Europe never appears in it. Nothing in the tool signals that the German press is missing; there is simply less coverage than there really was.

Spiegel is the clearest single case. Its German output is many times the English selection, and the stories that carry weight in German politics, industry and business are frequently German only. For anyone with operations, customers or regulatory exposure in Germany, missing that is not a rounding error.

Doing it properly means handling German text as German. Compound nouns defeat naive keyword matching, and umlauts get transliterated in URLs and slugs but not in body text. A brand or product name has to be matched with those variations in mind or the recall drops without anyone noticing. We build the matching around that rather than bolting a filter onto an English pipeline.

The second reason is the paid tier. A collector that does not flag SPIEGEL Plus quietly delivers openings as if they were articles, and every word count, sentiment score and model trained on that data inherits the error. Flagging it costs nothing and keeps the dataset honest.

The third is prominence. What Spiegel leads with tends to set the German agenda for the day, so front position collected over time is a genuinely predictive signal rather than a vanity metric.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your Der Spiegel feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, watches it as the site changes, and repairs it before your dashboard goes quiet.

You see a sample first, in your format, over the sections and editions you actually track, with the paywall flag visible so you can judge the real coverage before committing.

FAQ

Can you collect articles behind SPIEGEL Plus?

No. We do not use subscriber accounts and we do not work around the meter. What we collect is what an ordinary visitor sees: headline, standfirst, byline, section, both timestamps, front position and the opening text that renders publicly, with the article marked as paid. If you need full bodies, the route is a licence from the publisher and we will tell you that rather than deliver openings as if they were articles.

Is the English edition the same content as the German site?

No, and assuming it is will cost you most of the coverage. The English international edition is a curated selection, not a translation of the German output, and many stories exist only in German. We collect both, tag every row with its edition and language, and treat a story existing in one but not the other as a finding rather than a gap.

How do you handle German compound words and umlauts?

Deliberately, because naive matching fails on both. German compounds hide a brand name inside a longer word, and umlauts appear transliterated in URLs and slugs but not in body text. We build the matching for your specific brand and product names with those variants included, rather than running an English keyword filter over German text and quietly losing half the hits.

The site is huge. How do you keep the collection affordable?

By using the cheap discovery surfaces first. The headline overview lists recent output in publication order and the news sitemap covers what is new, so staying current does not require crawling the site. Full section passes are reserved for the desks that matter to you, and the cadence is set per source rather than uniformly.

Is scraping Der Spiegel legal?

Spiegel journalism is copyrighted and their terms restrict automated collection and reuse, so this dataset is for analysis rather than republication. We work only on public pages, never with subscriber credentials, we honour the crawl rules the site publishes and we pace requests. Media monitoring on published metadata and publicly rendered text is common practice, but the use case is yours: have your own counsel approve it before the project starts.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582