ZDF Scraper for Catalog Metadata and Availability

The crawl rules name scraping frameworks by tool, not just AI bots. The attached rules are narrow, and the signal about intent is not.

ZDF Scraper
Solutions

Managed broadcast metadata, run end to end by us

ScrapeIt runs the collection as a managed service. You name the genres, programme types or the full catalogue; we enumerate through the published sitemaps, resolve availability windows to absolute dates, type news separately, and hand back CSV, JSON, Excel or a push into your warehouse.

We pace conservatively on this source by deliberate choice rather than because a number was published, and we scope briefs to what is needed rather than sweeping the catalogue.

We collect published metadata only - never streams, never anything behind a sign in, no viewer data - and we do not use the material as model training data. Content is copyrighted, so the dataset is for analysis rather than republication, and your counsel should see the use case before the project starts.

ZDF fields in every export

Programme records carry the title, series and episode identifiers, description, genre, duration, first broadcast date, availability window with its expiry, and the address.

Expiry is resolved to an absolute date at collection. Where a page states a relative window, storing the relative form guarantees a meaningless row later, so the resolution happens while the context is present.

Content type separates drama, documentary, news and sport, since they follow completely different availability rules - news material in particular has much shorter windows than drama.

Series structure is preserved as a hierarchy so that completeness and ordering questions remain answerable.

Every row carries the collection timestamp and the sitemap or section it was discovered through, so coverage can be audited.

ZDF fields in every export
Departure feeds, structure comparison and limits

Departure feeds, structure comparison and limits

Departure feeds are the standard output of any time-limited catalogue: what is leaving in the next week or month, by genre and programme type, built from resolved expiry dates.

Structure comparison against the federated platform is the distinctive use here - two German public broadcasters with different institutional shapes, catalogued on one schema, which makes the effect of structure on output visible.

News volume analysis is worth separating out entirely. Short windows and high turnover make it a different dataset that swamps the rest if merged, and on its own it is a usable record of a public broadcaster's news output.

Limits: metadata only, never streams or anything behind a sign in, no viewer data. Content is copyrighted and public funding does not imply free reuse. Where a client's purpose is model training we treat the named-agent list as the broadcaster's position and decline.

A single national broadcaster with a time-limited shelf

ZDF is one of Germany's national public broadcasters, and its media library carries drama, documentaries, news, sport and a substantial current affairs output.

Unlike the federated platform run by the other side of German public broadcasting, this is a single institution with one commissioning structure, which makes its catalogue a cleaner object to analyse and a useful comparison against a federation.

Like all German public broadcasting it operates availability windows: material stays online for a defined period and then goes. The catalogue is therefore a set of time-limited rights rather than a permanent library, and a snapshot without expiry dates records only what exists today.

The crawl rules deserve a careful read. There is no general rule block at all; instead there is a named list of user agents - AI model crawlers alongside scraping frameworks and SEO tools - and the rules attached to that list are narrow, covering specific sports highlight paths. Sitemaps are published for discovery.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why a narrow rule with a broad signal deserves care

The crawl rules here are an unusual document and they are worth reading as communication rather than only as a technical constraint.

Technically, no rule applies to an ordinary crawler - there is no general block - and even the named agents are only restricted from a handful of sports highlight addresses. On a literal reading, almost everything is permitted.

What the file also does is name scraping frameworks explicitly, alongside AI crawlers and SEO tools. That is a public broadcaster saying something about how it views automated collection, even though the rules it attached are narrow. We take the literal permission and the expressed attitude together: we collect published metadata, we pace conservatively, we stay well inside anything that could look like load, and we would rather scope a brief down than test where the tolerance ends.

The analytical case for using the source at all is the availability windows and the comparison value. A single national broadcaster's catalogue set beside a federated one answers questions about how institutional structure shapes commissioning that neither answers alone.

The last point is that news material behaves differently from everything else here, with short windows and high volume, and a pipeline that treats it like drama will produce a catalogue dominated by items that vanish within days.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your ZDF feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, follows the published sitemaps, and repairs the collector before a departure feed misses the week that mattered.

You see a sample first, in your format, over the genres you actually track, with expiries resolved and news typed separately so you can judge the structure on real programmes.

FAQ

The rules name scraping tools - is collection allowed?

Literally, yes: there is no general rule block and the named agents are restricted only from a few sports highlight paths. But naming scraping frameworks says something about how the broadcaster views automated collection, so we pace conservatively and scope briefs down rather than testing the tolerance.

Why resolve availability windows at collection?

Because a relative window stored as written becomes meaningless the moment it is saved. The resolution to an absolute date has to happen while the context is present, and it is what makes a departure feed possible afterwards.

Why type news separately?

Because it has much shorter windows and far higher volume than anything else. Merged in, it dominates the catalogue with items that vanish within days; separated, it is a usable record of a broadcaster's news output.

What is worth comparing this against?

The federated German platform. Two public broadcasters with different institutional shapes, on one schema, makes visible how structure affects commissioning and availability - a question neither catalogue answers alone.

Can this be used as AI training data?

Not through us. The named-agent list includes model crawlers, and we treat that as the broadcaster's position on the question regardless of how narrow the attached rules are.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582