ARD Mediathek Scraper for a Federated Broadcast Catalog

One platform, nine regional broadcasters behind it. Drop the producer field and a federation looks like a single national channel.

ARD Mediathek Scraper
Solutions

Managed public broadcast metadata, run end to end by us

ScrapeIt runs the collection as a managed service. You name the producers, genres or the full catalogue; we enumerate from the published sitemaps, record the producing broadcaster and resolve availability windows to absolute dates, and hand back CSV, JSON, Excel or a push into your warehouse.

Where departures matter we run repeat collection so arrivals and expiries become records rather than a difference somebody has to reconstruct.

We collect published metadata within the general crawl rules and honour the reservation of rights: this material is not used as model training data, by us or on a client's behalf. Content is copyrighted and terms restrict reuse, so the dataset is for analysis rather than republication. Your counsel should see the use case before the project starts.

ARD Mediathek fields in every export

Programme records carry the title, series and episode identifiers, description, genre, duration, the producing broadcaster, first broadcast date, availability window and the address.

Producing broadcaster is a first class field. It is the federation's defining structure and it is what allows questions about regional commissioning, regional content share and how the institutions differ from each other.

Availability window is recorded with its expiry, because German public broadcasting operates time limits on how long material stays online and those limits are part of what the catalogue is.

Discovery source is recorded per row - which sitemap an address came from - so coverage can be audited rather than assumed complete.

Every row carries the collection timestamp, which is what makes an expiry date interpretable.

ARD Mediathek fields in every export
Regional commissioning, expiry tracking and limits

Regional commissioning, expiry tracking and limits

Regional commissioning analysis is the distinctive output: how much each institution contributes, in which genres, and how that has changed. It is a question about the structure of German public broadcasting and it needs the producer field to be answerable at all.

Expiry tracking produces the departure feed that every time-limited catalogue supports: what is leaving, when, by producer and genre. It requires windows resolved to absolute dates at collection.

Coverage auditing from the sitemap source field tells you honestly how much of the catalogue a collection actually reached, rather than presenting whatever was found as complete.

Limits: metadata only, never streams or anything behind a sign in, no viewer data, and no use of this material as model training data. Content is copyrighted and public funding does not imply free reuse.

A federation presented as one service

ARD Mediathek is the joint on demand platform of Germany's public broadcasting federation. What looks like a single service is a shared front end over the catalogues of the regional broadcasting institutions that make up the federation, each with its own commissioning and regional remit.

That structure is the thing to preserve. A programme made by one regional broadcaster for its own region is a different object from a federation-wide commission, and a catalogue that records only the platform loses the distinction entirely - which for any question about German public broadcasting is the distinction that matters.

Discovery works through published sitemaps rather than through a browsable index. The detail and video sitemaps are referenced in the crawl rules and the addresses they list respond directly, which makes enumeration straightforward once you read the rules rather than trying to crawl the interface.

The crawl rules also carry an explicit reservation of rights headed Nutzungsvorbehalt, directed at AI training and naming the major model crawlers individually.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the rights reservation needs reading precisely

The crawl rules here contain a reservation of rights, and reading it carelessly in either direction leads somewhere wrong.

It is headed Nutzungsvorbehalt - reservation of use - and its content is an AI training opt-out. It names model training crawlers individually and disallows them from the entire site. What it does not do is prohibit crawling generally: the rules for ordinary crawlers remain narrow, covering technical paths rather than editorial content.

So the honest reading has two parts. Collecting catalogue metadata for analysis is not what the reservation addresses, and we do it within the general rules. Using this material as training data for a model is exactly what the reservation addresses, and we will not do it - for ourselves or for a client who tells us that is the purpose.

That distinction is worth stating because it cuts both ways. Treating the reservation as a total prohibition would write off a legitimate source; treating it as decorative would ignore a clearly expressed position from a public broadcaster. Neither is acceptable.

Beyond that, the analytical value is the federated structure and the availability windows - the same time-limited catalogue problem that public broadcasters everywhere present, with an extra dimension because nine institutions feed the same shelf.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your ARD feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, follows the published sitemaps as the catalogue changes, and repairs the collector before an expiry record develops a gap that cannot be filled.

You see a sample first, in your format, over the producers and genres you actually track, with producer and expiry populated so you can judge the structure on real programmes.

FAQ

The site has a rights reservation - can you collect it at all?

The reservation is an AI training opt-out naming model crawlers individually. It does not prohibit crawling generally, and the rules for ordinary crawlers are narrow and technical. We collect catalogue metadata within those rules and do not use the material as training data.

What if my purpose is training a model?

Then the reservation applies to your use and we will not do it. That is the case the broadcaster has explicitly addressed, and taking the job anyway would mean ignoring a clearly expressed position from a public institution.

Why record which broadcaster produced a programme?

Because this is a federation, not a single channel. Regional commissioning and regional content share are the questions people actually have about German public broadcasting, and without the producer field none of them can be asked.

How do you find everything if there is no index?

Through the sitemaps the site publishes in its crawl rules - the documented route. We also record which sitemap each address came from, so coverage can be audited rather than presented as complete because nothing obvious was missing.

Does material expire?

Yes - German public broadcasting operates time limits on how long material stays online. We resolve those windows to absolute dates at collection, which is what makes a departure feed possible.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582