RTVE Play Scraper for Spanish Broadcast Archive Data

Most streaming archives go back a few years. A public broadcaster's goes back decades, and that depth is the only thing a commercial service cannot buy.

RTVE Play Scraper
Solutions

Managed archive metadata, run end to end by us

ScrapeIt runs the collection as a managed service. You name the periods, genres or programme types; we build the pipeline with broadcast date precision recorded, content types separated and sparse older entries kept rather than filtered, and hand back CSV, JSON, Excel or a push into your warehouse.

Where availability drift matters we run repeat collection and keep every observation, because material that became unavailable leaves no note saying so.

We collect published metadata only - never streams, never anything behind a sign in - honour the crawl rules and pace requests. Public funding does not imply free reuse, so terms are recorded rather than assumed and your counsel should see the use case before the project starts.

RTVE Play fields in every export

Programme records carry the title, content type, genre, description, original broadcast date where published, duration, season and episode structure, channel and language.

Original broadcast date is the field that makes the archive usable. A catalogue entry without it is a title floating free of the period it belongs to, and for archive research the period is the question. Where the site publishes both a broadcast date and a publication date we keep them apart.

Content type distinguishes series, news, documentary and radio material, because they have different structures and mixing them produces aggregates that describe none of them.

Language is recorded where published, since Spain has co-official languages and the broadcaster produces in more than one.

Every row carries the collection timestamp and the source address.

RTVE Play fields in every export
Archive research, availability drift and licensing

Archive research, availability drift and licensing

Archive research is the strongest use: how much of which decade remains available, by genre and channel, which for media historians and policy researchers is a concrete question about public memory that this catalogue can answer.

Availability drift over time - which archive material becomes unavailable and which returns - is only visible with repeat collection and is a direct read on how rights affect a public archive.

Programme structure analysis across decades shows how formats, episode lengths and season structures changed, which is straightforward once duration and structure are collected consistently.

On licensing: a public broadcaster's material is copyrighted and its reuse terms are its own. Public funding does not imply free reuse, and we record terms rather than assuming. The dataset is for analysis, and anything republishing content needs a conversation with the broadcaster.

A public archive that happens to stream

RTVE Play is the on demand service of Spain's public broadcaster, carrying current programming, news, documentaries, drama and a substantial archive of past broadcast material.

The archive is what distinguishes it. Commercial services carry what they licensed recently; a public broadcaster carries what it made, and in this case that runs back decades. For research into Spanish media, television history or public service broadcasting, that depth is the asset and there is no commercial substitute for it.

The catalogue also spans several content types that behave differently: series with episode structure, news programmes tied to dates, documentaries, and radio material from the same organisation.

Catalogue sections respond directly with substantial content. The crawl rules run to several hundred lines and are mostly a named blocklist of link checkers and harvesting tools; the general rules concern media files and site infrastructure rather than editorial content, and the file publishes a set of sitemaps.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why archive depth needs different handling

A pipeline built for a commercial streaming catalogue will work here and will waste most of what the source offers.

Commercial catalogue logic assumes a few thousand recent titles with consistent metadata. An archive spanning decades has inconsistent metadata by nature: older material was catalogued under different conventions, broadcast dates are recorded with varying precision, and some entries carry rich description while others carry a title and a year. Treating that unevenness as dirty data and cleaning it away destroys the information that the entry is old.

The better approach is to record precision explicitly - an exact broadcast date and a year-only date are different facts and should be distinguishable - and to keep the sparse entries rather than filtering them out for failing a completeness check.

The second reason is that archive material changes availability. Rights on older programming lapse and return, so what is streamable this year may not be next, and repeat collection turns that into a record of what a public broadcaster can actually keep available.

The third is that this is one national broadcaster. It is excellent for Spain and says nothing about the wider Spanish-language world, where the production centres are elsewhere.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your archive feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, handles the metadata unevenness that decades of cataloguing conventions produce, and repairs the collector before an availability record develops a gap.

You see a sample first, in your format, over the periods you actually research, with date precision recorded so you can see on real entries what a cleaned-up pipeline would have thrown away.

FAQ

Why keep entries with sparse metadata?

Because sparseness is information: older material was catalogued under different conventions, and an entry with only a title and a year is telling you it is old. Filtering it for failing a completeness check removes exactly the part of the archive that research is about.

How do you handle imprecise broadcast dates?

By recording the precision explicitly. An exact date and a year-only date are different facts, and a dataset that stores both as a date silently invents precision that was never there.

Does the archive stay available?

Not entirely. Rights on older programming lapse and return, so availability drifts. Repeat collection turns that into a record of what a public broadcaster can actually keep accessible, which is itself a research finding.

Does this cover Spanish-language media generally?

No. It is one national broadcaster and it is excellent for Spain. The wider Spanish-language production world sits elsewhere, and we scope briefs on that basis rather than letting one archive stand in for a language.

Is public funding the same as free reuse?

No, and it is a common assumption worth correcting. The material is copyrighted and reuse terms are the broadcaster's own. We record what applies rather than assuming, and republication needs a conversation with them.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582