ITVX Scraper for Catalogue, Episode and TV Schedule Data

ITVX prints the leaving date in the record itself. Every episode carries an availability window, so catalogue churn on a broadcaster's streaming service is readable without diffing snapshots.

ITVX Scraper
Solutions

From scope to delivery on an ITVX scrape

You set the scope - which categories, whether the schedule ships with the catalogue, and whether you want a single snapshot or a repeating run that turns into a change log of arrivals, tier moves and removals. We build the crawler, run it on your cadence and deliver CSV, JSON, XLSX or an endpoint. Field names and the refresh interval are agreed before the first run.

The engineering is ours. On an ITVX scraping project that means anti-bot handling, proxy rotation and CAPTCHA solving as part of the managed service, and UK exit points, since the service is served to a UK audience. What we will not do is just as firm: we never touch the video stream and go nowhere near content protection, we do not sign in, and we collect no account, profile or viewing records and no personal data. ITV terms forbid extracting data from the service - we say that plainly, and anything contested should go past your own counsel first.

Fields an ITVX scraper reads from a programme and episode record

Programme and episode records are open to a signed-out visitor - no account, no player, nothing held behind the paywall. This is what we extract from ITVX for each one.

  • Identifiers - programmeId and episodeId in all three encodings, productionId with a version suffix such as #003, the ccid plus brandCCId and titleCCId, a series id that appends the series number to the programme id, and nextProductionId pointing at the following episode.
  • Titles and text - title and titleSlug, episodeTitle where one exists, a one-line description, a longDescription, an EPG description, and on brands a long synopsis field written for the programme page.
  • Numbering - series and episode numbers, contentInfo rendered as S1:E1, a series label, numberOfAvailableSeries on the brand and numberOfAvailableEpisodes on each series, plus fullSeries and longRunning flags that say whether the boxset is complete or still running.
  • Timing - duration as 46m for display and notFormattedDuration as PT46M for machines, productionYear, and dateTime with broadcastDateTime carrying the first transmission.
  • Classification - genres and subgenres with stable ids such as DRAMA_AND_SOAPS and CRIME_AND_THRILLER, the originating channel, and a guidance line that spells the content warning out in words instead of a certificate.
  • Access - tier reading FREE or PAID, a premium boolean, isPaid on browse tiles, a premium tag on the artwork, an adRule value, and contentOwner naming the supplier when the stock is licensed in.
  • Accessibility - accessibility tags of ad, s and bsl, with audioDescribed, subtitled and visuallySigned as separate booleans and a distinct signed playlist where a BSL version exists.
  • Artwork - a template address with placeholders for class, aspect ratio, width, height and quality, alongside the breakpoint presets the site requests for itself.

Two fields carry more weight than the rest. availabilityFrom and availabilityUntil are stamped on every episode, so the day a programme leaves is published rather than guessed. And tier splits the free ad-funded library from ITVX Premium inside the same record, which is the division most buyers of ITVX data actually need to measure.

Fields an ITVX scraper reads from a programme and episode record
TV schedule data: the ITV guide, live channels and now-next

TV schedule data: the ITV guide, live channels and now-next

The broadcast schedule is a second dataset no subscription-only service can offer, and on ITVX it sits in the open beside the catalogue.

  • Day pages - the TV guide is addressed by date and covers a rolling fortnight: a week behind, today, a week ahead. Five channels are gridded - ITV1, ITV2, ITVBe, ITV3 and ITV4 - at roughly thirty slots each per day.
  • Slot fields - title, start and end in UTC, duration in seconds, the title CCID, genre, series number, and both an episode link and a programme link wherever catch-up exists.
  • The join - about half the slots carry a catch-up link; regional weather, promos and continuity carry none, which is exactly the boundary between what is broadcast and what enters the catalogue.
  • Live channels - eighteen streams: five simulcasts of the linear channels and thirteen FAST channels addressed as fast1 through fast22, among them Midsomer Murders, Vera, The Chase and ITVSigned.
  • Now and next - each channel publishes the current and following programme with start, end, duration, brand title, short synopsis, genre, a numeric age rating and, on films, the BBFC certificate, plus subtitle and signing flags.
  • Upcoming and collections - a short sitemap of /upcoming pages for announced programmes, and editorial rails addressed by slug and a 22-character id, Last Chance to Watch among them, which appear in no sitemap.

The homepage adds a Top 10 Most Watched rail: rank order only, no viewing figures.

How the ITVX catalogue sits on top of the ITV broadcast grid

ITVX is the streaming service of a British broadcaster, and that is why its data looks nothing like a subscription-only library. Most of the catalogue is free with advertising, ITVX Premium is a paid tier layered onto the same records, and behind both runs a real transmission schedule on ITV1, ITV2, ITVBe, ITV3 and ITV4. A programme therefore exists twice - as a slot in tonight's grid and as an entry in the on-demand catalogue - and the published data links the two.

Addressing follows that structure. A programme page is /watch plus the title slug plus a programme id; an episode page appends the episode id. The same number is carried in three encodings: 2a1926 in the address bar, 2_1926 in the ad parameters, and a slash form used internally. Films and one-off specials are addressed by a production-level id ending in the letter B, so a film has no separate episode layer at all. Every record also carries a seven-character CCID such as gsytf1j, and that CCID keys the artwork host, so images resolve by id and never by title.

Browse is flat and small. Nine categories cover the whole library: Documentaries and Lifestyle, Drama, Kids, Film, Sport, Comedy, News, Entertainment and Reality, and Signed - BSL. Each hub has an /all address that returns the entire category in a single payload, with no pagination and no scroll token. There is no A-Z index for the adult catalogue - the Shows A-Z link in the header belongs to the kids profile - and the search path is closed to crawlers in the robots file. The coverage plan falls out of that: nine /all listings for the programme set, the sitemaps for the episode set, and the dated day pages of the TV guide for the schedule.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

What rights teams, guides and analysts do with ITVX data

The people who pay to scrape ITVX need to prove where a programme was, on which day, and on what terms.

  • Rights holders and distributors - confirming that a licensed series surfaced on the date the contract named, and that it comes down when the window closes rather than lingering.
  • Rival broadcasters and streamers - reading what a public-service competitor keeps free with advertising and what it moves behind the ITVX Premium price, title by title.
  • TV guides and availability apps - where the grid is the product, and a slot pointing at the wrong catch-up episode is a support ticket the same evening.
  • Media analysts and investors - sizing the free library against the paid slate and measuring how fast a broadcaster's archive turns over.
  • Advertisers and media agencies - the free tier is ad-funded, so the ad-supported inventory is the catalogue itself, and the genre and channel mix is planning input.
  • Production and research teams - tracking which suppliers hold shelf space, because the content owner is named on the record when a title is licensed in.

None of them wants a one-off list of titles. They want the same record read again on a schedule, so that an arrival, a tier change and a removal each land as a dated row. That is the difference between a catalogue dump and a usable feed, and it is why the availability window matters more here than anywhere else.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

ScrapeIt as your supplier of ITVX data

ScrapeIt is a managed web scraping agency. You describe the dataset, we build and operate the ITVX scraper and hand back files or an API; there is no tool for you to run and no proxy bill to manage. A project here usually opens with a paid sample - a few hundred programmes across two or three categories, with their episode records and a fortnight of schedule - so you can test the fields and the availability dates against your own records before committing to a schedule. Delivery goes to S3, GCS, an SFTP drop, a webhook or a database you nominate.

FAQ

Does ITVX have a public API for catalogue or schedule data?

No. There is no developer portal, no documented ITVX API and no data product for buyers. The site is a single-page application talking to internal services, and those endpoints refuse callers from outside the player, so they are not a supply route. Everything we deliver comes from pages a signed-out visitor can open: programme and episode records, the nine category listings, the dated TV guide pages and the live channel page. What you get from us is an ITVX API of our own - an endpoint over the extracted dataset, with your field names and your refresh interval.

Does ITVX publish when a programme will leave the catalogue?

Yes, and it is the most valuable thing on the record. Every episode carries an availability window with both an opening and a closing timestamp, so departures are read rather than derived. The spread is wide: a daily news programme sits for about a week, a licensed film can run just thirty days, an acquired drama boxset two years, and a kids series can be dated out six years ahead. There is an editorial Last Chance to Watch rail as well, but it is curated and short - the dates on the records are the reliable source.

Can you tell free ITVX content apart from ITVX Premium?

Yes. Access is explicit in the data rather than inferred. Each programme and episode carries a tier of FREE or PAID, a premium boolean beside it, and browse tiles carry an isPaid flag plus a premium tag on the artwork. The free ad-funded library is what the nine category listings return; the paid slate surfaces through the home rails and collections, and includes archive comedy licensed in from other broadcasters. The ITVX Premium price starts at GBP 5.99 with no contract, and titles licensed from third parties name their content owner on the record.

How often can ITVX data be refreshed, and in what formats?

Daily, weekly or monthly, and the schedule side can go faster. The catalogue moves slowly enough that a daily run captures arrivals and removals cleanly; the TV guide is worth pulling every day because the fortnight window rolls forward and only the near days are firm. Output is CSV, JSON, XLSX or a REST endpoint, delivered to S3, GCS, SFTP, a webhook or straight into a database. Every row carries the snapshot date alongside the availability dates, so successive runs stack into a history instead of overwriting each other.

Is it legal to scrape ITVX, and what will you not touch?

Be clear-eyed. The ITV terms cover personal, non-commercial use and forbid extracting any data or metadata from the services, including by scraping, and forbid building an index or database from the content; the robots file names a list of crawlers and blocks them outright, and closes the search and account paths to everyone else. We work only on pages a signed-out visitor can reach, never sign in, never touch the video stream or content protection, and collect no account, profile or personal data. We are a contractor, not your legal adviser - take the terms and your intended use to your own counsel first.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582