CBC Gem Scraper for Canadian Titles and Origin Data

In a market with content quotas, where a programme was made is not trivia. It is a regulated category and the broadcaster publishes it.

CBC Gem Scraper
Solutions

Managed catalogue metadata, run end to end by us

ScrapeIt runs the collection as a managed service. You name the languages, genres or the full catalogue; we build the pipeline with origin, language and tier as fields and series structure intact, and hand back CSV, JSON, Excel or a push into your warehouse.

Where turnover matters we run repeat collection and keep every observation, since a title that has left leaves no trace of having been there.

We collect published metadata only - never streams, never anything behind a sign in, no viewer data - honour the crawl rules, use the published sitemaps for discovery and pace requests. Content is copyrighted, so the dataset is for analysis rather than republication, and your counsel should see the use case before the project starts.

CBC Gem fields in every export

Title records carry the title, original title, content type, genre, language, production country where published, release year, season and episode structure, description, maturity rating and availability tier.

Language and production origin are first class fields. In a bilingual public broadcaster operating under content rules, they are the two axes along which every interesting question is asked, and recovering them later from title text is guesswork.

Availability tier is recorded because the service has free and premium levels, and a catalogue that does not say which tier a title sits in cannot answer what a viewer gets without paying - usually the first question asked of a public broadcaster's streaming service.

Series structure is preserved as a hierarchy rather than flattened, so completeness and ordering questions stay answerable.

Every row carries the collection timestamp and, where repeat collection shows a change, the date a title appeared or disappeared.

CBC Gem fields in every export
Content share analysis, turnover and limits

Content share analysis, turnover and limits

Domestic content share analysis is the output that makes this source distinctive: the proportion of Canadian production by genre, language and tier, tracked over time. It is directly relevant to policy research and to producers arguing about commissioning.

Turnover analysis separates the standing catalogue from acquired content that rotates, which is only visible with repeat collection and is what tells you how much of the service is durable.

Bilingual comparison - how the English and French catalogues differ in size, genre and originals share - is a question about a national broadcaster that few datasets can answer.

Limits: metadata only, never streams, nothing behind a sign in and no viewer data. Playback is geo-restricted to Canada and that is the broadcaster's restriction to enforce, not ours to work around. Content is copyrighted, so the dataset is for analysis rather than republication.

A public broadcaster in a regulated market

CBC Gem is the streaming service of the Canadian Broadcasting Corporation, carrying the public broadcaster's television programming, originals, documentaries, news and acquired content, with both free and premium tiers.

Canada regulates domestic content in broadcasting, and that regulatory context makes this catalogue unusually informative. Where a programme was produced, and whether it qualifies as Canadian content, is not a stylistic detail here - it is a category with consequences, and it shows in what a public broadcaster commissions and carries.

The corporation is also bilingual, operating in English and French with a separate service for the French market. A catalogue collected without language as a field flattens two distinct commissioning worlds into one list.

Show and catalogue sections respond directly. Crawl rules are published and are unusually light - a single technical path disallowed - with a sitemap index published for discovery. Playback is geo-restricted to Canada; the metadata is not.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why origin and language carry the analysis

A catalogue list from a public broadcaster looks like a modest dataset until the regulatory context is applied, at which point it becomes one of the more informative sources about a national media system.

Content origin is the reason. In a market with domestic content requirements, the share of Canadian production in a public broadcaster's streaming catalogue is a measurable fact about how the system is working. Researchers, policy bodies and producers all want it and it requires origin to have been collected as a field rather than inferred.

Language is the second axis. English and French commissioning differ in volume, genre mix and budget, and a combined catalogue count tells you about neither. Collected separately, the comparison is direct.

The third reason is the free and premium split. For a publicly funded broadcaster, what sits behind a paywall is a question with a political dimension as well as a commercial one, and the tier field answers it.

The fourth is turnover. Acquired content arrives and leaves as licences change, and distinguishing the permanent commissioned catalogue from the rotating acquired one needs repeat collection - a snapshot cannot tell them apart.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your Canadian catalogue feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, follows the published sitemaps as titles are added, and repairs the collector before a turnover record develops a gap.

You see a sample first, in your format, over the languages and genres you actually track, with origin and tier populated so you can judge the fields on real titles.

FAQ

Why is production origin worth a column?

Because Canada regulates domestic content, so origin is a category with consequences rather than trivia. The share of Canadian production in a public broadcaster's catalogue is a measurable fact about how the system works, and it has to be collected as a field rather than inferred later.

Why separate English and French?

Because they are distinct commissioning worlds with different volume, genre mix and budgets. A combined catalogue count describes neither, and the comparison between them is one of the more interesting things this source supports.

Does the free and premium split matter?

Yes, and for a publicly funded broadcaster it carries a political dimension as well as a commercial one. The tier field answers what a viewer gets without paying, which is usually the first question asked of this kind of service.

Can you tell commissioned content from acquired?

Largely, and turnover is the tell: acquired content rotates as licences change while commissioned content stays. Distinguishing them needs repeat collection - a single snapshot cannot separate the two.

Playback is Canada-only. Does that block you?

No, because we collect metadata rather than streams. The geo-restriction is on playback and it is the broadcaster's to enforce, not ours to work around. Titles, descriptions, origin and tier are published openly.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582