The Telegraph Data: A 402 and a Crawl File That Misleads

The server literally says payment required. A crawl file that seems to wave everyone through does not outrank that, and reading it as permission would be a mistake.

The Telegraph Scraper
Solutions

Managed work within an agreement, run end to end by us

ScrapeIt runs the work as a managed service. If you hold a licence or syndication agreement, we build within its terms, record its reference on every row, and hand back CSV, JSON, Excel or a push into your warehouse.

If you do not, we say so before quoting and scope what can be answered from citations in other outlets.

We read the server's payment-required answer as the publisher's position and do not rely on the crawl file's apparent permission. Your counsel should see the agreement and intended use before the project starts.

What a Telegraph dataset can honestly carry

Where a client holds a licence or syndication agreement, article records carry the headline, standfirst, byline, section, publication timestamp and article identifier, within the agreement's scope and with its reference on every row.

Citation records are built without the site: where other outlets cite, quote or follow up Telegraph reporting, collected from sources that permit collection. They show which Telegraph stories drove wider coverage.

Access evidence is recorded with any project scoping on this source - the status code and the crawl rules as they stood on the date checked - because the gap between the two is exactly the kind of thing a later reviewer will ask about.

Every row carries the collection timestamp and its source type, so licensed content and citation observations are never mixed.

What a Telegraph dataset can honestly carry
Citation analysis, licensed work and limits

Citation analysis, licensed work and limits

Citation and follow-up analysis is the strongest product here: which Telegraph stories were picked up, by whom and how quickly, collected from outlets that permit collection.

Licensed or syndicated content work is ordinary structuring and delivery where a client holds the right agreement, with its reference attached throughout.

Access documentation is worth keeping for any project that touches this source. The mismatch between the server's answer and the crawl file is the sort of detail that matters in a later review, and recording it at the start costs nothing.

Limits: we do not collect the paywalled site without an agreement, we do not treat the crawl file's apparent permission as consent, and we do not attempt to get past the payment wall.

A subscription paper with an unusual status code

The Telegraph is a major British national newspaper, published online behind a subscription. It covers politics, business, sport and culture for a large UK readership.

Its answer to anonymous requests is unusual. Most paywalled sites return 401 or 403; this one returns 402, the HTTP status that literally means payment required. It is rarely used in the wild, and here it states the position about as directly as a status code can.

The crawl rules are the second unusual thing. The file names the major search engines - Google, Bing, Apple, DuckDuckGo and others - and disallows them from the whole site, then allows everything to all other crawlers. Read literally, it keeps search engines out and lets everyone else in.

No news publisher intends to hide from search engines while inviting scrapers, so the file almost certainly does not say what its authors meant. That makes this source a useful lesson in how to read access signals together rather than one at a time.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why a permissive-looking file is not permission

A crawler that reads only the crawl rules would conclude this site is open to it. A crawler that reads only the server's answer would conclude it needs to pay. The second reading is the right one, and the reasons generalise well beyond this paper.

Crawl rules express intent, and intent has to be read in context. A rule set that blocks every major search engine while welcoming unknown crawlers is not a coherent publishing policy for a newspaper; it is far more likely a configuration error, such as blocks written in the wrong order. Treating an obvious error as consent is not a defensible position, and it would not survive a conversation with the publisher's lawyers.

The server's answer, on the other hand, is unambiguous. It asks for payment. Content behind that answer is subscription content, and the route to it is a subscription or a licence, not a technicality in a text file.

What remains open without the paper's agreement is the paper's influence: how other outlets cite and follow its reporting. That is a useful dataset in its own right and needs nothing from telegraph.co.uk.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your coverage feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team reads access signals together rather than one at a time, keeps agreements attached to the data, and repairs collection when formats change.

You see a sample first, in your format. Without an agreement in place, the sample is a citation analysis from other outlets, so you can judge the value before any licensing conversation.

FAQ

The crawl rules allow everyone - so can you collect it?

No. The file blocks the major search engines and allows all other crawlers, which no newspaper intends and which is almost certainly a configuration error. The server itself answers 402 Payment Required. We read the unambiguous signal, not the misprint.

What does 402 mean?

Payment required. It is a rarely used HTTP status, and here it states the position plainly: this is subscription content, and the route to it is a subscription or a licence.

What can I get without a licence?

Citation and follow-up analysis: which Telegraph stories other outlets picked up, when and how. It is collected from sources that permit collection and needs nothing from telegraph.co.uk.

Why document the access signals?

Because the mismatch between the server's answer and the crawl file is the kind of detail a later review asks about. Recording both, as they stood on the day checked, costs nothing at the start and saves a difficult conversation later.

Would you work with syndicated Telegraph content?

Yes, within the syndication or licence terms you hold. We record the agreement's reference on every row so permitted use stays attached to the data after delivery.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582