Financial Times Data: Licences and the AI Prohibition

The FT wrote its AI position into the one file every crawler reads first. A licence for research does not quietly extend to training a model.

Financial Times Scraper
Solutions

Managed work within an FT licence, run end to end by us

ScrapeIt runs the work as a managed service. If you hold an FT licence, we build within its terms, record the licence reference and a use flag on every row, and hand back CSV, JSON, Excel or a push into your warehouse.

If you do not, we tell you before quoting, point you to the licensing contact the FT publishes, and scope what can be answered from citations elsewhere.

The FT prohibits AI and machine learning use of its content. We will not build anything that feeds FT content into model training without an FT licence that covers it. Your counsel should see the licence and intended use before the project starts.

What a licensed FT pipeline carries

Where a client holds an FT licence, article records carry the headline, standfirst, byline, section, publication and update timestamps and the article identifier, in the form and scope the licence permits.

A use flag is recorded on every row stating what the licence covers. Because the AI prohibition is explicit, a pipeline built for internal research should mark its output as not for model training, so the restriction survives when the data moves between teams.

The licence reference travels with the data, so any row can be traced back to the terms it was collected under.

Citation records are a separate table built without the FT site: where other outlets cite or quote FT reporting, collected from sources that permit it, showing which FT stories shaped wider coverage.

Every row carries the collection timestamp and its source type.

What a licensed FT pipeline carries
Citation analysis, use flags and what we refuse

Citation analysis, use flags and what we refuse

Citation analysis is the strongest product that needs no FT access: which FT reports are cited, by whom, how quickly and in what context, collected from outlets that permit collection.

Use flags on licensed data keep permitted and prohibited uses apart after delivery. On this source the flag is not bureaucracy; it is the difference between a compliant pipeline and a contract breach waiting to happen.

Licensed archive work, where a client holds the right terms, is ordinary structuring and delivery, with the licence reference attached throughout.

What we refuse: collection from ft.com without a licence, any attempt to get past the 403, and any use of FT content for machine learning or AI purposes unless the FT has licensed exactly that.

A paywalled paper with its position in writing

The Financial Times is one of the most influential business newspapers in the world, published behind a subscription and read closely by finance, policy and corporate audiences.

Its crawl rules carry two statements in plain words. All use of FT content is subject to its terms and copyright policy. And the paper expressly prohibits any use of its content or data for machine learning or artificial intelligence purposes, including the training or development of AI technologies or language models. The same file publishes a contact for licensing enquiries.

Requests to the site returned 403. Taken together, the position could hardly be clearer: access is by subscription or licence, and AI use is excluded unless the FT itself agrees otherwise.

That leaves a narrower but real space for data work, and it starts with reading the licence rather than the page.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the AI clause matters even with a licence

Most publishers leave their position on AI implied. The FT states it, and states it broadly: any machine learning or AI purpose, including training and development. That wording reaches further than people expect.

The practical consequence is that a licence obtained for one purpose does not silently cover another. A research team licensing FT content for analysts, then passing the corpus to a data science group building a model, has moved from a permitted use to an explicitly prohibited one unless the licence says otherwise. It is an easy line to cross by accident, which is why we mark it on the data rather than trusting memory.

The second point is that most FT data briefs are really about influence: which FT stories moved markets or set agendas. That question can be answered from how other outlets cite the FT, which needs no access to ft.com.

The third is plain: the site answers anonymous requests with 403, and the paper has put its terms in writing. We do not engineer around either.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps a licensed feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds pipelines that stay inside the licence, keeps use flags attached to the data, and repairs collection when formats change.

You see a sample first, in your format. Without a licence in place, the sample is a citation analysis from other outlets, so you can judge the value before any licensing conversation.

FAQ

Can you scrape the FT?

Not without a licence. Access is by subscription or licence, anonymous requests get 403, and the paper states its terms in its crawl rules. With a licence we build within its terms; without one we scope what can be answered from citations elsewhere.

What does the AI prohibition cover?

Any use of FT content or data for machine learning or artificial intelligence purposes, including training or developing AI technologies or language models. It is broad by design and stated in the file every crawler reads first.

Does a licence let me train a model?

Only if the licence says so explicitly. A licence for research or internal reading does not silently extend to model training, and passing a licensed corpus to a modelling team is an easy line to cross by accident - which is why we flag permitted use on every row.

How do I get a licence?

The FT publishes a licensing contact in the same crawl rules file. For most commercial uses that conversation is the first step, and we are happy to help scope what to ask for.

Can I measure FT influence without their content?

Yes. Citations of FT reporting in other outlets show which stories shaped coverage, how fast and where, collected from sources that permit it. That answers most influence questions without touching ft.com.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582