Companies House Scraper for UK Companies and Officers

Company records are corporate facts. Officer records are real people with birth months and correspondence addresses, and they need treating that way.

Companies House Scraper
Solutions

Managed company register data, run end to end by us

ScrapeIt runs the collection as a managed service. You name the companies, sectors or filters; we use the registrar's free interface where it covers the need, extract filed accounts where it does not, mark personal data fields explicitly, and hand back CSV, JSON, Excel or a push into your warehouse.

Where a company-level dataset would meet the brief we propose that version, because it is cheaper for you and carries no personal data obligations at all.

This is a statutory public register and we collect it politely. Officer and control records describe living individuals, and you are the controller of whatever you build from them. Your counsel should see the use case before the project starts - that is not a formality on this source.

Companies House fields in every export

Company records carry the company number, name, previous names, status, type, incorporation date, registered office address, SIC codes, accounts and confirmation statement dates.

The company number is the identifier everything joins on. Names change and are reused; the number does not, and a dataset matched on name will merge unrelated companies and split renamed ones.

Officer records carry the name as filed, role, appointment and resignation dates, nationality, occupation and the partial date of birth the register publishes. These describe identified living people and the dataset marks them as such, because what happens downstream to a personal data column should be a deliberate decision rather than an oversight.

Persons with significant control are collected as their own record type, with the nature of control as published.

Document records carry the filing type, date and address, with extracted content where a client needs the accounts as numbers rather than as files.

Companies House fields in every export
Accounts extraction, network analysis and limits

Accounts extraction, network analysis and limits

Accounts extraction is the substantial technical work here. Filed accounts exist as documents in varying formats and turning them into comparable figures across thousands of companies is a real pipeline, not a download.

Directorship network analysis - which people sit on which boards and how companies connect through shared officers - is the most requested analytical product and the one that most clearly needs the personal data conversation first, because a network of named individuals is exactly the kind of profiling that carries obligations.

Status and filing monitoring is the lighter product: which companies changed status, filed late or moved to dissolution, built from company-level fields with no personal data involved at all. Where a brief can be met this way we suggest it.

Limits stated plainly: we pace politely on a public service, we use the free interface where it fits, and we raise the personal data question at scoping every time rather than treating the statutory publication as the end of the matter.

A public register with a free interface

Companies House is the UK registrar of companies. Its public service publishes company records, filing histories, accounts, officer appointments and persons with significant control, for every registered company in the United Kingdom.

It also runs a free public programmatic interface with its own developer portal, which covers company profiles, officers, filing history and search. For most briefs that interface is the right route, and a scraping project that ignores it is solving a problem the registrar already solved.

Where collection still has a role is in the documents. Filed accounts and certain statements exist as documents rather than as structured fields, and extracting their contents into data is work the interface does not do for you.

The search and company pages respond directly. There is no robots file at the service address, which we take as no stated restriction rather than as an invitation, and we pace accordingly.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why officer data needs a decision before collection

This register is public by statute, and that is often where the thinking stops. It should not.

Company records are corporate facts and carry no personal data question. Officer records are different: a name, a role, a nationality, an occupation, a month and year of birth and a correspondence address, describing a living person. The register publishes them under a legal regime designed for transparency about who controls companies. That is not the same thing as a general permission to compile them into profiles for any purpose.

In practice the client is the controller of whatever they build, and the obligations that follow - purpose, retention, the rights of the people in the dataset - are theirs. Our part is to make the personal data visible as such in the schema rather than letting it arrive as just more columns, and to raise the question at scoping rather than after delivery.

The second reason to think before collecting is scope. A great many briefs that ask for all officers actually need officers of a specific set of companies, and the narrower version is both cheaper and much easier to defend.

The third is that the free interface already covers most of it, which means the conversation is usually about what a client should build rather than whether we can fetch it.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your register feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, maintains accounts extraction as filing formats change, and repairs it before a monitoring series misses a status change.

You see a sample first, in your format, over the companies you actually track, with personal data fields marked so the decision about them is one you make knowingly.

FAQ

Should I use the API instead of scraping?

For most briefs, yes, and we will say so. The registrar runs a free public interface covering profiles, officers, filings and search. Collection earns its place on the filed documents, which the interface hands you as files rather than as data.

The register is public - why raise data protection?

Because public by statute and free to compile for any purpose are different things. Officer records describe living people, and whatever you build makes you the controller with the obligations that follow. We surface it at scoping rather than after delivery.

Can you do directorship network analysis?

Yes, and it is the product that most needs the conversation above first. A network of named individuals across boards is precisely the kind of profiling that carries obligations, so we scope it deliberately rather than as a side effect.

Is there a version with no personal data?

Yes, and we often propose it: company-level fields only - status, filings, accounts, addresses, SIC codes. Many briefs asking for everything actually need this, and it is cheaper and carries no personal data obligations.

Why is extracting accounts hard?

Because filed accounts are documents in varying formats rather than structured fields. Turning them into figures comparable across thousands of companies is a real pipeline, and it is the part of this source where collection genuinely adds something.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582