ClinicalTrials.gov Scraper for Trial Records and Changes

The registry has a full public API, so the honest question is not how to scrape it. It is what the API cannot tell you, and the answer is when anything changed.

ClinicalTrials.gov Scraper
Solutions

Managed clinical trial data, run end to end by us

ScrapeIt runs the collection as a managed service. You name the conditions, interventions, sponsors or geographies; we build the pipeline, run it on your cadence and hand back CSV, JSON, Excel or a push into your warehouse, with facilities as rows, sponsors normalised and a change log rather than a snapshot.

Where the public API is the right route for part of the brief we use it and say so. Paying for a crawl that duplicates a supported endpoint is a waste, and we would rather charge for the part that is genuinely hard: keeping every observation, diffing it reliably and joining it to everything else you hold.

This is public registry data and it is meant to be read, but it is registry data about medical research and we treat it accordingly. We collect no patient level information, we do not attempt to identify participants, and we pace requests. If your use case touches regulatory submissions or safety reporting, bring it to your own counsel and to your regulatory team before the project starts.

Clinical trial fields in every export

The study record covers NCT identifier, brief title, official title, sponsor and collaborators, condition, intervention and intervention type, phase, study type, enrolment and whether it is actual or estimated, eligibility criteria, sex and age limits, primary and secondary outcome measures, and overall status.

Dates come as a set and each one is kept separately: first posted, last update posted, study start, primary completion and completion, each flagged as actual or estimated. Mixing an estimated completion date with an actual one is the fastest way to build a misleading timeline, and the registry marks the difference, so there is no excuse for losing it.

Facilities are delivered as their own rows rather than flattened into a string. Each site carries the facility name, city, state, country and recruitment status, which is what makes questions about geographic footprint answerable: how many sites in Germany, which of them stopped recruiting, when did the first site in Japan appear.

The change history is the field set that distinguishes this from an API pull. Every observation is stored, and the differences between observations are delivered as a change log: what field changed, from what value, to what value, and when we saw it. Status transitions, enrolment revisions, date shifts and site additions all fall out of that automatically.

Results and linked publications are joined where posted, along with the collection timestamp on every row.

Clinical trial fields in every export
Sponsors, results postings and joining to literature

Sponsors, results postings and joining to literature

Sponsor normalisation matters more than it sounds. The same pharmaceutical company appears under subsidiary names, historical names and abbreviations across thousands of records, and a portfolio view that does not resolve them undercounts badly. We normalise sponsors and collaborators to a single identifier and keep every surface form encountered, so a corporate family can be viewed whole or by entity.

Results postings are a separate and unevenly populated part of the registry. Where results exist we collect the outcome measures, participant flow and adverse event tables as structured rows rather than as text, which makes them usable in analysis. Where they do not exist, the absence is itself recorded with the date the study completed, because the gap between completion and posting is a compliance question people track.

Joining to literature is the most requested extension. A trial identifier appearing in a published paper connects the registry record to the reported result, and we run that join against literature metadata rather than leaving you to match strings. It is the same pipeline described on our PubMed page, and the two are usually bought together.

Historical depth is straightforward here: the registry retains records and identifiers are stable, so a retrospective across a therapeutic area is a one time job. The change history, however, only exists from the point where somebody starts observing, which is the argument for beginning a watch before you need it.

What the registry holds and what the public API returns

ClinicalTrials.gov is the United States registry of clinical studies and, in practice, the closest thing the world has to a single index of trials. Each record carries an NCT identifier, a brief and an official title, the sponsoring organisation, the condition studied, the intervention, phase, enrolment, eligibility criteria, status, and the list of facilities running it.

It also has a documented public API, and we will say so before quoting anyone for a crawl. The version 2 endpoint returns study records as structured JSON, and bulk downloads are offered as well. For anyone who needs the current state of the registry, that is the correct route, it is faster than any crawl, and it is supported.

What the API returns is a snapshot. It tells you what a study record says today. It does not, by itself, hand you a history of how that record has changed, and change is where most of the commercial value in this dataset sits. A study that moved from recruiting to active not recruiting, added twelve sites in one country, cut its enrolment target or quietly shifted its primary completion date has told you something that a current snapshot cannot.

The site also carries context around the record that a client often wants joined in: results postings, linked publications and the sponsor pages behind them. That joining is ordinary data work rather than scraping heroics, and it is usually where the project actually goes.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the change log is the part worth paying for

Competitive intelligence in pharmaceuticals runs on timing. A rival programme that slips its primary completion date by six months has told you something material, and it has told you quietly, by editing a field. Nobody announces it. If your dataset holds only the current state, the edit is invisible and the six months are gone by the time it shows up in a press release.

The same applies to enrolment. A trial that halves its target has usually run into recruitment trouble; a trial that raises it has often expanded an arm. Both are legible from a revision history and neither is legible from a snapshot.

Site expansion is the third signal and the most geographic. Watching facilities appear country by country traces where a sponsor is taking a programme next, often months ahead of any regulatory filing. Because facilities are separate rows with their own status, that reads directly off the data rather than requiring interpretation.

None of this requires fighting the registry. We pull through the public interfaces where they serve, on a schedule, and the value comes from keeping every observation rather than overwriting yesterday with today. That is a discipline rather than a trick, and it is the reason clients stop maintaining their own nightly API job and hand it over.

The fourth use is coverage of the wider landscape. A single registry is not the whole world of trials, and serious competitive work joins this to other registries and to literature. We build that join rather than delivering one source and calling it a landscape.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your clinical trial feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, watches it as the registry and its interfaces change, and repairs it before your dashboard goes quiet.

You see a sample first, in your format, over the conditions and sponsors you actually track, with the change log already populated from a real observation window so you can judge it on data rather than on a promise.

FAQ

The registry has a public API. Why would I pay for collection?

For the current state of a record, the API is the right route and we will tell you so. What it returns is a snapshot: what the record says today. It does not hand you a history of edits, and the commercial signal here is almost entirely in the edits - a completion date that slipped, an enrolment target that was cut, twelve sites that appeared in one country. We keep every observation and deliver the differences.

What exactly is in the change log?

Per study, per observation: which field changed, its previous value, its new value and when we saw the change. Status transitions, enrolment revisions, date shifts, site additions and removals, and eligibility amendments all fall out of that. It only covers the period from when observation started, which is why beginning the watch early matters on this source.

Are trial sites delivered as separate rows?

Yes. Each facility keeps its name, city, state, country and recruitment status as its own row linked to the study. That is what makes geographic questions answerable: how many active sites in a country, which ones stopped recruiting and when, where a sponsor opened first. Flattening sites into a single text field destroys all of that.

Can you join trials to published results?

Yes, and it is the most common extension. Trial identifiers appear in published papers, so the registry record can be linked to the literature record rather than matched by title. We run that join in the same pipeline, which is why clients usually buy this alongside literature metadata collection.

Is any of this patient data?

No. The registry publishes study level information: protocols, sponsors, sites, eligibility criteria and aggregate results where posted. There is no patient level data in it and we make no attempt to identify participants. If your use case touches regulatory submissions or safety reporting, take it to your own counsel and regulatory team before the project starts.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582