EMA Data Collection for Authorised Medicines and Reports

The agency publishes its medicine data for download. Quoting you for a crawl without checking that first would be taking money for work nobody needs.

EMA Scraper
Solutions

Managed regulatory data, run end to end by us

ScrapeIt builds and runs the pipeline as a managed service. You name the therapeutic areas, substances or holders; we use the agency's published downloads where they cover the need, collect the documents where they do not, extract assessment reports with structure preserved, and hand back CSV, JSON, Excel or a push into your warehouse.

We check the downloads first and say so in the quote. Charging for a crawl that reproduces a published export is not something we are willing to do.

This is public regulatory information, published for use, and it contains no patient data. If your use case feeds regulatory submissions or pharmacovigilance, take it to your own regulatory team before the project starts, since the tolerance for error there is different from ordinary analytics.

EMA fields in every export

The medicine record covers the medicine name, international non proprietary name, active substance, therapeutic area, ATC code where published, marketing authorisation holder, authorisation status, authorisation date and the procedure type.

Status history is kept rather than overwritten. A medicine that was authorised, then suspended, then reinstated has a history that a current status field destroys, and the transitions are precisely what regulatory and competitive teams watch for.

Document records carry the document type, title, publication date and the medicine it belongs to. Where a client needs the content, assessment report text is extracted with its section structure preserved, because the indication section and the discussion section answer different questions and a flat text blob answers neither well.

Safety information is collected as its own record type: referrals, their trigger, the committee involved and the outcome with its date.

Provenance is recorded per row: published download or page collection, because a mixed pipeline should always be able to say which.

EMA fields in every export
Assessment reports, shortages and joining to trials

Assessment reports, shortages and joining to trials

Assessment report extraction is the substantial piece of work. Reports are long documents with consistent structure, and extracting them section by section - indication, clinical efficacy, safety, conditions of authorisation - produces something an analyst can query rather than read. Doing it well takes account of the fact that the structure has changed over the years.

Shortage notices are a separate and time sensitive record type: which medicine, which countries affected, the reason given and the expected resolution. For supply chain and procurement teams in healthcare that is one of the more operationally useful things the agency publishes.

Joining to clinical trials is the standard extension. An authorisation rests on trials, and the trial identifiers frequently appear in the assessment report, which connects the regulatory record to the study record. That is the same pipeline as our clinical trials work and the two are usually bought together.

Historical depth is good: the register retains withdrawn and superseded medicines, so a retrospective on a therapeutic area is practical as a one time job.

What the regulator publishes and in what form

The European Medicines Agency evaluates medicines for the European Union. Its site carries the authorised medicine register, the assessment reports behind each authorisation, safety referrals, shortage notices and the committee outcomes that drive all of it.

It also publishes medicine data for download, on its own page, in bulk. That is the first thing to check on any EMA project and the first thing we check, because where the download covers a client's need there is no reason to build a crawler at all and we will say so.

What the downloads do not always carry is the documents. Assessment reports are substantial PDFs per medicine, and the useful content inside them - indications, the committee's reasoning, conditions attached to an authorisation - is not in a tabular export. That is where collection earns its place.

The register is structured around the medicine: name, active substance, therapeutic area, authorisation holder, status and the document set attached. Section pages respond and crawl rules are published.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the documents are the part worth paying for

Most of what people first ask for on this source already exists as a download. A list of authorised medicines with their holders and statuses is published in bulk, and building a crawler to reproduce it is work that benefits nobody.

What is not in the download is the reasoning. An assessment report explains what the committee concluded, on what evidence, and with what conditions attached, and for regulatory affairs and competitive intelligence that content is the entire point. It exists as documents per medicine, and turning several thousand of them into structured, queryable text is a real piece of work.

The second thing collection adds is history. A download is a snapshot of current state, and the interesting events here are transitions: an authorisation granted, a referral opened, a status changed, a shortage declared. Those exist only if somebody was observing, and once they exist they answer questions about timing that no current-state export can.

The third is joining outward. A medicine record connects to clinical trials, to literature and to national registers, and the value in a regulatory dataset usually sits in the join rather than in any single source. We build that rather than delivering one source and a problem.

The last is that this is public regulatory information, published to be used, which makes it one of the few sources in this section with no access argument to have at all.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your regulatory feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, watches it as the site and the document formats change, and repairs it before your regulatory monitoring goes quiet.

You see a sample first, in your format, over the therapeutic area you actually track, with assessment report extraction applied so you can judge the hard part on real documents.

FAQ

Why would I pay you if the agency publishes downloads?

For most of what people first ask for, you would not, and we check the downloads before quoting. What they do not carry is the assessment reports - the committee's reasoning, the evidence, the conditions attached - which exist as documents per medicine. Turning several thousand of those into queryable structured text is the work worth paying for.

Can you keep the history of a medicine's status?

Yes, as transitions rather than a current value. A medicine authorised, then suspended, then reinstated has a history that a current status field destroys, and those transitions are exactly what regulatory and competitive teams watch. They exist only from the point where observation started.

How do you extract the assessment reports?

Section by section - indication, clinical efficacy, safety, conditions of authorisation - so an analyst can query rather than read. The structure has changed over the years, so the extraction accounts for several document generations rather than assuming one layout.

Can you collect shortage notices?

Yes, as their own time sensitive record type: medicine, countries affected, reason given and expected resolution. For procurement and supply chain teams in healthcare it is one of the more operationally useful things the agency publishes.

Is any of this patient data?

No. This is regulatory information about medicines: authorisations, assessments, safety referrals and shortages. There is no patient data on these surfaces. If the output feeds regulatory submissions or pharmacovigilance, take it to your regulatory team first, since the error tolerance there differs from ordinary analytics.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582