DrugBank Database Download, Parsed, Versioned and Joined to Open Data

DrugBank sells its full database under license, and that is where we start: your licensed XML becomes clean tables, every release gets a diff, and the open drug data around it is joined in.

DrugBank Scraper
Solutions

Scraping around DrugBank: your license, our pipeline, open sources joined in

You keep the license and the download: DrugBank's terms forbid shared logins, so the file lands in your bucket or our job runs in your environment. We extract DrugBank data from the XML into tables for drugs, synonyms, products, interaction pairs, proteins, pathways and identifiers, diff each release against the one before, and deliver CSV, Parquet, SQL or an internal API inside your own cloud. Around it we collect the open sources your license does not cover - FDA labels, RxNorm, ChEMBL, PubChem, UniProt, trial registries - and join them on UNII, InChIKey, RxCUI and UniProt IDs, within what your subscriber order lets you merge.

Where an open source is guarded, we bring anti-bot handling, proxy rotation and CAPTCHA solving; none of it is ever pointed at DrugBank's sign-in, bot check or downloads. Nothing here touches personal data: these records describe molecules, proteins and products, not people.

Inside the DrugBank XML download: one drug record, element by element

The DrugBank full database XML is one file with a drug element per entry, described by the drugbank.xsd schema. Each drug carries a type of small molecule or biotech, created and updated dates, and a primary drugbank-id of DB plus five digits, followed by secondary IDs left by merged entries and older formats such as APRD, BIOD, BTD, EXPT and NUTR. Salts get DBSALT IDs and metabolites DBMET IDs.

Identity. Name, description, CAS number, UNII, average and monoisotopic mass, physical state, groups (approved, investigational, experimental, withdrawn, illicit, nutraceutical, vet_approved), synonyms tagged by language and coder, international brands with their company, mixtures, salts, a chemical taxonomy from kingdom to direct parent, categories with MeSH IDs, ATC codes that carry all four parent levels, and AHFS codes.

Pharmacology. Indication, pharmacodynamics, mechanism of action, toxicity, metabolism, absorption, half-life, protein binding, route of elimination, volume of distribution and clearance arrive as curated prose, not numbers. Calculated properties add logP, logS, pKa, polar surface area, SMILES, InChI, InChIKey and Rule of Five; experimental ones add melting point, water solubility and caco2 permeability.

Proteins. Targets, enzymes, carriers and transporters share one shape: name, organism, actions such as inhibitor or substrate, a known-action flag of yes, no or unknown, references, and a polypeptide with UniProt ID, gene name, cellular location, Pfam and GO terms and sequence. Enzymes add inhibition and induction strength. Pathways carry SMPDB IDs, reactions map a parent compound to its metabolite with the enzyme responsible, and SNP effects carry rs-ID, allele and PubMed ID.

Interactions and market. A drug-interaction is a partner DrugBank ID, its name and one sentence of description; food interactions are short advice lines. Products list labeller, NDC, DPD and EMA codes, marketing dates, dosage form, strength, route, FDA application number and generic or OTC flags, beside prices with currency, patents with expiry and pediatric extension, and external identifiers for PubChem, ChEMBL, ChEBI, KEGG, PharmGKB, BindingDB, ZINC and RxCUI.

Inside the DrugBank XML download: one drug record, element by element
DrugBank API: endpoints, keys, regions and how pricing is built

DrugBank API: endpoints, keys, regions and how pricing is built

The DrugBank API comes in two parts, and DrugBank issues the keys for both rather than offering a self-serve signup. The Clinical Intelligence API at api.drugbank.com/v1 answers in JSON, filtered by the regions in your license such as us, ca or eu, in pages of at most 50 records with Link and X-Total-Count headers. A production key counts calls against a quota, and calls above it still pass as overage; a development key is capped per month for testing; browser apps use short-lived tokens through a separate api-js host.

Its endpoints cover medication and drug name search with fuzzy autocomplete, products and product concepts including lookup by RxCUI, drug-drug interactions by DrugBank ID, product concept or local product ID, indications, conditions with ICD-10 search, allergies and cross-sensitivities, ATC and MeSH categories, adverse effects, boxed warnings, packages and current FDA labels with their history.

The Discovery API at /discovery/v1 serves the research side with no region filter, because it covers unmarketed experimental and illicit drugs too: drugs, products, drug-target bonds, SNPs, bio-entities, polypeptides and molecules.

The subscription is modular, and that is what sets DrugBank API pricing: a Drug Search base module, then add-ons for adverse effects, allergies, contraindications and boxed warnings, interactions with severity and management, conditions, indications, pharmacology and US labels, plus extra regions. Mappings to SNOMED CT, MedDRA and ICD-10 need their own licenses.

DrugBank dataset access: public preview, free account, licensed releases

DrugBank is a drug and drug-target knowledgebase that began at the University of Alberta and is run today by OMx Personal Health Analytics Inc., a Canadian company trading as DrugBank. A DrugBank dataset ties each molecule to the proteins it acts on, the pathways it runs through, the drugs it interacts with and the products it ships in, all under one DB accession number. The company now sells that knowledge as a platform: a knowledge graph, Biopharma OS, the Clinical Intelligence API, Data Library subscription packages, Snowflake sharing and an MCP server for AI tools.

Access comes in three layers. Signed out, go.drugbank.com shows a preview card at addresses such as /drugs/DB01076: name, groups, type, summary, mechanism, the opening of the primary indication, formula and weight, first approval by country, synonyms, code names, brands, identifiers from UNII and CAS to ChEMBL, PubChem, RxNorm and ATC, InChIKey, SMILES, and targets and transporters as UniProt accessions. Enzymes and carriers stay hidden, and interactions, trials, indications, contraindications, products and references show only as counts. A free account opens the full card for non-clinical research. Anything in bulk - XML releases, CSV and SQL packages, the API - needs a license: the free Academic License for eligible researchers, paid Academic+, or a commercial subscription.

The terms of use draw the line plainly. The free license covers internal, non-clinical research; copying the data, co-mingling it into derivative datasets for commercial use or building another database of drug information, interactions or targets needs a signed commercial agreement. Disputes fall under Alberta law, and academic work must cite the DrugBank 6.0 paper.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

DrugBank DDI dataset, target maps and release discipline

Buyers fall into four camps. Pharma and biotech teams use DrugBank data for target identification, repurposing and licensing due diligence. Bioinformatics groups build drug-target networks from the protein blocks. Machine learning teams train interaction and binding models. Clinical software vendors need medication search, interaction checks and allergy cross-sensitivity inside an EHR or a telehealth app.

The DrugBank DDI dataset is where most confusion starts, because its two formats differ. In the XML an interaction is a pair of IDs plus one sentence, with no severity field; the sentences follow a small set of templates, which is why DDI prediction work often turns them into interaction-type labels. Severity ratings and management advice live in the interaction module of the Clinical Intelligence API. A training set and a prescribing screen are two different purchases.

Releases need discipline. Versions follow MAJOR.MINOR.PATCH: a major bump means an incompatible schema change, a minor one new data, a patch every new export, and releases are kept by version, so a model or a paper can pin 5.1.22. Merged entries survive as secondary IDs, so joins must resolve old accessions to the primary one before a diff between two releases means anything.

It is also why a DrugBank scraper is the one thing we will not build: the preview card holds a fraction of the record, the site guards itself against automated traffic, and bulk copying is exactly what the license exists for.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

The team behind your DrugBank data pipeline

ScrapeIt works as your outsourced data team: we design the parsers and collectors, run them on your release calendar, and keep them working when a schema, a registry or a label format changes. Before we open a licensed file we sign a non-disclosure agreement at least as strict as your DrugBank terms, which is what those terms ask of contractors.

ScrapeIt has no tie to DrugBank or OMx. We never host, resell or redistribute DrugBank content, and every pipeline we build can run inside your own cloud.

FAQ

Is the DrugBank API free, and how is DrugBank API pricing set?

No. DrugBank runs two APIs, the Clinical Intelligence API for health software and the Discovery API for research data, and both use keys that DrugBank issues under a paid subscription; the development key is a second key capped per month for testing, not a free tier. Responses are paged JSON for one query at a time, not a copy of the release. Pricing follows the modules and regions you license, with calls over the agreed quota counted as overage. The free routes are the website with a free account and, when downloads are open, the Academic License files and the CC0 vocabulary.

Where does a DrugBank XML or CSV download come from?

From the releases page, under your own account. The 5.1.22 full database XML is a 204 MB zip, fetched over HTTP basic authentication with your account email and password. A DrugBank CSV download on the same page covers structure links, external drug links, protein identifiers and the CC0 vocabulary; the full data in CSV, JSON or SQL comes with Academic+ or a commercial package. In September 2026 every file there, open data included, is marked temporarily unavailable while DrugBank reworks distribution; commercial data moves through separate channels, the commercial downloads portal, Data Library packages and the API. We flatten whichever file you hold into CSV or Parquet.

Which tables do you build from a DrugBank release, and what keys link them?

Everything the licensed release carries, split into tables: the drug with its primary and secondary IDs, groups, synonyms, brands, taxonomy, ATC and MeSH categories and pharmacology prose; products with NDC, DPD and EMA codes; interaction pairs; food interactions; targets, enzymes, carriers and transporters with UniProt IDs and actions; pathways, SNP effects, patents and properties. Joins run on UNII and CAS for regulators, InChIKey for chemistry, RxCUI for clinical vocabularies and UniProt for proteins.

How often does DrugBank release new data, and how do you track changes?

Every export raises the patch number. The current release, 5.1.22, came out on 27 June 2026, and its notes read no significant changes; a minor bump would signal new data and a major one a schema change, which is the event that breaks parsers. Releases are kept by version, so you can pin one for a paper or a model. We diff each release against the previous one - new accessions, merged IDs, added or dropped interactions, changed product rows - and deliver the change set with the full snapshot.

Is scraping DrugBank legal, and can we use a copy from GitHub?

DrugBank's terms forbid copying the data, building derivative datasets for commercial use or another drug, interaction or target database without a signed subscriber order, so we do not scrape DrugBank pages. A copy of the DrugBank database on GitHub or Kaggle carries no more rights than the release it came from, at best the non-commercial academic license, and nothing proves which release it was or what was edited; we neither use nor supply such copies. Without a license, the open DrugBank alternatives are ChEMBL, PubChem, UniProt, RxNorm and FDA labels, joined on UNII, InChIKey and RxCUI.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582