MedlinePlus API and Data Extraction for Health Topics, Genetics and Medical Tests

MedlinePlus gives away its consumer health data in pieces: a daily XML file, a search API capped at 85 calls a minute, a code lookup capped at 100. We put the pieces together.

MedlinePlus Scraper
Solutions

Your own consumer health API, assembled from MedlinePlus sources

Scoping starts with layers: daily topic snapshots with a change file, Connect crosswalks for your code sets, genetics records completed with the page-only sections, test pages, recipes with nutrition panels, or all of them joined on MeSH, ICD-10-CM, SNOMED CT, RxCUI and LOINC. We extract MedlinePlus data from the official files and services first and keep MedlinePlus scraping for what those lack, inside NLM's published limits and with the tool and email parameters set. You receive CSV, JSON, Parquet, XML or a hosted endpoint, a consumer health API of your own, with source URL, retrieval date and the NLM credit line on each row. When a project also pulls guarded commercial sources, such as pharmacy price pages, anti-bot handling, proxy rotation and CAPTCHA solving are part of the service; NLM's hosts need none of it.

MedlinePlus data, record by record: topics, genes, tests, recipes

MedlinePlus data is more structured than the pages suggest, once you know where each section keeps it.

Health topics. The daily XML file holds every topic in both languages, 1,017 English and 1,016 Spanish in September 2026, about 29 MB raw or under 5 MB zipped. A record carries a numeric id, title, URL, language, creation date and meta description; the full summary as HTML; also-called synonyms and see references; MeSH descriptors with their D numbers; one or more of 44 groups; related topics; the matching topic in the other language; translations in over 50 languages; a primary NIH institute where one is assigned; and every curated link with its title, URL, organization, information category such as Start Here, Treatments and Therapies, Clinical Trials or Patient Handouts, and flags such as NIH, PDF, Video or Easy-to-Read.

Genetics. The MedlinePlus genetic conditions database currently covers 1,305 condition, 1,500 gene and 24 chromosome summaries plus mitochondrial DNA. A condition record gives name, synonyms, description, inheritance codes such as ad, ar, xr or m, related gene symbols, keys into OMIM, GTR, MeSH, SNOMED CT and ICD-10-CM, and reviewed and published dates; a gene record adds its normal function, NCBI Gene ID and the conditions that cite it. The legacy shows: the compendium still ships as ghr-summaries.xml, and the search service still calls this database ghr.

Medical tests. Around 300 test pages per language, each answering the same questions: what the test is used for, why it is ordered, what happens, how to prepare, risks, what results mean, then dated references.

Recipes, drugs and supplements. All 238 recipes in 16 categories show a full Nutrition Facts panel; 202 also carry schema.org Recipe markup with times, yield, ingredients, steps and a credited source, and the other 36 are parsed from the page. About 2,060 drug pages per language are indexed by monograph number, title, brand-name cross-references and last-revised date. Herbs and supplements are no longer hosted: the 79 English entries link out to NCCIH, the NIH Office of Dietary Supplements and the National Cancer Institute.

MedlinePlus data, record by record: topics, genes, tests, recipes
MedlinePlus Connect API: ICD-10-CM, SNOMED CT, RxNorm and LOINC lookups

MedlinePlus Connect API: ICD-10-CM, SNOMED CT, RxNorm and LOINC lookups

The MedlinePlus Connect API is the code-driven door, built for EHRs and patient portals on the HL7 Infobutton standard. Its base address, connect.medlineplus.gov/service, takes a code system OID and a code: ICD-10-CM, ICD-9-CM or SNOMED CT for problems, RxCUI or NDC for drugs, LOINC for lab tests, CPT or SNOMED CT for procedures. Answers come as Atom XML, JSON or JSONP, in English or Spanish; a Spanish drug request needs a code, not a name.

It returns a link layer rather than a corpus. A problem code brings back the matched topic or genetics page with title, URL, synonyms, full summary and attribution; a drug code brings the monograph title, URL and a short snippet; a LOINC code brings the test page and a snippet. Links carry utm tags that should be stripped before joining, and an unmatched code returns null.

Limits decide bulk use. Connect takes 100 requests a minute per IP address; past that it stops answering for 300 seconds or until the rate falls, whichever is later, and NLM asks that results be cached for 12 to 24 hours. A full crosswalk, every ICD-10-CM code or every RxCUI on a formulary mapped to its page, is tens of thousands of paced calls, rerun when mappings move. SNOMED CT coverage centers on the CORE Problem List Subset, and the SNOMED CT and ICD-10-CM keys in genetics records were chosen for Connect, so treat them as pointers.

The MedlinePlus database: NLM's consumer health library, not MEDLINE

MedlinePlus is the health information site of the U.S. National Library of Medicine, part of NIH, written for patients and families in English and Spanish. It is free, carries no advertising and endorses no product. The MedlinePlus database behind it has six public sections: health topics, drugs and supplements, genetics, medical tests, the A.D.A.M. medical encyclopedia and healthy recipes, plus easy-to-read material and translated handouts in more than 50 languages. No part of it needs an account.

The name invites a standing mix-up. MEDLINE is NLM's index of journal citations, the core of PubMed, built for researchers; MedlinePlus is plain-language reading for the public. The two meet in one place: almost every English health topic carries a Journal Articles link, a PubMed search that NLM librarians build on the topic's MeSH heading. Citations and abstracts are a PubMed project; this page is about the consumer layer.

Addresses are flat and stable. A topic is a slug such as medlineplus.gov/asthma.html, mirrored under /spanish/; a drug page is a monograph number such as /druginfo/meds/a696005.html, with an -es suffix in Spanish; genetics sits under /genetics/condition/, /gene/ and /chromosome/, inherited from Genetics Home Reference, which folded into MedlinePlus in October 2020; tests live under /lab-tests/ and encyclopedia articles under /ency/article/. Pages are static HTML, and one sitemap file lists roughly 25,600 addresses, most of them encyclopedia and drug pages.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Where the free health topics API stops and a MedlinePlus scraper starts

Every official channel covers one slice. The web service at wsearch.nlm.nih.gov searches health topics, and genetics under a separate database name, returning XML with ten hits per call by default and no more than 85 requests a minute per IP address. The XML files are complete, but only for health topics and a few glossaries. The genetics files carry only the description of a condition: frequency, causes, the inheritance narrative, the reference list with PubMed IDs and the related medical tests appear on the page alone. Medical tests and recipes have no data file at all. That is why teams scrape MedlinePlus pages even with every file in hand.

Dates are the next gap. The XML records when a topic was created, never when it last changed; the page states that in its DC.Date.Modified tag and Last updated line. The sitemap does not help, because its lastmod values are rebuild dates: a topic untouched for two years shows this month's date.

Then volume and parity. Each XML release is a full snapshot, not a delta, so finding which of roughly 53,000 curated links were added or retired means diffing releases. Spanish is not a clean mirror: MeSH descriptors exist only on English records, genetics reaches Spanish only as the Help Me Understand Genetics handbook, and the Spanish herbs index is about a third of the English one. And there is no MedlinePlus price field anywhere, so price monitoring for the drugs it describes needs a pharmacy source joined on RxCUI or NDC.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

What ScrapeIt runs on MedlinePlus, and what it leaves alone

ScrapeIt does the engineering: we write the collectors, give each MedlinePlus section its own schedule, compare every run with the one before and repair the pipeline when NLM moves a template, a file or an endpoint. You receive data, not software to babysit.

We have no link to NLM, NIH or MedlinePlus and never imply their endorsement. Licensed text stays with its publishers: ASHP monographs reach you as titles, monograph numbers, revision dates and links, the A.D.A.M. encyclopedia only as the titles and links NLM's own XML attaches to topics. MedlinePlus holds no patient data, and we collect no personal data.

FAQ

What is MedlinePlus, and is it the same as MEDLINE?

MedlinePlus is the National Library of Medicine's health information site for patients and families, in plain English and Spanish: health topics, drug information, genetics, medical tests, an encyclopedia and recipes. It is not MEDLINE. MEDLINE is NLM's index of biomedical journal citations, searched through PubMed and written for researchers. The two touch at one point: English health topics link to a librarian-built PubMed search. If your project is citations, abstracts or MeSH-indexed literature, it is a PubMed collection, and we scope it separately.

Does MedlinePlus have a public API, and what does it leave out?

Yes, several, all free with no key or registration. The web service searches health topics and genetics by keyword and answers in XML, up to 85 requests a minute; MedlinePlus Connect maps ICD-10-CM, SNOMED CT, RxCUI, NDC, LOINC and CPT codes to pages in XML or JSON, up to 100 a minute; the genetics API serves each condition, gene and chromosome page as XML or JSON; daily XML files hold every health topic. None of them returns full medical test pages, recipes, genetics sections beyond the description, a topic's last-updated date, or the licensed encyclopedia and drug monograph text.

Is MedlinePlus reliable enough to build a product on?

As patient education, yes. Health topic links are chosen against published selection criteria and checked daily for breakage, the site carries no advertising, genetics summaries are reviewed by genetics experts before posting, medical tests get a full review at least every three years, and encyclopedia and drug pages show a review or revision date. Summaries are written for lay readers, so they suit portals, apps and education products rather than clinical decision support. The practical risk is staleness on your side, which is why every row we deliver keeps the page's own modified date.

How often does MedlinePlus data change, and how is it delivered?

It depends on the section. The health topic XML and the web service refresh once a day, Tuesday to Saturday; drug pages are added and revised monthly; medical tests run on a three-year review cycle with fixes in between; genetics pages change with expert review. We follow those clocks: topics are diffed on every release, other sections are polled on their own cadence, and each run lands as a full MedlinePlus dataset plus a change file of added, edited and retired records. Formats are CSV, JSON, Parquet or XML, delivered to S3, SFTP, a database or an API endpoint.

Can MedlinePlus content be reused commercially, and what will you not collect?

The government-written parts can: health topic summaries, medical tests, genetics summaries, NLM illustrations, recipes and videos are public domain in English and Spanish. NLM asks for a credit such as Source: MedlinePlus, National Library of Medicine, with no logo and no implied endorsement. The A.D.A.M. encyclopedia and ASHP drug monographs are licensed for display on MedlinePlus only: they may not be loaded into an EHR or portal without the publisher's own license, and the encyclopedia forbids automated extraction and AI use without Ebix's written consent, so we never collect that text.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582