ScienceDirect Scraper: Journal, Book and Article Data Extraction

A ScienceDirect record answers to two identifiers, a DOI and a PII, and the abstract page standing in front of the paywall already carries more usable fields than most buyers expect.

ScienceDirect Scraper
Solutions

How ScrapeIt Runs a ScienceDirect Scraping Project End to End

ScrapeIt builds the ScienceDirect scraping pipeline, hosts it, watches it and rebuilds the parts Elsevier changes. Anti-bot handling, proxy rotation and CAPTCHA solving are part of the managed service, and this project needs all three: the platform closes itself to unnamed automated clients by default and leaves only named search engines a narrow path.

We scope a run the way the catalogue is scoped - content class, then title, then volume and issue - because a title list with ISSN and coverage years is a cheaper seed than any search page. Volatile fields get their own schedule: citing-article counts and open access status move, a page range never does. A sample extract comes first, so field names, coverage and null rates can be checked before the full run. Delivery is CSV, JSON, JSONL, XLSX or an endpoint you call, pushed to S3, SFTP or your warehouse.

Every Field in a ScienceDirect Article Record, Under Its Own Name

The shape we extract ScienceDirect data into, under the platform's own field names:

  • Two identifiers. A DOI, and a PII - a seventeen-character string such as S0092867420302294 - which is the identifier the site's own URLs are built from. A normalised PII and an EID ride with it.
  • Title block. Article or chapter title, first author, the full author list with affiliation strings, and the source title of the journal, book series, handbook or reference work.
  • Abstract and keywords. Abstract text, a short teaser abstract, author keywords and publisher index terms, plus research highlights and a graphical abstract where used.
  • Where it sits. ISSN for serials, ISBN and edition for books, volume, issue number and name, first and last page, page range and article number.
  • Four dates, not one. Cover date as YYYY-MM-DD, cover display date in the publisher's wording, the date the item first went online and the separate date the version of record did.
  • Content type. Journal, book series, handbook series, ebook or reference work.
  • Article type. Research article, review, mini review, short communication, data article, case report, replication study, software publication, video article, correspondence, editorial, erratum, practice guideline, book chapter or encyclopedia entry.
  • Open access block. Open access and open archive flags, the open access type, the sponsor that paid for it and the user licence - CC BY, CC BY-NC, CC BY-NC-ND or the Elsevier user licence - beside the copyright line and publisher.
  • Cross-references. Scopus ID and Scopus EID, PubMed ID, and the reference list, whose entries link out to Crossref, Scopus and Google Scholar.
  • Attachments. Figures at thumbnail, downsampled and high-resolution sizes, plus tables, video, spreadsheets and 3D models, each with filename, mime type and size.
  • Metrics. The Scopus citing-article count and PlumX metrics split into usage, captures, mentions, social media and citations.

Author names are a bibliographic field here, as in any reference list. We deliver the byline and its affiliation and build no researcher profiles or contact lists.

Every Field in a ScienceDirect Article Record, Under Its Own Name
The Elsevier API Route: Keys, Institutional Tokens, Quotas and Open Mirrors

The Elsevier API Route: Keys, Institutional Tokens, Quotas and Open Mirrors

There is a documented ScienceDirect API, and who you are decides what it returns. Elsevier's developer portal issues a key on self-registration, free for non-commercial work by academic and public-sector researchers; commercial use needs a licence and a subscription. Full access goes only to people affiliated with a subscribing organisation, and entitlement is resolved from the caller's IP address.

Three ways to authenticate follow. A plain key rides on the institution's IP range. An authtoken settles the case where one address maps to several accounts and expires two hours after issue. An institutional token, sent as its own header, is minted by Elsevier's integration team for customers or partners acting for one, must stay server-side, never appears in browser code or an address bar, and can be revoked without notice.

Quotas sit per endpoint and reset every seven days. Search currently allows 20,000 calls a week at two per second and stops at 6,000 results per query; article retrieval currently allows 50,000 a week at ten per second. The search payload is deliberately thin - no abstract, no keywords - so a complete record costs a second call per item.

Open mirrors close part of the gap. Crossref holds over 25 million Elsevier DOIs with the PII as an alternative identifier, plus volume, issue, pages, funders and references, but almost no abstracts and almost no affiliations. OpenAlex adds ORCID, ROR-linked institutions, topics and open access colour to roughly 24.4 million Elsevier works, of which about 3.3 million carry an abstract.

How ScienceDirect Is Shelved: Journals, Book Series, Handbooks and Reference Works

ScienceDirect is Elsevier's full-text platform, and its shelves are sorted before its search box is. Five content classes carry everything the site publishes: journals, book series, handbook series, ebooks and major reference works. Elsevier's own query language codes them JL, BS, HB, BK and RW, and a separate switch splits the catalogue into serial and nonserial. A ScienceDirect scraper is scoped by content class first and by subject second, because that is the order the platform itself uses.

Addresses follow the same split. Journals answer at /journal/ plus a title slug, book series at /bookseries/, handbook series at /handbook/, single books at /book/ and encyclopedias at /referencework/. Articles have two front doors: /science/article/pii/ for the document and /science/article/abs/pii/ for the abstract-only page a reader without entitlement lands on. Term pages sit under /topics/, the catalogue browse under /browse/, and every family carries a sitemap of its own.

Counted in September 2026 the platform currently carries more than 3,000 peer-reviewed journals, 24 million articles and book chapters, up to 50,000 books from more than 50 imprints and more than 350,000 Topics pages. Browsing narrows that by domain and subdomain, publication type, journal status - whether the title still accepts submissions - access type and first letter.

Publication stage matters as much as subject. An item appears first in Articles in Press as an uncorrected proof or a journal pre-proof, becomes a corrected proof, and only then is assigned to a volume and issue as the version of record. The first-online date carries over, so identity survives the move while volume, issue and pages arrive late.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Open Access, Open Archive and Where the ScienceDirect Paywall Line Falls

The access line is drawn per document, not per journal, and there are five positions on it. Open access content is free to everyone. Open archive content becomes free after an embargo that runs from 12 to 48 months depending on the journal, under the Elsevier user licence. Subscribed content opens for an entitled reader. Non-subscribed content shows an abstract page and nothing further. A fifth case shows neither, because some items carry no separate abstract at all, and ScienceDirect prices those and the rest for single purchase through a shopping cart.

That geometry is the reason to scrape ScienceDirect metadata rather than chase full text. The bibliographic layer - identifiers, titles, abstracts, keywords, affiliations, dates, licences, open access status - stands in front of the line for a very large share of the catalogue, and 4.2 million validated open access articles sit entirely in the open. To be plain about it: we do not collect subscription full text, and we work with metadata and open access content.

What that buys is concrete. A competitive intelligence team maps which groups publish where, at what cadence, under which licence and with which funder acknowledged. A publisher measures how fast a rival journal moves from acceptance to version of record. A pharmaceutical library watches one therapeutic area across journals, book chapters and reference works at once. A discovery vendor needs an accurate title list with ISSN, coverage years and access type before it can promise a library anything.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Working With ScrapeIt on Elsevier and ScienceDirect Data

ScrapeIt is a managed data collection agency, and scholarly publishing sits among the markets we cover. You describe the slice of the catalogue you need, we build and host the collector, keep it alive through redesigns, and hand back clean files on a schedule. No crawler code and no proxy pool on your side.

We also say no. If a published title list, a Crossref pull or an OpenAlex query answers the question more cheaply, you hear that on the first call. We do not sign in with anyone's institutional credentials, we do not reach for subscription full text, and we do not assemble researcher dossiers or contact lists out of author bylines.

FAQ

Does ScienceDirect have a public API, and what does it return?

Elsevier runs a documented API family for ScienceDirect: search, article metadata, article retrieval, object retrieval for figures and multimedia, serial and nonserial title metadata, entitlement checks and holdings reports. A key comes from the developer portal on self-registration and is free for non-commercial academic use; commercial use needs a licence and a subscription. It is not an open feed. Entitlement follows the caller's IP address, search returns a thin record with no abstract, results stop at 6,000 per query, and full text is returned only where the account is entitled or the item is open access.

Which fields come back in a ScienceDirect article record?

DOI and PII, plus a normalised PII and an EID; title, first author and full author list with affiliations; source title with ISSN or ISBN; volume, issue, issue name, page range and article number; abstract, teaser abstract, author keywords and index terms; cover date, cover display date, first-online date and version-of-record date; content type and article type; the open access block with type, sponsor and user licence; copyright and publisher; Scopus and PubMed identifiers; the reference list; figures and supplementary files; and the Scopus citing-article count with PlumX metrics.

What can you collect without a subscription, and what stays behind the paywall?

Open access articles and chapters are fully public, and open archive content joins them once its embargo of 12 to 48 months has passed. For everything else the abstract page is the public surface: title, authors, affiliations, source, volume and issue, dates, keywords, licence and open access status are readable there, and on some items even the abstract is withheld. The full text of subscription articles, the PDF and the reader view stay behind the paywall. We work with metadata and open access content and do not go past that line.

How often should a ScienceDirect data feed refresh, and in what formats?

Cadence follows the field rather than the calendar. Articles in Press turn over constantly and a daily or weekly pass keeps them current, since a pre-proof gains its volume, issue and pages days or weeks later. Citing-article counts, open access status and access type deserve a weekly or monthly pass. Titles, ISSNs, editorial boards and coverage years move a few times a year. Pricing a ScienceDirect project follows the same split, because re-collecting a stable field costs what collecting a volatile one costs. Output ships as CSV, JSON, JSONL or XLSX, to S3, SFTP, a warehouse or an endpoint you call.

Is it legal to extract ScienceDirect data, and what will you not collect?

Elsevier states a machine-readable mining reservation on its pages and publishes a policy that permits mining for scientific research under the EU copyright exception and asks for consent otherwise. Its own provisions cover academic mining through the API, ask that the resulting output carry an attribution notice, and require the underlying dataset to be deleted once the project ends. So the terms shape the job rather than decorate it. We collect what is published to an ordinary visitor, we do not sign in or use institutional credentials, we do not take subscription full text, and named individuals appear only as the author byline printed on the item.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582