Empik Scraper for Book Metadata, Prices and Seller Offers

Empik Scraper
Solutions

How the data reaches you

An Empik scraper we run is a scheduled job, not a one-off export. You set the scope - a publisher list, an ISBN list, a set of categories or the full book department - and the cadence. We run the crawl, normalise the fields, diff each run against the previous one and deliver CSV, JSON, XLSX, a Google Sheet, an S3 drop or a REST API your systems read on their own.

Change feeds are usually more useful than full dumps. We can send only the rows where price, seller count, buy box holder or availability moved since the last run, which keeps daily files small enough for a spreadsheet or a BI tool with no pipeline in between.

Volume, cadence and field list are set per project. Empik price tracking across a few thousand ISBNs runs differently from a full catalogue sweep, and we size the job to the one you need.

What we extract from Empik

We extract the fields below from Empik by default and add more on request. Polish labels are quoted as they appear on the page, because those strings are what a crawler matches.

  • Identifiers - the numeric product id from the URL plus "ISBN" and "EAN". On the book pages we checked, both fields carried the same 13-digit value.
  • Bibliographic detail - "Autor" (author), "Tłumaczenie" (translation), "Wydawnictwo" (publisher), "Seria" (series), "Numer wydania" (edition number), "Data premiery" (release date), "Liczba stron" (page count), "Oprawa" (binding), "Format", "Język wydania" (language of the edition) and "Język oryginału" (original language).
  • Price - the "cena" (price) charged in zł, the "cena sugerowana przez wydawcę" (publisher's suggested price) printed beside it, and the "najniższa cena z 30 dni" (lowest price in the last 30 days) line on listing tiles.
  • Badges - "Megacena", "Gwarancja najniższej ceny" (lowest price guarantee), "Promocja" (promotion), "Nowość" (new) and "Tylko w empiku" (only at Empik).
  • Dostępność (availability) - "Dostępny w salonie" (available in a salon), "Odbierz za 2 godziny" (collect in 2 hours), "Przewidywany czas wysyłki" (expected dispatch time), "Zapowiedzi" (announced titles) and "Produkt niedostępny do zakupu przez internet" (not available for purchase online).
  • Seller - the merchant line, the rival offer counter, and the cheapest competing price where the page shows one.
  • Opinie (reviews) - the rating as a decimal written with a comma, the review count in brackets, and review text where present.
  • Media and measurements - image URLs, the description block, and "Wysokość", "Szerokość" and "Głębokość" (height, width, depth).

Collection covers public product and pricing data. No customer data is collected.

What we extract from Empik
What Empik scraping has to handle

What Empik scraping has to handle

Empik publishes hub pages that most retailers do not. The sitemap index at empik.com/sitemap.xml points at separate files for author pages, publisher pages and publishing series pages, next to category and filter pages. Those hubs are a cheap way to enumerate a publisher's whole Empik presence, or to follow a series, without walking the entire catalogue.

Ranked bestseller lists are public. The books chart runs to a hundred positions with the rank printed beside each title, so a daily crawl builds a real time series of chart movement instead of a proxy assembled from sales estimates.

Pagination is worth stating plainly. Listing pages advance with a start parameter that steps in 60s, and a page parameter returns 404. On 31 August 2026 a listing reporting 220745 "produktów" (products) offered a pager running to 3680 pages, which matches that total at 60 per page, so the result set is walkable rather than truncated. Sorting is set with "Trafność: największa" (best match), "Cena: od najniższej" (price, lowest first), "Cena: od najwyższej" and "Popularność: największa" (most popular). One limit is worth knowing up front: Empik Premium prices appear only after a shopper signs in, so a public crawl records the public price and nothing else.

How Empik works as a data source

Empik is a Polish retailer built around books, music, film and press, with a salon network across the country and a large online catalogue at empik.com. That catalogue now reaches well past media into toys, electronics, fashion, stationery, home goods and beauty, but the book side still shapes the data. Empik reported 370 salons in Poland at the end of 2024 and set a target close to 400 for the year after.

Two kinds of offer share the same site. Empik sells its own stock, and EmpikPlace hosts third-party merchants who list against the same product records. A product page names its merchant outright, with a line reading "Sprzedaje firma Empik" (sold by the company Empik) when Empik itself is the seller. Where other merchants list the same item, the page carries a counter such as "Produkt sprzedają inni sprzedawcy (15)" (this product is sold by other sellers, 15 of them) together with the cheapest alternative offer. Listing pages expose a "Sprzedawca" (seller) facet that names Empik alongside individual marketplace shops.

Product URLs keep a fixed shape: a slug, a comma, the letter p with a numeric product id, then a category suffix. A book reads ,p1597338509,ksiazka-p and a toy reads ,p1431212244,zabawki-p. That id is stable and makes a clean join key for anything you store.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why an ISBN-keyed catalogue is worth crawling

An ISBN is the reason to scrape Empik rather than treat it as one more Polish shop. The same 13 digits identify the same edition at a distributor, in a library catalogue, in a publisher's own list and at every competing bookseller. Match on it and Empik data joins cleanly to whatever else you hold, with no fuzzy title matching and no guessing about which of five reprints you are looking at. Very little of general retail hands you a key like that.

The publisher's suggested price carries the second argument. Empik prints "cena sugerowana przez wydawcę" next to the price it actually charges, so a book page holds the list price and the street price side by side. Discount depth per title, per publisher or per series falls straight out of that pair, and you never have to rebuild a reference price from history.

The third is the seller mix. Because the Empik offer and marketplace offers share one product record, a single crawl shows where Empik holds the buy box, how many merchants sit behind it and what the cheapest of them asks. For a publisher or a distributor that reads as a channel map: which titles resellers are undercutting, and where unexpected stock has surfaced.

Related Case Studies

The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days

The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days

Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.

Learn More about The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days
The Lowest Allegro Prices from 150K Eans Collected

The Lowest Allegro Prices from 150K Eans Collected

Daily scraping of lowest prices for 150K products on Allegro.pl to support marketplace pricing and margin optimization.

Learn More about The Lowest Allegro Prices from 150K Eans Collected
Ralph Lauren Monitoring on Amazon, 8 Markets Scanned Into One Clean Dataset

Ralph Lauren Monitoring on Amazon, 8 Markets Scanned Into One Clean Dataset

Regular monitoring of Ralph Lauren clothing, footwear, and accessories sold across Amazon subdomains: AE, DE, ES, FR, IT, NL, PL, UK.

Learn More about Ralph Lauren Monitoring on Amazon, 8 Markets Scanned Into One Clean Dataset
Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

6 E-Commerce Sites Like eBay to Scrape in 2026

6 E-Commerce Sites Like eBay to Scrape in 2026

If you sell online, run a marketplace, or advise e-commerce clients, you already know why eBay matters: it’s one of the few places where big retailers compete side by side with thousands of small merchants and private sellers.

Top 8 E-commerce Websites to Scrape in 2026 (From Amazon to 1688)

Top 8 E-commerce Websites to Scrape in 2026 (From Amazon to 1688)

E-commerce teams do not just need “some” competitor data anymore. They need a continuous stream of real prices, discounts, stock levels, reviews, and seller behavior from the platforms that actually shape their markets.

How to Scrape Amazon Data: Benefits, Challenges & Best Practices

How to Scrape Amazon Data: Benefits, Challenges & Best Practices

Amazon provides valuable information gathered in one place: products, reviews, ratings, exclusive offers, news, etc. So scraping data from Amazon will help solve the problems of the time-consuming process of extracting data from e-commerce.

scrapeit logo

Who runs the crawler

ScrapeIt is a managed web scraping agency. We write the crawlers, host them, watch them and repair them when Empik changes its markup, so nobody on your side maintains scraping code. You receive data on a schedule, in the format your team already works with.

Scope, fields and delivery are agreed before we start and can be adjusted later without a rebuild. Tell us what you need from Empik and we will scope it.

FAQ

Does Empik have a public API for product and price data?

No. There is no public catalogue or pricing API for reading empik.com. The API that does exist sits on the seller side: EmpikPlace runs on Mirakl, and a merchant with a seller account can generate an API key to push offers and pull their own orders. It is scoped to that merchant's own listings, so it answers nothing about competitors, other sellers or the catalogue at large. Empik's robots.txt also disallows the site's own GraphQL gateway paths, /gateway/api/graphql/products and /gateway/api/graphql/cart, so those are not a route either. Public product and listing pages are what our crawlers read.

Can you scrape Empik from a list of ISBNs?

Yes, and it is the usual way these projects start. You send ISBNs, we resolve each one to its Empik product page, and we report back the ones with no match so your list stays honest. Formats are kept apart: a printed edition, an ebook and an audiobook of the same title sit on separate product pages with separate ids, so we flag which format each row belongs to rather than collapsing them into one line.

Does Empik show whether a book is available in a physical store?

Product pages carry "Dostępny w salonie" (available in a salon) and, for qualifying items, "Odbierz za 2 godziny" (collect in 2 hours). That is a site-level signal about collection being offered, not a per-branch stock count. Choosing a specific salon happens inside checkout, which we do not enter. If per-store detail matters to you, say so at scoping and we will tell you plainly what the public pages support.

How do you tell Empik's own offers apart from marketplace sellers?

The product page states it. When Empik is the merchant the page reads "Sprzedaje firma Empik", and when third parties list the same item it shows a counter such as "Produkt sprzedają inni sprzedawcy" with the lowest alternative price. Listing pages carry a "Sprzedawca" (seller) facet naming Empik and individual marketplace shops. We store the seller name on every row, so you can filter to Empik's own assortment, to marketplace supply, or to both.

How often can you refresh Empik prices, and in what format do we get the data?

Cadence is a scoping decision: daily suits price and availability work, weekly or monthly suits catalogue and metadata work, and a narrow ISBN watchlist can run more often than a full department sweep. We agree it against your scope rather than quoting a fixed number. Delivery is CSV, JSON, XLSX, a Google Sheet, an S3 or SFTP drop, or a REST API, with a change feed if you would rather receive only the rows that moved.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582