The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn MoreOne eight-digit id carries the same product across eleven Lookfantastic storefronts and three currencies. We scrape Lookfantastic shade by shade and hand back a feed that already lines up.
This runs as a managed service, not a script handover. Coverage is planned from the category tree with its page parameter at forty products a page, then topped up from the product sitemap, which currently lists about sixteen thousand four hundred addresses, so nothing hides on a shelf that no menu links to. For markdown work we sweep the offer branch, roughly 4,800 discounted lines deep, sorted by discount so the deepest cuts land first. Anti-bot handling, proxy rotation and CAPTCHA solving are part of the service, and so is the retry logic that keeps a shade-level crawl internally consistent. Output arrives as CSV, JSON, JSONL or Parquet over S3, GCS, SFTP or a webhook, daily or hourly, with a change feed so you read only what moved. There is no public Lookfantastic API for buyers, so we deliver the feed such an API would have given you.
The record here is unusually full, and an extract Lookfantastic product data job can take almost all of it from public pages, with no account and no basket.
Coverage is honest rather than uniform. Description and directions sit on nearly every line, the INCI list on most of them, the key-ingredient block on roughly half, and accreditations on a minority. We report the fill rate per field for your slice instead of promising a full grid, and we flag shades that sit in the option list but read as sold out, because those vanish and return without warning.
Beauty retail sells bundles, and bundles break naive price comparison. Around forty addresses sit under the beauty box branch, and the title states the contents value out loud: a monthly box selling at GBP 13 is billed as worth over GBP 55, and the seasonal calendars are billed as worth over GBP 665 and over GBP 469. Set those beside a single moisturiser and the comparison is nonsense. A usable feed carries a set flag, the declared contents value and the selling price as three separate columns and leaves the arithmetic to the buyer.
Subscriptions add a second trap. The monthly box keeps one product id while its slug and its contents change every month, so a naive price history for that id reads as a wildly volatile product when nothing about it is comparable month to month. The same id also sells as one, three, six and twelve month terms at different totals: a term ladder, not a size ladder.
Promotions stack. The saving percentage is measured against RRP, not against yesterday's price. Some cuts only land when a code is typed at checkout, so the shelf price and the paid price part company. Offer landing pages are cut by channel, with separate affiliate, email, paid search and student versions, and tiered spend thresholds and choose-your-gift pages sit on top. We record which mechanic produced each price instead of flattening them into one number.
Two things are not published and we do not invent them: period after opening, and batch or expiry dates, which live only on the pack. Reviews carry buyer display names, which is personal data, so by default we keep score, date and text and drop the author.
Lookfantastic is a British beauty retailer inside the THG group, selling skincare, make-up, haircare, fragrance, bodycare, men's grooming and beauty electricals. For anyone who needs to scrape Lookfantastic, what matters is how plainly the shop is addressed: two trees hold the whole catalogue, and one number holds every product.
Category pages sit under /c/health-beauty/ and fan out into face, make-up, hair, body, fragrance, men, suncare, nails, electrical, gifting and Korean beauty, three or four levels deep. Brand pages sit under /c/brands/, and the brand directory carries roughly 400 brand hubs, most with shelves of their own: Acqua di Parma alone splits into fragrances, eau de parfum, hand and body, and gift sets.
Products live at /p/slug/id/ behind an eight-digit id. The slug is decoration. Put the wrong slug in front of the right id and the site quietly redirects to the canonical address, so a crawl keyed on the id survives every marketing rewrite of a title. Adding a variation parameter swaps the page to one shade or one size while the canonical still points at the parent.
Skincare adds an axis ordinary retail does not have: browsing by active. A Skincare Ingredients hub fans out into AHA and BHA, azelaic acid, ceramides, collagen, ectoin, glycolic acid, hyaluronic acid, niacinamide, peptides and retinal. Those shelves are where a lot of Lookfantastic data demand actually starts.
Get a QuoteDevelopers
Customers worldwide
Pages extracted
Hours saved for our clients
€199 / one-time
setup fee - included
€169 / mo
setup fee €499
€229 / mo
setup fee €499
€349 / mo
setup fee €499
€549 / mo
setup fee €799
Eleven Lookfantastic shop fronts share one address scheme and one product id. The UK site prices in GBP. The Irish, German, Austrian, Dutch, Italian, Spanish, French and Polish sites and the pan-European storefront price in EUR. The Gulf site prices in AED. The path is identical down to the slug, so Lookfantastic price data joins across borders on the id alone, with no fuzzy matching, no title cleaning and no brand dictionary.
What differs is everything after the join. One foundation carries thirty-four shades on the UK, German and Spanish sites and twenty-six in Ireland. The euro ladder is not one ladder either: Irish and pan-European prices for the same shade sit above the German and Italian ones. French and Polish storefronts simply do not hold every line the UK holds, and that gap is a clean assortment signal rather than a fault. Even the filters admit the split, because the price and savings facets are keyed per currency, so each storefront sifts against its own money.
Australia is the exception worth knowing before anyone quotes a project. The Australian shop is a separate build on a different platform, with its own product handles, vendor part codes and AUD prices, and no shared id at all. It has to be matched by barcode and title. That single fact decides whether a global Lookfantastic scraping brief runs a week or a month.
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn More
Daily scraping of lowest prices for 150K products on Allegro.pl to support marketplace pricing and margin optimization.
Learn More
Regular monitoring of Ralph Lauren clothing, footwear, and accessories sold across Amazon subdomains: AE, DE, ES, FR, IT, NL, PL, UK.
Learn MoreLearn how to use web scraping to solve data problems for your organization
If you sell online, run a marketplace, or advise e-commerce clients, you already know why eBay matters: it’s one of the few places where big retailers compete side by side with thousands of small merchants and private sellers.
E-commerce teams do not just need “some” competitor data anymore. They need a continuous stream of real prices, discounts, stock levels, reviews, and seller behavior from the platforms that actually shape their markets.
Amazon provides valuable information gathered in one place: products, reviews, ratings, exclusive offers, news, etc. So scraping data from Amazon will help solve the problems of the time-consuming process of extracting data from e-commerce.
ScrapeIt is a managed web scraping agency. We build the crawler, run it on our own infrastructure, watch it when the site changes, and send data instead of code. Beauty is a shape we know: shade and size variants, INCI blocks, unit pricing and set logic get the same careful treatment on every retailer we cover. Scope, sample and schema come before any invoice, and the sample is real rows pulled from the live shop, never a mock.
No. Lookfantastic runs an affiliate programme that pays commission on links, but there is no documented public API that hands a buyer product, price, stock or ingredient data. The storefront talks to an internal service that is not offered to third parties and is not stable ground to build on. Our crawler reads the same public pages a shopper sees and returns a versioned feed with a fixed schema, so no field renames itself underneath your pipeline.
Yes, and for Europe and the Gulf it is cheap, because the UK, Irish, German, Austrian, Dutch, Italian, Spanish, French, Polish, pan-European and UAE storefronts all address a product with the same eight-digit id. Cross-border comparison joins on that id, and we return one row per storefront with its own currency, selling price, RRP, saving and stock state. Australia is the exception: it is a separate build with its own handles and vendor codes, so those rows are matched by barcode and title and priced in AUD.
The Full Ingredients List is published as INCI on most lines and comes back as raw text plus a parsed array, so you can search for retinal, niacinamide or a preservative across the catalogue. The key-ingredient block, which names an active and says what it does, sits on roughly half. Vegan and cruelty-free are not a simple flag: they appear as named third-party accreditations such as The Vegan Trademark or PETA Cruelty-Free, on a minority of lines, and we return them exactly as named rather than collapsing them to true or false.
Daily is the usual cadence for a full catalogue sweep. Price and stock on a watchlist can run hourly, which matters here because coded offers and tiered promotions turn over inside a day. Cost tracks shade count rather than product count: a foundation with thirty-four shades is thirty-four price points, not one. We normally combine a nightly full pass with frequent passes over the offer branch and a chosen brand list, and ship a change feed so you only process what actually moved.
We stay on public pages. No account, no basket, no checkout, and no paths the site asks crawlers to leave alone, which on this shop covers site search, facet-filtered listings and sort-order URLs. We read those shelves through the plain category tree instead. Reviewer display names are personal data and are dropped by default, leaving score, date and text. Anything the shop does not publish, such as period after opening or per-store stock, we say is unavailable rather than estimating it.
Step 1 - Make a Request
You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.
Step 2 - Configuring Custom Web Crawlers
Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.
Step 3 - Collect and Deliver
Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.
Step 4 - Maintain and Support
Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.
Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582