The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn MoreEvery alternate.de address ends in a numeric Art.-Nr, the same number opens the same item on eleven Alternate storefronts, and one mainboard record can run to fifty-eight rows of specification.
ScrapeIt runs this as a managed service. We build the Alternate scraper, host it, watch it and repair it when the site moves; you receive files, not code. Anti-bot handling, proxy rotation and CAPTCHA solving are part of the service.
Coverage follows the site's shape: product addresses from the gzipped sitemaps, category structure from the listing sitemap, detail pages refreshed on the cadence you pick. Listings page twenty-four items at a time and an overrun page quietly serves the first page again, so we walk to the stated result count rather than page until something breaks. Delivery is CSV, JSON, XLSX, a Google Sheet, a direct database load or an Alternate API endpoint we host and you call. We stay on public catalogue pages, never touch cart, checkout or account areas, and leave reviewer names and review text out by default because they are personal data; the average rating and the count come through as ordinary fields.
A product address has four parts: the brand, a slugified name, the literal segment html/product, and the number. The number is the record key. It is printed on the page as Art.-Nr., repeats as the sku in the structured data, and it survives the slug being rewritten when marketing copy changes, so the crawler keys on it and treats the slug as decoration. Older lines carry a seven digit number, newer ones a nine digit number; both are live.
Each page carries a Product block and a breadcrumb list in schema.org markup. From those we extract Alternate product data in a stable shape:
The Details table is where this catalogue earns its place in a components dataset. It is a three column grid of group, attribute and value, and it goes deep. A mid range mainboard record ran to fifty-eight rows: socket, chipset, form factor, memory slot count and type, channels, ranks and supported speed grades, the PCIe slot arrangement, M.2 count with the PCIe generation of each slot and the module lengths accepted in millimetres, SATA ports, usable RAID levels, fan and addressable RGB header counts, audio codec, LAN controller, board width and height, and a named list of every processor the board accepts down to the silicon codename. A graphics card record answers what decides a build instead: card length, occupied slot count, power connector type, average draw and minimum power supply rating.
Listing cards repeat the useful part: image, brand logo, name, a short bullet list of the specs that decide that category, price, availability wording and any promotional badge.
The PC-Konfigurator is the part of alternate.de a components dataset cannot get elsewhere. It walks the base parts in a fixed order - processor, mainboard, cooler, memory, graphics card, SSD, case, case fans, power supply and operating system - then optional parts and services, and it checks the combination before letting the order through. It also publishes a 3DMark based estimate: expected frames per second and benchmark figures for the processor, memory and graphics card chosen, averaged from real user runs. And it prices the same parts two ways: as an assembled system or as loose components. Read over time it yields a compatibility graph and an assembly premium. Two more configurators sit beside it, for Macs and for balcony solar plants, and a small own brand ALTERNATE PCs range is sold prebuilt with processor, graphics, memory and chipset spelled out on the card.
Used and open box stock is a separate slice in two flavours that should not be merged. Outlet holds B-Ware and Restposten (open box and clearance) in a mirrored category tree, and those offers are marked used in the structured data while new stock is marked new, so the split is trivial to honour. Refurbished is a different programme with a four step grade: G1 Wie neu (like new), G2 Exzellent, G3 Sehr gut and G4 Gut, applied mainly to smartphones, tablets, notebooks and monitors.
Regulated goods carry the EU energy class as a card bullet linking to the label and the product fiche. An llms.txt file, expressly permitted in the robots file, lists the whole catalogue as titled links - a cheap discovery pass for anything the sitemaps leave out.
ALTERNATE GmbH runs alternate.de from Linden in Hesse, and it is a component shop first: mainboards, processors, graphics cards, memory, power supplies and cases sit at the centre of the catalogue, with finished notebooks, smartphones and televisions arranged around them. Every offer on the site names ALTERNATE GmbH as the seller. There is no third party marketplace layer, so one product number means one seller, one stock position and one price, and an Alternate scraper never has to pick between rival sellers on the same listing.
The physical footprint is small and easy to state. Everything sits on one site at Linden: Store 1 for computing, television and audio, a separate Grill-Store, a service counter for repairs and upgrades, and pickup warehouses at Linden and neighbouring Pohlheim. Alternate is not a branch chain, so prices are not split by store and the website is the shop.
The catalogue reaches well past electronics. Beside Hardware, PC, Notebook, Smartphone, TV-Audio, Gaming, Server and Workstations and Smart Home sit Werkzeug (power tools), Garten (garden), Grill, Spielzeug (toys), Photovoltaik (solar) and Elektroinstallation (electrical fittings), plus a SimRacing department selling wheelbases, pedals, shifters and rigs.
Eleven storefronts share the same catalogue: alternate.de, .at, .ch, .nl, .be in Dutch and in French, .lu, .fr, .it, .es and .dk. The product number is identical on all of them, so one address path opens the same board in four countries while the figure moves: 139.90 euro in Germany, 141.90 in Austria, 129.00 in the Netherlands and 120.90 Swiss francs, each with its own shipping line beneath it. That turns Alternate price data into a cross border comparison keyed on a number instead of on model names.
Get a QuoteDevelopers
Customers worldwide
Pages extracted
Hours saved for our clients
€199 / one-time
setup fee - included
€169 / mo
setup fee €499
€229 / mo
setup fee €499
€349 / mo
setup fee €499
€549 / mo
setup fee €799
Nobody shops for a mainboard by name. They shop by socket, chipset, form factor, capacity, length and wattage, and that is what lifts Alternate price data above a plain price list: the Details table is structured enough to build compatibility on. Put a board's socket and its supported processor list beside the CPU catalogue and you have a validated pairing table. Put card length and occupied slots beside case clearance and you can answer whether the card fits. Put average draw and the minimum supply rating beside the power supply catalogue and you can size a build without guessing.
The commercial jobs we are asked for:
Most teams begin with two or three component categories, settle the field mapping against their own catalogue, then widen into finished goods.
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn More
Daily scraping of lowest prices for 150K products on Allegro.pl to support marketplace pricing and margin optimization.
Learn More
Regular monitoring of Ralph Lauren clothing, footwear, and accessories sold across Amazon subdomains: AE, DE, ES, FR, IT, NL, PL, UK.
Learn MoreLearn how to use web scraping to solve data problems for your organization
If you sell online, run a marketplace, or advise e-commerce clients, you already know why eBay matters: it’s one of the few places where big retailers compete side by side with thousands of small merchants and private sellers.
E-commerce teams do not just need “some” competitor data anymore. They need a continuous stream of real prices, discounts, stock levels, reviews, and seller behavior from the platforms that actually shape their markets.
Amazon provides valuable information gathered in one place: products, reviews, ratings, exclusive offers, news, etc. So scraping data from Amazon will help solve the problems of the time-consuming process of extracting data from e-commerce.
ScrapeIt is a managed web scraping agency, not a toolkit you have to operate. We already run retail crawls across Europe, so a German component feed arrives in the same shape as the rest of your price monitoring stack and joins on the same EAN and manufacturer part number keys you use elsewhere.
You get a named engineer, a sample file before anything is signed and a fixed monthly figure. Tell us the categories, the fields and the cadence, and we will send a sample drawn from alternate.de so you can test the mapping against your own catalogue first.
No. There is no developer programme, no documented endpoint and no partner product feed for buyers of data. The affiliate arrangement runs through the AWIN network and hands out banners and tracked links, not a catalogue. The robots file also puts the site's own internal service path and its JSON responses off limits to crawlers. What you can have is an Alternate API we operate for you: we crawl the public catalogue, normalise the fields and expose them as a REST endpoint or a scheduled file drop, so your systems call one stable interface even when the shop changes its own.
A nightly full sweep plus an hourly watchlist is the usual arrangement. Component pricing moves daily, and the TagesDeals section runs on a countdown that shows both the time left and how much of the promotional quantity remains, so a deal can close on quantity long before its clock runs out. For a watchlist of a few thousand numbers we go hourly; for a full catalogue sweep, nightly. Availability wording is worth the same cadence as price, because for parts it is the scarcer signal of the two.
Yes, and it is the cheapest win in the project. The German, Austrian, Swiss, Dutch, Belgian in two languages, Luxembourgish, French, Italian, Spanish and Danish sites share the catalogue and the product number, so the same path resolves on each. We return one row per storefront with its own price, currency, shipping rate and availability, joined on the number, which gives you a clean cross border comparison without any name matching. Not every line is carried on every storefront, so we mark absence explicitly rather than leaving a gap.
As deep as the page goes. The Details table is delivered as group, attribute and value triples rather than flattened into one text blob, so socket, chipset, memory channels, PCIe generation per M.2 slot, accepted module length in millimetres, RAID levels, header counts, board dimensions, card length, occupied slots and minimum power supply rating all arrive as separate addressable fields. Where a row holds a list, such as every processor a board supports, we split it into items. Units are normalised on request so millimetres, watts and gigabytes compare across brands.
We collect public catalogue pages only. No account is created, no login is used, and the paths the robots file excludes are not touched. We gather factual product data - numbers, prices, specifications, conditions and availability - rather than reproducing editorial copy or imagery wholesale. Customer review text and the display names attached to it are personal data and stay out by default; if you need sentiment we deliver the average rating and the review count instead. Grading and condition labels come through as plain fields.
Step 1 - Make a Request
You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.
Step 2 - Configuring Custom Web Crawlers
Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.
Step 3 - Collect and Deliver
Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.
Step 4 - Maintain and Support
Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.
Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582