The Entire Iherb Supplements Catalog, Captured End-to-End in 3 Days
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn MoreA Baker Creek packet costs 3 to 5 dollars whatever is inside it, and inside it may be 15 seeds or 250. We pull the rareseeds.com catalog into one table you can sort.
We build and run the Baker Creek scraper as a managed job. You name the crops, the categories or the whole seed catalog; we design the schema, write the crawl, run it weekly, daily or by season, and hand back clean rows. Anti-bot handling, proxy rotation and CAPTCHA solving are part of the service, so access is our problem rather than yours. Delivery is CSV, JSON, Excel, a feed into your warehouse or an endpoint we host. The usual shape is one row per variety with the growing bullets split into their own columns, a computed cost per seed, and a collection date on every row so availability and price history accumulate instead of overwriting themselves. We take only what any visitor sees, and reviewer names and the growing zones posted beside them are personal data we leave out by default.
Names come in two parts with the crop first: Tomato Seeds, Abe Lincoln Original; Potato Tubers, Russet Burbank; Garlic Bulbs, Elephant (1 lb). Pack quantity, where there is one, is written into the name and the address instead of being an option, so garlic-inchelium-red-4-bulbs is a separate item, not a size. Here is what we extract for Baker Creek data:
What is not there: germination percentage, lot number, packing year, organic certification. The seed is untreated but not certified organic, so no organic flag exists to collect.
Sold out is a state, not a deletion. The variety keeps its address, its SKU, its price and its reviews, the availability flag turns over, and the buy button becomes a subscribe form for an in-stock notice. On listing pages the difference is structural: an orderable variety is drawn as an add-to-cart form carrying SKU and price, while a sold-out one is a plain block holding only the id, the name and a link, labelled Out of Stock. Read the cart forms and you lose the varieties worth watching.
Seasonality shows in the counts. Tomato Seeds currently holds 119 varieties with 5 gone, the 2026 Introductions shelf holds 125 with 9 gone, garlic has 4 of 5 orderable, and all 8 potato and sweet potato items are out because the tuber season is shut. A pre-order is announced inside the product name, as in PRE-ORDER 2027 The Whole Seed Catalog. There is no rare, exclusive or best-seller badge anywhere; the only structural handle on novelty is that introductions shelf.
Addressing is flat, and it decides the coverage plan. Categories live under /store/plants-seeds/, a variety sits at the root as /tomato-abe-lincoln-original with no prefix and no numeric id, and guides share that same root, so an address will not tell a variety from an article. Some slugs invert the name: dester-tomato, pork-chop-tomato. Listing pages take ?p=N and accumulate rather than replace, so page N returns the first N times 24 tiles and any page past the end returns the whole category in one document. The item map is incomplete as well, with live tomato varieties missing from it, so coverage is built from the 193 category pages and from search.
Baker Creek Heirloom Seed Company trades at rareseeds.com out of Mansfield, Missouri, and it sells one kind of product: open-pollinated seed. No hybrids, no treated seed, nothing you are obliged to buy again next year. Jere Gettle printed the first catalog in 1998, and the house line since then is that seed belongs to the people and every packet can be saved, shared and replanted. That single rule is what makes Baker Creek data unlike a garden-center feed. A row here is a lineage with a person, a place and a decade attached, not a product code with a colour swatch.
The store currently lists about 1,496 items, and the company puts its own count at more than 1,350 heirloom varieties. Vegetable Seeds holds roughly 918 of them, Flower Seeds about 463, Herb Seeds 109 and Bulbs 38, with 43 more in gifts and supplies. The tree runs three levels deep - Plants and Seeds, then Vegetable, Flower, Herb or Bulbs, then one category per crop - and beans, peppers and squash split again into a fourth level, so Bean Seeds carries Common, Fava, Garbanzo, Hyacinth, Lima, Long, Runner, Soy and Winged underneath it.
Around the shop sit the parts a seed buyer recognises: the printed Whole Seed Catalog, spring and fall festivals, the seed store in Mansfield, the Petaluma Seed Bank in California and Comstock-Ferre in Wethersfield, Connecticut, trial grounds in Jamaica, California and Kenya, and roughly 406 articles split into Growing Guides and Seed Stories. Shipping inside the United States is free, most overseas orders are capped at ten packets, and Australia, Chile, New Zealand and the EU are not served at all, which is why this catalog reads as a United States assortment.
Get a QuoteDevelopers
Customers worldwide
Pages extracted
Hours saved for our clients
€199 / one-time
setup fee - included
€169 / mo
setup fee €499
€229 / mo
setup fee €499
€349 / mo
setup fee €499
€549 / mo
setup fee €799
This catalog prices almost everything between 3.00 and 5.00 dollars, and the sticker barely moves between a common lettuce and a variety a handful of growers keep alive. What moves is the packet. Little Gem lettuce is 3.00 dollars for at least 250 seeds, near 1.2 cents each. Moonshadow hyacinth bean is 4.50 for 15 seeds, about 30 cents each. Glass Gem popping corn is 5.00 for 75. The same shelf carries a twenty-five-fold spread in cost per seed, and the site never prints that column.
Baker Creek price monitoring is therefore arithmetic before it is anything else. Divide the price by the minimum seed count and the catalog sorts itself into what is cheap to trial and what is being rationed. Watch those two columns week to week and a packet count cut without a price change becomes visible, which is a quiet increase nobody announces.
For a seed company measuring itself against the heirloom shelf, the questions are about assortment. Which varieties exist here and not on your list. Where a rival puts 250 seeds in a packet and you put 100 for the same money. Which crops sit at the top of the band and which never leave the floor. Which of your own lines have no counterpart here at all, and which of theirs your buyers keep asking for. An absent row is a finding, and it is the finding a category manager can act on inside one season.
Scraping supplement products from iHerb.com with full details, including descriptions and packaging variations.
Learn More
Daily scraping of lowest prices for 150K products on Allegro.pl to support marketplace pricing and margin optimization.
Learn More
Regular monitoring of Ralph Lauren clothing, footwear, and accessories sold across Amazon subdomains: AE, DE, ES, FR, IT, NL, PL, UK.
Learn MoreLearn how to use web scraping to solve data problems for your organization
If you sell online, run a marketplace, or advise e-commerce clients, you already know why eBay matters: it’s one of the few places where big retailers compete side by side with thousands of small merchants and private sellers.
E-commerce teams do not just need “some” competitor data anymore. They need a continuous stream of real prices, discounts, stock levels, reviews, and seller behavior from the platforms that actually shape their markets.
Amazon provides valuable information gathered in one place: products, reviews, ratings, exclusive offers, news, etc. So scraping data from Amazon will help solve the problems of the time-consuming process of extracting data from e-commerce.
ScrapeIt is a managed web scraping agency. We have been running production crawls for retailers, marketplaces, pricing teams and analysts for years, and the job is the same every time: agree the fields, prove them against the live site, keep the crawl alive when the site changes, and deliver on a schedule you can plan around. You do not maintain code, rotate proxies or chase a broken selector at midnight. You get a sample first, then the dataset, then the same dataset again on the day you asked for it.
Baker Creek publishes no documented data API and runs no developer program. The storefront sits on a Magento base, so a GraphQL layer exists and it answers questions about the category tree, giving names, paths and how many products sit on each shelf, which is genuinely useful for planning coverage. It does not hand product records to outside callers, and the merchant REST endpoints are closed behind a token. The search box has a suggest service returning a thin JSON slice for a few matches: title, link, price, a stock flag and a rating score. None of that is a catalog feed, so the variety table has to be assembled from the public pages.
Variety name, crop, Latin name, SKU, numeric store id, page address, category path, price, minimum seed count, a computed cost per seed, stock state, and the growing bullets split into columns: days to maturity, days to sprout, sun hours, seed depth, plant spacing, ideal temperature, frost hardy, growth habit, height at maturity and life cycle. We also take the description and the seed-saving notes as text, the star rating, the review count and the image addresses. Fields the site does not publish, such as germination rate, lot number, packing year and organic certification, stay empty rather than guessed.
Coverage is built from the category tree rather than the item map, because the map is missing live varieties. We walk all 193 category pages down to the fourth level, take each category in a single deep request since listing pages accumulate rather than replace, and reconcile the tile count against the count the shelf itself reports. Search fills the last gaps, mostly items whose address does not begin with the crop name. Every run is compared with the previous one, so a variety that appears, disappears or changes its packet count is reported rather than silently overwritten.
Weekly suits most assortment work, daily makes sense from January through spring when the shelves move fastest, and a late-summer run catches the new introductions. Sold-out varieties are recorded, not dropped: the page stays up with its price and SKU, and we keep the row with the stock state set and the collection date attached. That is how you get a stock-out history, showing how long a variety was gone, whether it came back, and which crops empty out every year at the same point in the calendar.
We collect public catalog pages that any visitor can open, at a polite rate, and we honour what the site asks for. There is no separate terms-of-use page here; the robots file is the stated position, and it allows general crawling while refusing the use of content for training AI models and blocking a named list of AI crawlers. We respect both, so your dataset is for assortment and pricing work rather than model training. We stay out of checkout, account, wishlist and review areas, and buyer names and the growing zones posted with their reviews are left out of the delivery by default. If your use case needs a legal review, we say so before the first run.
Step 1 - Make a Request
You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.
Step 2 - Configuring Custom Web Crawlers
Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.
Step 3 - Collect and Deliver
Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.
Step 4 - Maintain and Support
Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.
Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582