How to Choose a Web Scraping Vendor

The fastest way to separate a real scraping provider from a reseller is to ask what happens when the target site changes its layout on a Friday night. A real provider answers with a process - who gets alerted, how fast someone looks, whether you get a partial delivery or a late one. A reseller answers with a sentence about robust infrastructure. Everything below is a variation on that question.

The questions that separate a provider from a reseller

Most of what is sold as "web scraping" is one of three things: a proxy or unblocker API, a self-serve tool you operate, or a managed service where someone else owns the outcome and you never touch infrastructure or code. All three are legitimate, and vendors are vague about which one they sell, because "we handle it" sounds better than "you handle it with our proxies."

  • The Friday night question. A layout change ships at 22:00 on a Friday. Who notices - a person, a monitor, or you on Monday when the file is empty? What arrives in the meantime?
  • Who writes the parser? If the answer is "you configure it in our dashboard", you are buying a tool and the maintenance is yours. That may be fine, but it should not be a surprise.
  • Have you scraped this specific site? Not this category - this site. A vendor with a real catalog answers immediately and shows you the output.
  • What breaks most often on a site like mine? Anyone who has run crawlers at volume has opinions: pagination caps, geo-gated pricing, lazy-loaded fields, A/B tests that change the DOM for part of the traffic. A vendor with no opinions has no scars.
  • What happens on partial failures? The dangerous outcome is not a crawler that dies loudly. It is one that quietly returns a fraction of the rows with no error, and nobody notices for weeks.

Standardizing five European property portals into one schema at 230,000 rows a day is the scale at which those answers get tested.

Pricing models, and why per-row pricing misleads

Per-row and per-request pricing looks like the fairest model on the page and is usually the least predictable. Three reasons:

  • Row counts are not under your control. The site adds listings, changes pagination, or duplicates inventory, and your bill moves without your business changing.
  • Retries and blocks are where the cost lives. A hard site burns far more requests per delivered row than an easy one. Per-row pricing either hides that in a high unit price or exposes you to it in the fine print.
  • Maintenance is not a row. The hour spent rewriting a parser after a redesign has no unit to bill against, so per-row quotes omit it - and it arrives later as a change request, or not at all.

Ask instead for a fixed monthly figure tied to a defined scope - these sites, these fields, this frequency - with maintenance inside it, and a stated rule for what triggers a new quote. ScrapeIt plans start from EUR 199. If one quote is much cheaper than the others, find out what is excluded before you call it a bargain - how web scraping cost actually works covers what usually sits outside it.

What a sample must prove, and how to test a vendor cheaply

Every vendor will send a sample. Most prove only that clean rows can be pulled from the easiest pages on the site. Make the sample do real work:

  • You pick the URLs. Include the ugly cases - a listing with missing fields, a page in a secondary language, something deep in pagination.
  • Check the fields you will join on. IDs, timestamps, currency, units. A price column mixing "1.299,00" and "1299.00" costs you more downstream than a missing column, because it fails silently.
  • Ask what a null means. Field absent, field present but empty, extraction failed - three different things that must not collapse into one blank cell.
  • Ask for the same URLs twice, a day apart. Rows that appear and vanish without the source changing are the signature of an undertuned crawler.
  • Confirm the delivery format is the production one - CSV, JSON, XLSX, straight into your database, or via API. Sample-by-email and production-into-your-warehouse are different problems.

You can look at real extracted data before talking to anyone. Then skip the three-month evaluation and buy one small, real deliverable:

  1. Pick one site and a narrow field set, then check a few dozen rows against the source by hand. It finds what no vendor documentation will.
  2. Set a delivery date, then say nothing more about it. Whether they warn you before the deadline that something slipped tells you the most.
  3. Change something mid-flight - a field, the frequency, the format. How a small change is handled predicts how the big ones will be.
  4. Then scale. Swiss private listings into a client CRM via API at 85K rows a day went live in 7 days, and 250,000 iHerb products were delivered end to end in 3 days. Ask a vendor what its equivalent is.

SLAs, and what they are worth without a maintenance commitment

An uptime SLA on a scraping contract is close to meaningless on its own. The vendor's servers being up says nothing about whether the target site changed and the parser is writing empty strings into a live pipeline. Delivered, correct data is the metric. What is worth negotiating:

  • Delivery window - the file lands by a stated time on a stated schedule, and someone tells you when it will not.
  • Break-fix commitment - a target-site change is fixed within a named window, and that work is included, not quoted separately. This is the clause that matters most and the one most often missing.
  • Completeness checks - agreed thresholds on row counts and required fields, with an alert when a run falls outside them.

Layout changes are not an edge case, they are the ongoing cost of the service - what maintenance involves sets out the month-to-month reality.

The red flags

  • Vague answers on legality. "Everything on the internet is public" is not a compliance position. A serious vendor can tell you what it collects and what it refuses to collect. Ours is written down on the compliance page: public data only, nothing personal beyond what the site publishes itself, GDPR and CCPA posture, confidentiality agreements, and a list of what we will not do.
  • No named contact. If you cannot name the person who answers at 09:00 on the day a delivery is wrong, you do not have a vendor, you have a form.
  • No sample, or a sample only on their chosen URLs. Refusal here usually means the hard pages were never solved.
  • Pricing that never mentions maintenance. The quote is for a build, not a service, whatever the proposal calls it.
  • Promises to bypass logins or paywalls. Treat this as disqualifying. A vendor offering to defeat authentication will take your instructions over the law and its own contracts, and the exposure lands on you as the buyer. That is a liability dressed as a capability. Data behind an account is not always off the table, but it is a question of who holds a legitimate right of access, not of whether the vendor can break in.

When hiring a vendor is the wrong call

Sometimes you should not outsource this, and a vendor who never says so is selling rather than advising.

  • One stable site, a few thousand rows, once a month. A script and an afternoon beats a contract. Revisit when the number of sites, not the number of rows, starts growing.
  • Scraping is the product. If extraction is your core differentiator, the expertise belongs in-house even if it is slower and costlier at first.
  • You do not know what you need yet. A spec that changes weekly is badly suited to fixed-scope delivery. Prototype it yourself, then hand over the stabilized version.
  • You have idle in-house capacity and a tolerant deadline. The honest calculation is engineering time against monthly fee, and it does not always favor buying. Build vs buy works through that comparison.

Where outsourcing wins is breadth and endurance: many sites, one schema, running for years without your team owning the breakage. If that is your situation, put the same test site and field list in front of two or three vendors, ask us for a quote too, and bring the Friday night question with you.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582