Do Web Scraping Vendors Sign NDAs and DPAs

Yes - a serious scraping vendor will sign a mutual NDA, and will sign a DPA when the engagement actually involves personal data. The harder part of procurement is not the signature. It is knowing which document you need, what it should cover in a scraping engagement, and whether the vendor's answers on subprocessors, retention and deletion survive legal review.

What the NDA is actually protecting

Most buyers assume the NDA protects the scraped data. Usually that is not the valuable part. Public product prices, public listings, public job postings - nobody can be prevented from knowing this data exists, because anyone with a browser can see it. The confidential asset in a scraping engagement is almost always your specification.

The target list tells a competitor which markets you are watching. The field spec tells them what you consider a signal. The frequency tells them how fast you react. A brief that says "these marketplaces, EAN-level, refreshed daily, lowest price only" is a readable summary of a pricing strategy. When we monitor 150K EANs daily on Allegro.pl for lowest price, the interesting thing to an outsider is not the prices - it is the EAN list.

The NDA clauses that matter most here:

  • Scope of confidential information - make sure it names the target site list, the field schema, sample outputs, and the delivery cadence, not just "data provided by the Client". A target list you gave verbally may not be covered.
  • Mutuality - the vendor shares its own methods and architecture with you during scoping, so a one-way NDA often stalls at the vendor's legal review. Mutual is faster.
  • Publicity and case studies - if you do not want to be named, say so in the agreement, not in an email. Some vendors publish client names by default.
  • Survival period - a target list stays commercially sensitive long after the contract ends. Two years is usually too short for a competitive-intelligence use case.
  • Return or destruction on termination - and specify what "destruction" covers: working copies, intermediate storage, logs, and backups each have different lifecycles.

ScrapeIt's compliance page states the confidentiality position and the willingness to sign agreements, but a page is the starting point, not the contract. If your legal team has a template, send it.

When you need a DPA - and when you do not

A data processing agreement is required when the vendor processes personal data on your behalf. That is the entire test. It does not turn on whether the work is called "scraping", how much data there is, or how nervous the project makes people feel.

A large share of commercial scraping does not involve personal data at all:

  • Product catalogs and prices - iHerb supplements, Amazon listings, marketplace EANs. Product attributes are not personal data.
  • Business listings and company directories - a company name, a business address, a registration number. Careful: a sole trader's business contact details can still be personal data in the EU, where the identifier points to a natural person.
  • Property listings without agent contact details - address, price, size, photos. When we standardize 230,000 rows a day from five European property portals, the schema is the listing, not the person who posted it.
  • Travel inventory - hotel rates, flight fares, availability. No data subjects involved.

You do need a DPA when the field spec includes named agent or seller contacts, reviewer profiles and review text tied to a person, or job postings naming the hiring manager. If you are uncertain, look at the field list rather than the site: the same real estate portal is a no-personal-data project or a personal-data project depending on whether "agent_name" and "agent_phone" are columns you asked for.

The honest answer for many procurement reviews: the fastest route to approval is to remove the personal data fields from the spec. If nobody in your organization can say what the agent phone number is for, drop the column and the DPA question resolves itself. ScrapeIt's stated position is public data only, with no personal data beyond what is publicly published, which narrows the surface without eliminating the question - "publicly published" and "not personal data" are different things.

The questions to ask before you sign

These distinguish a vendor who has thought about this from one who has not.

  • Who are your subprocessors, and where are they? Scraping engagements involve infrastructure providers and often proxy providers. Ask for the list by name and jurisdiction, and what each one touches: a proxy network that only sees outbound requests is a different risk from a cloud provider holding your output files.
  • How do you handle subprocessor changes? A notice period and a right to object are standard. Get them in writing.
  • How long do you keep our data, and what counts as "our data"? Distinguish deliverables, raw intermediate storage, and operational logs. Ask for a retention period per category, not one number.
  • What does deletion mean operationally? Ask whether backups are included and the maximum lag before a deletion propagates. "We delete on request" without a backup answer is incomplete.
  • Where is data processed and stored? If you need EU-only processing, say it during scoping, not after the crawlers are built.
  • Who on your side can see our target list? Access control on the spec matters as much as access control on the output.
  • What happens to the crawler code if we leave? More commercial than legal, but it belongs in the same conversation - see build vs buy for how that trade-off plays out.
  • How do you deliver? Delivery into your own database or via API keeps the dataset inside your perimeter and shortens the retention conversation - one Switzerland private-listings project pushes 85K rows a day into the client's CRM. CSV, JSON and XLSX are available too.

What a vendor should refuse to do

A vendor who says yes to everything is a liability, because the risk does not disappear - it moves onto your balance sheet. The refusals are a quality signal.

  • Data behind a login the client does not have a right to use. Credentialed access changes the legal picture entirely. Where a members-only catalog is involved, as in the BestSecret project running 200K rows a day, access has to be agreed in advance.
  • Personal data beyond what is publicly published. Enrichment, joining against other datasets to build profiles, or reconstructing contact details that the source did not publish.
  • Circumventing protections in ways the client cannot defend. If you would not want the method described in a deposition, do not buy it.
  • Guaranteeing that a target site will never change its terms or its markup. Nobody can promise that. What a vendor can promise is that they fix it when it breaks - see what maintenance involves.
  • Implying ownership of trademarks or site names. Site names in a proposal are descriptive references to targets, not endorsements or partnerships. That same compliance page carries the full list of what will not be done.

Getting it through your own review faster

The pattern that stalls deals is legal reviewing a scraping contract as if it were a SaaS contract with a full personal-data footprint, when the project is public product prices. Front-load the facts and the review gets short.

What to prepare before the first legal call:

  • The field list, column by column. This single document determines whether a DPA is needed.
  • The target list and the legal basis you rely on for each - public pages, or access you already hold.
  • Delivery destination and who inside your organization will hold the data.
  • Retention: how long you actually need each delivery, which for a daily refresh is usually shorter than assumed.

A concrete sample helps more than a description: if your reviewers want to see the shape of the output first, the data sample shows real extracted rows. Once the field list is settled, request a quote - plans start from EUR 199, and how that scales with volume and site count is covered in how pricing actually works.

One case where you should not hire a vendor at all: if the field spec itself cannot leave the company and you have engineers to spare, building in-house removes the contract question. That is a real answer for some teams. For most, the NDA and - where it applies - the DPA are routine, and the field list decides whether the review takes weeks or months.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582