Parcels App Scraper for Carrier Detection and Formats

Given only a tracking number, which carrier is it? The answer is a pattern library, and this is one of the few places it is written down.

Parcels App Scraper
Solutions

Managed carrier format data, run end to end by us

ScrapeIt runs the collection as a managed service. You name the regions or take the full directory; we build the pipeline with format rules expressed as applicable patterns, overlaps recorded, postal conventions flagged, and hand back CSV, JSON, Excel or a push into your warehouse.

Formats change slowly, so this is usually a one-time build with occasional refresh rather than a monitored feed, and we quote it that way instead of selling a subscription for static data.

We collect published reference data only. No tracking events for third party shipments, no recipient information. Crawl rules are honoured and requests paced. Your counsel should see the use case before the project starts if the product will handle tracking data for real consignees.

Parcels App fields in every export

Carrier records carry the carrier name, country, carrier type, website and the service descriptions published for it.

Format records are the heart of the dataset: the pattern as documented, expressed both as published and as a regular expression, with length, prefix, suffix and any check digit convention as separate fields so the rule can be applied rather than only read.

Ambiguity is recorded rather than resolved away. Where a pattern is shared by several carriers, the overlap is a field, because a detection system that does not know a number is ambiguous will confidently route it to the wrong carrier.

Universal postal formats are flagged separately, since a number following the international postal convention encodes an origin country and belongs to a class of its own rather than to one operator.

Every row carries the collection timestamp and the source address.

Parcels App fields in every export
Detection testing, postal coverage and scope

Detection testing, postal coverage and scope

Detection testing is worth building alongside the dataset: a set of known numbers per carrier to check that the rules actually match what they claim. Rules documented in prose frequently disagree with reality at the edges, and testing is how that gets found before a customer does.

Postal operator coverage mapping answers which national posts are documented and which are not, which for a cross-border product is a direct list of where manual handling will still be needed.

Overlap analysis across the whole rule set shows where detection is inherently ambiguous, and it is the input to a sensible fallback design rather than an afterthought.

Scope is the same as for any tracking source. We collect carrier and format reference data. We do not collect tracking events for shipments that are not the client's own, because those describe parcels moving toward named recipients and that is personal data whatever route it arrived by.

A directory built around number formats

Parcels App is a shipment tracking service covering a wide set of postal operators and couriers internationally, with a public carrier directory describing each one.

What makes it worth collecting separately from other aggregators is emphasis. The directory documents how carriers' tracking numbers are structured - prefixes, lengths, check digit conventions, the country codes embedded in postal formats - which is the raw material for a problem every shipping product eventually hits: a customer pastes a number and gives no carrier.

Carrier detection from a number alone is a pattern matching exercise with real ambiguity, because formats overlap and several carriers can plausibly match the same string. A useful dataset therefore has to support ranked candidates rather than a single answer, and that shape has to be decided before collection.

The carrier directory responds directly with substantial content. Crawl rules are published, allow the site generally, and point at a sitemap.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why detection needs candidates, not an answer

Carrier detection is one of those problems that looks solved until it is deployed, and the failure mode is quiet.

Formats genuinely overlap. A thirteen character string ending in a country code fits the international postal convention and could belong to any of dozens of operators. A numeric string of a common length matches several couriers. A system returning one carrier for such a number is not detecting, it is guessing with a confident interface, and the customer gets a tracking page for the wrong company.

Building it properly means returning ranked candidates with the matched rule attached, so the calling application can ask the customer, try several, or apply its own priors about which carriers it actually uses. That design decision has to be made at collection, because a dataset that recorded only the first matching carrier cannot be repaired later.

The second reason to use this source is coverage of postal operators. Express couriers are documented in many places; national posts in smaller countries are documented in few, and those are exactly the carriers a cross-border e-commerce product struggles with.

The third is that this is reference data with a long shelf life. Formats change rarely, so a well built dataset stays useful for years rather than needing constant refresh.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your format rule set

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, tests the rules against known numbers rather than trusting the documentation, and refreshes the set when carriers change their conventions.

You see a sample first, in your format, over the carriers you actually encounter, with overlaps recorded so you can see immediately where detection will need a fallback.

FAQ

Can you tell me the carrier from a tracking number?

We can give you the rule set that does it, returning ranked candidates rather than one answer. Formats genuinely overlap, and a system returning a single carrier for an ambiguous number is guessing with a confident interface - the customer gets the wrong company's tracking page.

Why record overlaps instead of resolving them?

Because resolving them requires context we do not have and you do - which carriers you actually use. Recorded as overlaps, your application can ask the customer, try several, or apply its own priors. Resolved silently, it just gets them wrong.

Do the documented formats actually match reality?

Mostly, and not always at the edges, which is why we recommend building a test set of known numbers alongside. Rules written in prose disagree with practice often enough that testing is how you find it before a customer does.

How often does this need refreshing?

Rarely. Formats change slowly, so this is usually a one-time build with occasional refresh. We quote it that way rather than selling a subscription for data that is essentially static.

Can you track actual parcels for me?

Not for shipments that are not yours. A tracking event describes a parcel moving toward a named recipient, which is personal data whatever route it arrived by. For your own shipments, the platform's API is the right tool.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582