TrackingMore Scraper for Carrier Coverage and Statuses

A thousand carriers, a thousand ways to say the parcel is at the depot. Normalising that vocabulary is the entire engineering problem.

TrackingMore Scraper
Solutions

Managed carrier reference data, run end to end by us

ScrapeIt runs the collection as a managed service. You name the regions or take the full directory; we build the pipeline with carrier records, coverage and status vocabulary mapped to a normalised lifecycle, raw statuses kept alongside, and hand back CSV, JSON, Excel or a push into your warehouse.

The directory changes steadily rather than constantly, so monthly collection with change records suits most briefs and is what we quote.

We collect published reference data only. No tracking events for third party shipments, no recipient data, no addresses. Crawl rules are honoured and requests paced. Your counsel should see the use case before the project starts, particularly if the product will handle tracking data for real consignees.

TrackingMore fields in every export

Carrier records carry the carrier name, its identifier on the platform, country or countries of operation, carrier type such as postal, express or last mile, website, and the tracking number formats it uses where documented.

Status vocabulary is collected per carrier as its own rows: the raw status text the carrier emits, and the normalised lifecycle stage we map it to. Both are kept, because the normalisation is an interpretation and a user who disagrees with a mapping needs the original to work from.

The normalised lifecycle is a small deliberate set - information received, in transit, out for delivery, delivered, exception, returned - chosen because a finer scheme cannot be populated consistently across a thousand carriers and a coarser one loses the distinctions people act on.

Coverage fields record which countries and lanes a carrier serves, which is the question most e-commerce operators bring to this data.

Every row carries the collection timestamp, since carrier coverage and naming change more often than people expect.

TrackingMore fields in every export
Coverage overlap, format rules and scope

Coverage overlap, format rules and scope

Coverage overlap analysis answers which carriers are redundant and which are the only option into a market, which is directly useful when negotiating or when planning where to add support.

Tracking number format rules, where documented, support carrier detection - working out which carrier a number belongs to without asking the customer. It is a pattern problem and the documented formats are the raw material.

Carrier change tracking over time shows which operators are added and which disappear, which is a read on consolidation in last mile delivery that is hard to see any other way.

Scope is firm. No tracking events for third party shipments, no recipient information, no addresses. Carrier reference data contains none of that, and a brief that needs live tracking should use the platform's API for the client's own shipments rather than a scraper for everyone else's.

A directory of carriers and the words they use

TrackingMore is a shipment tracking aggregator connecting to a very large number of carriers worldwide - national posts, international express companies, regional couriers and last mile operators - and presenting their tracking under one interface.

The publicly useful asset is the carrier directory. It lists which carriers are supported, where they operate, what their tracking looks like and how each one is identified, and it is substantial enough to be a reference dataset in its own right for anyone building shipping or e-commerce logistics.

The interesting problem underneath is vocabulary. Every carrier describes shipment progress in its own words and its own granularity - one says out for delivery, another says with courier, a third emits a depot scan code with no plain language at all. Mapping those onto a single lifecycle is what an aggregator does, and it is what a dataset built here has to do too.

The carrier section responds directly with a great deal of content. Crawl rules are published and carry no restrictions for general crawlers beyond pointing at a set of sitemaps.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the directory is worth more than the tracking

The instinct is to want tracking events. It is the wrong target for two reasons, one practical and one about people.

The practical reason is that tracking events belong to specific shipments and are retrieved per tracking number, which means bulk collection is neither possible nor sensible - the aggregator's own API exists for exactly that and using it is the correct route.

The reason about people is more important. A tracking event is a record of a parcel moving toward a named recipient at an address. Collecting those in bulk builds a dataset about individuals receiving goods, which is personal data regardless of how it was obtained. We do not collect tracking events for shipments that are not the client's own.

What the directory supports is entirely different and entirely legitimate: which carriers serve which countries, how coverage overlaps, which carriers a marketplace or platform needs to support, and how status vocabulary differs across them. For anybody building shipping software that is the reference data the project actually needs.

The fourth point is that vocabulary normalisation is reusable. Built once against a large carrier set, the mapping serves any tracking integration the client builds afterwards, which is usually worth more than any single dataset.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your carrier reference feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, maintains the status normalisation as carriers change their wording, and repairs the collector before your coverage map goes out of date.

You see a sample first, in your format, over the regions you actually ship to, with raw and normalised statuses side by side so you can judge the mapping rather than inherit it.

FAQ

Can you collect tracking events in bulk?

No, on two grounds. Practically, events are retrieved per tracking number and the platform's API exists for that. More importantly, a tracking event is a record of a parcel moving toward a named recipient, so bulk collection builds a dataset about individuals receiving goods.

Why keep both raw and normalised statuses?

Because the normalisation is an interpretation. A user who disagrees with a mapping needs the carrier's original wording to work from, and without it they would have to re-derive the whole thing from scratch.

Why such a small set of lifecycle stages?

Because a finer scheme cannot be populated consistently across a thousand carriers - many emit depot codes with no plain language at all. A coarser one loses the distinctions people act on. Six stages is where it holds together.

What is the carrier directory actually for?

Coverage decisions. Which carriers serve which countries, where coverage overlaps, which one is the only option into a market, and which a platform needs to support. For anyone building shipping software that is the reference data the project needs.

Can I use the format rules to detect carriers?

That is one of the better uses. Where formats are documented they support working out which carrier a number belongs to without asking the customer - a pattern problem for which the documented formats are the raw material.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582