FreightWaves Scraper for Freight News and Index Values

One domain, two datasets: a news corpus and a numeric series. Merge them and you get a table that is bad at being either.

FreightWaves Scraper
Solutions

Managed freight media data, run end to end by us

ScrapeIt runs the collection as a managed service. You name the modes, topics or the full feed; we build the pipeline with articles and extracted figures as separate record types, every figure carrying its source article, and hand back CSV, JSON, Excel or a push into your warehouse.

News moves quickly and gets corrected, so we collect frequently and keep revisions rather than overwriting.

We collect published editorial content, honour the crawl rules and pace requests. Content is copyrighted: the dataset is for measurement and analysis, not republication, and anything surfacing article text to end users is a licensing conversation with the publisher. Your counsel should see the use case before the project starts.

FreightWaves fields in every export

Article records carry the headline, publication timestamp, modification timestamp where available, author, section, topic tags, mode, word count and the address.

Index and figure records are separate, carrying the index or metric name, the value, the unit, the period it describes, the lane or region if specified, and the article it appeared in. That last field is what makes a number checkable rather than something a pipeline asserted.

Mode is a first class field - trucking, rail, ocean, air - because freight modes move independently and a combined view averages markets that have nothing to do with each other.

Entities mentioned, such as carriers, ports and shippers, are extracted where a client needs them, which turns the corpus into something queryable by company rather than by keyword.

Every row carries the collection timestamp, and articles carry a revision count where re-collection shows the text changed.

FreightWaves fields in every export
Event timelines, entity monitoring and rights

Event timelines, entity monitoring and rights

Event timelines are what the corpus does best: how coverage of a disruption builds, peaks and decays, which for anyone running an early warning process is more actionable than any single article.

Entity monitoring by carrier, port or shipper turns the corpus into a competitive intelligence feed, and needs entity extraction rather than keyword matching to avoid missing the mentions written differently.

Numeric series assembled from reporting are useful as a secondary check against a primary provider rather than as a replacement for one, and we scope them that way rather than presenting journalism as a market data subscription.

On rights we are direct: the content is copyrighted. Measurement and structuring is analysis; republishing text is not. A product surfacing article bodies to end users is a licensing conversation with the publisher, and we say so at scoping.

News and numbers on the same domain

FreightWaves is a freight industry news organisation that also produces market data. Its coverage spans trucking, rail, ocean, air and logistics technology, published at high volume, and its reporting regularly carries index values and market figures.

That combination is the structural fact for a data project. There is a news corpus - articles with headlines, authors, timestamps, topics - and there is a numeric layer of index values and market figures that appear inside the reporting. They are different data with different uses, and one table holding both is bad at each.

The news corpus supports coverage and sentiment work, event timelines and competitive monitoring. The numeric layer supports market analysis but only if the values are extracted with their periods and their source articles attached.

Article sections respond directly with substantial content. Crawl rules are published and disallow search query addresses and the content API paths, leaving editorial content open.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the numbers need their articles attached

Extracting figures from reporting is straightforward to do and easy to do irresponsibly, and the difference is one field.

An index value lifted out of an article and put in a spreadsheet loses its qualification. The article may have said the figure was preliminary, or covered a partial week, or was the publisher's own estimate rather than a measured value. Keeping the source article and the surrounding claim text on the row means a user can check, and the cases where checking matters are precisely the ones where the number is surprising.

The second reason is the update problem, sharper in freight news than most. Figures get revised, articles get corrected, and a corpus with only a publication timestamp mixes the original claim with the corrected one. Recording modification times and keeping revisions is what stops a dataset from quietly rewriting its own history.

The third is that coverage volume in trade media is a real signal. A sustained rise in articles about a port, a lane or a carrier usually precedes measurable disruption, and that analysis needs article metadata rather than the numeric layer.

The fourth is scope. This is journalism, not a data feed. It reports on markets and it is not the market, and where a figure originates with another provider the right source for a serious series is that provider.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your freight media feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, retunes figure extraction as reporting formats change, and repairs the collector before an event timeline loses the days that mattered.

You see a sample first, in your format, over the modes you actually follow, with figures extracted and their source articles attached so you can judge the extraction on real reporting.

FAQ

Why keep articles and figures separate?

Because they are different data with different uses. A news corpus supports coverage analysis and event timelines; a numeric series supports market analysis. One table holding both is bad at each, and the merge is not reversible afterwards.

Can I use these index values as market data?

As a secondary check, not as a replacement for a primary provider. A figure in an article may be preliminary, partial or the publisher's estimate, which is why we keep the source article and claim text on every row so you can see what was actually asserted.

What happens when an article is corrected?

We record modification times and keep revisions rather than overwriting. Figures get revised and articles get corrected, and a corpus with only a publication timestamp quietly rewrites its own history.

Is coverage volume actually useful?

Yes, and it is one of the better signals here. A sustained rise in articles about a port, lane or carrier usually precedes measurable disruption, and that analysis runs off article metadata rather than the numeric layer.

Can I republish the articles?

No. The content is copyrighted and the dataset is for measurement and analysis. Surfacing article text to your end users is a licensing conversation with the publisher, not a scraping project, and we raise it at scoping rather than leaving it to become your problem.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582