Forbes Scraper Separating Staff Reporting From Contributors

Two articles on this site can carry the same masthead and mean entirely different things: one is staff reporting, the other is a contributor's opinion.

Forbes Scraper
Solutions

Managed Forbes collection, run end to end by us

ScrapeIt runs the collector as a managed service. You name the sections, companies or topics; we build the pipeline, run it on your cadence and hand back CSV, JSON, Excel or a push into your warehouse, with author type on every row and ranking lists as their own records.

Cadence is per section. Fronts during an active story justify frequent passes; archive work runs slower. Requests are paced within the crawl rules the site publishes.

Articles are copyrighted and the terms restrict reuse, so the dataset is metadata, entity extraction and publicly rendered text for analysis rather than material for republication. Bring the use case to your own counsel before the project starts.

Forbes fields in every export

The article record covers canonical URL, headline, standfirst, section, byline, publication timestamp, last updated timestamp, item type and the article identifier.

Author type is the field that makes this source usable: staff, contributor or other, taken from what the page states rather than inferred from the name. It sits next to the byline so a client can filter, weight or separate at will, and so that any total can be recomputed either way without going back to the source.

List and ranking items are typed separately, with the list name, edition year, position and the entity named at each position. Those compilations are among the most cited things the site publishes and they do not fit an article schema at all.

Front observations run alongside, recording which items sat on a section front, in what order and when. On a site with heavy contributor volume, promotion is a much better proxy for editorial attention than a count of published items.

Then the usual context: topic tags where published, publicly rendered body text and word count, outbound links, lead image URL, and the collection timestamp on every row.

Forbes fields in every export
Lists, contributor networks and cross-outlet comparison

Lists, contributor networks and cross-outlet comparison

Ranking lists are worth their own collection. Name, edition year, position and the entity at each position gives a clean time series, and for the companies and people who appear on them it is one of the more valuable single facts a media dataset can carry.

Contributor network analysis falls out of typed bylines. Which contributors write about a sector, how often, and whether coverage of a company concentrates in a handful of authors is a question communications teams ask constantly and cannot answer from an untyped feed.

Cross outlet comparison is the usual extension and the reason typing matters beyond this source. A share of voice number across several business titles is only meaningful if each outlet's contributor or sponsored content is identified consistently, and outlets label it differently. We normalise that during collection.

Historical backfill is practical because article URLs are stable and section pagination reaches back. It runs as a one time job separate from the ongoing feed.

One brand, two very different kinds of article

Forbes publishes staff journalism and, alongside it, a large volume of material from an external contributor network. Both appear under the same brand, in the same templates, with the same design, and to a reader skimming a feed they look identical.

They are not. Staff reporting goes through an editorial process; contributor pieces are written by outside authors under a looser arrangement, and the site labels them as such on the page. For a communications team measuring coverage, a staff investigation and a contributor's opinion column are not comparable events, and averaging them produces a number that overstates attention.

Section fronts such as business and innovation respond directly and carry ordered selections, and crawl rules are published. Volume is substantial, and a large share of it comes from the contributor side, which is exactly why the distinction has to be in the data rather than in a caveat.

The site also runs lists and rankings that behave differently again: annual compilations with their own structure, entity rich and heavily cited, which are worth collecting as their own type rather than as articles.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why the byline type decides whether the report is honest

A client asks how much coverage they got on Forbes. Untyped, the answer counts an investigative piece by a staff reporter and a contributor's enthusiastic column as one each, and the resulting number is presented to a board as attention from a major business title.

That is not a rounding error. Contributor volume can dominate the count for a given company, particularly one that is actively pitching, and a report built on the raw total tells the board something that is not true.

Typed, the same data becomes defensible: two staff articles, eleven contributor pieces, of which nine came from three authors. That is a picture somebody can act on, and it usually changes what they do next.

The second reason is source credibility weighting, which anyone building a media dataset for analysis eventually needs. Feeding a model a corpus in which opinion columns and reported journalism are indistinguishable teaches it that they are the same, and every downstream sentiment or salience score inherits that.

The third is the lists. A company appearing in an annual ranking is a specific, dateable, highly citable event that has nothing in common with an article mention, and it deserves its own row rather than being buried as one more item in a coverage count.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your Forbes feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, watches it as templates and labelling change, and repairs it before your coverage numbers start counting the wrong things.

You see a sample first, in your format, over the companies or topics you actually track, with author type applied so you can see the staff to contributor split before committing.

FAQ

Why does staff versus contributor matter so much?

Because they are not comparable events and the count hides it. Contributor volume can dominate the total for a company that is actively pitching, so a report built on the raw number tells a board it received attention from a major business title when most of what it got was opinion columns. Typed apart, the same data becomes defensible.

Where does the author type come from?

From what the page states, not from inference on the name. The site labels the distinction, so we read it rather than guessing, and we keep it next to the byline so any total can be recomputed either way without going back to the source.

Can you collect the ranking lists?

Yes, as their own type: list name, edition year, position and the entity at each position. Those compilations are among the most cited things the site publishes, they do not fit an article schema, and for a company that appears on one it is a far more valuable fact than another article mention.

How does this help with cross-outlet share of voice?

By making the comparison honest. Several business titles carry contributor or sponsored content and each labels it differently, so a share of voice number is only meaningful if that material is identified consistently across all of them. We normalise the typing during collection rather than leaving it to a spreadsheet later.

Can I republish the articles?

No. The journalism is copyrighted and the terms restrict reuse. What you get is metadata, entity extraction and publicly rendered text for analysis: coverage measurement, author network analysis, ranking histories. Have your own counsel approve the use case before the project starts.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582