What the Times shows before the paywall, and what it does not
The New York Times runs a metered paywall, and the first thing to settle on any project here is what that leaves visible. Section fronts are public and complete: the world, business, technology and politics indexes list what the desk is running, in order, with headlines and summary lines. Article pages render the headline, the summary, the byline, the section, both timestamps and usually an opening portion of the text before the wall appears. The rest is not ours to take and we do not take it.
Machine readable surfaces exist alongside. The Times publishes a news sitemap for recently published items, and article and section pages carry structured markup identifying the headline, publisher and dates. There is also a public developer programme with documented APIs covering article search, top stories and most popular items, which for some briefs removes the need for collection entirely.
Structurally the site is organised by section and by topic, with topic pages that gather everything on a subject and reach considerably further back than the fronts. Newsletters, live coverage and interactive features sit alongside conventional articles and behave differently enough that they are typed rather than lumped together.
The practical consequence is that a Times dataset is a metadata dataset with a text sample attached. For media monitoring, share of voice, section level prominence and revision tracking, that is sufficient. For anything that needs full article bodies, the answer is a licence from the publisher, not a scraper.
Get a Quote