MLB Scraper for Statcast Pitches, Game Feeds, Rosters and MiLB

MLB publishes its game on two public sites - mlb.com with its club and MiLB sections, and Baseball Savant - and both share one set of game and player ids. We turn them into one dataset.

Plans from €169/month · Free project assessment · Reply within 1 business day

MLB Scraper
Solutions

Managed MLB data scraping, from backfill to game-day loads

ScrapeIt runs MLB data scraping as a managed service. You name the levels, seasons and tables; we build one pipeline to extract MLB data across mlb.com, the club sections and Baseball Savant, and deliver CSV, JSON, Parquet or a load into your warehouse on the cadence you set. The mlb.com edge turns away many automated clients, so anti-bot handling, proxy rotation and CAPTCHA solving are part of the job, along with the slicing, retries and row-count checks that keep a Savant backfill complete. No logins and no paywalls: MLB.TV, MiLB.TV and fan accounts stay out. Player data is professional sporting information about public figures; contact details, family and biography text are not collected.

What an MLB Statcast dataset holds, pitch by pitch

A Savant pitch row is the densest record baseball publishes: well over a hundred columns per pitch, keyed on game_pk, at_bat_number and pitch_number, with the batter, the pitcher and fielder_2 to fielder_9 as MLBAM ids. It is the MLB pitching dataset most buyers mean by Statcast.

  • Release - pitch_type and pitch_name, release_speed and effective_speed, release_spin_rate and spin_axis, release_extension, arm_angle and the release point in three axes.
  • Flight and location - pfx_x and pfx_z movement, velocity and acceleration from vx0 to az, plate_x and plate_z, zone, and sz_top and sz_bot for the batter's zone.
  • Contact - exit velocity as launch_speed, launch angle as launch_angle, hit_distance_sc, hc_x and hc_y, bb_type and a six-step launch_speed_angle class from Weak to Barrel.
  • Value - expected stats (xBA, xwOBA, xSLG), woba_value, delta_run_exp and delta_home_win_exp.
  • Bat tracking - bat_speed, swing_length, attack_angle, attack_direction and swing_path_tilt.
  • Context - balls, strikes, outs_when_up, inning, on_1b to on_3b as runner ids, infield and outfield alignment, scores before and after, times through the order and days of rest.

The Gameday view and the box score add what Savant leaves out: pitch break and a play id per pitch; weather, attendance, first pitch, game length and delays; the umpire crew and official scorer; mound visits, replay challenges and, from 2026, ABS challenges won, lost and remaining. Win probability, win probability added and leverage index run play by play, and Gameday replays any moment of a game as it stood then. MLB standings add clinch codes, elimination and magic numbers, run differential, an expected won-lost record and sixteen split records, from one-run games to grass and turf.

Two definitions changed in 2026: plate_x and plate_z are measured at the middle of the plate rather than the front, and sz_top and sz_bot follow the ABS zone rather than an operator's mark. Every row we deliver says which applies.

What an MLB Statcast dataset holds, pitch by pitch
MLB player stats, transactions, the injured list and the minors

MLB player stats, transactions, the injured list and the minors

The stats pages on mlb.com split into player and team, hitting and pitching, and Standard, Expanded and Statcast views, with splits and qualifier filters, but no export button, so exportable MLB stats come from us as CSV or JSON. Behind the tables sit season, career, game log, split, head-to-head and sabermetric views with WAR, expected statistics and ZiPS projections, and we build an MLB player dataset across all of them.

Roster moves form a transaction log with date, effective date, club, player and one of forty-odd type codes: Optioned, Recalled, Designated for Assignment, Outrighted, Claimed Off Waivers, Rule 5 Selection, Status Change. Injured list moves read as clubs announce them - 10-day for position players, 15-day for pitchers, 60-day, retroactive dates - with a one-line injury description, and the 40-man roster marks each player active, reassigned or injured. We keep that line and nothing beyond it. The draft table adds round labels for competitive balance and compensation picks, slot value and signing bonus.

Minor League Baseball runs on the same ids, Triple-A through Rookie ball as their own sport ids and leagues. Tracking is uneven: Triple-A has full Statcast from 2023 and the Florida State League from 2021, while Double-A pitch data has calls but no velocity or spin. MiLB.com hubs now forward into mlb.com/milb.

MLB.com and Baseball Savant: two public sites, one set of MLB ids

Major League Baseball publishes its game through MLB Advanced Media, and the record reaches the public on two sites: the pages on mlb.com with the thirty club sections, and Baseball Savant, the Statcast research site. An MLB scraper worth paying for treats the two as one system, because they share one set of keys. The pages carry MLB results, standings, box scores, rosters, transactions and the Gameday view of every pitch; Savant carries the tracking data underneath.

The game key is gamePk, a plain integer. The same number opens the Gameday page, the box score and the game_pk column of every Savant pitch row, so a pitch, its box score and the weather at first pitch join without matching names. Players carry an MLBAM person id that follows them from a /player/ page to the Savant batter and pitcher columns, the draft table and the transaction log. MLB team ids run from 108 to 158 with gaps and survive moves: 133 has been the Philadelphia, Kansas City and Oakland Athletics and is now simply the Athletics, coded ATH.

Levels are sport ids - 1 for the majors, 11 Triple-A, 12 Double-A, 13 High-A, 14 Single-A, 16 Rookie ball - and game types run from S for spring training and R for the regular season to F, D, L and W for the Wild Card Series, Division Series, League Championship Series and World Series. Schedules, standings, rosters and transactions all carry these ids, so every table joins on the same keys.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

MLB Scraping Plans and Pricing

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why teams buy an MLB API built from public pages

Buyers range from media and analytics teams who need MLB live scores within seconds to clubs, agencies and researchers who want a decade of pitches in one clean table. A managed feed earns its fee in four places.

  • Joining the sources - Savant pitches, mlb.com game context and the minors land in one schema on gamePk and person ids, with doubleheaders, postponements and split-squad games resolved.
  • History - a full season runs past half a million pitches, so we collect it in more than two dozen date slices and reconcile each one against the game's pitch count before it reaches your MLB database.
  • Snapshots - standings, rosters and their injury notes can be rebuilt for past dates, but probable pitchers, posted lineups and depth charts are overwritten. Only a timed capture shows what was known before first pitch.
  • Delivery - incremental loads into your MLB data warehouse, versioned schemas and an alert the day Savant adds a column.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

The team that keeps your MLB database current, spring to October

ScrapeIt is a managed scraping agency, not a tool you configure. A named engineer scopes the levels and tables with you, builds the collectors and watches them through spring training, the 162-game season and October, the weeks a new Savant column or a reshaped page would otherwise break your loads. The price of an MLB feed follows the plan you pick and the levels in scope, not the number of pages loaded. A sample comes first, in your format, so you can check it against games you know.

FAQ

Do you offer an MLB stats API?

Yes. The ScrapeIt MLB API returns schedules, live game state and final box scores, standings, rosters with injured list status, transactions and Statcast pitch rows, all keyed on gamePk and MLBAM person ids. You get JSON on the schedule you choose - every few seconds during games, daily for rosters - plus CSV and XLSX exports, and every field is described in one data dictionary.

How do you collect a full season of Baseball Savant pitches?

A season runs past half a million pitches, more than any single Statcast search returns, so we slice it by date, check each slice against the pitch count of its games and stitch the rows on game_pk, at_bat_number and pitch_number before they reach you. The minors have their own search, covering Triple-A and the Florida State League, and join on the same person ids. Every row carries its tracking source and the 2026 plate-location flag.

How often can MLB data refresh during a game?

During games we refresh live state every ten seconds or so, and any moment we missed is rebuilt from the game timeline. Schedules, probable pitchers, rosters and transactions run on daily and pre-game passes, and Statcast pitch rows are collected after games go final. Official scoring changes can land days after a game, so recent weeks are re-read before a season is frozen.

How do you scrape MLB history when the tracking systems changed?

Column by column. Box scores and line scores reach back to the early 1900s, play-by-play fills in for later decades, pitch velocity arrives in 2008 with PITCHf/x and batted-ball tracking in 2015 with Statcast, and Savant rescales the 2008-16 velocities to the out-of-hand point so they sit on one scale. Bat tracking is newer again. Team ids hold across moves: 133 is the Philadelphia, Kansas City and Oakland Athletics and now the Athletics, while 120 was the Montreal Expos before Washington. We flag the tracking source and the 2026 plate-location change in every row.

Is it legal to scrape MLB?

We collect only publicly available data - everything a visitor can see on MLB.com and Baseball Savant - and we collect it legally. No logins, no paywalls: MLB.TV, MiLB.TV and fan accounts stay out of scope. Player records are professional sporting information about public figures, and in line with GDPR we keep personal data to that record: no contact details, no family or biography text, and injury data stops at the injured list line the club announced.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582