edX Scraper for Courses, Programs and University Partners

edX sells university credentials rather than courses, and its catalogue is organised around that. Session dates and enrolment tracks matter more here than the price tag.

edX Scraper
Solutions

Managed edX collection, run end to end by us

ScrapeIt runs the collector as a managed service. You name the subjects, partner institutions, credential levels and the countries you need prices for; we build the pipeline, run it on your cadence and hand back CSV, JSON, Excel or a push into your warehouse, with programs as parents, sessions as rows and tracks kept separate from price.

Cadence is per purpose. Session and deadline data justifies frequent passes because it goes stale in days; catalogue structure and syllabus content run monthly.

We collect what the catalogue publishes openly, honour the crawl rules the site publishes and pace requests. Course content belongs to the platform and its partner universities, so the dataset is metadata for analysis rather than material for republication, and we do not touch lecture material. Bring the use case to your own counsel before the project starts.

edX fields in every export

The catalogue record covers course or program identifier, title, product type and credential level, canonical URL, partner institution, subject and subcategory, level, stated effort per week, total duration, language and available transcripts.

Track and pricing fields are kept apart: whether a free audit track exists, whether audit access is time limited, the verified or paid track price, the currency, and the country the price was collected for. One price column would answer a question nobody asked.

Session data is collected as its own set of rows where the course runs in dated sessions: session start date, enrolment deadline, end date, and whether the run is current, upcoming or archived. This is the field set that keeps a catalogue snapshot honest about what a learner can actually enrol in today.

Program structure is preserved rather than flattened. A MicroMasters or professional certificate is delivered as a parent record with its component courses linked, in order, so the credential can be analysed as a whole or by part.

Then the usual context: instructors, prerequisites where stated, syllabus outline where published, ratings and enrolment figures where shown, and the collection timestamp with the country on every row.

edX fields in every export
Programs, archived runs and cross-platform mapping

Programs, archived runs and cross-platform mapping

Program level collection is the part most clients end up wanting once they see it. A parent credential with its component courses linked lets you total the cost, total the hours, and see whether the components are shared with other programs, which is common and which nobody notices from a flat export.

Archived runs are worth keeping rather than discarding. A course that ran three times and stopped tells you something about demand, and a subject where every run is archived is a retreat. We keep archived sessions with their dates and mark them, instead of only reporting the current state.

Cross platform mapping is the usual extension, and it is where the subject taxonomy problem appears. Every learning platform uses its own subject tree, so comparing coverage across two or three of them requires a mapping layer. We build that during collection rather than leaving three incompatible taxonomies for the client to reconcile in a spreadsheet.

Change tracking runs on the same basis as elsewhere: new listings, retired listings, price movements, new sessions and deadline changes delivered as their own records so a client sees what moved without diffing full exports.

How edX packages courses, programs and credentials

edX grew out of university partnerships and the catalogue still shows it. Alongside individual courses sit MicroMasters programs, professional certificates, XSeries bundles, bootcamps and full degrees, most of them attached to a named university or company rather than to an independent instructor.

That changes what the data has to carry. On an instructor marketplace the interesting entity is the course; here it is often the program, with courses as components, and a client tracking competitive coverage cares which institutions are offering what at which credential level.

Enrolment tracks are the second structural feature. Many courses can be taken on a free audit basis with limited access, while a verified track adds graded work and a certificate for a fee. Audit access is also frequently time limited. So availability, price and what you actually get are three separate attributes of the same listing.

Session scheduling matters more here than on an always open marketplace. A good deal of the catalogue runs in dated sessions with start dates, enrolment deadlines and archived runs, and a catalogue snapshot that ignores those dates will show courses as available that nobody can currently join. The site publishes a search surface and a sitemap, both of which respond, so enumerating the catalogue is practical.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why credential level and session dates carry the signal

Comparing this platform with an instructor marketplace on price alone produces nonsense. A short course from a named university at one price and a fifty hour course from an independent instructor at another are not competing on the same axis, and any benchmark that treats them as comparable will mislead whoever reads it.

The axis that does compare is the credential. What level of qualification is being offered, by which institution, in which subject, and at what total cost across a multi course program. That requires program structure to be preserved, component courses to be linked and prices to be summable, and it is the reason program data is worth more here than course data.

Session dates are the second thing people underestimate. A catalogue that lists a course as available when its enrolment deadline passed three weeks ago is describing a shop window rather than a shop. Collecting sessions separately, with their deadlines and status, is what makes availability analysis real.

The third is institutional coverage. Because listings attach to universities and companies, this source answers questions about which institutions are moving into online credentials, in which subjects, and how quickly. For a university planning its own online portfolio that competitive map is the brief.

The fourth is track economics. Audit availability, time limits on audit access and the verified track price together describe how a platform converts free learners into paying ones, and tracking how those settings change across the catalogue over time is a genuine strategic signal.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and keeps your edX feed

ScrapeIt is a managed extraction company, not a tool you have to learn. Our team builds the pipeline, watches it as the catalogue changes, and repairs it before your coverage map goes stale.

You see a sample first, in your format, across the subjects and institutions you actually track, with programs and their components already linked so you can judge the structure on real records.

FAQ

Can you collect programs, not just individual courses?

Yes, and on this platform that is usually the point. A MicroMasters or professional certificate is delivered as a parent record with its component courses linked in order, so the credential can be costed and measured as a whole. Components are frequently shared between programs, which is invisible in a flat course export.

Do you capture session dates and enrolment deadlines?

Yes, as their own rows: start date, enrolment deadline, end date and whether the run is current, upcoming or archived. Without them a catalogue snapshot lists courses as available whose deadlines passed weeks ago, which turns an availability analysis into fiction.

How do you handle free audit versus paid tracks?

As separate attributes rather than one price. We record whether a free audit track exists, whether audit access is time limited, the paid track price, the currency and the country collected for. How a platform converts free learners into paying ones is itself a signal, and it is only visible if those fields stay apart.

Can you compare this platform against others?

Yes, and it needs a subject mapping layer because every platform uses its own taxonomy. We build that during collection so you receive one comparable dataset rather than three incompatible exports. It is also worth noting that credential level, not price, is the axis that compares meaningfully between a university platform and an instructor marketplace.

How often can the feed refresh, and in what formats?

Session and deadline data frequently, since it goes stale in days; catalogue structure and syllabus content monthly, since re-reading them buys nothing. Delivery is CSV, Excel, JSON, JSONLines or XML over FTP, SFTP, Amazon S3, Google Cloud Storage, Dropbox, Google Drive or email, or written straight into your database.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582