Duolingo Data: Course Pairs, Subjects and Public Plan Pages
A Duolingo scraper starts from the way the app organizes teaching: not one course per language, but one course per language pair - the language you learn and the language you learn it from. Duolingo, the Pittsburgh company behind the app, publishes that grid in its public catalog, shows a learner count on every language course and lets a visitor filter the list by base language under I speak. As of September 2026 the all-languages view listed 285 language pairs covering 42 taught languages, from English for Spanish speakers at 54M learners down to English for Kannada speakers at 4.47K, next to three subjects that are not languages: Math, Music and Chess.
The grid is uneven, and that is what makes it worth collecting. English is taught from 32 base languages, while Spanish, French, German, Italian, Portuguese, Japanese, Korean and Chinese (Simplified) are taught from 27 each. Most other languages, from Welsh and Navajo to High Valyrian and Klingon, exist only for English speakers; Catalan is taught only from Spanish, and Cantonese only from Chinese. The website interface runs in more than 30 languages on subdomains such as de.duolingo.com.
Around the catalog sit other public layers: the plan pages and help articles for Super Duolingo, the Family Plan and Duolingo Max, the Duolingo for Business seat calculator, app store listings with in-app prices, the Duolingo English Test site with its accepting institutions, and the company's own reports. Lessons and the learning path sit behind sign-in and stay out of scope. To scrape Duolingo data well you pick the layer first, because each one is published and refreshed differently.
Get a Quote