Reddit Scraper for Subreddits, Posts, Comment Threads and Votes

Reddit sorts public conversation into communities with their own rules and flair. We turn subreddits, posts and whole comment threads into rows with score, upvote ratio, flair and reply depth.

Plans from €169/month · Free project assessment · Reply within 1 business day

Reddit Scraper
Solutions

Managed Reddit Scraping by Community, Keyword or Thread List

ScrapeIt runs Reddit scraping as a managed service, not as a Reddit scraping tool you install and babysit. You name the scope - subreddits, brand and product keywords, or a list of thread URLs - and we build the Reddit crawler, schedule it and fix it when Reddit changes its pages. Reddit screens automated traffic hard, so anti-bot handling, proxy rotation and CAPTCHA solving are part of the service. We collect public and restricted communities with no logins, so private, quarantined and 18+ communities stay out. Every row keeps Reddit's IDs and permalinks, so runs merge without duplicates. Delivery is CSV, JSON Lines, XLSX or Parquet to S3, SFTP or your warehouse, or through the ScrapeIt Reddit scraper API, hourly, daily, weekly or once.

Reddit Post Scraper: Fields for Each Community and Post

To extract Reddit data we read three levels - community, post and comment - and every row keeps Reddit's own IDs, so a comment always leads back to its post and its subreddit. Comments get their own section below; the first two levels carry these fields.

  • Subreddit. Name and t5_ ID, display title, public description, creation date, community type, the status line moderators set next to the name, bookmarks such as the wiki, and the subreddit rules in their order, each with its title and description.
  • Activity. Weekly visitors and weekly contributions, the two figures Reddit now shows in place of member totals. Weekly visitors count unique visitors over the past seven days on a 28-day rolling average, bots excluded; weekly contributions count posts and comments of the past seven days, removed ones excluded. The old member total survives only in the code of post pages, and we keep it while it lasts.
  • Post flair. The flair text and background color on each post. A community can define up to 350 flair templates, used as categories - Routine Help, Meme/Macro, Cables - or as status marks such as Expired on a dead deal.
  • Post. t3_ ID, permalink, title, body text of text posts, post type, linked URL and its domain, creation and edit times, score, upvote ratio, comment count, award count and the language code Reddit assigns. Image, gallery and video URLs are recorded too, so the same feed works as a Reddit media scraper.
  • Status marks. Archived, locked, pinned as one of up to six community highlights, spoiler, the green moderator shield on distinguished posts, and the Brand Affiliate tag that authors add to incentivized or commercial posts.

The author of a post or comment is kept as the public username, hashed or dropped, as you prefer.

Reddit Post Scraper: Fields for Each Community and Post
Reddit Comment Scraper: Whole Threads as Text and Structure

Reddit Comment Scraper: Whole Threads as Text and Structure

A Reddit comment scraper is only as good as its thread reconstruction. Each comment row carries its t1_ ID, its parent ID - another comment or the post itself - its depth and position among siblings, score, creation time and permalink, so the comment tree can be rebuilt exactly or flattened into rows. To scrape Reddit threads completely, we open the "more replies" and "Continue this thread" branches that Reddit folds away, and we record the sort order the thread was read in.

Deleted and removed comments keep their place: "Comment deleted by user" and "Comment removed by moderator" stay as empty markers, so reply chains do not break. Reddit's own comment count still includes removed comments, so it often runs higher than the replies you can see, and we deliver both numbers. Labels add context: OP on replies by the original poster, MOD and ADMIN on moderators and staff, the automated-account label on bots such as AutoModerator, and a collapsed flag on comments Reddit folds as likely spam.

Scores move. Reddit hides vote counts during a post's first hours, and a community can hide comment scores for up to 24 hours, so we store dated snapshots rather than one value. Where a community archives old posts, threads older than six months stop taking votes and comments, and Reddit archived posts get one final pass. Reddit also machine-translates posts and comments into languages such as French, German, Spanish, Portuguese, Italian, Hindi and Thai; we keep the original text with its language code.

Subreddits, Posts and Votes: How Reddit Data Is Organized

A Reddit scraper has to follow the way Reddit itself is built: around communities rather than people. Each subreddit, from r/buildapcsales to r/SkincareAddiction, is created and run by volunteer moderators who write its rules, set its post flair and decide what may be posted. As of June 2026 Reddit reports more than 100,000 active communities and over 130 million daily active uniques.

Communities come in four types. Public ones are open to read, restricted ones are open to read but limit who may post or comment, and private and premium-only communities stay closed to visitors. Communities labeled Mature (18+) sit behind sign-in and an age check, and quarantined ones behind an explicit opt-in.

Inside a community every post has a short base-36 ID with the t3_ prefix and an address of the form /r/community/comments/ID/title/; every comment has its own t1_ ID and a permalink under its post, and the subreddit itself has a t5_ ID. A post can be text, a link, images or a gallery, a video, a poll or a repost of another community's post, and each carries a title, a vote score, an upvote ratio and a comment count. Comments nest into reply trees of any depth. Feeds sort by Best, Hot, New, Top and Rising, with Top limited to the past hour, day, week, month, year or all time, and comment sections sort by Best, Top, New, Controversial, Old or Q&A.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Reddit Scraping Plans and Pricing

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why Teams Scrape Reddit Data: Brand Mentions, Sentiment, Demand

Reddit threads are where people ask what to buy, report what broke and argue over alternatives, in their own words and with a vote score attached. Three uses come up most.

  • Brand monitoring. Used as a Reddit keyword scraper, the feed tracks brand, product and competitor names across the subreddits you choose, with the matching passage, community, score and time. Posts their authors tag as Brand Affiliate stay apart from organic brand mentions, and Reddit reviews - threads where owners compare mattresses, graphics cards or moisturizers - arrive with every reply.
  • Topic and sentiment analysis. A Reddit dataset for sentiment analysis needs more than text. Score, upvote ratio and reply depth show whether an opinion was endorsed or disputed, the Controversial sort surfaces split threads, and post flair gives a ready topic label for training and evaluation.
  • Market research. Weekly visitors and contributions compare the pull of rival communities over time. Deal communities such as r/buildapcsales put the category and the price in the title - [Handheld] ... - $423.99 - and dead deals get the Expired flair, so Reddit price data on promotions comes with the crowd's verdict attached.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

What is the Scraping Web Data for Sentiment Analysis & How it Helps Marketers and Data Scientists

What is the Scraping Web Data for Sentiment Analysis & How it Helps Marketers and Data Scientists

The use of sentiment analysis tools in business benefits not only companies but also their customers by allowing them to improve products and services, identify the strengths and weaknesses of competitors' products, and create targeted advertising.

Empower Your Business with Google Maps Data

Empower Your Business with Google Maps Data

Using web scraping to extract Google Maps data will help you quickly and efficiently find businesses in any industry, city, state, or region. And the extracted contact information, such as phone or email address, social networks, or links to web pages, will help to contact them.

How to Generate Business Leads Using Web Scraping

How to Generate Business Leads Using Web Scraping

Lead scraping is a reliable way to get the right customer contacts to market your product, saves organizations time, and helps you better understand the audience you want to attract. Contacts include emails, phone numbers, or social media profiles.

scrapeit logo

Your Reddit Dataset, Scoped, Sampled and Maintained by ScrapeIt

ScrapeIt is a managed web scraping agency, so instead of a Reddit web scraper to maintain you get a delivered dataset. Tell us the question - every mention of your product in 40 subreddits, a month of posts from five hobby communities with full threads, or a daily feed of new deals - and an engineer designs the collection and keeps it running. A sample comes first, so you check fields, thread depth and fill rates before committing. The price of a Reddit dataset depends on how many communities it covers, how deep the threads go and how often it refreshes.

FAQ

Do you offer a Reddit API?

Yes. The ScrapeIt Reddit API returns the dataset we collect for you as JSON: subreddits with description, rules, flair list and weekly visitors and contributions; posts with ID, permalink, title, body, type, link, media URLs, flair, score, upvote ratio, comment count and timestamps; and comments with parent ID, depth, score and text. It updates on the schedule you set, and every delivery is also available as CSV or XLSX.

Can you scrape Reddit comments with the full reply structure?

Yes. Every comment row carries its own ID, its parent's ID, its depth and its position, so the thread rebuilds exactly as Reddit shows it or loads flat into a spreadsheet. Branches folded behind "more replies" and "Continue this thread" are expanded, deleted and removed comments stay as empty placeholders, and the sort order is recorded. Images, GIFs and videos posted inside comments come through as media URLs. As a Reddit thread scraper it also takes a plain list of thread URLs, and very long threads can be taken in full or capped at a depth or comment count you choose.

How do you scrape Reddit posts that mention a brand or keyword?

Two ways, usually combined. For the subreddits you name, we read the New feed on a schedule and match titles, post text and comments against your keyword list, so mentions in those communities are caught as they appear. To find new places, we run Reddit search for each keyword across posts and comments, sorted by New or Relevance with a time filter from the past hour to all time, and add communities that keep turning up to the watch list. Each match keeps the passage, post, community, score and time.

How often can Reddit data be refreshed, and in what formats?

As often as the question needs: hourly or daily for brand monitoring, weekly or once for research. Scores and comment counts keep changing after a post goes up, and Reddit hides vote counts during a post's first hours, so we re-read posts at set ages - for example after 1 hour, 24 hours and 7 days - and store each snapshot with its time. Files come as CSV, JSON Lines, XLSX or Parquet in S3, SFTP or your warehouse, or through our API.

Is it legal to scrape Reddit?

We collect only publicly available data - everything a visitor can see on Reddit - and we collect it legally. That covers public communities with their rules and flair, posts and comment threads, with no logins and nothing from private communities. Usernames are pseudonyms, yet GDPR still treats them as personal data, so we never build user dossiers: no profile pages and no post history of any one person.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582