SITC Scraper for Annual Meeting Sessions and Abstracts

SITC publishes its annual meeting as a filterable schedule and a searchable abstract database. Our SITC scraper delivers both as tables of sessions, titles and institutions.

Plans from €169/month · Free project assessment · Reply within 1 business day

SITC Scraper
Solutions

A SITC dataset built and maintained for you

The meeting site is assembled in the browser: the schedule and the abstract database fill in after the page loads. We render the pages, read every session and every abstract entry, and deliver them as rows. You choose the meeting year, the sections and the moment - for example once when the regular titles are released and again when the late-breaking ones join them.

We build the crawlers, run them and re-fit them when the site is rebuilt for the next meeting; rendering, pacing and proxy rotation are part of the service. Member areas behind the login are not collected. Names of speakers and authors are personal data and stay out of the dataset by default - organisations, titles, numbers and categories go in.

SITC data we extract from the schedule and abstracts

Two record types make up most of a SITC extract: the session and the abstract. Each has its own key on the site - sessions carry a number in the title, such as Session 106a, and every abstract has an abstract number that also decides its poster day.

  • Sessions - session number and title, date, start and end time with the time zone, and the room given as building, level and hall.
  • Session type - the label the schedule prints above each entry, among them Pre-Conference Program, Regular Session, Concurrent Session, Keynote / Plenary, Poster Session, Sponsored Symposium and Tech Talk.
  • Session schedule - inside a session, each presentation with its time slot, its title and the speaker's institution.
  • Abstract titles - abstract number, title, presentation type (Poster or Oral) and the day it presents.
  • Classification - Primary Category, Abstract Type, the keywords list and the Top 150 marker.
  • Organisations - the presenting author institution as written, which may be a university, a cancer centre, a hospital or a company; an entry with more than one affiliation keeps them all.
  • Industry sessions - for sponsored symposia and tech talks, the company behind the slot, the talk title, time and room.
  • Exhibitors and supporters - company name with booth number from the floor plan, and supporter names with their tier and the programme they support.

Speaker, chair and author names are shown on the site but are personal data, so the default dataset leaves them out and keeps the organisation. Every row carries the meeting year and the address it was read from.

SITC data we extract from the schedule and abstracts
SITC abstracts over time: embargo dates, past meetings, guidelines

SITC abstracts over time: embargo dates, past meetings, guidelines

Meeting content is released in steps, and a dataset is only as complete as the step it was taken at.

  • Release steps. Titles and author information of regular abstracts appear first, about a month ahead; late-breaking titles join in the last week. The embargo on regular abstracts runs until the eve of the meeting, on late-breaking abstracts until the first poster day. We time the runs to those dates and mark each row with the step it reflects.
  • Filters as data. The abstract database counts entries per Primary Category, per presenting institution, per presentation type and per abstract type. Those counts are a ready summary of who presents what, and we deliver them as their own table.
  • Past meetings. Earlier meetings stay online under their year, and the archive page lists past editions with their dates and city. Running the same extraction over several years gives a comparable series of categories, titles and presenting organisations.
  • Schedule changes. The site notes that session details are added as they are finalised. Repeated runs keep each version, so a moved or withdrawn item is visible.
  • Other public sections. The clinical practice guidelines index, the events calendar with dates, time zones and places, and the news posts with their dates can join the same project.

How the SITC annual meeting site is organised

SITC is the Society for Immunotherapy of Cancer, a non-profit medical society based in Milwaukee, Wisconsin. Its website, sitcancer.org, serves members and the wider field: education programmes, career development, the society's journal JITC and its guidelines, news, and the annual meeting, which is where a data project usually starts.

Each meeting has its own area under the year - the current one is SITC 2026, held at the Phoenix Convention Center with pre-conference programmes before the main days. Inside it the menu splits into Schedule, Abstracts, Exhibit and Support, Attendee Resources and Press. The schedule can be narrowed to a single day and sorted by time, title, session type or room. The abstract section, Titles and Publications, is a search with filters for category, presenting author, institution, presentation type and abstract type. Exhibitors sit on an interactive floor plan with booth numbers, supporters on a page grouped by tier, and Industry Presentations lists sponsored symposia and tech talks by company.

Not everything is open. The member directory, community discussions and on-demand recordings ask for a login, and none of that is part of what we collect. What remains is the programme itself: what is presented, when, in which room, under which category and from which organisation.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

SITC Scraping Plans and Pricing

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Delivery: CSV, JSON or XLSX files, our API or a JSON feed. MCP server on request.

Full plan details →

Why scrape SITC meeting data

The abstract database of SITC 2026 alone holds more than a thousand titles, and the cancer immunotherapy field reads it through filters, a page at a time. A table changes that.

  • Competitive intelligence. Biopharma teams filter abstract titles by category and keyword to see which modalities - cellular therapies, checkpoint blockade, immune cell engagers - other companies bring this year, and from which institutions.
  • Conference planning. Medical affairs and business development build agendas from sessions, rooms and times without copying them by hand, and spot clashes between concurrent sessions.
  • Market mapping. Exhibitors and supporters show which tool makers, contract research organisations and drug developers invest in the field, and at what tier.
  • Research analytics. Counts per Primary Category and per institution, compared across meeting years, show where the field's attention is moving.
  • Content and alerting. Publishers and analysts prepare coverage from titles on the release date and update it when the embargo lifts.

The dataset describes the programme, not the medicine: it carries titles, numbers, categories and organisations and offers no clinical guidance.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

About ScrapeIt

ScrapeIt is a managed data collection agency. For scientific and medical sources we do the same job as for any other site - build the crawler, run it, deliver documented tables - with stricter defaults: no patient data, no member areas, and names and contact details of researchers left out by default. Personal data is handled under GDPR, and we do not sell contact lists of people.

Society meeting sites like SITC change with every edition, so the work includes re-fitting the extraction each year and keeping the columns stable across meetings.

FAQ

Do you offer a SITC API for annual meeting data?

Yes. The ScrapeIt SITC API returns session number, title, session type, time and room, and for abstracts the number, title, Primary Category and presenting institution, as JSON from an endpoint we host, refreshed on the schedule you set. The same tables are also sent as CSV or XLSX files, through a JSON feed or into your database.

Is it legal to scrape SITC?

Yes. The schedule, abstract titles, exhibitor floor plan and supporter list are open pages of the meeting site. We collect only publicly available data - everything a visitor can see on SITC without logging in - and we collect it legally. Member-only areas are never entered, and personal names are left out by default.

When are SITC abstracts released, and can runs follow those dates?

Titles and author information of regular abstracts are published about a month before the meeting, and late-breaking titles in the last week before it. The embargo on regular abstracts lifts on the eve of the meeting, and on late-breaking abstracts on the first poster day. We schedule a run for each release, so the dataset follows the listing as it grows.

Can you scrape SITC abstracts and sessions from past meetings?

Yes. Recent meetings remain on the site under their own year with their abstract database and programme pages, and the archive lists past editions with their dates and city. One extraction across several years returns a single table with a meeting-year column.

Are speaker and author names included?

Not by default. Names are personal data, so the standard dataset keeps the organisation - university, hospital or company - and drops the person. Sessions, titles, abstract numbers, categories and keywords are delivered in full.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

SITC data from €169/month. Free project assessment, reply within 1 business day.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582