Geoportal Scraper for Cadastral Parcels, Buildings and Ortho Layers

Poland does not keep its cadastre in one table. It keeps it in 380 county registers behind one map window, so Geoportal scraping is really a stitching job - and stitching is the part we do.

Geoportal Scraper
Solutions

Scoping a Geoportal scraping project: layers, territory, join key

A scoping call answers three questions. Which layers - cadastral parcels and buildings, address points and streets, administrative boundaries, BDOT10k classes, ortho or elevation tiles, utility networks, planning layers, transaction records. Which territory - the whole country, a set of voivodeships, or a list of TERYT codes. And what it should be joined on: parcel identifier, TERYT code, address point or geometry.

From there we build the collectors, keep the county roster current, and run each layer on its own cadence. Where a county portal throttles or challenges automated clients, anti-bot handling, proxy rotation and CAPTCHA solving are part of the service rather than a surcharge. We extract Geoportal data only from services open to any visitor, at polite rates. Delivery is CSV, GeoJSON, GeoPackage, Shapefile, Parquet, a PostGIS drop or a Geoportal API of your own.

What a Geoportal parcel record carries: identifier, area, land use, buildings

The parcel identifier. Every parcel carries a TERYT-based identifier shaped WWPPGG_R.OOOO.NR. In 141215_1.0051.2/7 the parts are voivodeship, county, commune, commune type, cadastral district and parcel number; the slash records a split, so 2/7 is the seventh piece cut from former parcel 2. Buildings extend the same string with a building number and the suffix _BUD. Because the type digit separates a town from the countryside around it, an urban-rural commune appears as two codes, 141204_4 and 141204_5.

What a parcel lookup returns. ULDK, the parcel location service, answers sixteen request types over plain GET: parcel by identifier, by coordinates, by district name and number, plus building, district, commune, county and voivodeship lookups, vertex snapping and parcel merging. A result parameter picks the fields - id and teryt, parcel number, region, commune, county, voivodeship, function for buildings, geom_wkt or geom_wkb, geom_extent, and datasource, naming the county service that answered.

Attributes from the map layer. A point query on the KIEG parcel layer returns the parcel identifier, voivodeship, county, commune and cadastral district names, parcel number, registered area in hectares, register group code, land-use designation, classification contour code such as Bi, and publication date. Sibling layers add buildings, contours and an LPIS fill-in layer.

Bulk county packages. Per-county GeoPackages harvested from those services carry parcels with id_dzialki, numer_dzialki, nazwa_obrebu, numer_jednostki, nazwa_gminy, pole_ewidencyjne, klasouzytki_egib and grupa_rejestrowa, and buildings with id_budynku, a one-letter type code spanning residential, garage, industrial, office, transport and health, and storey counts above and below ground. One mid-sized county holds about 268,000 parcels and 99,000 buildings.

Boundaries and addresses. The register of boundaries publishes state, voivodeship, county, commune, town, cadastral unit and cadastral district polygons, statistical districts, and the service areas of courts, police and tax offices. Each feature carries JPT_KOD_JE as its TERYT code, JPT_NAZWA_ as its name, JPT_POWIER as area, plus INSPIRE identifiers and validity dates; address points and streets come from the same register.

What a Geoportal parcel record carries: identifier, area, land use, buildings
Beyond the cadastre: ortho vintages, elevation, utility networks and plans

Beyond the cadastre: ortho vintages, elevation, utility networks and plans

Orthophoto. Coverage is indexed one layer per acquisition year, forty-nine of them from 1957 to the current season, each feature carrying the sheet code, flight date, pixel size, colour mode, coordinate system, file size and a download link. The standard product is 25 cm, towns are flown at 10 cm or better, and the sharpest sheets reach 3 cm. True ortho, which removes the lean of buildings, is a separate service.

Elevation. The terrain model is a 1 m grid, dropping to 5 m where only stereo measurement exists; the surface model covers the same sheets. Raw laser scanning comes as LAS and LAZ with classes for ground, low, medium and high vegetation, buildings, water, noise and overlap, at 4 to 20 points per square metre; coverage services return the same rasters as ASCII Grid or GeoTIFF.

Topography and buildings. BDOT10k holds roughly 60 million objects in nine classes - watercourses, road and rail networks, utility lines, land cover, protected areas, administrative units, buildings, development complexes and other objects - at 1:10,000 detail. Building solids come as CityGML 2.0, LoD1 nationwide in several vintages and LoD2 for 236 counties.

Networks, plans and zones. Utility networks split by medium: water, sewer, power, gas, district heating, telecom, special and unidentified, plus devices. Planning layers carry local plans as raster and vector, general commune plans adopted and drafted, and landscape resolutions. Around the edges sit forgotten layers such as wind turbine exclusion zones, carrying mast height beside the polygon.

Inside geoportal.gov.pl: one map window over 380 county registers

Geoportal.gov.pl is the national geoportal of the Polish spatial information infrastructure, run by the Head Office of Geodesy and Cartography (GUGiK) under the Surveyor General. Four hosts do four jobs: the editorial site carries the catalogue, mapy.geoportal.gov.pl runs the map client and most OGC endpoints, integracja.gugik.gov.pl the aggregate services, opendata.geoportal.gov.pl the bulk files. The interface is Polish with a full English mirror, and layer names stay Polish in both.

The structure that matters to anyone planning to scrape Geoportal is administrative rather than technical. The land and building register, EGiB, is kept by the head of each county, which means 380 separate registers instead of one national table. GUGiK papers over that with aggregate services: KIEG for parcels and buildings, KIUT for utility networks, KIN for address points and streets, KIMP for local plans, KICN for transactions, KISOG for control networks. Each is a proxy that fans a request out to the county servers behind it and merges what comes back, so national coverage depends on which counties have joined.

Every registered dataset appears in EZiUDP, the register of spatial data sets and services, under an identifier such as PL.PZGiK.1, with the reporting authority, its REGON business-register number, the TERYT code of the area, INSPIRE theme codes and the URLs of its discovery, view and download services. There were about 11,900 such datasets in September 2026, 385 of them EGiB entries pointing at roughly 368 separate county hosts.

Nothing sits behind a login. The site terms state that material published there falls outside copyright protection under Article 4(2) of the Polish copyright act, and that republication is allowed provided the source is named.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

Why national coverage is harder to assemble than the map window suggests

The first problem is arithmetic. Cadastral parcels are not published once, they are published 385 times, by 385 authorities, across roughly 368 hosts running four or five different vendor stacks. The aggregate services hide that while a map is being drawn, and only while a map is being drawn. The moment you want attributes in rows rather than pixels in a tile, you are back to per-county endpoints with per-county behaviour.

The second is format. The national download services speak GML and only GML, in three flavours, with no GeoJSON option, so every harvest needs a parser before it needs a database. Paging is a bounding box plus a start index and a count, which makes the collection plan a tiling plan: cut the country into boxes small enough to stay under the response cap and keep a cursor for each one. The national grid also puts northing before easting, and half the sample calls in circulation have the pair the wrong way round.

The third is drift. Published example calls stop working when a service is reconfigured, and an endpoint that answers a capabilities request can still refuse a feature request. Elevation exists in two vertical datums, and the same sheet is published in both. Transaction dates arrive with plainly mistyped years. Parcel numbers containing slashes break naive delimiter handling. That is the case for buying Geoportal data as a maintained feed instead of building one: we hold the county roster, the tiling plan, the datum handling and the field mapping.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who builds and runs your Geoportal data feed

ScrapeIt runs the whole job as a managed service. We design the collectors, schedule them, watch the runs and repair them when a county changes supplier, a layer is renamed or a field quietly disappears. Your side receives files or an endpoint and keeps no scraping engineers of its own.

We are not affiliated with GUGiK or any Polish public authority. Details of who owns a parcel are not part of the public layers and we do not collect them: we work with geometry, identifiers, addresses and object characteristics.

FAQ

Does Geoportal have a public API, and what does it actually return?

Not in the way most buyers mean. The page labelled API service describes a JavaScript library for embedding a map in a web page, not a data endpoint. What does exist is a set of real machine interfaces: the ULDK parcel locator, which answers plain GET requests and returns geometry as well-known text or binary; the EZiUDP register, which serves JSON and semicolon-separated CSV with filters for TERYT, dataset, INSPIRE theme and service type; and the OGC stack of view, download, coverage, catalogue and predefined-download services. There is no single endpoint that hands over the whole cadastre, which is exactly the gap we fill.

Which Geoportal fields do you extract for parcels and buildings?

Identifiers first, because they are what turn separate layers into one graph: parcel identifier and its TERYT parts, cadastral district number and name, cadastral unit, commune, county and voivodeship, building identifier with its _BUD suffix. Then the payload: registered area in hectares, land-use and soil-class designations, classification contour code, register group, building type code, storeys above and below ground, and the source service with the moment it was harvested. Boundary and address layers add the TERYT code, official name, area and validity dates. If your model needs something we do not see published, we say so rather than inventing a column.

How often does Geoportal data change, and how do you keep a feed current?

Each layer moves on its own clock. County cadastral services publish continuously and the national parcel and building packages are re-harvested every few days, so a parcel edit in one county surfaces without waiting for the rest. Boundaries and the address register change with administrative acts. Orthophoto and elevation arrive as whole vintages, town by town, on a multi-year rotation. We poll each layer at its own cadence, diff against the previous run on stable identifiers, and deliver changes rather than full reloads, so a difference between two runs is real movement and not re-delivery.

Can we get Geoportal price data from the property transaction register?

Partly, and the nuance matters. The register of real estate prices is kept by each county from notarial deeds, and the portal states that its view and download services do not expose the transaction price or detailed property characteristics, with fuller access charged under the statutory tariff. What is published openly are per-county transaction packages split into parcels, buildings and premises, carrying gross price and VAT, deed date, market type, transaction type, property type, land or usable area, land-use and plan designation, room count and floor, with buyer and seller recorded as a category rather than a name. We take what is open and flag the rest.

Is scraping Geoportal legal, and do you touch owner information?

The site terms say that material published there is not covered by copyright protection under Article 4(2) of the Polish copyright act, and that republication is permitted as long as the source is named. Orthophoto, elevation, laser scanning, topographic objects, boundaries and geographical names are stated on their own pages to be free of charge and usable for any purpose, and parcel and building geometry with basic attributes is likewise open. On people: owner details are not in the public layers, transaction parties are stored as category codes such as natural person or State Treasury, and we neither collect nor attempt to reconstruct anything about individuals.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582