A Polish Law Text Database Scraped from LexLege, Article by Article

LexLege does the tedious part of Polish law: every code cut into addressable articles, every amendment dated, the next one flagged. We turn that into rows and tell you when ISAP must have the last word.

LexLege Scraper
Solutions

Scoping a LexLege scraping project: acts, depth, refresh

A scoping call settles three things. Which acts: the whole catalogue, the codes and tax statutes, or a watchlist of named laws. Which depth: current text by article, then per-article history with redlines, then the upcoming-changes calendar. Which refresh: a one-off corpus, a weekly pass or a daily diff driven by the change feed. Those answers set the price of a LexLege feed; page counts do not. On request each act also carries its ISAP entry and publication date.

The site sits behind Cloudflare, and anti-bot handling, proxy rotation and CAPTCHA solving are part of the service rather than a surcharge. We extract LexLege data only from pages open to any visitor, at polite rates: no login, no subscriber case law, no paid templates. Output comes as JSON or JSONL for retrieval pipelines, CSV or Parquet for analysts, or an API of your own, with source URL, base citation and collection time on every row.

LexLege data fields: act, article, paragraph, point and version

A LexLege scraper works on nested levels, and each level has an address of its own, so every row traces back to the page it came from.

  • Act. Short title, full title such as Ustawa z dnia 26 czerwca 1974 r. - Kodeks pracy, the abbreviation the site uses (KP, KEA, pr. poczt.), the base citation in the form Dz.U.YYYY.0.POS with a t.j. suffix when it is a consolidated text, the slug, the archive flag and the PDF export. A table of contents lists every DZIAL, Rozdzial and Oddzial with its article range, and each division has a page of its own, as in /kodeks-pracy/rozdzial-iv-regulamin-pracy/260. Regulations open with a preamble quoting their legal basis, which ties an implementing act to its parent statute.
  • Article. One URL per unit: /kodeks-pracy/art-30, art-22-2 for article 22 with superscript 2, letter suffixes such as art-119zzl in Ordynacja podatkowa, and paragraf-1 for units of regulations and ethics codes. Each carries its number, an editorial title such as Rozwiazanie umowy o prace - a heading the official text does not have - and a body split into paragraphs (paragraf in the codes, ustep elsewhere) and numbered points, with repealed units kept as uchylony.
  • Cross-references. Citations inside the text are live links to the cited article. Kodeks pracy holds 573 article units and 334 such links, 21 of them into other acts such as the trade union act, the sickness benefits act and GDPR.
  • Article version. Historia jednostki lists the earlier states of an article back to 2011 with Stan prawny na, the amending act as Dz.U.2023.0.641.zm, Obowiazuje od and the full wording of that version, removed and inserted words marked separately - a redline, not just a date.
  • Act version. Wersje aktu gives one row per version: date, status (Obowiazujacy stan prawny, Nadchodzace zmiany or Archiwalna), the amending Dz.U. position and the articles it touched. Ordynacja podatkowa alone runs to 152 versions, from 2011 to changes scheduled for 2027.

Around the text sit templates linked to each article and Prawo w praktyce commentary from the same branch of law. We map both as references rather than copying them.

LexLege data fields: act, article, paragraph, point and version
Polish consolidated acts on LexLege: change feed, archive flags, court rulings

Polish consolidated acts on LexLege: change feed, archive flags, court rulings

The change feed. Zmiany w prawie runs two tabs, Ostatnie zmiany and Najblizsze zmiany, both grouped by effective date, each act listed with the number of changes it takes that day and a link to its version table. A daily pass only has to re-read what moved.

Labels that mislead. Each act header prints Stan prawny aktualny na dzien with the current date, and each page title promises an aktualny tekst jednolity - on repealed acts, on ministerial regulations with no consolidated text and on EU regulations alike. Repeal shows only as the AKT ARCHIWALNY prefix in the title and slug. The base citation can trail the journal: at the end of September 2026 Kodeks pracy still named Dz.U. 2025 poz. 277, days after Dz.U. 2026 poz. 1245 appeared. An address for an article that does not exist still renders an article page with an empty body, so every unit is checked for real text before it enters a dataset.

Court rulings. The home page title promises orzeczenia, yet the open search returns articles, acts and documents only. Case-law lines sit in the paid LexLege SIP package, just as rulings once sat in the LexLege PRO tier of the old lexlege.pl.

Templates. The wzory pism library claims over 2,000 documents at a flat 25 zl each; we index titles, categories and types, never the paid files.

LexLege vs ISAP: a private reading layer over Dziennik Ustaw

Before buying LexLege data, know what LexLege is: a private legal information service, not an official publisher. It grew up in the ArsLege group of Prawomaniacy sp. z o.o., an Olsztyn company known for ArsLege law exam tests, and in 2026 it moved under Puls Biznesu, the Bonnier Business Polska daily. The old lexlege.pl address now forwards to lexlege.pb.pl, where the service calls itself a System Informacji Prawnej. The interface is Polish only.

The statute text is not LexLege's own work. Act pages name the Dziennik Ustaw position they are built on - for Kodeks pracy the official tekst jednolity Dz.U.2025.0.277.t.j. - and LexLege's editors fold every later amending act into that base. The site calls the result ujednolicone teksty, the Polish term for an unofficial consolidation, and that is exactly what it is. The first-hand record is ISAP, the Sejm Chancellery's register of acts from Dziennik Ustaw and Monitor Polski; anyone who needs the complete Polish legal acts database should start there. LexLege is a reading layer on top of it.

What the layer adds is structure. ISAP records each act with its text as published and its consolidations as PDF files; LexLege serves the wording in force article by article, every article at its own address, with an editorial title, a dated change log and a calendar of changes not yet in force. The catalogue is curated rather than complete: about 1,290 documents in its act list in September 2026 - roughly 590 statutes, 340 regulations and 300 repealed acts kept under an AKT ARCHIWALNY label. It also carries texts the official journals never print, such as the ethics codes of advocates, legal advisers, notaries and bailiffs and the Warsaw exchange rulebooks for NewConnect and Catalyst.

Get a Quote
dev_w
25

Developers

customers
500+

Customers worldwide

pages
1 500 000 000+

Pages extracted

stime
15000+

Hours saved for our clients

Plans

Airplane

€199 / one-time

setup fee - included

Data limits100,000
Frequencyone-time
Run timeup to 5 days
Data storing7 days

Helicopter

€169 / mo

setup fee €499

Data limits250,000
Frequencymonthly
Run timeup to 5 days
Data storing14 days

Glasses

€229 / mo

setup fee €499

Data limits1,000,000
Frequencyweekly
Run timeup to 5 days
Data storing30 days

DNA

€549 / mo

setup fee €799

Data limits3,000,000
Frequency3 times daily
Run timesame day
Data storing90 days

What a LexLege scraper feeds: article search, model training, amendment alerts

Teams buy LexLege data for the shape of the law more than for the words, which the journal already publishes. Three uses keep coming back.

Retrieval that cites the article. Most questions about Polish law land on one provision - art. 30 of Kodeks pracy, art. 14 of the PIT act - and LexLege already cuts the law at that grain, with a stable address for every article and its paragraphs and points in order. A legal-tech assistant built on it can quote the unit it answers from, the editorial titles double as a subject index the statute lacks, and the in-text links give a citation graph without writing a parser.

Before-and-after pairs for models. The article history pairs each superseded wording with its replacement, the amending act and the in-force date, deletions and insertions marked. That is ready material for models that summarise a nowelizacja or answer as of a given day. Statute text is free of copyright; LexLege's titles and commentary are not, so they stay out of a training corpus unless you hold the rights.

A calendar of what changes next. Payroll, tax and compliance teams care about the day a rule starts to apply. LexLege dates each change by entry into force and lists changes still in vacatio legis, so a watchlist of acts becomes a list of affected articles weeks ahead. It will not replace the official record: the catalogue covers a fraction of the journal, and final wording still gets checked against Dziennik Ustaw through ISAP.

Our Blog

Reads Our Latest News & Blog

Learn how to use web scraping to solve data problems for your organization

How Artificial Intelligence Is Used In Web Scraping

How Artificial Intelligence Is Used In Web Scraping

Leveraging advances in technology, the AI-powered web scraper has skyrocketed in demand and is helping to expand capabilities by automating tedious daily tasks and speeding up data collection from thousands of websites several times over.

What is Web Scraping and What is it Used For?

What is Web Scraping and What is it Used For?

Web scraping is a method of obtaining web data by extracting it from pages of web resources with the help of a program, that is, in automatic mode. It is used to syntactically convert web pages into more usable forms.

Web Scraping for Machine Learning

Web Scraping for Machine Learning

If you specialize in machine learning, you need to feed large amounts of data to the algorithms. Web scraping is the easiest and the most efficient method of collecting the data from all over the Internet.

scrapeit logo

Who runs the LexLege feed, and what stays out of it

ScrapeIt runs the job as a managed service. We design the collectors, schedule them, watch every run and repair them when LexLege renames a block, changes a URL pattern or moves host again, as it did in 2026. You receive files or an endpoint and keep no scraping engineers of your own.

We are not affiliated with LexLege, Puls Biznesu, Bonnier Business Polska, ArsLege or the Chancellery of the Sejm. Commentary bylines, author e-mail addresses and the questions readers send to the site's legal advice form stay out of every delivery: we collect law, not people.

FAQ

Does LexLege have an API or a bulk download?

No. LexLege publishes no public API, data feed, bulk export or sitemap. What it offers a machine is HTML pages at stable addresses and a Pobierz PDF button that prints the current version of one act at a time, with superscript article numbers flattened, so art. 3 with superscript 1 and the repealed art. 31 read the same. If you need a Polish legislation API in the official sense, the Sejm's ELI service behind ISAP is that route, covered on its own page. For LexLege we deliver a LexLege API of your own: acts, articles, versions and changes as JSON.

Is LexLege an official source of Polish law?

No. Polish law binds in the wording published in the official journals, Dziennik Ustaw and Monitor Polski, and a consolidated text counts only once it is itself announced there. LexLege publishes unofficial consolidations, adds article titles of its own and can take days to switch to a newly announced tekst jednolity. That makes it excellent for reading, searching and tracking, but for a contract clause, a court filing or a compliance opinion the wording should be checked against the official publication, which ISAP links for every act.

Does LexLege include Polish court rulings data?

Not on the open site. Its page title mentions rulings, but public search covers only articles, acts and document templates, and the case-law lines LexLege does hold are part of the paid LexLege SIP package, behind a subscription we do not enter. Before the move to Puls Biznesu, rulings of the common and administrative courts also sat behind a paid PRO tier. For a Polish court rulings dataset, scope the courts' own portals instead: orzeczenia.ms.gov.pl for common courts and orzeczenia.nsa.gov.pl for administrative ones.

How current are the Polish consolidated acts on LexLege, and how often can a feed refresh?

For most major codes and the PIT, CIT and VAT acts the base citation matches the newest consolidated text in Dziennik Ustaw, and every change carries the date it enters into force, with upcoming changes listed in advance. Two caveats: a freshly announced tekst jednolity can take days to replace the base citation, and the date in the act header is simply today's date. A daily diff driven by the Zmiany w prawie feed keeps a dataset in step; weekly or monthly passes suit research corpora.

Is it legal to scrape LexLege data?

Statutes and regulations are not protected works under art. 4 of the Polish copyright act, so the law text itself is free to reuse. LexLege's own layer is a different matter: article titles, version tables, commentary and templates belong to the publisher, and the Bonnier terms covering pb.pl claim copyright and database rights and limit reuse beyond reading. We therefore scope LexLege projects for internal analysis and monitoring, keep every row traceable to its source, leave paid and subscriber content alone, and point you to ISAP for text you plan to republish. We collect no personal data.

How does it Work?

Step 1 - Make a Request

You share your needs, expectations, and desired timeframe. We’ll suggest the best solution based on your request and budget.

Step 2 - Configuring Custom Web Crawlers

Our specialists configure the crawlers and extract a sample dataset for your review before proceeding with the full-scale extraction.

Step 3 - Collect and Deliver

Once you approve the sample, we launch the project and start full data collection. We gather, filter, and structure the data for easy use, delivering it on time in your preferred format.

Step 4 - Maintain and Support

Our team manages ongoing processes, monitors website changes, and supports all data extraction cycles. We can also help integrate data into your systems or create dashboards to simplify analysis.

Request a Quote

Tell us more about you and your project information.
Which sites, which fields, how often. A couple of lines is enough.

We reply within 1 business day. No obligation.

scrapiet

Scrapeit Sp. z o.o.
10/208 Legionowa str., 15-099, Bialystok, Poland
NIP: 5423457175
REGON: 523384582