ParseForge Scrapers

dane.gov.pl Scraper - Poland Open Data Portal API

parseforge/poland-open-data-scraper

BusinessNews & media

Scrape dane.gov.pl, Poland's national open data portal: 26,922 datasets, 1.6M resources, 7,471 institutions, showcases, change history and table rows over the public JSON:API.

Run this scraper See the API call
Total runs
26
Bookmarked
0
Last modified
8 days ago

This scraper was last updated on .

What does the dane.gov.pl Scraper - Poland Open Data Portal API return?

ParseForge

dane.gov.pl Scraper - Poland Open Data Portal API

Pull all 26,989 datasets, 1.6 million resources and 7,476 publishing institutions out of dane.gov.pl, Poland's national open data portal, in one run. Every dataset row carries 49 fields: Polish title and description, publisher with its REGON and address, EU DCAT themes, licence and every reuse condition, update frequency, openness score, file formats, view and download counts and four timestamps. No login, no API key, no rate limit. Exports to CSV, JSON, Excel and XML.

The portal has a public JSON:API, and that sounds like it settles the matter until you try to use it. The listing endpoint ignores a filter name it does not know and hands back the entire catalogue with HTTP 200, so a typo looks like a successful unfiltered query. sort behaves the same way. The tag filter is in the official spec and returns HTTP 400. Licence and update frequency filter on the search index and nowhere else. The resource index stops paging at exactly 40,000 rows with all shards failed, the audit log at 10,000. Single requests hang for a full minute at random. This scraper knows all of that, measures every filter against the unfiltered total before it sends it, paginates around the walls, and gives you a flat table instead of nested JSON:API envelopes.

Who uses it What they scrape dane.gov.pl for
Data engineers Building an open data catalog mirror with formats, licences and refresh cadence per dataset
Property analysts The 6,635 developers legally obliged to publish flat price lists, and their daily updates
Compliance teams High value datasets under the Polish open data act and the EU HVD list
Civic tech and journalists Which public bodies publish what, how often, and which of their links are already dead
AI and RAG builders Polish public sector text and real table rows as grounded training and retrieval material

What it does

  • 📚 Ten collections in one Actor - datasets, resources, institutions, showcases, change history, the broken links report, news, knowledge base, cross-model search, and the actual rows inside a table.
  • 🇵🇱 49 fields on every dataset - title and description in plain text and HTML, publisher with type, city and website, DCAT themes with EU codes, keywords, regions, formats, delivery types, licence, reuse conditions, openness score, update frequency, views, downloads, harvest source and four timestamps.
  • 📊 Real table rows, with real headers - the data API returns col1, col2, col3; this Actor restores the publisher's own column names, so you get Nazwisko aktualne and Liczba, not numbered placeholders.
  • 🔗 Live link checking - one HEAD request per resource records the HTTP status, content type and byte size of the file itself, which finds dead downloads the portal has not flagged yet.
  • 🏛️ Full publisher profiles - REGON, postal address, email, telephone, fax, ePUAP inbox, electronic delivery address, website and the external catalogues the body is harvested from.
  • 🧮 Catalogue breakdowns - one row per format, theme, category, delivery type, openness score and publisher, with counts and share of whatever slice you filtered to.
  • 🇬🇧 English or Polish - English translates the controlled vocabularies and links to the English portal. Titles and descriptions are written in Polish by the publishers, so they stay Polish either way.
  • ⏱️ Paging that respects the walls - the resource index dies past 40,000 rows, the audit log and any single table past 10,000; the run stops cleanly at each ceiling and logs exactly how many rows are out of reach instead of failing. A page that times out is retried, never dropped in silence.

What you can do with dane.gov.pl data

Mirror Poland's open data catalog. Pull all 26,989 datasets with their formats, licences, openness scores and update frequencies, then re-run weekly with a Modified from date to catch only what changed. The 49-field row is flat, so it lands in a warehouse table without any JSON flattening step.

Track the property market at source. Polish developers must publish their flat price lists as open data under the 2021 developer act, which is why 6,635 of the 7,476 registered institutions are developers and why 21,858 of the datasets sit under Economy and finance. Filter to Property developer, turn on the file check and the table preview, and you get the price lists with their live download URLs.

Audit a public body's transparency. Point it at one institution ID, turn on the change history block, and you get every dataset that body publishes, when each was last refreshed against its promised cadence, and the full audit trail of who edited what.

Feed a RAG index with Polish public sector text. Datasets, news and knowledge base articles all carry full descriptions, and the table rows collection gives you the underlying data itself. It is public domain or CC BY in almost every case, so the licence question is already answered in the row.

Why choose this scraper

What you get
Every filter measured Each filter was sent against the unfiltered total before it shipped. The ones this API accepts and then silently ignores are applied after the fetch instead, on the full row, before anything is billed
Nothing always empty The five sparse licence-condition columns are folded into one readable list, and no field ever comes back as a literal null
The walls documented 40,000 rows on resources, 10,000 on the audit log, 10,000 per table on the data API, 26,989 datasets fully reachable, each one bisected and logged at run time
Ten record types Not just the catalogue: the files, the publishers, the reuse showcases, the audit log, the dead-link report, the editorial content and the table contents
Pay for what returned An enrichment block that came back empty is not charged, and the run asks the charging manager before it writes a batch rather than after
No proxy, no key The API is open. There is no proxy cost in your bill and no credential to manage

How it compares

There is one other dane.gov.pl Actor in the Apify Store, benthepythondev/poland-dane-gov-scraper, at $2.50 per 1,000 rows against our $7.00. It is cheaper and it has two billing events against our nineteen, which is the honest summary of the difference: it charges for a start and a row, and that is what it does. Ours costs 2.8x more per dataset row and in exchange covers ten record types instead of one, carries 49 fields on the primary row, restores real column headers on table data, checks whether files still download, and knows where each index stops paging. If you want one flat pass over the dataset list and nothing else, the cheaper one is the right call.

This Actor benthepythondev/poland-dane-gov-scraper The portal's own API
Price per 1,000 datasets $7.00 $2.50 Free
Billing events 19 2 n/a
Record types 10 1 20 endpoints, JSON:API envelopes
Fields on a dataset row 49 not published 37 raw attributes, nested
Table rows with real headers Yes No col1, col2, col3
Live file checking Yes No No
Paging walls handled Yes Not published Fails with all shards failed
Users in the last 30 days new 0 n/a

What a dane.gov.pl dataset looks like

One unedited row from the verified platform run fW4CEt1RpHtSsx2t1:

{
  "datasetId": "1",
  "slug": "dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
  "title": "Dane liczbowe dot. kontroli prowadzonych przez WIIH w 2014 r.",
  "description": "Dane liczbowe dot. kontroli prowadzonych przez wojewódzkich inspektorów Inspekcji Handlowej (IH) w 2014 r.",
  "descriptionHtml": "<p>Dane liczbowe dot. kontroli prowadzonych przez wojewódzkich inspektorów Inspekcji Handlowej (IH) w 2014 r.</p>",
  "url": "https://dane.gov.pl/pl/dataset/1,dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
  "apiUrl": "https://api.dane.gov.pl/1.4/datasets/1,dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
  "externalUrl": "Not Disclosed",
  "institutionId": "26",
  "institutionName": "Urząd Ochrony Konkurencji i Konsumentów",
  "institutionType": "Administracja rządowa",
  "institutionCity": "Warszawa",
  "institutionWebsite": "https://uokik.gov.pl/index.php",
  "dcatThemes": [{ "id": "148", "code": "GOVE", "title": "Rząd i sektor publiczny" }],
  "dcatThemeCodes": [
    "GOVE"
  ],
  "legacyCategory": "Administracja Publiczna",
  "keywords": [
    "uokik",
    "inspekcja handlowa"
  ],
  "regions": [
    "Polska"
  ],
  "formats": [
    "xlsx"
  ],
  "resourceTypes": [
    "file"
  ],
  "visualizationTypes": [],
  "resourceCount": 1,
  "showcaseCount": 0,
  "licenseName": "CC0 1.0",
  "licenseCode": "N/A",
  "licenseDescription": "Other (Public Domain)",
  "licenseConditions": [],
  "updateFrequency": "Rocznie",
  "opennessScore": 2,
  "opennessScores": [
    2
  ],
  "hasHighValueData": "No",
  "hasEuHighValueData": "Not Disclosed",
  "hasDynamicData": "No",
  "hasResearchData": "No",
  "isPromoted": "No",
  "viewsCount": 673,
  "downloadsCount": 138,
  "harvestedFromTitle": "Not Disclosed",
  "harvestedFromUrl": "Not Disclosed",
  "harvestedFromType": "Not Disclosed",
  "harvestedLastImport": "Not Disclosed",
  "supplements": [],
  "imageUrl": "Not Disclosed",
  "imageAlt": "Not Disclosed",
  "createdAt": "2016-04-19T07:39:49Z",
  "modifiedAt": "2024-03-19T19:14:09Z",
  "verifiedAt": "2016-04-19T09:41:11Z",
  "resourceModifiedAt": "2022-12-05T11:40:46Z",
  "scrapedAt": "2026-08-27T20:21:21.768Z"
}

Configure the run

Pick one or more collections, set how many rows you want, and add filters. Max Items caps the whole run, and with several collections selected each gets an even share of it, with anything one collection does not use rolling forward to the next. Everything else has a working default. The catalogue filters run inside the portal index; the post-fetch refinements are the fields the index accepts and then ignores, so they cost more upstream reads and are labelled as such in the schema.

Fresh CSV datasets from central government, newest first:

{
  "collections": ["datasets"],
  "formats": ["csv"],
  "minOpennessScore": "3",
  "createdFrom": "2026-01-01",
  "sortBy": "-created",
  "maxItems": 5000
}

Every file a single publisher offers, with a live download check and the table schema:

{
  "collections": ["resources"],
  "institutionIds": ["26"],
  "includeTabularPreview": true,
  "includeFileProbe": true,
  "maxRowsPerResource": 5,
  "maxItems": 2000
}

The most common Polish surnames, straight out of the PESEL register table:

{
  "collections": ["tabular-rows"],
  "resourceIds": ["28052"],
  "maxRowsPerResource": 1000,
  "maxItems": 1000
}

That table holds 298,692 rows and the data API pages 10,000 of them, so ask for more than that and the run stops at the ceiling and says so.

Free users

Free Apify accounts get 10 rows per run as a preview, enough to check the field set and the filters before committing. Paid plans lift the cap to 1,000,000 rows. Upgrade here.

Run it

  1. Create a free Apify account. It comes with $5 of platform credit, no card needed.
  2. Open the Actor and pick your collections. Leave everything else alone for a first run.
  3. Set Max Items and press Start.
  4. Export from the Storage tab as CSV, JSON, Excel or XML, or pull it over the API.

Use with AI agents (MCP)

claude mcp add apify --transport sse https://mcp.apify.com/sse --header "Authorization: Bearer YOUR_APIFY_TOKEN"

Then ask in plain language:

  • "Get me every CSV dataset published on dane.gov.pl since January, with its publisher and licence."
  • "Which Polish public institutions have broken download links on the open data portal right now?"
  • "Pull the 500 most common Polish surnames out of the PESEL register dataset."

Troubleshooting

No results at all. The most likely cause is a format filter in the wrong case. This index stores formats lower case and CSV returns zero while csv returns 10,698. Narrow date ranges do the same thing quietly.

Fewer rows than I asked for. Three ceilings are real and all three are logged: the resource index stops at 40,000 rows, the audit log at 10,000, and a single table on the data API at 10,000, after which the portal answers all shards failed. Flip Sort by to reach the other end of an index, add a filter to cut the result set below the wall, or use Table row search to reach deeper rows of one table.

A field I expected is empty. Almost every field on this portal is optional. harvestedFrom* is filled on 39% of datasets, externalUrl on 31%, the legacy category on 17%. Empty means the publisher left it blank, and the row says Not Disclosed rather than hiding it.

My licence or update frequency filter did nothing. Those two only filter on the cross-model search index. In the dataset collection they are applied after the fetch, under Post-fetch refinements, which is slower but correct. Use the Search collection with Search: licence for the server-side version.

The run is slower than I expected. Single requests to this API hang for up to a minute at random, roughly one in five. The Actor times each attempt out at 25 seconds and retries against the clock, and runs six pages in parallel, which is why a 5,000 dataset run measured 63 rows per second even though one request averages three seconds. Turning on the opt-in blocks costs one extra call per row and drops that to around 8 rows per second.

FAQ

Question Answer
Do I need an API key for dane.gov.pl? No. The portal's JSON:API is fully anonymous and this Actor uses no proxy and no credentials.
How many datasets are there? 26,989 datasets, 1,613,952 resources, 7,476 institutions, 119 showcases and 269,811 audit log entries when this was measured. The catalogue grows daily, so treat these as a floor.
Can I download the actual CSV files? The row gives you the direct downloadUrl for each resource, and the table rows collection returns the parsed contents of tabulated resources without you downloading anything.
Why do so many publishers look like property companies? Polish developers are legally required to publish flat price lists as open data, so 6,635 of the 7,476 institutions are developers. Filter Institution type to Central government for the ministries.
Is the data in English? The controlled vocabularies are, when you set Language to English. Titles and descriptions are written in Polish by the publishers and are never translated by the portal.
Can I get only what changed since last week? Yes. Use Modified from in the dataset collection, or the Change history collection with a Changed from date for the audit trail itself.
What licence is the data under? 24,098 datasets are CC0 1.0 and 2,825 are CC BY 4.0, with 66 across the NonCommercial and ShareAlike variants. Each row carries the licence name and its reuse conditions.
How fast is it? 63 rows per second, measured on a 5,000 dataset run on the Apify platform that finished in 79 seconds. At that rate the whole catalogue takes about seven minutes. The opt-in blocks are one HTTP call each and cut it to roughly 8 rows per second.
What is the openness score? The portal's 1 to 5 star rating on Tim Berners-Lee's scale. 3 means a non-proprietary format like CSV, 4 adds URIs, 5 adds links out to other data. 13,990 datasets are 3 stars or better.
Does it work on the English portal? Yes. Setting Language to English switches the vocabularies and points every url at dane.gov.pl/en/....

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial scraper, not affiliated with or endorsed by the Ministry of Digital Affairs, dane.gov.pl or the Polish government. It reads only data the portal publishes anonymously to anyone. Institution records contain business contact details of public bodies, published by those bodies as open data; handle any personal data you extract in line with GDPR, CCPA and PIPL, and keep the licence terms carried in every row.

What input does the dane.gov.pl Scraper - Poland Open Data Portal API accept?

FieldTypeWhat it doesDefault
collections array Which dane.gov.pl records to collect. Datasets is the catalogue itself (26,989 records). Resources are the downloadable files and APIs inside them (1.6M). Institutions are the 7,476 publishing bodies. Showcases are the 119 apps built on the data. Change history is the portal audit log. Broken links is the portal's own dead-link report. News and Knowledge base are the editorial content. Search runs one cross-model query. Table rows pulls the actual rows out of a tabulated resource. ["datasets"]
searchTerm string Free text query applied to whichever collections are selected. On the dataset, resource and institution indexes diacritics are folded, so 'budzet' works. The cross-model search does NOT fold them: use 'budżet' there. not set
datasetIds array Numeric dane.gov.pl dataset IDs, for example 1681. Given here, only these datasets are fetched, and the Resources collection is limited to their files. not set
resourceIds array Numeric resource IDs. Required target for the Table rows collection; leave empty there and the ten most viewed tabulated resources are used instead. not set
institutionIds array Numeric institution IDs, for example 26 for UOKiK. Filters datasets to those publishers, and the Institutions collection to those bodies. not set
maxItems integer Free users: limited to 10 items (preview). Paid users: up to 1,000,000. This is the cap for the whole run: with several collections selected each gets an even share and whatever one does not use rolls forward to the next. 10
formats array Keep only records offering at least one of these formats. The index is case sensitive and stores formats in lower case. not set
dcatThemes array The 14 EU data themes the portal tags datasets with. Multiple values are OR'd. not set
legacyCategories array The 13 original dane.gov.pl categories. Only about 1,156 datasets still carry one, so this is a narrow filter by design. not set
datasetTypes array How the data is delivered: a downloadable file, a live API, or a link to an external website. not set
visualizationTypes array Keep only records the portal can render. Table means the rows are queryable through the data API, which is what the Table rows collection needs. not set
keywords array Keyword tags exactly as the portal stores them, for example 'uokik'. Multiple values are OR'd. Note the API's own 'tag' filter is documented but returns HTTP 400, so keywords are the working vocabulary. not set
titleContains string Full text match against the title only, instead of the whole record. not set
notesContains string Full text match against the description only. not set
minOpennessScore string (5 options) The portal's 1 to 5 star openness rating, using Tim Berners-Lee's scale. 3 means a non-proprietary format such as CSV, 4 adds URIs, 5 adds links to other data. not set
createdFrom string Only records first published on or after this date, as YYYY-MM-DD. not set
createdTo string Only records first published on or before this date, as YYYY-MM-DD. not set
onlyHighValueData boolean Only datasets flagged as high value data under the Polish open data act (17,536 datasets). false
onlyEuHighValueData boolean Only datasets on the European Commission's high value datasets list (118 datasets). false
onlyDynamicData boolean Only datasets the publisher marks as dynamic, meaning near real time (3,950 datasets). false
onlyResearchData boolean Only datasets flagged as research data (1,413 datasets). false
onlyPromoted boolean Only datasets the portal editors promote on the front page (5 datasets). false
licenseNames array Keep only datasets under one of these licences. not set
updateFrequencies array Keep only datasets the publisher refreshes at this cadence. Values are the Polish labels the portal stores. not set
modifiedFrom string Only datasets last modified on or after this date, as YYYY-MM-DD. not set
modifiedTo string Only datasets last modified on or before this date, as YYYY-MM-DD. not set
minViews integer Only datasets viewed at least this many times on the portal. not set
minDownloads integer Only datasets downloaded at least this many times. not set
minResources integer Only datasets carrying at least this many resources. not set
searchModels array Which record types the cross-model search returns. Leave empty for all of them. not set
searchLicenseCodes array Licence filter, server side, available only on the cross-model search. not set
searchUpdateFrequencies array Update frequency filter, server side, available only on the cross-model search. not set
searchDateFrom string Cross-model search only. Records dated on or after this day, as YYYY-MM-DD. not set
searchDateTo string Cross-model search only. Records dated on or before this day, as YYYY-MM-DD. not set
institutionTypes array Which kind of publishing body to keep. Developer is by far the largest group because property developers are legally obliged to publish their price lists here. not set
institutionCity string Exact city name as the register stores it, for example Warszawa (1,324 institutions) or Kraków (465). not set
institutionRegon string Polish REGON statistical number, which resolves to exactly one institution. not set
historyAction string (3 options) Which kind of audit log entry to keep. UPDATE covers 231,131 of the 269,810 entries. not set
historyChangedFrom string Only audit log entries recorded on or after this date, as YYYY-MM-DD. not set
historyChangedTo string Only audit log entries recorded on or before this date, as YYYY-MM-DD. not set
tabularQuery string Full text query inside a resource's own table, for the Table rows collection. Searching a name inside the PESEL register works this way. not set
maxRowsPerResource integer How many rows to pull out of each tabulated resource before moving to the next one. Also caps the table preview block. 100
language string (2 options) English translates the controlled vocabularies (themes, update frequency) and links to the English portal. Titles and descriptions are only ever written in Polish by the publishers, so they stay in Polish either way. pl
sortBy string (20 options) Order the records come back in. An unsupported value falls back to the collection's default with a warning, because this API silently ignores a sort field it does not know. not set
includeDatasetResources boolean Attach every resource of the dataset with its format, byte size, openness score, download URL and data date. One extra API call per dataset. false
maxResourcesPerDataset integer How many resources to attach when the block above is on. The API caps a page at 100. 10
includeDatasetShowcases boolean Attach the showcases, meaning the public services and apps, that reuse this dataset. Only fetched when the dataset declares at least one. false
includeDatasetHistory boolean Attach the portal audit log for the dataset: every update with its timestamp, the editor account and which fields changed. false
includeInstitutionProfile boolean Attach the publishing body's full record: REGON, postal address, email, phone, fax, ePUAP inbox, website and its harvest sources. false
includeTabularPreview boolean For resources, attach the column schema with real header names and the first rows of the table. Only tabulated resources have any; the rest return nothing and are not billed. false
includeFileProbe boolean Send one HEAD request per resource to record the HTTP status, content type and byte size of the actual file, which is how you find dead links the portal has not noticed. false
includeFacets boolean After the records, emit one row per facet bucket: every format, theme, category, delivery type, openness score and top publisher with its dataset count and share of the filtered catalogue. false

How much does the dane.gov.pl Scraper - Poland Open Data Portal API cost?

from $6.23 per 1,000 results

Charged forWhat it isPrice each
Actor Start Charged when the Actor starts running. Number of events charged depends on Actor memory (one event per GB, minimum one event). $0.0178 to $0.02
Catalogue page scanned One page of up to 50 records read from a dane.gov.pl index. This is the fixed cost of paging through the catalogue, spread across every row that page yields. $0.00356 to $0.004
Dataset One dane.gov.pl dataset with its full 49-field record: Polish title and description in text and HTML, portal and API URLs, publishing institution with its type, city and website, EU DCAT themes with their codes, the legacy portal category, keywords, geographic regions, every file format and delivery type it offers, resource and showcase counts, licence name, code, plain-English description and every reuse condition, update frequency, openness score, the high value, EU high value, dynamic, research and promoted flags, view and download counts, the external catalogue it was harvested from, supplementary documents, and the created, modified, verified and resource-modified timestamps. $0.00623 to $0.007
Resource One distribution: the downloadable file, API or service inside a dataset, with its format, media type, byte size, direct download URL, CSV and JSON-LD conversions, openness score, language, data date, the protected-data and high-value flags, the special-sign legend the publisher used for missing cells, and the dataset and institution it belongs to. $0.00267 to $0.003
Institution One of the 7,476 publishing bodies with its REGON statistical number, institution type, postal address, city, email, telephone, fax, ePUAP inbox, electronic delivery address, website, dataset and resource counts, and the external catalogues it is harvested from. $0.00712 to $0.008
Showcase One of the 119 public services, websites and apps built on dane.gov.pl data, with its author, external URL, licence type, platform, mobile and desktop store links, keywords, screenshots and how many datasets it reuses. $0.00534 to $0.006
Change history entry One entry from the portal audit log: the action, the timestamp, the editor account, the record it touched, which fields changed and the full before-and-after difference the portal recorded. $0.00267 to $0.003
Broken link One row from the portal's own dead-link report: the institution, the dataset, the resource and the URL that stopped resolving, with the date the report was last rebuilt. $0.00356 to $0.004
News item One portal news article with its title, full text, keywords, author and publication date. News is only reachable through the cross-model search index, never through a listing endpoint. $0.00356 to $0.004
Knowledge base article One knowledge base article: the portal guidance, reports and training material for data publishers, with its category, keywords and full text. $0.00356 to $0.004
Search result One hit from the cross-model search index, which is the only place the portal filters on licence, update frequency and date, and the only way to query datasets, resources, institutions, showcases, news and knowledge base in a single pass. $0.00267 to $0.003
Table row One real row of data out of a resource's own table, with the publisher's column headers restored instead of the API's col1, col2, col3, plus the row number and the date that row was last refreshed. $0.00178 to $0.002
Catalogue breakdown One aggregation bucket: a format, EU theme, portal category, delivery type, openness score or publisher with its dataset count and its share of the filtered catalogue. This is the shape of the catalogue rather than its contents. $0.00178 to $0.002
Files of a dataset Opt in. Every file inside one dataset attached to its row, each with its format, byte size, openness score, download URL, data date and visualisation types. Billed only when the dataset actually returned files. $0.00356 to $0.004
Apps reusing a dataset Opt in. The public services and apps that reuse one dataset, attached to its row with their authors, URLs and licence types. Billed only when the dataset actually has one. $0.00356 to $0.004
Change history of a dataset Opt in. The portal audit trail for one dataset attached to its row: every update with its timestamp, the editor account and the fields that changed, plus the total change count. Billed only when the audit log had entries. $0.00356 to $0.004
Full publisher profile Opt in. The publishing body's complete record attached to the row: REGON, postal address, email, telephone, fax, ePUAP inbox, electronic delivery address, website, dataset and resource counts and harvest sources. Billed only when the profile came back. $0.00445 to $0.005
Table preview Opt in. The column schema with the publisher's real header names and types, the total row count, and the first rows of a resource's table attached to its row. Only tabulated resources have any, and the rest are not billed. $0.00356 to $0.004
Live file check Opt in. One HEAD request against the actual download URL recording the HTTP status, content type and byte size, which is how you find files that stopped resolving before the portal notices. Billed only when the check completed. $0.00267 to $0.003

Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.

How do I call the dane.gov.pl Scraper - Poland Open Data Portal API API?

One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.

POST · run and get results
curl -X POST "https://api.apify.com/v2/acts/parseforge~poland-open-data-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "collections": [
      "datasets"
    ],
    "maxItems": 10,
    "onlyHighValueData": false,
    "onlyEuHighValueData": false,
    "onlyDynamicData": false
  }'

What example inputs can I use?

Use these inputs to see how a run is configured.

input.json
{
  "collections": [
    "datasets"
  ],
  "maxItems": 10,
  "onlyHighValueData": false,
  "onlyEuHighValueData": false,
  "onlyDynamicData": false
}

What do users say about the dane.gov.pl Scraper - Poland Open Data Portal API?

No reviews yet. Be the first.

How do I report an issue with the dane.gov.pl Scraper - Poland Open Data Portal API?

We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.

Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.

What related scrapers can I use?

Run dane.gov.pl Scraper - Poland Open Data Portal API on Apify All scrapers