ParseForge Scrapers

Netherlands Open Data Scraper - data.overheid.nl API

parseforge/data-overheid-netherlands-open-data-scraper

Business

Scrape all 20,380 datasets from data.overheid.nl, the national open data portal of the Netherlands: full DCAT-AP-DONL metadata, distributions, 4,667 data requests, registers and live link checks.

Run this scraper See the API call
Total runs
24
Bookmarked
0
Last modified
8 days ago

This scraper was last updated on .

What does the Netherlands Open Data Scraper - data.overheid.nl API return?

ParseForge

Netherlands Open Data Scraper - data.overheid.nl CKAN API

Scrape every one of the 20,380 datasets on data.overheid.nl, the national open data portal of the Netherlands, with all 86 metadata fields the CKAN API carries. Each row is a flat DCAT-AP-DONL record: the data owner as an OWMS authority, the EU themes in Dutch and English, licence, access rights, update frequency, contact point, legal foundation and every file with its format. No login, no API key, no rate limit. Export to CSV, JSON, Excel, or XML.

The portal's own search shows you ten results a page and hides most of the record behind them, and its government open data API answers only in Dutch URIs: a licence comes back as http://creativecommons.org/licenses/by/4.0/deed.nl and a theme as .../owms/terms/Natuur_en_milieu. This Actor pages the whole open data catalogue in one run, resolves every URI to a readable name, and adds four registers the CKAN API does not expose at all, including the 4,667 public data requests Dutch citizens have filed asking the government to open a dataset.

Who uses it What they scrape data.overheid.nl for
Data engineers A complete inventory of Netherlands government datasets to pipe into a warehouse
Open data researchers Licence, access-rights and update-frequency coverage across 187 public bodies
Civic technologists Live download links by format, checked, before building on top of them
Journalists The 4,667 data requests and the government's own answer to each one
Compliance and policy teams The 229 EU High Value Datasets and the 34 statutory basisregistraties

What it does

This Actor reads the CKAN open data API behind data.overheid.nl and the five community registers on the site itself, and returns each entry as a flat row tagged with rowType:

  • ๐Ÿ‡ณ๐Ÿ‡ฑ Datasets: all 20,380, with 86 fields each: title, description, national identifier, data owner and publisher, harvest organization, source catalogue, OWMS themes in Dutch and English, keywords, licence with an open-licence flag, access rights, availability status, update frequency, contact point with email, phone and website, issue and modification dates, temporal coverage, spatial extent and the legal foundation with its wetten.overheid.nl link.
  • ๐Ÿ“ฆ Files: around 74,600 distributions, 3.66 per dataset, with access and download URLs, the EU file-type URI resolved to a plain name, a machine-readable flag, media type, preview URL and licence.
  • ๐Ÿ“ฃ Data requests: all 4,667 public requests to open a dataset, with the phase they reached in Dutch and English, the body that was asked, and optionally the full request and the government's written answer.
  • ๐Ÿ›๏ธ Registers: 1,671 public bodies with their governance layer (Rijk, Gemeente, Provincie, Waterschap), 91 curated groups, 36 source catalogues and the registered applications built on the data.
  • ๐Ÿ“š Directories: 187 data owners, 99 themes, 40 formats, 8 licences, 23 source catalogues and 10,403 keywords, each with the dataset count and a working link to its slice of the portal.
  • โœ… Flags the portal buries: 115 datasets flagged high value and 229 carrying an EU High Value Dataset category, 34 basisregistraties, 6,909 nationally-scoped datasets, 51 reference-data sets, and the 2,020 datasets catalogued but not openly downloadable.
  • ๐Ÿ”Ž Optional checks: fetch each file to see whether it still downloads (9.0 percent are not), and read its first bytes to see whether it really is what it claims (10.0 percent are not: 24 files declared JSON in a 469 file sample served an HTML page instead).

Results export to CSV, JSON, Excel, or XML, or stream from the API.

What you can do with data.overheid.nl data

๐Ÿ—‚๏ธ Build a complete inventory of Netherlands government data. One run with no filters returns all 20,380 datasets with their owners, licences, formats and update frequencies. Tick the file rows and you also get every download URL in the country's open data catalogue, ready to schedule.

๐Ÿ”— Find the dead links before your users do. data.overheid.nl harvests 23 separate catalogues and records a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. Turn on the link check and each file comes back with its real HTTP status, redirect target, content type and size. In a 647 file sample, 9.0 percent were dead: mostly WFS and WMS services answering 500, plus spreadsheets behind a 403.

โš–๏ธ Audit open data policy compliance. Filter to the 229 datasets in an EU High Value Dataset category, or to the 2,020 datasets whose access rights are not public, or to the 1,376 published under a closed licence, and you have the evidence for a coverage report in one export.

๐Ÿ“ฃ Read what the public is asking for. The data request register is 4,667 requests from citizens, journalists and companies asking a named public body to open a dataset, each with the phase it reached and the answer it got. 336 were refused as not public and 1,970 ended with the government pointing at data that already existed.

Why choose this scraper

What you get
Full coverage All 20,380 datasets, 86 fields each, plus five registers the CKAN API does not expose
Readable, not URIs Every OWMS and EU URI resolved to a name, themes in Dutch and English
Filters measured against the live totals 28 filters, each one checked against the catalogue count, with the real number in the dropdown
Clean rows ISO dates, real numbers, Yes/No booleans, "Not Disclosed" where the portal withheld a value, zero always-empty columns
A register nobody else carries 4,667 public data requests with the government's own answer
Priced per row You pay for rows written, and the optional checks only when they returned something

How it compares

One other Actor covers this exact source, and two more cover CKAN portals generically. The generic ones are cheaper per row and give you what a bare CKAN call gives you: an id, a name, a few tags. The named data.overheid.nl Actor lists six fields in its own store description. This Actor costs more per row than either, and the difference buys the other 80 fields, the URI-to-name resolution the portal makes you do yourself, the five community registers, and filters that were each measured against the live catalogue instead of copied from CKAN documentation.

This Actor benthepythondev/netherlands-data-overheid-scraper Generic CKAN Actors The portal's own API
Fields per dataset 86 6, per its own description Whatever CKAN returns raw Raw URIs only
Filters 28, measured against totals None documented Free text You write the Solr query
Data requests 4,667 No No Not exposed
Registers 5 No No Only organizations
Link checking Yes, optional No No No
Users in the last 30 days New 1 1 each n/a

What a Netherlands open dataset looks like

One real row from the verified run, unedited apart from truncating the description:

{
  "rowType": "dataset",
  "url": "https://data.overheid.nl/dataset/wpozittenblijvers-v1",
  "apiUrl": "https://data.overheid.nl/data/api/3/action/package_show?id=wpozittenblijvers-v1",
  "id": "7cc95d10-bb52-4211-bbe8-8a18bb6e0f0d",
  "slug": "wpozittenblijvers-v1",
  "title": "Zittenblijvers in het bo en sbo per vestiging",
  "description": "Dit document beschrijft de bestanden over leerlingen in het Primair Onderwijs die via een API (A ...",
  "identifier": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
  "alternateIdentifiers": [],
  "organization": "dienst-uitvoering-onderwijs",
  "organizationTitle": "DUO",
  "organizationId": "3b966f6a-4bd6-4f63-95b4-76adcdf31657",
  "dataOwner": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
  "dataOwnerName": "Dienst Uitvoering Onderwijs",
  "publisher": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
  "publisherName": "Dienst Uitvoering Onderwijs",
  "sourceCatalog": "https://onderwijsdata.duo.nl",
  "sourceCatalogName": "Dienst Uitvoering Onderwijs (DUO)",
  "themes": [
    "http://standaarden.overheid.nl/owms/terms/Onderwijs_en_wetenschap"
  ],
  "themeLabels": [
    "Onderwijs en wetenschap"
  ],
  "themeLabelsEn": [
    "Education and science"
  ],
  "keywords": [
    "Leerlingen"
  ],
  "keywordCount": 1,
  "licence": "http://creativecommons.org/licenses/by/4.0/deed.nl",
  "licenceLabel": "CC-BY (4.0)",
  "licenceUrl": "http://creativecommons.org/licenses/by/4.0/deed.nl",
  "isOpenLicence": "Yes",
  "accessRights": "http://publications.europa.eu/resource/authority/access-right/PUBLIC",
  "accessRightsLabel": "Public",
  "datasetStatus": "http://data.overheid.nl/status/beschikbaar",
  "datasetStatusLabel": "Available",
  "restrictionsStatement": "Not Disclosed",
  "updateFrequency": "http://publications.europa.eu/resource/authority/frequency/ANNUAL",
  "updateFrequencyLabel": "Annual",
  "languages": [
    "http://publications.europa.eu/resource/authority/language/NLD"
  ],
  "languageLabels": [
    "Dutch"
  ],
  "metadataLanguage": "http://publications.europa.eu/resource/authority/language/NLD",
  "highValueDataset": "No",
  "hvdCategories": [],
  "hvdCategoryLabels": [],
  "basisRegister": "No",
  "nationalCoverage": "No",
  "referenceData": "No",
  "datasetQuality": "N/A",
  "contactName": "Informatieproducten",
  "contactTitle": "Not Disclosed",
  "contactEmail": "informatieproducten@duo.nl",
  "contactPhone": "Not Disclosed",
  "contactWebsite": "Not Disclosed",
  "contactAddress": "Not Disclosed",
  "author": "Informatieproducten",
  "authorEmail": "informatieproducten@duo.nl",
  "maintainer": "Not Disclosed",
  "maintainerEmail": "Not Disclosed",
  "version": "1.0.0",
  "versionNotes": "Not Disclosed",
  "metadataCreated": "2020-04-02T22:20:24.932693",
  "metadataModified": "2026-08-27T08:10:20.178083",
  "modified": "2026-01-09T07:43:36",
  "issued": "Not Disclosed",
  "datePlanned": "Not Disclosed",
  "temporalStart": "Not Disclosed",
  "temporalEnd": "Not Disclosed",
  "temporalLabel": "Not Disclosed",
  "spatialValues": [],
  "spatialSchemes": [],
  "legalFoundationLabel": "Not Disclosed",
  "legalFoundationRef": "Not Disclosed",
  "legalFoundationUrl": "Not Disclosed",
  "provenance": [],
  "documentation": [],
  "samples": [],
  "sources": [],
  "conformsTo": [],
  "relatedResources": [],
  "landingPage": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
  "syncChecksum": "Not Disclosed",
  "extraFields": [],
  "changeType": "updated",
  "state": "active",
  "isPrivate": "No",
  "fileCount": 1,
  "fileFormats": [
    "http://publications.europa.eu/resource/authority/file-type/CSV"
  ],
  "fileFormatLabels": [
    "CSV"
  ],
  "files": [
    {
      "name": "Aantal zittenblijvers bo en sbo per schoolvestiging",
      "url": "https://onderwijsdata.duo.nl/dataset/f0c79c94-6ffb-44bf-a3d5-9a05aef4ade4/resource/df7297a9-70a3-4a3a-b4a6-d289c7d16014/download/brin6_zittenblijvers.csv",
      "format": "CSV",
      "position": 0
    }
  ],
  "scrapedAt": "2026-08-27T19:34:28.746Z"
}

Configure the run

Leave everything empty and the Actor sweeps the whole open data catalogue, newest change first. Every dropdown carries the live dataset count next to each value, so you can see what a filter is worth before you run it. Filters combine with AND; several values inside one filter combine with OR.

Every dataset, newest change first, with its files:

{
  "maxItems": 20000,
  "sortBy": "metadata_modified desc",
  "includeDistributions": true
}

Environmental datasets published as CSV under an open licence, checked to see whether they still download:

{
  "themes": ["http://standaarden.overheid.nl/owms/terms/Natuur_en_milieu"],
  "formats": ["http://publications.europa.eu/resource/authority/file-type/CSV"],
  "licences": ["http://creativecommons.org/licenses/by/4.0/deed.nl"],
  "includeDistributions": true,
  "includeLinkCheck": true,
  "maxResourcesPerDataset": 5,
  "maxItems": 2000
}

The data request register with the government's full answer to each request:

{
  "datasetIds": [],
  "includeDataRequests": true,
  "includeDataRequestDetails": true,
  "dataRequestPhases": ["Data is niet openbaar"],
  "maxItems": 400
}

Free users

Free Apify accounts get 10 rows per run as a preview, which is enough to see every column and decide. Any paid plan lifts the cap to whatever you set in Max Items, up to 1,000,000. Upgrade here.

Run it

  1. Create a free Apify account. New accounts come with $5 in free credit, which is about 700 datasets.
  2. Open the Actor, leave the input empty for a full sweep or pick a theme, a data owner or a format from the dropdowns.
  3. Click Start. A 10-row preview finishes in about 2 seconds; a full 20,380-dataset sweep with files runs at roughly 267 rows per second.
  4. Download the results as CSV, JSON, Excel or XML, or pull them from the dataset API.

Use with AI agents (MCP)

claude mcp add apify --transport http https://mcp.apify.com --header "Authorization: Bearer YOUR_APIFY_TOKEN"

Then ask your agent in plain language:

  • "List every Netherlands open dataset about air quality that is published as CSV under an open licence."
  • "Which Dutch public bodies publish the most datasets, and how many of theirs are not openly downloadable?"
  • "Show me the data requests where the government refused because the data is not public."

Troubleshooting

No results at all. A filter value the portal does not know returns zero rows rather than an error. Every dropdown here only offers values that exist, so the usual cause is combining two filters that have no overlap, for example a theme and a data owner that never publish together. Clear one filter and run again.

Fewer rows than I asked for. Max Items is a single budget shared by every row type. If you tick the extra registers, the datasets are collected first and can use the whole budget before the registers start. Raise Max Items or run a register on its own with no dataset filters.

A field is empty on every row. Most DCAT fields are optional upstream and the portal leaves them blank. legalFoundationLabel is filled on 3 percent of datasets, datasetQuality on 9 percent, spatialValues on under 1 percent. "Not Disclosed" means the portal withheld it; "N/A" means it does not apply. No column is empty on every dataset in the catalogue.

No community column on my rows. The Communities filter works but does not export: the portal indexes the community list for search and never returns it on the dataset record.

The run is slow. Plain dataset collection runs at about 267 rows per second, so 8,000 rows take 28 seconds. The link check and the byte probe fetch files from 23 different government servers, some of which take seconds to answer: 1,500 rows with both on and three files checked per dataset took 560 seconds. Lower "Files to check per dataset" to speed it up.

FAQ

Question Answer
Do I need an API key for data.overheid.nl? No. The open data API is anonymous, and so is every page this Actor reads.
Is this the official CKAN API? It reads the portal's CKAN API and its own website. It is not run by or affiliated with data.overheid.nl.
How many datasets are there? 20,380 as of 27 August 2026, across 26 harvest organizations and 23 source catalogues.
Is the data in Dutch? The records are, since the portal is Dutch-first: 18,956 are in Dutch, 1,389 in English, 35 in German. Themes, statuses, access rights, licences, formats and data request phases are also given in English.
Can I get only the machine-readable files? Yes. Filter by format, and every file row also carries a machineReadable flag.
What is a basisregistratie? One of the ten statutory Dutch national registers. 34 datasets are flagged as belonging to one, and there is a filter for them.
What are EU High Value Datasets? Categories from the EU Open Data Directive that member states must publish for free. 229 datasets carry one, in five categories.
Does it check that the downloads work? Optionally. The portal itself records a link status on only 8 percent of files; the link check does the other 92 percent, and found 9.0 percent of a 647 file sample dead. OGC services are asked for their GetCapabilities, so a live WFS is not reported as broken.
How fresh is it? Live. Every run reads the portal at that moment, and 5,825 catalogue records changed in the last 30 days.
Can I scrape one dataset by URL? Yes. Paste data.overheid.nl dataset or community URLs into Dataset or community URLs, or put CKAN slugs into Dataset IDs.

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial tool, not affiliated with or endorsed by data.overheid.nl, KOOP or the government of the Netherlands. It collects only data that is already published openly, without logging in and without circumventing any access control. Records may contain names and contact details of public officials acting in their professional capacity; if you process them, GDPR, CCPA and PIPL obligations are yours as the data controller.

What input does the Netherlands Open Data Scraper - data.overheid.nl API accept?

FieldTypeWhat it doesDefault
searchTerms array Free-text queries run against title, description and keywords. Dutch works best: verkeer, bevolking, milieu, parkeren. Leave empty to sweep the whole catalogue of 20,380 datasets. []
startUrls array Specific data.overheid.nl pages, for example https://data.overheid.nl/dataset/wpozittenblijvers-v1 or https://data.overheid.nl/community/datarequest/schengen-entry-ban. When you supply these, only these pages are scraped and the filters below are ignored. []
datasetIds array CKAN dataset names or UUIDs, one per line, for example wpozittenblijvers-v1. Same effect as Dataset URLs, without the URL. []
maxItems integer Free users: limited to 10 items (preview). Paid users: up to 1,000,000. 10
themes array Dutch government OWMS themes. Sub-themes are listed under their parent. The catalogue stores the leaf theme only, so picking a parent does NOT include its children: tick both if you want both. []
dataOwners array The public body legally responsible for the data (OWMS authority). All 188 values the portal offers are listed with their dataset count; the catalogue index itself resolves 187 of them. []
organizations array The portal account a dataset was loaded under, which is usually the source portal rather than the data owner. 26 exist. []
sourceCatalogs array The regional, municipal or ministerial portal a dataset was harvested from. 23 feed the national catalogue. []
formats array Keep only datasets that publish at least one file in these formats. The index stores the full EU file-type URI, so the short spelling CSV matches nothing and every value here is the URI the portal really uses. []
licences array All 8 licences the portal records. CC-BY 4.0, CC-0 and Public domain together cover 18,349 of the 20,380 datasets. []
statuses array Whether the data behind the record is actually available. 20,108 datasets are Available; the other three states are rare and worth isolating. []
accessRights array DCAT access rights. 2,020 datasets are catalogued but not openly downloadable, which is exactly what a data-availability audit is looking for. []
updateFrequencies array DCAT accrual periodicity. Populated on 30 percent of datasets; the rest declare nothing and are excluded by any choice here. []
languages array Language the metadata is written in. The portal is Dutch-first: 18,956 records are Dutch, 1,389 English, 35 German. []
hvdCategories array Categories from the EU Open Data Directive. Only 229 datasets carry one, and they are the ones the Directive obliges the Netherlands to publish for free. []
communities array The four thematic communities data.overheid.nl runs. This one filters but does not export: the portal indexes the community list for search and never returns it on the dataset record, so there is no community column on the output rows. Counts here are the ones the index really matches, which are lower than the site shows. []
keywords array Keyword tags exactly as the portal stores them, one per line, for example verkeer, natuur, bodem or water. 10,403 exist. Tick "Keywords directory" below to export the full list with counts. []
minResources integer Keep only datasets that ship at least this many distributions. 6,754 datasets have 3 or more. not set
modifiedFrom string Only datasets whose catalogue record changed on or after this date (YYYY-MM-DD). 5,825 changed in the last 30 days. not set
modifiedTo string Only datasets whose catalogue record changed on or before this date (YYYY-MM-DD). not set
createdFrom string Only datasets first indexed on data.overheid.nl on or after this date (YYYY-MM-DD). 8,531 were added since 2024. not set
createdTo string Only datasets first indexed on data.overheid.nl on or before this date (YYYY-MM-DD). not set
issuedFrom string The publisher's own issue date, not the harvest date (YYYY-MM-DD). Populated on 49 percent of datasets. not set
issuedTo string Upper bound for the publisher's own issue date (YYYY-MM-DD). not set
onlyHighValue boolean Keep only the 115 datasets flagged as high value under the EU Open Data Directive. false
onlyBasisRegister boolean Keep only the 34 datasets that belong to a Dutch basisregistratie, the ten statutory national registers. false
onlyNationalCoverage boolean Keep only the 6,909 datasets that cover the whole country rather than one municipality or province. false
onlyReferenceData boolean Keep only the 51 datasets marked as reference data (code lists and authoritative vocabularies). false
customFilterQuery string Raw CKAN fq clause ANDed with everything above, for example num_resources:[10 TO *] AND -organization:nationaalgeoregister-nl. Wrap any OR group in parentheses: a bare top-level OR returns the whole catalogue upstream. A field name that does not exist returns zero rows rather than an error. not set
sortBy string (6 options) Only these six orders actually work upstream. Any other value is silently ignored by the portal and falls back to relevance. metadata_modified desc
dataRequestSearch string Free text over the data request register, for example verkeer. The register ignores any other search parameter without saying so, and 160 of the 4,667 requests mention verkeer. not set
dataRequestPhases array Where the request ended up. "Data is beschikbaar" means the government pointed at data that already exists; "Data is niet openbaar" means it refused. []
dataRequestAuthorityKinds array Which layer of government the request was aimed at. []
includeDistributions boolean Emit one extra row per downloadable file or service, with its own URL, EU format URI, media type, licence, preview URL and the link status the portal itself last recorded. Datasets average 4.1 files each. false
includeDataRequests boolean Emit the 4,667 public data requests: what was asked for, which body was asked, the phase it reached and the status. Nothing else on the Store carries this register. false
includeApplications boolean Emit the 7 registered applications built on Netherlands open data, with their type, maker, live URL and contact. false
includeCatalogs boolean Emit the 36 source catalogues the portal harvests, each with its own portal URL, description and dataset count. false
includeGroups boolean Emit the 91 curated dataset groups, for example the ten basisregistraties or the municipal High Value Data list, with their descriptions. false
includeOrganizations boolean Emit the 1,671 public bodies in the portal register with their governance layer (Rijk, Gemeente, Provincie, Waterschap), description, logo and dataset count. This is a large register: raise Max Items before ticking it. false
includeDataOwners boolean Emit one row per data owner that appears in your filtered result set, with the number of datasets it holds. false
includeThemes boolean Emit all 99 OWMS themes and sub-themes with their Dutch and English names, their parent theme and their dataset counts. false
includeFormats boolean Emit every file format in the matched result set with its canonical name, machine-readable flag and dataset count. false
includeLicences boolean Emit every licence in the matched result set with its plain-English label, whether it is open, and how many datasets declare it. false
includeKeywords boolean Emit the keyword tags used by the datasets that match your filters, with counts, ranked. The catalogue holds 10,403 of them. false
includeSourceCatalogs boolean Emit the source portals that feed the matched result set, with how many datasets each contributes. false
includeDataRequestDetails boolean Open every data request and add the 15 fields the listing hides: the exact data asked for, the format and period wanted, the intended use and impact, what the requester already tried, the full government reply, the resolution and the body named as the possible data owner. false
includeLinkCheck boolean Request each file and record the HTTP status, redirect target, content type, byte size and latency. data.overheid.nl harvests 23 catalogues and stores a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. false
includeFileProbe boolean Download the first 64 KB of each file and report what the bytes really are, whether that matches the declared format, the CSV delimiter and the column headers. false
includeAuthorityProfile boolean Add the owning body's governance layer (Rijk, Gemeente, Provincie, Waterschap), description, logo and total dataset count to every dataset row. Each owner is fetched once per run and reused. false
maxResourcesPerDataset integer Upper bound on how many files the two checks above touch per dataset. Only used when one of them is on. 5

How much does the Netherlands Open Data Scraper - data.overheid.nl API cost?

from $2.67 per 1,000 results

Charged forWhat it isPrice each
Actor Start Charged when the Actor starts running. Number of events charged depends on Actor memory (one event per GB, minimum one event). $0.04806 to $0.054
Catalogue page scanned One page of up to 500 datasets read from the data.overheid.nl CKAN index. This is the fixed cost of paging through the catalogue of 20,380 datasets, spread across every row that page yields. $0.00356 to $0.004
Dataset One Netherlands open dataset with its full DCAT-AP-DONL record: title, description, national identifier, data owner and publisher as OWMS authorities, harvest organization, source catalogue, OWMS themes in Dutch and English, keywords, licence with an open-licence flag, access rights, availability status, update frequency, EU High Value Dataset category, the basisregistratie, national-coverage and reference-data flags, contact point with email, phone and website, issue and modification dates, temporal coverage, spatial extent, legal foundation with its wetten.overheid.nl link, provenance, documentation, samples, related resources, and every file summarised with its format. $0.00623 to $0.007
File One downloadable file or OGC service: its access and download URLs, EU file-type URI with a plain name, machine-readable flag, media type, MIME type, distribution type, preview URL, licence, release date, and the link status the portal itself last recorded, plus the identity, owner and organization of the dataset it belongs to. $0.00267 to $0.003
Register page scanned One listing page of a community register. The registers paginate at a fixed 10 entries a page, so this is the paging cost of the 4,667 data requests and the 1,671 organizations, billed per page read rather than per row. $0.00356 to $0.004
Data request One public request from a citizen, journalist or company asking a Dutch public body to open a dataset: the title, what was asked for, the request number, the phase it reached in Dutch and English, the status, the body named as data owner, the theme and the date. 4,667 exist and no other Actor on the Store carries this register. $0.00534 to $0.006
Data request detail read Optional. The data request page opened in full, adding the 15 fields the listing hides: the exact data and format and period requested, the intended use and impact, what the requester had already tried and why the existing data did not do, the government reply in full, the resolution, how the request arrived and the organization named as the possible data owner. $0.00534 to $0.006
Application One of the applications registered as built on Netherlands open data, with its type, whether a public body or a company made it, its live URL, its data owner and the contact behind it. $0.00534 to $0.006
Source catalogue One of the 36 regional, municipal or ministerial portals data.overheid.nl harvests, with its own portal URL, description and how many datasets it contributes. $0.00445 to $0.005
Group One curated dataset group, for example the ten statutory basisregistraties or the municipal High Value Data list, with its full description and dataset count. 91 exist. $0.00445 to $0.005
Organization One of the 1,671 public bodies in the portal register, with its governance layer (Rijk, Gemeente, Provincie, Waterschap), description, logo, permanent link and dataset count. Includes the bodies registered with no datasets, which is what a coverage audit is looking for. $0.00445 to $0.005
Data owner One data owner that appears in the filtered result set, with its OWMS authority URI, plain name, dataset count and a working link to its datasets on the portal. 187 publish at least one dataset. $0.00267 to $0.003
Theme One of the 99 OWMS themes and sub-themes with its Dutch and English name, its parent theme, its URI and the number of datasets the catalogue index really returns for it. $0.00267 to $0.003
Format One file format in the matched result set with its EU file-type URI, canonical name, machine-readable flag and dataset count. The catalogue uses 40 of them, from CSV and JSON to OGC WMS and WFS services. $0.00267 to $0.003
Licence One licence in the matched result set with its URI, short code, the portal's own label, whether it is an open licence and how many datasets declare it. $0.00267 to $0.003
Keyword One keyword from the catalogue vocabulary with the number of datasets carrying it, ranked, honouring whatever filters the run set. The catalogue holds 10,403 of them. $0.00178 to $0.002
Source catalogue count One source portal in the matched result set with how many datasets it contributes, so you can see at a glance which of the 23 feeder catalogues a filtered slice really comes from. $0.00267 to $0.003
File link checked Optional. One file URL fetched from the publisher's own server to see whether it is actually alive: HTTP status, redirect target, content type, byte size and latency. data.overheid.nl harvests 23 catalogues and records a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. $0.00356 to $0.004
File bytes probed Optional. The first 64 KB of a file read with a range request, returning what the bytes really are, whether that contradicts the declared format, the CSV delimiter, the column headers and a sample row. Not charged when the file could not be read. $0.01068 to $0.012
Data owner profile attached Optional. The owning body's governance layer, description, logo and total dataset count added to a dataset row. Each owner is fetched once per run and reused across every dataset it published. $0.00534 to $0.006

Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.

How do I call the Netherlands Open Data Scraper - data.overheid.nl API API?

One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.

POST ยท run and get results
curl -X POST "https://api.apify.com/v2/acts/parseforge~data-overheid-netherlands-open-data-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "searchTerms": [],
    "startUrls": [],
    "datasetIds": [],
    "maxItems": 10,
    "themes": []
  }'

What example inputs can I use?

Use these inputs to see how a run is configured.

input.json
{
  "searchTerms": [],
  "startUrls": [],
  "datasetIds": [],
  "maxItems": 10,
  "themes": []
}

What do users say about the Netherlands Open Data Scraper - data.overheid.nl API?

No reviews yet. Be the first.

How do I report an issue with the Netherlands Open Data Scraper - data.overheid.nl API?

We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.

Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.

What related scrapers can I use?

Run Netherlands Open Data Scraper - data.overheid.nl API on Apify All scrapers