ParseForge Scrapers

OpenAIRE Publications Scraper

parseforge/openaire-scraper

EducationOtherLead generation

Scrapes open access research publications from OpenAIRE by search query and optional year range. Returns each publication as a flat row with title, authors, DOI, date, and access status.

Run this scraper See the API call
Total users
2
Monthly active
1
Total runs
100
Bookmarked
0
Rating
Not rated yet
Last modified
12 days ago

Overview

ParseForge

OpenAIRE Publications Scraper

Scrape open access research publications from OpenAIRE by keyword, year range, or both, up to a million per run. Every record comes with its title, authors, DOI, publication date, and access status. No API key or registration. Export to CSV, JSON, Excel, or XML.

OpenAIRE aggregates millions of open access publications from repositories, journals, and aggregators across Europe and beyond. Its official API requires registration and has rate limits. This Actor reads the public search results directly, filtered by search term and publication year, and returns each match in one fixed schema.

Who uses it What they scrape OpenAIRE for
Academic researchers Building a literature review dataset for a specific topic
Librarians Monitoring new open access publications in their institution's field
Data analysts Tracking publication trends over time for reporting
Grant managers Finding open access outputs from funded projects

What it does

This Actor collects research publications from OpenAIRE by search query and optional year range, and returns each one as a flat row.

  • ๐Ÿ” Keyword search: any term, phrase, or boolean query OpenAIRE supports, e.g. 'machine learning' or 'climate change'.
  • ๐Ÿ“… Year range filter: set fromYear and toYear to narrow results to a specific period.
  • ๐Ÿ“ฆ Bulk collection: collect up to 1,000,000 records per run with a single input.
  • ๐Ÿ“„ Flat output: each publication is one row with title, authors, DOI, date, and access status.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with OpenAIRE data

๐Ÿ“š Build a literature review dataset.

A PhD student enters a research topic and collects all matching open access publications from the last five years to seed their reference manager.

๐Ÿ“ˆ Track publication trends.

A data analyst runs the Actor monthly with a fixed query and year range to count new publications and spot emerging topics.

๐Ÿ”Ž Monitor open access compliance.

A librarian checks which publications from their institution's researchers are openly available in OpenAIRE.

๐ŸŒ Map research output by region.

A policy researcher collects publications mentioning a country name and analyzes the geographic distribution of authors.

Why choose this scraper

What you get
No API key No registration or OAuth, a search term
Open access focus Returns publications that are freely available to read
Year filtering Limit results to a specific publication window
Scalable Collect up to a million records per run

How it compares

No other Store actor targets OpenAIRE the same way, so the honest comparison is with the alternatives teams actually weigh.

OpenAIRE Publications Scraper Build it in-house By hand
Setup Run it now, zero config Days of engineering None, but hours per pull
When OpenAIRE changes Maintained for you You fix it You re-learn the page
Proxies, retries, anti-bot Built in Your problem Browser only
Output Fixed JSON schema, CSV/Excel export Whatever you build Copy-paste
Cost Pay per result Engineering time Analyst hours

Configure the run

Drive the Actor with a search query and optional year range, and filters run as each publication is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

{
 "searchQuery": "machine learning",
 "maxItems": 10
}

A larger pull:

{
 "searchQuery": "machine learning",
 "maxItems": 200
}

Pricing

Pay-per-result: $0.021 per result collected. You pay only for the results written to your dataset.

Results collected Approximate cost
100 results $2.10
1,000 results $21.00
10,000 results $210.00

New Apify accounts start with $5 in free credit.

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the OpenAIRE Publications Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to OpenAIRE through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/openaire-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your search query for typos or overly specific terms. Try a broader keyword or remove the year filter to see if any records exist.

Why are results missing some fields?

OpenAIRE metadata varies by source. Some publications may not have a DOI or author list. The Actor returns whatever is available.

Why did the run stop before reaching maxItems?

OpenAIRE may not have enough matching records. The Actor collects all available results up to your limit.

Why is the run slow?

Large result sets require pagination through OpenAIRE's search. Reduce maxItems or narrow your query to speed things up.

Can I search for an exact phrase?

Yes. Put the phrase in double quotes, like "climate change adaptation".

FAQ

Question Answer
What is OpenAIRE? OpenAIRE is a European open access infrastructure that aggregates metadata for millions of research publications from repositories, journals, and other sources.
Do I need an API key to use this Actor? No. The Actor reads the public OpenAIRE search interface directly, so no registration or key is required.
What data does each result include? Each row includes the publication title, authors, DOI, publication date, access status, and other metadata available from OpenAIRE.
Can I filter by publication year? Yes. Set the fromYear and toYear inputs to restrict results to a specific range.
How many results can I collect? You can set maxItems up to 1,000,000 per run.
What search queries does OpenAIRE support? You can use simple keywords, phrases in quotes, and boolean operators like AND, OR, and NOT.
Does this Actor return only open access publications? OpenAIRE focuses on open access content, but some records may be metadata-only or have restricted access. The access status field tells you.
Can I export the results? Yes. You can export to CSV, JSON, Excel, or XML from the Apify dataset.
Is this Actor free to use? The Actor itself is free, but you need an Apify account and may use free tier credits. Large runs may require a paid plan.
What is the difference between this and the OpenAIRE API? The official API requires registration and has rate limits. This Actor handles pagination and rate limiting for you and returns a clean dataset.

Related actors

  • google-scholar-scraper: Use this if you need citation counts and a broader scholarly index, including paywalled articles.
  • crossref-scraper: Use this if you need DOI metadata from Crossref, including funding information and references.

Browse the full ParseForge collection for more scrapers.

๐Ÿ†˜ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by OpenAIRE A.M.K.E. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

Input

FieldTypeWhat it doesDefault
searchQuery string Keywords to search for (e.g. 'climate change', 'machine learning', 'COVID-19') machine learning
maxItems integer How many research products to collect per run. 10
fromYear integer Filter publications from this year (e.g. 2020) not set
toYear integer Filter publications up to this year (e.g. 2024) not set

Pricing

from $19.00 per 1,000 results

Charged forWhat it isPrice each
result Single result in the default dataset. $0.019 to $0.021

Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.

API

One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.

POST ยท run and get results
curl -X POST "https://api.apify.com/v2/acts/parseforge~openaire-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "searchQuery": "machine learning",
    "maxItems": 3
  }'

Examples

Input that runs as-is.

input.json
{
  "searchQuery": "machine learning",
  "maxItems": 3
}

Reviews

No reviews yet. Be the first.

Issues

We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.

Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.

Related scrapers

Run OpenAIRE Publications Scraper on Apify All scrapers