ParseForge Scrapers

Crossref Academic Paper Metadata Scraper

parseforge/crossref-academic-paper-scraper

EducationBusinessDeveloper tools

Scrapes academic paper metadata from Crossref by search query or publication type. Returns DOI, title, authors, journal, citation counts, and dates.

Run this scraper See the API call
Total users
2
Monthly active
1
Total runs
3
Bookmarked
0
Rating
Not rated yet
Last modified
16 hours ago

Overview

ParseForge

Crossref Academic Paper Metadata Scraper

Scrape Crossref academic paper metadata for any DOI, search query, or publication type, up to a million records per run. Each record includes title, authors, journal, DOI, citation counts, and publication dates. No API key required. Export to CSV, JSON, Excel, or XML.

The Actor queries the public Crossref REST API, filters by publication type, search query, and sort order, and returns each matching scholarly work as one flat row. It covers journal articles, books, dissertations, datasets, and more.

Who uses it What they scrape Crossref for
Academic researchers Building a bibliography for a literature review
Librarians Verifying citation metadata for institutional repositories
Data scientists Gathering a corpus of paper metadata for analysis
Journal editors Checking DOI registration and metadata completeness
PhD students Collecting references for a thesis

What it does

This Actor collects academic paper metadata from Crossref by search query or publication type, and returns each work as a flat row with DOI, title, authors, journal, citation counts, and dates.

  • ๐Ÿ”Ž Search query: free-text search across titles and other metadata fields.
  • ๐Ÿ“š Publication type filter: journal-article, book-chapter, book, proceedings-article, dissertation, report, standard, dataset, or posted-content.
  • ๐Ÿ“… Sort order: by publication date or relevance to the search query.
  • ๐Ÿ“ฆ Bulk export: up to 1,000,000 records per run for paid users.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Crossref data

๐Ÿ“š Build a literature review bibliography.

A researcher enters a search query like "climate change migration" and exports all matching journal articles with full metadata to CSV for reference management.

๐Ÿ” Verify DOI metadata for a journal issue.

An editor filters by journal-article and searches for a specific title to confirm authors, volume, issue, and page numbers before publication.

๐Ÿ“Š Analyze publication trends.

A data scientist fetches all journal articles sorted by published date for a given year and uses citation counts to identify influential papers.

๐ŸŽ“ Collect references for a thesis.

A PhD student searches for a topic, filters to dissertations and journal articles, and exports the metadata to build a reference list.

Why choose this scraper

What you get
No API key Uses the public Crossref REST API without authentication
Rich metadata Returns DOI, title, authors, journal, citation counts, and dates
Flexible filtering Search by text or filter by scholarly work type
Scalable Fetch up to a million records per run

How it compares

This Actor focuses on flexible search and filtering of Crossref metadata with no API key required, similar to other Crossref scrapers but with a simpler input schema.

Feature ParseForge CrossRef Academic Metadata Scraper Crossref Academic Paper Search Crossref Api Scraper
Search by text query Yes Yes Yes Yes
Filter by publication type Yes Not listed Not listed Yes
Sort by relevance Yes Not listed Not listed Not listed
Fetch up to 1,000,000 records Yes Not listed Not listed Not listed
No API key required Yes Not listed Not listed Yes

What a Crossref record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
 "DOI": "10.1157/13053466",
 "type": "journal-article",
 "title": "Comentario: Prevenciรณn de los factores de riesgo de los trastornos de la conducta alimentaria en adolescentes:",
 "containerTitle": "Atenciรณn Primaria",
 "shortContainerTitle": "Aten Primaria",
 "publisher": "Elsevier BV",
 "issue": "7",
 "volume": "32",
 "page": "408-409",
 "publishedPrintDate": "2203-10",
 "issuedDate": "2203-10",
 "indexedDateTime": "2025-05-30T05:22:24Z",
 "indexedTimestamp": 1748582544347,
 "createdDateTime": "2003-11-17T16:44:58Z",
 "createdTimestamp": 1069087498000
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor with a search query or leave it empty to fetch the most recently published journal articles. Filter by publication type and sort by published date or relevance. The Input tab lists every parameter.

A first run with the defaults:

{
 "maxItems": 10,
 "filterType": "journal-article",
 "sortOrder": "published"
}

A larger pull:

{
 "maxItems": 200,
 "filterType": "journal-article",
 "sortOrder": "published"
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account.
  2. Open the Crossref Academic Paper Metadata Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Crossref through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/crossref-academic-paper-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your search query for typos or overly specific terms. Try a broader query or leave the search field empty to fetch recent articles.

Why are some fields empty?

Crossref metadata depends on what publishers deposit. Some fields may be missing for certain records.

Why is my run limited to 10 items?

Free users are limited to 10 items as a preview. Upgrade to a paid plan to fetch up to 1,000,000 records.

Why does sorting by relevance not work?

Relevance sorting requires a search query. If the query is empty, results are sorted by publication date.

FAQ

Question Answer
Do I need a Crossref API key? No, this Actor uses the public Crossref REST API without authentication. You can start scraping immediately.
What metadata fields are returned? Each record includes DOI, type, title, containerTitle, shortContainerTitle, publisher, issue, volume, page, publishedPrintDate, issuedDate, indexedDateTime, indexedTimestamp, createdDateTime, createdTimestamp, depositedDateTime, depositedTimestamp, isReferencedByCount, referencesCount, URL, ISSN, issnPrint, issnElectronic, licenseUrl, licenseContentVersion, licenseDelayInDays, linkUrl, linkContentType, linkContentVersion, linkIntendedApplication, resourcePrimaryUrl, source, member, prefix, score, alternativeId, journalIssue, authors, scrapedAt, error.
Can I search by DOI? Yes, you can enter a DOI as a search query to retrieve metadata for a specific paper.
What publication types are supported? Journal articles, book chapters, books, proceedings articles, dissertations, reports, standards, datasets, and posted content.
How many records can I fetch? Free users are limited to 10 items as a preview. Paid users can fetch up to 1,000,000 records per run.
Can I sort results? Yes, sort by publication date (newest first) or by relevance to your search query.
Is the data from Crossref complete? Crossref contains metadata for over 150 million scholarly works, but some records may have missing fields depending on what publishers deposited.
Can I export to Excel? Yes, you can export results to CSV, JSON, Excel, or XML.
Does this Actor get abstracts? The dataset does not include an abstracts field.
How do I cite the data? You should cite the original papers, not the metadata. Crossref metadata is provided under a CC0 license.

Related actors

Browse the full ParseForge collection for more scrapers.

๐Ÿ†˜ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Crossref. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

๐Ÿ’ฐ How much does it cost to scrape Crossref Academic Paper Metadata?

This Actor uses pay-per-result pricing: $0.004 per result collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.

Input

FieldTypeWhat it doesDefault
maxItems integer Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000 10
searchQuery string A text query to search for in paper titles, abstracts, and other metadata fields. Leave empty to fetch the most recently published journal articles. not set
filterType string (9 options) Filter results to a specific scholarly work type. The default is journal articles. journal-article
sortOrder string (2 options) The order in which results are returned. Published order sorts by publication date, relevance sorts by match to the search query. published

Pricing

from $3.62 per 1,000 results

Charged forWhat it isPrice each
result Single result in the default dataset. $0.00362 to $0.004

Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.

API

One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.

POST ยท run and get results
curl -X POST "https://api.apify.com/v2/acts/parseforge~crossref-academic-paper-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "helloWorld": 123
  }'

Examples

Input that runs as-is.

input.json
{
  "helloWorld": 123
}

Reviews

No reviews yet. Be the first.

Issues

We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.

Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.

Related scrapers

Run Crossref Academic Paper Metadata Scraper on Apify All scrapers