Crossref Academic Paper Metadata Scraper
parseforge/crossref-academic-paper-scraper
EducationBusinessDeveloper tools
Scrapes academic paper metadata from Crossref by search query or publication type. Returns DOI, title, authors, journal, citation counts, and dates.
- Total users
- 2
- Monthly active
- 1
- Total runs
- 3
- Bookmarked
- 0
- Rating
- Not rated yet
- Last modified
- 16 hours ago
Overview
Crossref Academic Paper Metadata Scraper
Scrape Crossref academic paper metadata for any DOI, search query, or publication type, up to a million records per run. Each record includes title, authors, journal, DOI, citation counts, and publication dates. No API key required. Export to CSV, JSON, Excel, or XML.
The Actor queries the public Crossref REST API, filters by publication type, search query, and sort order, and returns each matching scholarly work as one flat row. It covers journal articles, books, dissertations, datasets, and more.
| Who uses it | What they scrape Crossref for |
|---|---|
| Academic researchers | Building a bibliography for a literature review |
| Librarians | Verifying citation metadata for institutional repositories |
| Data scientists | Gathering a corpus of paper metadata for analysis |
| Journal editors | Checking DOI registration and metadata completeness |
| PhD students | Collecting references for a thesis |
What it does
This Actor collects academic paper metadata from Crossref by search query or publication type, and returns each work as a flat row with DOI, title, authors, journal, citation counts, and dates.
- ๐ Search query: free-text search across titles and other metadata fields.
- ๐ Publication type filter: journal-article, book-chapter, book, proceedings-article, dissertation, report, standard, dataset, or posted-content.
- ๐ Sort order: by publication date or relevance to the search query.
- ๐ฆ Bulk export: up to 1,000,000 records per run for paid users.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with Crossref data
๐ Build a literature review bibliography.
A researcher enters a search query like "climate change migration" and exports all matching journal articles with full metadata to CSV for reference management.
๐ Verify DOI metadata for a journal issue.
An editor filters by journal-article and searches for a specific title to confirm authors, volume, issue, and page numbers before publication.
๐ Analyze publication trends.
A data scientist fetches all journal articles sorted by published date for a given year and uses citation counts to identify influential papers.
๐ Collect references for a thesis.
A PhD student searches for a topic, filters to dissertations and journal articles, and exports the metadata to build a reference list.
Why choose this scraper
| What you get | |
|---|---|
| No API key | Uses the public Crossref REST API without authentication |
| Rich metadata | Returns DOI, title, authors, journal, citation counts, and dates |
| Flexible filtering | Search by text or filter by scholarly work type |
| Scalable | Fetch up to a million records per run |
How it compares
This Actor focuses on flexible search and filtering of Crossref metadata with no API key required, similar to other Crossref scrapers but with a simpler input schema.
| Feature | ParseForge | CrossRef Academic Metadata Scraper | Crossref Academic Paper Search | Crossref Api Scraper |
|---|---|---|---|---|
| Search by text query | Yes | Yes | Yes | Yes |
| Filter by publication type | Yes | Not listed | Not listed | Yes |
| Sort by relevance | Yes | Not listed | Not listed | Not listed |
| Fetch up to 1,000,000 records | Yes | Not listed | Not listed | Not listed |
| No API key required | Yes | Not listed | Not listed | Yes |
What a Crossref record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{
"DOI": "10.1157/13053466",
"type": "journal-article",
"title": "Comentario: Prevenciรณn de los factores de riesgo de los trastornos de la conducta alimentaria en adolescentes:",
"containerTitle": "Atenciรณn Primaria",
"shortContainerTitle": "Aten Primaria",
"publisher": "Elsevier BV",
"issue": "7",
"volume": "32",
"page": "408-409",
"publishedPrintDate": "2203-10",
"issuedDate": "2203-10",
"indexedDateTime": "2025-05-30T05:22:24Z",
"indexedTimestamp": 1748582544347,
"createdDateTime": "2003-11-17T16:44:58Z",
"createdTimestamp": 1069087498000
}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor with a search query or leave it empty to fetch the most recently published journal articles. Filter by publication type and sort by published date or relevance. The Input tab lists every parameter.
A first run with the defaults:
{
"maxItems": 10,
"filterType": "journal-article",
"sortOrder": "published"
}
A larger pull:
{
"maxItems": 200,
"filterType": "journal-article",
"sortOrder": "published"
}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account.
- Open the Crossref Academic Paper Metadata Scraper.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to Crossref through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/crossref-academic-paper-scraper"
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check your search query for typos or overly specific terms. Try a broader query or leave the search field empty to fetch recent articles.
Why are some fields empty?
Crossref metadata depends on what publishers deposit. Some fields may be missing for certain records.
Why is my run limited to 10 items?
Free users are limited to 10 items as a preview. Upgrade to a paid plan to fetch up to 1,000,000 records.
Why does sorting by relevance not work?
Relevance sorting requires a search query. If the query is empty, results are sorted by publication date.
FAQ
| Question | Answer |
|---|---|
| Do I need a Crossref API key? | No, this Actor uses the public Crossref REST API without authentication. You can start scraping immediately. |
| What metadata fields are returned? | Each record includes DOI, type, title, containerTitle, shortContainerTitle, publisher, issue, volume, page, publishedPrintDate, issuedDate, indexedDateTime, indexedTimestamp, createdDateTime, createdTimestamp, depositedDateTime, depositedTimestamp, isReferencedByCount, referencesCount, URL, ISSN, issnPrint, issnElectronic, licenseUrl, licenseContentVersion, licenseDelayInDays, linkUrl, linkContentType, linkContentVersion, linkIntendedApplication, resourcePrimaryUrl, source, member, prefix, score, alternativeId, journalIssue, authors, scrapedAt, error. |
| Can I search by DOI? | Yes, you can enter a DOI as a search query to retrieve metadata for a specific paper. |
| What publication types are supported? | Journal articles, book chapters, books, proceedings articles, dissertations, reports, standards, datasets, and posted content. |
| How many records can I fetch? | Free users are limited to 10 items as a preview. Paid users can fetch up to 1,000,000 records per run. |
| Can I sort results? | Yes, sort by publication date (newest first) or by relevance to your search query. |
| Is the data from Crossref complete? | Crossref contains metadata for over 150 million scholarly works, but some records may have missing fields depending on what publishers deposited. |
| Can I export to Excel? | Yes, you can export results to CSV, JSON, Excel, or XML. |
| Does this Actor get abstracts? | The dataset does not include an abstracts field. |
| How do I cite the data? | You should cite the original papers, not the metadata. Crossref metadata is provided under a CC0 license. |
Related actors
Browse the full ParseForge collection for more scrapers.
๐ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
โ ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Crossref. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
๐ฐ How much does it cost to scrape Crossref Academic Paper Metadata?
This Actor uses pay-per-result pricing: $0.004 per result collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.
Input
| Field | Type | What it does | Default |
|---|---|---|---|
| maxItems | integer | Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000 | 10 |
| searchQuery | string | A text query to search for in paper titles, abstracts, and other metadata fields. Leave empty to fetch the most recently published journal articles. | not set |
| filterType | string (9 options) | Filter results to a specific scholarly work type. The default is journal articles. | journal-article |
| sortOrder | string (2 options) | The order in which results are returned. Published order sorts by publication date, relevance sorts by match to the search query. | published |
Pricing
from $3.62 per 1,000 results
| Charged for | What it is | Price each |
|---|---|---|
| result | Single result in the default dataset. | $0.00362 to $0.004 |
Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.
API
One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.
curl -X POST "https://api.apify.com/v2/acts/parseforge~crossref-academic-paper-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"helloWorld": 123
}' Examples
Input that runs as-is.
{
"helloWorld": 123
} Reviews
No reviews yet. Be the first.
Issues
We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.
Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.
Related scrapers
Run Crossref Academic Paper Metadata Scraper on Apify All scrapers
