ParseForge Scrapers

openHPI Courses Scraper

parseforge/openhpi-courses-scraper

Developer toolsOther

Scrapes openHPI course metadata from the public catalog. Returns each course as a flat row with title, language, instructors, and description. Filter by language or keyword.

Run this scraper See the API call
Total users
2
Monthly active
1
Total runs
68
Bookmarked
0
Rating
Not rated yet
Last modified
8 days ago

Overview

ParseForge

openHPI Courses Scraper

Scrape openHPI courses from the public catalog, up to a million per run. Each course comes with its title, language, instructors, and description. No login or API key. Export to CSV, JSON, Excel, or XML.

openHPI's course catalog is spread across paginated web pages that are tedious to collect by hand. This Actor reads the public course list directly, filters by language or a search term, and returns each match in one fixed schema. It is ideal for tracking new MOOCs from the Hasso Plattner Institute.

Who uses it What they scrape openHPI for
EdTech analysts Monitor which topics the Hasso Plattner Institute is teaching this quarter
Content curators Build a searchable index of openHPI courses for a learning portal
Researchers Study the evolution of MOOC offerings in German and English
Marketing teams Find openHPI courses that align with a product launch or campaign

What it does

This Actor collects openHPI courses from the public catalog and returns each one as a flat row.

  • ๐Ÿ” Search term filter: keep only courses whose title, instructor, or description contains your keyword.
  • ๐ŸŒ Language filter: restrict results to English or German courses.
  • ๐Ÿ“Š Flat row output: every course is returned as one record, ready for CSV or JSON export.
  • โš™๏ธ Max items control: set a hard limit from 1 to 1,000,000 courses per run.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with openHPI data

๐Ÿ“š Build a course directory.

A learning platform scrapes all openHPI courses once a week and publishes a searchable catalog for its users.

๐Ÿ“ˆ Track MOOC trends.

An analyst runs the Actor monthly with no filters to see which subjects the Hasso Plattner Institute is adding or retiring.

๐Ÿ”Ž Find courses by topic.

A content marketer searches for 'Python' to list relevant openHPI courses in a newsletter.

๐ŸŒ Compare language offerings.

A researcher filters by German and English separately to compare course availability across languages.

Why choose this scraper

What you get
No API key Reads the public openHPI catalog directly, no registration or OAuth
Fixed schema Every course returns the same fields, so downstream processing is predictable
Language aware Filter to English or German courses with one dropdown
Scalable Collect up to a million courses in a single run

How it compares

No other Store actor targets openHPI the same way, so the honest comparison is with the alternatives teams actually weigh.

openHPI Courses Scraper Build it in-house By hand
Setup Run it now, zero config Days of engineering None, but hours per pull
When openHPI changes Maintained for you You fix it You re-learn the page
Proxies, retries, anti-bot Built in Your problem Browser only
Output Fixed JSON schema, CSV/Excel export Whatever you build Copy-paste
Cost Pay per result Engineering time Analyst hours

Configure the run

Drive the Actor with an optional search term and language filter, and set a maximum number of courses to collect per run. The Input tab lists every parameter.

A first run with the defaults:

{
 "maxItems": 10
}

A larger pull:

{
 "maxItems": 200
}

Pricing

Pay-per-result: $0.005 per result collected. You pay only for the results written to your dataset.

Results collected Approximate cost
100 results $0.50
1,000 results $5.00
10,000 results $50.00

New Apify accounts start with $5 in free credit.

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the openHPI Courses Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to openHPI through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/openhpi-courses-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your search term and language filter. If both are set, they are combined with AND, so a course must match both. Try clearing one filter.

The Actor returns fewer courses than expected.

Make sure maxItems is high enough. Also, the openHPI catalog may have fewer courses matching your filters.

Can I scrape a specific course by URL?

No, this Actor only reads the catalog list. It does not accept direct course URLs.

The run takes a long time.

If you set maxItems to a very large number, the Actor will paginate through the entire catalog. Lower maxItems or add filters to speed it up.

FAQ

Question Answer
Does this Actor require an openHPI account or API key? No. It reads the public course catalog directly, so no login or key is needed.
What data does each course row include? Each row includes the course title, language, instructors, and description, along with other catalog fields. The exact fields are shown in the sample output.
Can I filter courses by language? Yes. Use the language dropdown to select English, German, or any language.
Can I search for a specific topic? Yes. Enter a keyword in the search term field, and only courses whose title, instructor, or description contains it will be returned.
How many courses can I collect in one run? You can set the maximum from 1 to 1,000,000 courses per run.
What export formats are supported? The Actor outputs data in CSV, JSON, Excel, or XML, depending on your Apify dataset settings.
Does this Actor scrape course content or videos? No. It only collects metadata from the course catalog, not the actual course materials.
How often should I run this Actor? You can schedule it daily, weekly, or monthly. The openHPI catalog changes as new courses are added, so a weekly run is common.
Is this Actor affiliated with openHPI or the Hasso Plattner Institute? No. This is an independent scraper built by the Apify community.
Can I get only German courses? Yes. Set the language filter to German and leave the search term empty.

Related actors

Browse the full ParseForge collection for more scrapers.

๐Ÿ†˜ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Hasso Plattner Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

Input

FieldTypeWhat it doesDefault
maxItems integer How many courses to collect per run. 10
language string (3 options) Filter the openHPI catalog by course language. not set
query string Optional keyword. Keeps only courses whose title, instructor, or description contains it. not set

Pricing

from $4.52 per 1,000 results

Charged forWhat it isPrice each
result Single result in the default dataset. $0.00452 to $0.005

Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.

API

One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.

POST ยท run and get results
curl -X POST "https://api.apify.com/v2/acts/parseforge~openhpi-courses-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "helloWorld": 123
  }'

Examples

Input that runs as-is.

input.json
{
  "helloWorld": 123
}

Reviews

No reviews yet. Be the first.

Issues

We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.

Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.

Related scrapers

Run openHPI Courses Scraper on Apify All scrapers