NY Business Entity Scraper
parseforge/apps-scraper
AutomationDeveloper toolsOther
Scrapes New York business entity records by name, DOS ID, or assumed name. Returns entity type, status, and filing details as flat rows for CSV, JSON, Excel, or XML export.
- Total users
- 4
- Monthly active
- 0
- Total runs
- 245
- Bookmarked
- 0
- Rating
- Not rated yet
- Last modified
- 12 days ago
Overview
NY Business Entity Scraper
Scrape New York business registry records by name, DOS ID, or assumed name, up to a million per run. Every record comes with its entity type, status, and filing details. No login or API key. Export to CSV, JSON, Excel, or XML.
The New York Department of State business entity registry holds every corporation, LLC, and partnership filed in the state, but its public search is slow and manual. This Actor queries the registry directly by entity name, DOS ID, or assumed name, filters by status and type, and returns each match as a flat row. No official API access is needed.
| Who uses it | What they scrape New York Business Entity Registry for |
|---|---|
| Compliance officers | Verify that a business is active and in good standing before signing a contract |
| Market researchers | Build a list of all LLCs in a given industry or county for outreach |
| Journalists | Trace corporate ownership and registration history for investigative stories |
| Sales teams | Enrich lead lists with official entity names and DOS IDs for accurate CRM records |
What it does
This Actor collects New York business entity records by name, DOS ID, or assumed name, and returns each one as a flat row with entity type, status, and filing details.
- ๐ Search by name, DOS ID, or assumed name: match contains, begins with, or base word for precise results.
- ๐ Filter by entity type: corporation, LLC, limited partnership, or limited liability partnership.
- ๐ข Filter by status: active, inactive, suspended, or all statuses.
- ๐ Flat row output: every record is one row, ready for CSV, JSON, Excel, or XML export.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with New York Business Entity Registry data
๐ Build a prospect list of New York LLCs.
A sales team searches for all active LLCs with 'consulting' in the name and exports the results to CSV for their CRM.
๐ Verify a vendor's registration status.
A compliance officer enters a DOS ID and checks that the entity is active before approving a payment.
๐ฐ Investigate corporate filings.
A journalist searches by assumed name to uncover related businesses and their registration history.
๐๏ธ Enrich internal records with official data.
A data analyst matches a list of company names against the registry to add DOS IDs and entity types.
Why choose this scraper
| What you get | |
|---|---|
| No official API needed | Queries the public registry directly, so you skip registration and rate limits |
| Flexible search | Search by entity name, DOS ID, assumed name, or assumed name ID |
| Filtered results | Narrow by entity type and status before data is collected |
| Scalable | Collect up to one million records per run |
How it compares
No other Store actor targets New York Business Entity Registry the same way, so the honest comparison is with the alternatives teams actually weigh.
| NY Business Entity Scraper | Build it in-house | By hand | |
|---|---|---|---|
| Setup | Run it now, zero config | Days of engineering | None, but hours per pull |
| When New York Business Entity Registry changes | Maintained for you | You fix it | You re-learn the page |
| Proxies, retries, anti-bot | Built in | Your problem | Browser only |
| Output | Fixed JSON schema, CSV/Excel export | Whatever you build | Copy-paste |
| Cost | Pay per result | Engineering time | Analyst hours |
Configure the run
Drive the Actor with a business name, DOS ID, or assumed name, and filters run as each record is read so only matches reach your dataset. The Input tab lists every parameter.
A first run with the defaults:
{
"maxItems": 10,
"searchValue": "APPLE"
}
A larger pull:
{
"maxItems": 200,
"searchValue": "APPLE"
}
Pricing
Pay-per-result: $0.01 per result collected. You pay only for the results written to your dataset.
| Results collected | Approximate cost |
|---|---|
| 100 results | $1.00 |
| 1,000 results | $10.00 |
| 10,000 results | $100.00 |
New Apify accounts start with $5 in free credit.
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account with $5 in credit.
- Open the NY Business Entity Scraper.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to New York Business Entity Registry through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/apps-scraper"
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check your search value and match mode. If you use Begins With, try Contains. Also ensure your entity type and status filters are not too restrictive.
Why is the run slow?
The registry can be slow for broad searches. Try narrowing your search value or reducing the maximum records.
Why are some fields empty?
Not all entities have every field populated. For example, an assumed name may not exist for every entity. Empty fields are normal.
Can I search without a search value?
Yes. Leave the search value empty to run a broad prefix scan. This can return many records, so set a reasonable maximum.
FAQ
| Question | Answer |
|---|---|
| What is the New York Business Entity Registry? | It is the official database of corporations, LLCs, and partnerships filed with the New York State Department of State. It includes entity names, DOS IDs, status, and filing history. |
| Do I need an API key or login? | No. This Actor queries the public registry directly, so no registration or authentication is required. |
| Can I search by DOS ID? | Yes. Set Search By to DOS ID and enter the exact ID. You can also search by assumed name or assumed name ID. |
| What does 'Match Mode' mean? | It controls how your search value is matched: Contains finds the value anywhere in the field, Begins With matches from the start, and Base Word matches whole words. |
| Can I filter by entity type? | Yes. Choose one or more from corporation, LLC, limited partnership, and limited liability partnership. |
| Can I filter by status? | Yes. Select active, inactive, suspended, or all statuses. |
| How many records can I get? | You can set Maximum records up to 1,000,000 per run. The actual number depends on how many entities match your search. |
| What output formats are supported? | The Actor exports to CSV, JSON, Excel, and XML. You can choose the format when you download the dataset. |
| Is the data current? | The Actor reads the registry live at the time of the run, so you get the latest publicly available information. |
| Can I run this on a schedule? | Yes. You can set up a recurring run in Apify to monitor changes or keep your dataset up to date. |
Related actors
Browse the full ParseForge collection for more scrapers.
๐ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
โ ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by New York State Department of State. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
Input
| Field | Type | What it does | Default |
|---|---|---|---|
| maxItems | integer | How many business registry records to collect per run. | 10 |
| searchValue | string | Name, DOS ID, or assumed name to search. Leave empty to run a broad prefix scan. | APPLE |
| searchByTypeIndicator | string (4 options) | Field used for matching your search value. | EntityName |
| searchExpressionIndicator | string (3 options) | How the value should match: contains, starts with, or base word. | Contains |
| entityStatusIndicator | string (4 options) | Filter by registration status. | AllStatuses |
| entityTypeIndicator | array | Entity types to include in results. | ["Corporation","LimitedLiabilityCompany" |
Pricing
from $4.52 per 1,000 results
| Charged for | What it is | Price each |
|---|---|---|
| Actor Start | Charged once when the run starts. | $0.005 |
| Result Item | Charged once per result collected. | $0.00452 to $0.005 |
| result details | Detailed result with additional fields. | $0.00905 to $0.01 |
Tiered: the lower figure is the price on a higher Apify plan. Billing and the free credit live on Apify.
API
One POST returns the dataset directly. Same shape for every scraper in the library, so swapping the slug is the only change.
curl -X POST "https://api.apify.com/v2/acts/parseforge~apps-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"helloWorld": 123
}' Examples
Input that runs as-is.
{
"helloWorld": 123
} Reviews
No reviews yet. Be the first.
Issues
We build and maintain this scraper, so a problem with it comes to us. Report it on the Apify listing and the thread stays attached to the scraper where the next person can find it: open an issue.
Broken and urgent, or you would rather not post in public? Write to parseforge@protonmail.com and it reaches the people who wrote it.
