Categories:Scraping Use Cases,Scraping ToolsView as Markdown
How to Build a Bulk Web Scraper in Activepieces with Scrape.do (Low Code)

Growth
Activepieces is an open-source automation platform, the self-hostable answer to Zapier and Make. It has an HTTP piece, a Code piece and loops, which means it can technically fetch any web page you point it at.
Technically. In practice, point a bare HTTP request at a real e-commerce or listings site and you'll get a 403, a CAPTCHA page, or an HTML shell with no content because the page renders in JavaScript. Your flow runs from one server with one IP address, and target sites are very good at noticing that.
That's the wall most Activepieces web scraping attempts hit. The way around it: import a ready-made Activepieces flow that sends a list of URLs through Scrape.do, loops over them, and extracts structured fields from each page. There's no custom piece to install and no scraping code to maintain, and the finished flow file is in the Downloads section at the end of this post. Building in n8n or Zapier instead? Start with Scrape.do's n8n and Zapier integrations.
What you'll build
A four-step flow:
- A Schedule trigger runs the scrape on a cadence (daily by default).
- A Code step, Build URL List, holds your target URLs (or swap it for a Google Sheets step).
- Loop Over URLs iterates the list.
- Inside the loop, the HTTP step Scrape.do Request scrapes each page and the Code step Extract Fields pulls out the title, H1, status and page size.
Schedule → Build URL List → Loop Over URLs → (Scrape.do Request → Extract Fields)
Everything runs on the free plan of Activepieces Cloud or on your own self-hosted instance.
Why route Activepieces requests through Scrape.do
Scrape.do is a web scraping API: send a target URL, get the page content back. It handles server-side what a plain HTTP piece can't:
- Rotating proxies, datacenter by default, residential and mobile with
super=truefor hard targets. - Anti-bot bypass and CAPTCHA handling.
- JavaScript rendering with
render=true, where a headless browser renders the page before it's returned. - You're charged only for successful requests. Failed ones, such as a
502or a429, cost nothing. A404or400from the target site counts as a result and is charged. - Free plan: 1,000 successful API credits a month. A plain request costs 1 credit. Sign up here.
Because it's a plain GET request, no custom Activepieces piece is needed. The built-in HTTP piece is enough.
Step 1: Import the Activepieces flow template
- Download the template:
scrapedo-activepieces-flow.json. - Open cloud.activepieces.com (or your self-hosted instance) and sign in.
- Go to Flows → Import Flow and select the file.
Activepieces validates the JSON against the pieces available on your instance and drops the trigger and action tree straight into a new draft. All four steps appear wired up.
Step 2: Add your Scrape.do token to the HTTP step
Open the Scrape.do Request step. Under Query Params you'll see two entries:
| Key | Value |
|---|---|
token |
YOUR_SCRAPEDO_TOKEN |
url |
{{step_2.item.url}} |
Replace the placeholder with your token from the Scrape.do dashboard. That's the only credential in the entire flow. The template ships with a placeholder, so it's safe to share. Once your real token is in the step, put the placeholder back before you export the flow for anyone else.
Leave the url mapping alone: it pulls the current item from the loop.
Step 3: Set the URLs to scrape
Open the Build URL List step. It returns a simple array:
export const code = async (inputs) => {
const urls = [
'https://books.toscrape.com/catalogue/page-1.html',
'https://books.toscrape.com/catalogue/page-2.html',
'https://books.toscrape.com/catalogue/page-3.html',
];
return urls.map((url) => ({ url }));
};
Swap in your own URLs. Each object becomes one loop item, so the shape { url: '...' } is what the HTTP step expects.
Want the list to come from a spreadsheet instead? Delete this step and add Google Sheets → Get Rows in its place, then point the loop at that step's output and map the URL column in the HTTP step. Same flow, dynamic input.
Step 4: Test the flow and publish it
Click Test Flow. Watch the loop iterate: each pass fires one Scrape.do request and returns a clean object:
{
"url": "https://books.toscrape.com/catalogue/page-1.html",
"status": 200,
"title": "All products | Books to Scrape - Sandbox",
"h1": "All products",
"html_length": 51294,
"scraped_at": "2026-08-18T09:14:22.117Z"
}
Happy with the output? Hit Publish and the schedule takes over. That's web scraping automation on a schedule, with no cron job or scraping script of your own to maintain.
How the flow works under the hood
- The HTTP piece calls
https://api.scrape.do/. The API responds only at the root path/. There is no/scrapeendpoint; invented paths return an access-denied error. - The target URL travels as a query-string parameter. The HTTP piece URL-encodes query params automatically. That matters, because a target the API can't read as a URL is rejected with "Your target 'URL' is not valid!"
- Scrape.do picks a proxy, gets past the site's anti-bot layer, optionally renders JavaScript, and returns the final HTML in the response body.
- The HTTP step's timeout is set to 120 seconds, the longest request timeout Scrape.do accepts.
- The Extract Fields step parses that HTML with plain regex: no dependencies, no
npm install. For heavier parsing, add a package to the step'spackageJson(Cheerio is the usual pick).
Need JavaScript rendering or a stubborn target? Add query params in the HTTP step: render → true for a headless browser, super → true for residential and mobile proxies. Both consume extra credits (5 for render, 10 for super, 25 for both, against 1 for a plain request, per the request costs table), so enable them per target rather than globally.
Five gotchas when scraping with Activepieces
1. Keep input and output separate. If you switch to a Google Sheets source, write results to different columns than the URL column. Output leaking back into input is the classic bulk-scraping bug: stale HTML gets glued onto a URL and every subsequent request fails as invalid.
2. Transient 502 ROTATION_FAILED. A proxy hop occasionally fails on the Scrape.do side. It's not charged. The HTTP step ships with No Error on Failure (failsafe) turned on, plus Retry on Failure and Continue on Failure, so one bad URL never kills the whole loop: the failed request comes back as the step's output and the loop moves on to the next URL. Persistent 502s on one domain mean the target is hard: switch it to super=true.
3. Don't store full HTML if you don't need it. The Extract Fields step deliberately returns only parsed fields plus a length counter. Passing 50 KB or more of raw HTML through every loop iteration bloats run logs and slows the flow. Extract first, store second.
4. Google Sheets triggers can be unreliable. Activepieces' "New Row" trigger has a long history of missed rows in the community forum. That's exactly why this template uses a Schedule trigger + explicit row fetch rather than a row-based trigger. Polling the sheet on a cadence is far more predictable than waiting to be told a row appeared.
5. Piece versions differ between instances. Self-hosted instances lag Cloud. If an imported step shows a version warning, open it and re-select the action: Activepieces re-resolves the piece to the version you actually have.
Run order
- Import the flow JSON.
- Paste your Scrape.do token into the HTTP step.
- Edit the URL list (or replace it with a Sheets step).
- Test Flow → verify output → Publish.
Downloads:
scrapedo-activepieces-flow.json- the importable Activepieces flow (template JSON)
Start scraping in Activepieces
Everything above runs on free tiers: the free plan of Activepieces Cloud (or your own self-hosted instance, which costs nothing but a server), and 1,000 free credits a month from Scrape.do, with proxy rotation, anti-bot bypass, CAPTCHA handling and JavaScript rendering included, charged only on successful requests.
The same pattern works in the other tools of this series: the web scraper in Windmill runs the requests in parallel from TypeScript, and the web scraper in Retool writes the results back to a database.
Get your free Scrape.do API token and turn any URL list into structured data, inside an automation platform you own.

Growth

