# How to Build a Bulk Web Scraper in Activepieces with Scrape.do (Low Code) > Source: https://scrape.do/blog/activepieces-web-scraper/ Published: 2026-10-05 · Updated: 2026-10-05 · Authors: Bugrahan Saka · Categories: Scraping Use Cases, Scraping Tools Activepieces is an open-source automation platform, the self-hostable answer to Zapier and Make. It has an HTTP piece, a Code piece and loops, which means it can technically fetch any web page you point it at. Technically. In practice, point a bare HTTP request at a real e-commerce or listings site and you'll get a 403, a CAPTCHA page, or an HTML shell with no content because the page renders in JavaScript. Your flow runs from one server with one IP address, and target sites are very good at noticing that. That's the wall most Activepieces web scraping attempts hit. The way around it: import a ready-made Activepieces flow that sends a list of URLs through Scrape.do, loops over them, and extracts structured fields from each page. There's no custom piece to install and no scraping code to maintain, and the finished flow file is in the Downloads section at the end of this post. Building in n8n or Zapier instead? Start with [Scrape.do's n8n and Zapier integrations](/documentation/integrations/). ## What you'll build A four-step flow: 1. A **Schedule** trigger runs the scrape on a cadence (daily by default). 2. A Code step, **Build URL List**, holds your target URLs (or swap it for a Google Sheets step). 3. **Loop Over URLs** iterates the list. 4. Inside the loop, the HTTP step **Scrape.do Request** scrapes each page and the Code step **Extract Fields** pulls out the title, H1, status and page size. **Schedule → Build URL List → Loop Over URLs → (Scrape.do Request → Extract Fields)** Everything runs on the free plan of Activepieces Cloud or on your own self-hosted instance. ## Why route Activepieces requests through Scrape.do Scrape.do is a web scraping API: send a target URL, get the page content back. It handles server-side what a plain HTTP piece can't: - **Rotating proxies**, datacenter by default, [residential and mobile](/documentation/proxies/super/) with `super=true` for hard targets. - **Anti-bot bypass and CAPTCHA handling.** - **JavaScript rendering** with `render=true`, where a [headless browser](/blog/how-does-a-headless-browser-help-with-web-scraping-and-data-extraction/) renders the page before it's returned. - **You're charged only for successful requests.** Failed ones, such as a `502` or a `429`, cost nothing. A `404` or `400` from the target site counts as a result and is charged. - **Free plan: 1,000 successful API credits a month.** A plain request costs 1 credit. [Sign up here](https://dashboard.scrape.do/signup). Because it's a plain GET request, no custom Activepieces piece is needed. The built-in HTTP piece is enough. ## Step 1: Import the Activepieces flow template 1. Download the template: [`scrapedo-activepieces-flow.json`](/uploads/blog/scrapedo-activepieces-flow.json). 2. Open [cloud.activepieces.com](https://cloud.activepieces.com) (or your self-hosted instance) and sign in. 3. Go to **Flows → Import Flow** and select the file. Activepieces validates the JSON against the pieces available on your instance and drops the trigger and action tree straight into a new draft. All four steps appear wired up. ## Step 2: Add your Scrape.do token to the HTTP step Open the **Scrape.do Request** step. Under Query Params you'll see two entries: | Key | Value | |---|---| | `token` | `YOUR_SCRAPEDO_TOKEN` | | `url` | `{{step_2.item.url}}` | Replace the placeholder with your token from the [Scrape.do dashboard](https://dashboard.scrape.do). That's the only credential in the entire flow. The template ships with a placeholder, so it's safe to share. Once your real token is in the step, put the placeholder back before you export the flow for anyone else. Leave the `url` mapping alone: it pulls the current item from the loop. ## Step 3: Set the URLs to scrape Open the **Build URL List** step. It returns a simple array: ```javascript export const code = async (inputs) => { const urls = [ 'https://books.toscrape.com/catalogue/page-1.html', 'https://books.toscrape.com/catalogue/page-2.html', 'https://books.toscrape.com/catalogue/page-3.html', ]; return urls.map((url) => ({ url })); }; ``` Swap in your own URLs. Each object becomes one loop item, so the shape `{ url: '...' }` is what the HTTP step expects. **Want the list to come from a spreadsheet instead?** Delete this step and add **Google Sheets → Get Rows** in its place, then point the loop at that step's output and map the URL column in the HTTP step. Same flow, dynamic input. ## Step 4: Test the flow and publish it Click **Test Flow**. Watch the loop iterate: each pass fires one Scrape.do request and returns a clean object: ```json { "url": "https://books.toscrape.com/catalogue/page-1.html", "status": 200, "title": "All products | Books to Scrape - Sandbox", "h1": "All products", "html_length": 51294, "scraped_at": "2026-08-18T09:14:22.117Z" } ``` Happy with the output? Hit **Publish** and the schedule takes over. That's web scraping automation on a schedule, with no cron job or scraping script of your own to maintain. ## How the flow works under the hood - The HTTP piece calls `https://api.scrape.do/`. The API responds **only at the root path** `/`. There is no `/scrape` endpoint; invented paths return an access-denied error. - The target URL travels as a query-string parameter. The HTTP piece URL-encodes query params automatically. That matters, because a target the API can't read as a URL is rejected with *"Your target 'URL' is not valid!"* - Scrape.do picks a proxy, gets past the site's anti-bot layer, optionally renders JavaScript, and returns the final HTML in the response body. - The HTTP step's timeout is set to 120 seconds, the longest request timeout Scrape.do accepts. - The Extract Fields step parses that HTML with plain regex: no dependencies, no `npm install`. For heavier parsing, add a package to the step's `packageJson` (Cheerio is the usual pick). Need JavaScript rendering or a stubborn target? Add query params in the HTTP step: `render` → `true` for a headless browser, `super` → `true` for residential and mobile proxies. Both consume extra credits (5 for `render`, 10 for `super`, 25 for both, against 1 for a plain request, per the [request costs table](/documentation/request-costs/)), so enable them per target rather than globally. ## Five gotchas when scraping with Activepieces **1. Keep input and output separate.** If you switch to a Google Sheets source, write results to *different* columns than the URL column. Output leaking back into input is the classic bulk-scraping bug: stale HTML gets glued onto a URL and every subsequent request fails as invalid. **2. Transient `502 ROTATION_FAILED`.** A proxy hop occasionally fails on the Scrape.do side. It's **not charged**. The HTTP step ships with **No Error on Failure** (`failsafe`) turned on, plus **Retry on Failure** and **Continue on Failure**, so one bad URL never kills the whole loop: the failed request comes back as the step's output and the loop moves on to the next URL. Persistent 502s on one domain mean the target is hard: switch it to `super=true`. **3. Don't store full HTML if you don't need it.** The Extract Fields step deliberately returns only parsed fields plus a length counter. Passing 50 KB or more of raw HTML through every loop iteration bloats run logs and slows the flow. Extract first, store second. **4. Google Sheets triggers can be unreliable.** Activepieces' "New Row" trigger has a long history of missed rows in the community forum. That's exactly why this template uses a **Schedule trigger + explicit row fetch** rather than a row-based trigger. Polling the sheet on a cadence is far more predictable than waiting to be told a row appeared. **5. Piece versions differ between instances.** Self-hosted instances lag Cloud. If an imported step shows a version warning, open it and re-select the action: Activepieces re-resolves the piece to the version you actually have. ## Run order 1. Import the flow JSON. 2. Paste your Scrape.do token into the HTTP step. 3. Edit the URL list (or replace it with a Sheets step). 4. **Test Flow** → verify output → **Publish**. **Downloads:** - [`scrapedo-activepieces-flow.json`](/uploads/blog/scrapedo-activepieces-flow.json) - the importable Activepieces flow (template JSON) ## Start scraping in Activepieces Everything above runs on free tiers: the free plan of Activepieces Cloud (or your own self-hosted instance, which costs nothing but a server), and **1,000 free credits a month** from Scrape.do, with proxy rotation, anti-bot bypass, CAPTCHA handling and JavaScript rendering included, charged only on successful requests. The same pattern works in the other tools of this series: the [web scraper in Windmill](/blog/windmill-web-scraper/) runs the requests in parallel from TypeScript, and the [web scraper in Retool](/blog/retool-web-scraper/) writes the results back to a database. [Get your free Scrape.do API token](https://dashboard.scrape.do/signup) and turn any URL list into structured data, inside an automation platform you own.