# How to Build a Bulk Web Scraper in Tray.ai with Scrape.do (Low Code) > Source: https://scrape.do/blog/tray-ai-web-scraper/ Published: 2026-10-05 · Updated: 2026-10-05 · Authors: Bugrahan Saka · Categories: Scraping Use Cases, Scraping Tools Tray.ai (formerly Tray.io) is built to move data between apps that have APIs. Its HTTP Client connector can call any of them, so web scraping in Tray.ai usually starts there: paste a product page URL into an HTTP Client step, click **Run step**, and expect the page back. From a protected site, what comes back is a block page, a CAPTCHA, or an HTML shell that only fills in once JavaScript runs in a browser. The HTTP Client sends a plain server request with no proxy rotation and no browser, and the pages worth scraping are built to turn those away. You'll fix that without leaving Tray. The same HTTP Client step sends each URL to Scrape.do instead of to the site, and around it you'll build a workflow that reads a list of URLs from Google Sheets, scrapes each one, parses the response with AI or a script, and writes structured rows to a results sheet. The finished workflow is a downloadable template, and either end can be swapped for Snowflake, Postgres, Salesforce, or any other Tray connector. ## What you'll build One Tray project holding a six-step workflow: 1. **Manual trigger**, which you can switch to Scheduled or Webhook later. 2. **`google-sheets-1`** reads the URL list from a source sheet. 3. **`loop-1`** runs the next three steps once per URL. 4. **`http-client-1`** sends the URL to Scrape.do and gets the page back. 5. **`script-1`** turns the page into fields. For sites with different layouts, an AI step plus a script does this instead. 6. **`google-sheets-2`** appends one row per URL to a results sheet. Your Scrape.do token and both sheet IDs live in the project's config data, so no step carries a hard-coded value. The example target is an Amazon product page, and the script pulls its ASIN, name, price and rating. The [Amazon scraping guide](/blog/amazon-scraping/) covers that target in more depth outside Tray. ## Why Scrape.do Scrape.do is a [web scraping API](/products/web-scraping-api/): you send it a target URL as a query parameter and get the page back. The parts the HTTP Client can't do on its own happen on Scrape.do's side: - **Rotating proxies.** Datacenter proxies by default, residential and mobile proxies with `super=true` for protected targets. - **JavaScript rendering.** `render=true` loads the page in a headless browser before returning it. - **Geo-targeting.** `geoCode=us` (or `gb`, `de` and many more) sends the request from that country. - **Markdown output.** `output=markdown` returns a compact Markdown version of the page that an AI step can read. - **Billing on success.** A plain request costs 1 credit, `render=true` costs 5, `super=true` costs 10, and both together cost 25. Failed requests are free. The free plan includes 1,000 successful API credits a month, enough to build and test this workflow. There is nothing to install in Tray either: every call is a plain GET request. ## Step 1: Create a project and add config data Every workflow in Tray belongs to a project. Config data lives at the project level, which is where your API token belongs so it never gets written into a step. - From your Tray dashboard, click **Create new** > **Project** and give it a name. - Open the project and go to the **Config data** tab. - Click **Add** and create the following keys: * **scrape_do_token** - Your API token from the [dashboard](https://dashboard.scrape.do) * **source_sheet_id** - The Google Sheets ID holding your URL list * **results_sheet_id** - The Google Sheets ID where results will be written - Click **Save**. Anywhere in the workflow you can now reference these with a jsonpath such as `$.config.scrape_do_token`. > Config data keeps the token out of every individual step, so the workflow definition holds no secret. It does not keep the token out of exports: Tray's project export carries config values along with their keys. Clear `scrape_do_token` before you share an exported project. The template in the Downloads section below ships with all three values empty, so whoever imports it fills in their own token. ## Step 2: Prepare your sheets Create two spreadsheets in Google Drive. **Source sheet** with a single header row: ```plaintext url ``` Add a few URLs to test with, for example `https://us.amazon.com/dp/B0BLRJ4R8F`. **Results sheet** with these headers: ```plaintext URL | ASIN | Name | Price | Rating | Status | Scraped ``` If your URLs already live in Postgres, Snowflake, or Salesforce, use that connector instead. Only the first and last steps of the workflow change. ## Step 3: Create a workflow, add a trigger and read the URL list - Inside your project, click **Create workflow**. - Choose a trigger: * **Manual** - For testing and on-demand runs * **Scheduled** - For automated hourly or daily scraping * **Webhook** - To trigger scraping from external sources - For this guide, select **Manual**. Now read the URL list. - Drag a **Google Sheets** connector onto the canvas. It will be named `google-sheets-1`. - Create or select your Google Sheets authentication. - Configure the step: * **Operation** - Select **Get rows** * **Spreadsheet** - Map to `$.config.source_sheet_id` * **Worksheet** - Enter `Sheet1` * **Has headers** - Toggle ON - Click **Run step** and confirm your URLs come back under `$.steps.google-sheets-1.rows`. ## Step 4: Loop through the URLs A single HTTP request scrapes one page. To scrape a list, wrap the request in a Loop. - Drag a **Loop** connector onto the canvas. It will be named `loop-1`. - Configure the step: * **Operation** - Select **Loop list** * **List** - Map to `$.steps.google-sheets-1.rows` * **Parallel processing** - Toggle ON * **Concurrency** - Set to your Scrape.do plan concurrency Inside the loop, the current row is available at `$.steps.loop-1.value`, so the URL for each iteration is `$.steps.loop-1.value.url`. ## Step 5: Send each URL to Scrape.do from the HTTP Client - Inside the loop, drag an **HTTP Client** connector. It will be named `http-client-1`. - Configure the step: * **Operation** - Select **GET** * **URL** - Enter `https://api.scrape.do/` - Under **Query parameters**, click **Add** and enter these one by one: * **token** - Map to `$.config.scrape_do_token` * **url** - Map to `$.steps.loop-1.value.url` * **output** - `markdown` for AI extraction, or `raw` to parse HTML yourself * **render** - `false` by default, set to `true` for JavaScript heavy websites * **super** - `true` to use residential and mobile proxies on protected targets, `false` for datacenter proxies * **geoCode** - Country code for proxy location, for example `us`, `gb`, `de`. View available locations [here](/documentation/proxies/geo-code/) - Set **Parse response** to OFF so the raw body comes through untouched. - Under **Error handling**, set **On error** to **Continue** with 2 retries. A single blocked URL should not kill the whole run. Which `output` to pick depends on step 6: `markdown` goes with the AI route (Option A), `raw` with the script route (Option B). The downloadable template uses Option B, so it ships with `output` set to `raw`. > Do not build the query string by hand. Tray encodes each query parameter for you, so mapping the target URL into the **url** field is enough. Concatenating it into the URL field yourself will break on any target containing `?` or `&`. - Click **Run step** and confirm you get a 200 with page content in `$.steps.http-client-1.response.body`. > If you cannot get a successful response, visit the [playground](https://dashboard.scrape.do/playground) and experiment with **Super**, **Render JavaScript**, and **Block Resources** until the request succeeds, then copy those parameters back into the step. ## Step 6: Extract data from the response You now have raw HTML or Markdown for each URL. There are two ways to turn that into structured fields. ### Option A: Use AI to extract data Best for scraping different websites with varying structures. AI returns a consistent format regardless of the source layout. Set **output** to `markdown` in the HTTP Client step for this route: the Markdown version of a page is a fraction of the HTML's size. - Inside the loop, after the HTTP Client step, drag an **Anthropic** or **OpenAI** connector. - Create or select your AI authentication. - Configure the step: * **Operation** - Select **Create message** or the equivalent chat operation * **Model** - Select a model such as `claude-sonnet-4-5-20250929` * **Prompt** - Use the structure below ```plaintext {$.steps.http-client-1.response.body} Analyze the markdown data in this document and extract ASIN, Product Name, Product Price, Review Rating, and Review Count as a structured JSON object. Return only the JSON, with no explanation and no code fences. ``` - Add a **Script** connector after it to turn the model output into clean fields: ```javascript exports.step = function (input) { let parsed = {}; try { parsed = JSON.parse(input.raw); } catch (e) { parsed = {}; } return { url: input.source_url, asin: parsed.ASIN || '', name: parsed['Product Name'] || '', price: parsed['Product Price'] || '', rating: parsed['Review Rating'] || '', status: Object.keys(parsed).length ? 'ok' : 'failed', scraped_at: new Date().toISOString() }; }; ``` Map `raw` to the AI step output and `source_url` to `$.steps.loop-1.value.url` in the step's **Variables** section. The script returns the same field names as Option B, so the mappings in step 7 work for either route as long as this step is named `script-1`. ### Option B: Use a script for extraction If you are scraping a single site with a consistent layout, a script is faster and consumes no AI credits. It also handles responses too large for a model context window. This route needs `output` set to `raw` in the HTTP Client step, because the script matches HTML tags. - Inside the loop, drag a **Script** connector. It will be named `script-1`. - Under **Variables**, add: * **source_url** - Map to `$.steps.loop-1.value.url` * **response** - Map to `$.steps.http-client-1.response.body` * **status** - Map to `$.steps.http-client-1.response.status_code` - Paste the extraction logic: ```javascript exports.step = function (input) { const body = input.response || ''; const ok = input.status === 200 && body.length > 0; const pick = (re) => { const m = ok ? body.match(re) : null; return m ? m[1].replace(/\s+/g, ' ').trim() : ''; }; return { url: input.source_url, name: pick(/]*>([\s\S]*?)<\/span>/), asin: pick(/"asin":"([A-Z0-9]{10})"/), price: pick(/\$([0-9.,]+)<\/span>/), rating: pick(/(\d+\.?\d*)\s*out of/), status: ok ? 'ok' : 'failed', scraped_at: new Date().toISOString() }; }; ``` - Click **Run step** and verify the fields come back populated. ## Step 7: Write results to Google Sheets - Still inside the loop, drag a second **Google Sheets** connector. It will be named `google-sheets-2`. - Configure the step: * **Operation** - Select **Add row** * **Spreadsheet** - Map to `$.config.results_sheet_id` * **Worksheet** - Enter `Sheet1` - Map each column to the script output. A Script step's return value sits under `result` in its output, so every path goes through `$.steps.script-1.result`: * **URL** - `$.steps.script-1.result.url` * **ASIN** - `$.steps.script-1.result.asin` * **Name** - `$.steps.script-1.result.name` * **Price** - `$.steps.script-1.result.price` * **Rating** - `$.steps.script-1.result.rating` * **Status** - `$.steps.script-1.result.status` * **Scraped** - `$.steps.script-1.result.scraped_at` ## Step 8: Test and deploy Before publishing, run the workflow end to end. - Click **Run workflow** at the top of the builder. - Open the **Logs** tab and verify that: * The loop ran once per row rather than once in total * Every HTTP Client step returned 200, with failures marked `failed` instead of stopping the run * The results sheet has one new row per URL with populated fields To skip steps 3 to 7, import `scrapedo-tray-workflow.json` from the Downloads section below into your own workspace. > Importing does not bring authentications with it. Tray will prompt you to select or create a Google Sheets authentication during the first import, and you will need to fill in the three config data values from step 1 before the workflow can run. - Once the run is clean, you can: * Switch the trigger from **Manual** to **Scheduled** and set an interval * Replace Google Sheets with Snowflake, Postgres, or Salesforce on either end * Export the project to json and promote it from your dev workspace to production * Wrap the scraping steps in a **Callable workflow** so other workflows in the project can reuse them ## How it works under the hood Each loop iteration makes one HTTP Client call, and Tray assembles it from the query parameters you mapped. For the example row, the request is equivalent to this one, which you can run from a terminal to test a URL outside Tray: ```bash curl "https://api.scrape.do/?token=&url=https%3A%2F%2Fus.amazon.com%2Fdp%2FB0BLRJ4R8F&output=raw&render=false&super=false&geoCode=us" ``` - **The target goes in a query parameter.** Scrape.do answers at the root, `https://api.scrape.do/`, and reads the page to fetch from `url`. That is why the HTTP Client's URL field never changes between iterations. - **Tray does the encoding.** Tray's HTTP Client docs say values passed as query parameters are escaped automatically, so the mapped target URL reaches Scrape.do intact, like the `https%3A%2F%2F...` form above. Paste the URL into the URL field yourself and the target's own `?` and `&` get read as Scrape.do parameters. - **`render` and `super` are opt-in** because they multiply the credit cost. Leave both `false` until the [playground](https://dashboard.scrape.do/playground) shows a target needs them. - **The script decides what counts as success.** It treats anything other than a 200 with a non-empty body as `failed` and still returns a row, so one bad URL becomes a visible row in the results sheet instead of a stopped run. ## Gotchas **1. Markdown output and the HTML script don't mix.** Option B matches HTML tags (``, `a-offscreen`). Feed it Markdown and name, ASIN and price come back empty while the row is still marked `ok`. On the example URL (tested 2026-10-05), only the rating matched. Pair `raw` with Option B and `markdown` with Option A. **2. Full product pages are big.** The example Amazon page came back as about 2.5 MB of raw HTML and about 145 KB of Markdown (2026-10-05). Tray's technical limits page allows 6 MB of data between two steps, and its HTTP Client docs mention a 1 MB page limit on responses. If the HTTP Client step errors on a large page, switch to `output=markdown` and the AI route. For Amazon in particular, the [Amazon Scraper API](/products/ready-api/amazon-scraper/) returns product details (ASIN, title, price, ratings) as structured JSON, so there is no page to parse at all. **3. Check where your token shows up.** Tray redacts authentications from workflow input logs and recommends referencing API tokens through `$.auth`. Config data is a separate store, so after your first run, open the HTTP Client step's input in the **Logs** tab. If the token is visible and other people can read those logs, turn on Log Masking for that step or move the token into a Tray authentication. **4. `render` and `super` multiply the bill.** A 500-URL run costs 500 credits on plain requests, 5,000 with `super=true`, and 12,500 with both `super=true` and `render=true`. Some domains get a proxy or render profile applied server-side, so read the `Scrape.do-Request-Cost` response header to see what a request actually cost. **5. Concurrency above your plan fails the excess requests.** Scrape.do answers requests over your plan's concurrency limit with a 429, which costs no credits, and the script records those rows as `failed`. Keep the Loop's concurrency at or below your plan's limit and rerun the failed rows. ## Run order 1. Create the project, add the three config data keys, and create the two sheets. 2. Build steps 3 to 7, or import the template and connect your Google Sheets authentication. 3. **Run workflow**, check **Logs** and the results sheet, then switch the trigger to **Scheduled**. **Downloads:** - [`scrapedo-tray-workflow.json`](/uploads/blog/scrapedo-tray-workflow.json) - the importable Tray.ai workflow (Google Sheets in, Scrape.do request, parse script, Google Sheets out) ## Start scraping Tray moves the data between your apps, and Scrape.do gets the page. The free plan's 1,000 successful credits a month cover 1,000 plain requests, with residential proxies, geo-targeting and JavaScript rendering there for the targets that need them. The same pattern runs in other tools too: see the [Retool web scraper](/blog/retool-web-scraper/) and the [Power Automate web scraper](/blog/power-automate-web-scraper/) guides. [Get your free Scrape.do API token](https://dashboard.scrape.do/signup), paste it into `scrape_do_token`, and point the source sheet at the pages you need.