Categories:Scraping Use Cases,Scraping ToolsView as Markdown

How to Build a Bulk Web Scraper in Tray.ai with Scrape.do (Low Code)

Clock13 Mins Read
calendarCreated Date: October 05, 2026
calendarUpdated Date: October 05, 2026

Tray.ai (formerly Tray.io) is built to move data between apps that have APIs. Its HTTP Client connector can call any of them, so web scraping in Tray.ai usually starts there: paste a product page URL into an HTTP Client step, click Run step, and expect the page back.

From a protected site, what comes back is a block page, a CAPTCHA, or an HTML shell that only fills in once JavaScript runs in a browser. The HTTP Client sends a plain server request with no proxy rotation and no browser, and the pages worth scraping are built to turn those away.

You'll fix that without leaving Tray. The same HTTP Client step sends each URL to Scrape.do instead of to the site, and around it you'll build a workflow that reads a list of URLs from Google Sheets, scrapes each one, parses the response with AI or a script, and writes structured rows to a results sheet. The finished workflow is a downloadable template, and either end can be swapped for Snowflake, Postgres, Salesforce, or any other Tray connector.

What you'll build

One Tray project holding a six-step workflow:

  1. Manual trigger, which you can switch to Scheduled or Webhook later.
  2. google-sheets-1 reads the URL list from a source sheet.
  3. loop-1 runs the next three steps once per URL.
  4. http-client-1 sends the URL to Scrape.do and gets the page back.
  5. script-1 turns the page into fields. For sites with different layouts, an AI step plus a script does this instead.
  6. google-sheets-2 appends one row per URL to a results sheet.

Your Scrape.do token and both sheet IDs live in the project's config data, so no step carries a hard-coded value. The example target is an Amazon product page, and the script pulls its ASIN, name, price and rating. The Amazon scraping guide covers that target in more depth outside Tray.

Why Scrape.do

Scrape.do is a web scraping API: you send it a target URL as a query parameter and get the page back. The parts the HTTP Client can't do on its own happen on Scrape.do's side:

  • Rotating proxies. Datacenter proxies by default, residential and mobile proxies with super=true for protected targets.
  • JavaScript rendering. render=true loads the page in a headless browser before returning it.
  • Geo-targeting. geoCode=us (or gb, de and many more) sends the request from that country.
  • Markdown output. output=markdown returns a compact Markdown version of the page that an AI step can read.
  • Billing on success. A plain request costs 1 credit, render=true costs 5, super=true costs 10, and both together cost 25. Failed requests are free.

The free plan includes 1,000 successful API credits a month, enough to build and test this workflow. There is nothing to install in Tray either: every call is a plain GET request.

Step 1: Create a project and add config data

Every workflow in Tray belongs to a project. Config data lives at the project level, which is where your API token belongs so it never gets written into a step.

  • From your Tray dashboard, click Create new > Project and give it a name.
  • Open the project and go to the Config data tab.
  • Click Add and create the following keys:
    • scrape_do_token - Your API token from the dashboard
    • source_sheet_id - The Google Sheets ID holding your URL list
    • results_sheet_id - The Google Sheets ID where results will be written
  • Click Save.

Anywhere in the workflow you can now reference these with a jsonpath such as $.config.scrape_do_token.

Config data keeps the token out of every individual step, so the workflow definition holds no secret. It does not keep the token out of exports: Tray's project export carries config values along with their keys. Clear scrape_do_token before you share an exported project. The template in the Downloads section below ships with all three values empty, so whoever imports it fills in their own token.

Step 2: Prepare your sheets

Create two spreadsheets in Google Drive.

Source sheet with a single header row:

url

Add a few URLs to test with, for example https://us.amazon.com/dp/B0BLRJ4R8F.

Results sheet with these headers:

URL | ASIN | Name | Price | Rating | Status | Scraped

If your URLs already live in Postgres, Snowflake, or Salesforce, use that connector instead. Only the first and last steps of the workflow change.

Step 3: Create a workflow, add a trigger and read the URL list

  • Inside your project, click Create workflow.
  • Choose a trigger:
    • Manual - For testing and on-demand runs
    • Scheduled - For automated hourly or daily scraping
    • Webhook - To trigger scraping from external sources
  • For this guide, select Manual.

Now read the URL list.

  • Drag a Google Sheets connector onto the canvas. It will be named google-sheets-1.
  • Create or select your Google Sheets authentication.
  • Configure the step:
    • Operation - Select Get rows
    • Spreadsheet - Map to $.config.source_sheet_id
    • Worksheet - Enter Sheet1
    • Has headers - Toggle ON
  • Click Run step and confirm your URLs come back under $.steps.google-sheets-1.rows.

Step 4: Loop through the URLs

A single HTTP request scrapes one page. To scrape a list, wrap the request in a Loop.

  • Drag a Loop connector onto the canvas. It will be named loop-1.
  • Configure the step:
    • Operation - Select Loop list
    • List - Map to $.steps.google-sheets-1.rows
    • Parallel processing - Toggle ON
    • Concurrency - Set to your Scrape.do plan concurrency

Inside the loop, the current row is available at $.steps.loop-1.value, so the URL for each iteration is $.steps.loop-1.value.url.

Step 5: Send each URL to Scrape.do from the HTTP Client

  • Inside the loop, drag an HTTP Client connector. It will be named http-client-1.
  • Configure the step:
    • Operation - Select GET
    • URL - Enter https://api.scrape.do/
  • Under Query parameters, click Add and enter these one by one:
    • token - Map to $.config.scrape_do_token
    • url - Map to $.steps.loop-1.value.url
    • output - markdown for AI extraction, or raw to parse HTML yourself
    • render - false by default, set to true for JavaScript heavy websites
    • super - true to use residential and mobile proxies on protected targets, false for datacenter proxies
    • geoCode - Country code for proxy location, for example us, gb, de. View available locations here
  • Set Parse response to OFF so the raw body comes through untouched.
  • Under Error handling, set On error to Continue with 2 retries. A single blocked URL should not kill the whole run.

Which output to pick depends on step 6: markdown goes with the AI route (Option A), raw with the script route (Option B). The downloadable template uses Option B, so it ships with output set to raw.

Do not build the query string by hand. Tray encodes each query parameter for you, so mapping the target URL into the url field is enough. Concatenating it into the URL field yourself will break on any target containing ? or &.

  • Click Run step and confirm you get a 200 with page content in $.steps.http-client-1.response.body.

If you cannot get a successful response, visit the playground and experiment with Super, Render JavaScript, and Block Resources until the request succeeds, then copy those parameters back into the step.

Step 6: Extract data from the response

You now have raw HTML or Markdown for each URL. There are two ways to turn that into structured fields.

Option A: Use AI to extract data

Best for scraping different websites with varying structures. AI returns a consistent format regardless of the source layout. Set output to markdown in the HTTP Client step for this route: the Markdown version of a page is a fraction of the HTML's size.

  • Inside the loop, after the HTTP Client step, drag an Anthropic or OpenAI connector.
  • Create or select your AI authentication.
  • Configure the step:
    • Operation - Select Create message or the equivalent chat operation
    • Model - Select a model such as claude-sonnet-4-5-20250929
    • Prompt - Use the structure below
{$.steps.http-client-1.response.body}

Analyze the markdown data in this document and extract ASIN, Product Name,
Product Price, Review Rating, and Review Count as a structured JSON object.
Return only the JSON, with no explanation and no code fences.
  • Add a Script connector after it to turn the model output into clean fields:
exports.step = function (input) {
  let parsed = {};
  try {
    parsed = JSON.parse(input.raw);
  } catch (e) {
    parsed = {};
  }
  return {
    url: input.source_url,
    asin: parsed.ASIN || '',
    name: parsed['Product Name'] || '',
    price: parsed['Product Price'] || '',
    rating: parsed['Review Rating'] || '',
    status: Object.keys(parsed).length ? 'ok' : 'failed',
    scraped_at: new Date().toISOString()
  };
};

Map raw to the AI step output and source_url to $.steps.loop-1.value.url in the step's Variables section. The script returns the same field names as Option B, so the mappings in step 7 work for either route as long as this step is named script-1.

Option B: Use a script for extraction

If you are scraping a single site with a consistent layout, a script is faster and consumes no AI credits. It also handles responses too large for a model context window. This route needs output set to raw in the HTTP Client step, because the script matches HTML tags.

  • Inside the loop, drag a Script connector. It will be named script-1.
  • Under Variables, add:
    • source_url - Map to $.steps.loop-1.value.url
    • response - Map to $.steps.http-client-1.response.body
    • status - Map to $.steps.http-client-1.response.status_code
  • Paste the extraction logic:
exports.step = function (input) {
  const body = input.response || '';
  const ok = input.status === 200 && body.length > 0;

  const pick = (re) => {
    const m = ok ? body.match(re) : null;
    return m ? m[1].replace(/\s+/g, ' ').trim() : '';
  };

  return {
    url: input.source_url,
    name: pick(/<span id="productTitle"[^>]*>([\s\S]*?)<\/span>/),
    asin: pick(/"asin":"([A-Z0-9]{10})"/),
    price: pick(/<span class="a-offscreen">\$([0-9.,]+)<\/span>/),
    rating: pick(/(\d+\.?\d*)\s*out of/),
    status: ok ? 'ok' : 'failed',
    scraped_at: new Date().toISOString()
  };
};
  • Click Run step and verify the fields come back populated.

Step 7: Write results to Google Sheets

  • Still inside the loop, drag a second Google Sheets connector. It will be named google-sheets-2.
  • Configure the step:
    • Operation - Select Add row
    • Spreadsheet - Map to $.config.results_sheet_id
    • Worksheet - Enter Sheet1
  • Map each column to the script output. A Script step's return value sits under result in its output, so every path goes through $.steps.script-1.result:
    • URL - $.steps.script-1.result.url
    • ASIN - $.steps.script-1.result.asin
    • Name - $.steps.script-1.result.name
    • Price - $.steps.script-1.result.price
    • Rating - $.steps.script-1.result.rating
    • Status - $.steps.script-1.result.status
    • Scraped - $.steps.script-1.result.scraped_at

Step 8: Test and deploy

Before publishing, run the workflow end to end.

  • Click Run workflow at the top of the builder.
  • Open the Logs tab and verify that:
    • The loop ran once per row rather than once in total
    • Every HTTP Client step returned 200, with failures marked failed instead of stopping the run
    • The results sheet has one new row per URL with populated fields

To skip steps 3 to 7, import scrapedo-tray-workflow.json from the Downloads section below into your own workspace.

Importing does not bring authentications with it. Tray will prompt you to select or create a Google Sheets authentication during the first import, and you will need to fill in the three config data values from step 1 before the workflow can run.

  • Once the run is clean, you can:
    • Switch the trigger from Manual to Scheduled and set an interval
    • Replace Google Sheets with Snowflake, Postgres, or Salesforce on either end
    • Export the project to json and promote it from your dev workspace to production
    • Wrap the scraping steps in a Callable workflow so other workflows in the project can reuse them

How it works under the hood

Each loop iteration makes one HTTP Client call, and Tray assembles it from the query parameters you mapped. For the example row, the request is equivalent to this one, which you can run from a terminal to test a URL outside Tray:

curl "https://api.scrape.do/?token=<your_token>&url=https%3A%2F%2Fus.amazon.com%2Fdp%2FB0BLRJ4R8F&output=raw&render=false&super=false&geoCode=us"
  • The target goes in a query parameter. Scrape.do answers at the root, https://api.scrape.do/, and reads the page to fetch from url. That is why the HTTP Client's URL field never changes between iterations.
  • Tray does the encoding. Tray's HTTP Client docs say values passed as query parameters are escaped automatically, so the mapped target URL reaches Scrape.do intact, like the https%3A%2F%2F... form above. Paste the URL into the URL field yourself and the target's own ? and & get read as Scrape.do parameters.
  • render and super are opt-in because they multiply the credit cost. Leave both false until the playground shows a target needs them.
  • The script decides what counts as success. It treats anything other than a 200 with a non-empty body as failed and still returns a row, so one bad URL becomes a visible row in the results sheet instead of a stopped run.

Gotchas

1. Markdown output and the HTML script don't mix. Option B matches HTML tags (<span id="productTitle">, a-offscreen). Feed it Markdown and name, ASIN and price come back empty while the row is still marked ok. On the example URL (tested 2026-10-05), only the rating matched. Pair raw with Option B and markdown with Option A.

2. Full product pages are big. The example Amazon page came back as about 2.5 MB of raw HTML and about 145 KB of Markdown (2026-10-05). Tray's technical limits page allows 6 MB of data between two steps, and its HTTP Client docs mention a 1 MB page limit on responses. If the HTTP Client step errors on a large page, switch to output=markdown and the AI route. For Amazon in particular, the Amazon Scraper API returns product details (ASIN, title, price, ratings) as structured JSON, so there is no page to parse at all.

3. Check where your token shows up. Tray redacts authentications from workflow input logs and recommends referencing API tokens through $.auth. Config data is a separate store, so after your first run, open the HTTP Client step's input in the Logs tab. If the token is visible and other people can read those logs, turn on Log Masking for that step or move the token into a Tray authentication.

4. render and super multiply the bill. A 500-URL run costs 500 credits on plain requests, 5,000 with super=true, and 12,500 with both super=true and render=true. Some domains get a proxy or render profile applied server-side, so read the Scrape.do-Request-Cost response header to see what a request actually cost.

5. Concurrency above your plan fails the excess requests. Scrape.do answers requests over your plan's concurrency limit with a 429, which costs no credits, and the script records those rows as failed. Keep the Loop's concurrency at or below your plan's limit and rerun the failed rows.

Run order

  1. Create the project, add the three config data keys, and create the two sheets.
  2. Build steps 3 to 7, or import the template and connect your Google Sheets authentication.
  3. Run workflow, check Logs and the results sheet, then switch the trigger to Scheduled.

Downloads:

Start scraping

Tray moves the data between your apps, and Scrape.do gets the page. The free plan's 1,000 successful credits a month cover 1,000 plain requests, with residential proxies, geo-targeting and JavaScript rendering there for the targets that need them.

The same pattern runs in other tools too: see the Retool web scraper and the Power Automate web scraper guides.

Get your free Scrape.do API token, paste it into scrape_do_token, and point the source sheet at the pages you need.