Categories:Scraping Use Cases,Scraping ToolsView as Markdown
How to Build a Custom Web Scraping Connector in Airbyte with Scrape.do

Growth
Airbyte moves data from APIs and databases into your warehouse, but its catalog has no general web scraping source you can point at a list of URLs. So when the data you need lives on web pages (competitor prices, product catalogs, listings), it ends up in a script that runs outside the pipeline, or in a custom connector that sends plain HTTP requests from a single IP address until the target starts blocking it.
Scrape.do fills that gap, the same way it plugs into n8n, Zapier and other workflow tools. With Airbyte's Connector Builder, you can turn the Scrape.do API into a proper source connector that takes a list of URLs, scrapes each page, and loads one row per page into Snowflake, BigQuery, Postgres, or any other destination in your stack.
Everything below happens in the Airbyte UI. No local development environment, no Docker, no Python.
What you'll build
An Airbyte custom connector with a single stream, scraped_pages:
- You give it your Scrape.do token, a list of URLs, a residential proxy toggle and a proxy country.
- Airbyte sends each address through Scrape.do, one request per URL. Scrape.do renders the page in a headless browser and returns it as JSON.
- Each URL becomes one row: the rendered HTML in
content, the target'sstatusCode, and two provenance fields,source_urlandscraped_at.
You import it from a ready-made YAML manifest, test it, publish it to your workspace, and sync it on a schedule like any built-in source. The columns you actually want (title, price, rating) come out of content with a few lines of SQL in the warehouse.
Why Scrape.do
Scrape.do is a web scraping API: send a target URL, get the page back. Server-side it handles what a plain HTTP request from a connector can't:
- Rotating proxies: datacenter by default, residential and mobile with
super=truefor protected targets. - JavaScript rendering with
render=true: a headless browser loads the page before it's returned. - JSON output with
returnJSON=true: the rendered page comes back inside a JSON object, which is the shape Airbyte reads. - Geo targeting with
geoCode, for prices and stock that change by country. - You're charged only for successful responses. Failed requests cost nothing.
- Free plan: 1,000 credits a month. Sign up here.
It's a single GET request with query parameters, which is exactly the kind of HTTP API Connector Builder turns into a source.
Step 1: Get your Scrape.do API token
- Sign up at dashboard.scrape.do if you do not have an account.
- Copy your API token from the dashboard. You'll paste it into Airbyte in step 4.
Step 2: Import the manifest into Connector Builder
- From your Airbyte workspace, go to Builder in the left sidebar.
- Click New custom connector.
- Choose Import a YAML manifest.
- Upload scrapedo-airbyte-manifest.yaml. It's in the Downloads section at the end of this guide.
The connector opens in the visual editor with everything already wired up. You can switch between the form view and the YAML view at any time using the toggle in the top right.
If you prefer to build it yourself, choose Start from scratch instead and follow steps 3 and 4 as configuration instructions rather than as a description of what the manifest already contains.
Step 3: Review the stream configuration
The manifest defines a single stream called scraped_pages. Open it in the builder and you'll see:
- Base URL -
https://api.scrape.do - Path -
/ - HTTP method -
GET - Authentication - No Auth, because Scrape.do takes the token as a query parameter rather than a header
- Query parameters -
token,url,render,returnJSON,super,geoCode
render and returnJSON are fixed to true. That pair is what turns a web page into a record Airbyte can store, and the under-the-hood section below explains why the connector needs both.
Step 4: Fill in the testing values
Open the Testing values panel on the right and enter:
- API Token - Paste the token from step 1. It is stored encrypted and never written into the YAML.
- URLs to Scrape - Add two or three addresses to test with, for example
https://us.amazon.com/dp/B0BLRJ4R8F - Use Residential Proxies - Leave off unless the target is heavily protected
- Proxy Country -
usby default; use another two-letter code such asgbordefor other markets
Step 5: Test the connector
- Click Test in the top right.
- Open the scraped_pages stream in the testing panel.
Verify that:
- The request returns a 200 status
- One record appears per URL in your list
contentholds the page HTML andstatusCodeis200source_urlmatches the URL that produced each row
If the response comes back with a structure you did not expect, adjust the Record Selector in the stream configuration. The manifest assumes the response body is the record itself, which is what returnJSON sends: one JSON object per page.
Step 6: Publish to your workspace
- Click the Publish chevron.
- Select Publish to workspace.
- Give the connector a version number and a short description.
The connector now appears in your source list like any built-in Airbyte source.
Step 7: Create a connection and sync
- Go to Connections and click Create connection.
- Select Scrape.do as the source and enter your token and URL list.
- Select your destination, for example Snowflake, BigQuery, or Postgres.
- Choose a sync frequency. Hourly and daily both work well for price monitoring and catalog tracking.
- Click Sync now and check the destination table.
Step 8: Turn the HTML into columns
Airbyte stores rows, not documents. Each row from this connector carries the whole rendered page in content, so the last step is pulling out the fields you want, and the right place for that is the warehouse the rows land in. For the Amazon product pages above, in BigQuery:
SELECT
source_url,
scraped_at,
REGEXP_EXTRACT(content, r'id="productTitle"[^>]*>\s*([^<]+?)\s*<') AS title,
REGEXP_EXTRACT(content, r'class="a-offscreen">([^<]+)<') AS price,
REGEXP_EXTRACT(content, r'data-hook="rating-out-of-text"[^>]*>([^<]+)<') AS rating
FROM your_dataset.scraped_pages
WHERE statusCode = 200;
That gives you title, price, and rating columns alongside source_url and scraped_at. Snowflake (REGEXP_SUBSTR) and Postgres (substring ... from) have equivalent functions, and putting the query in a view or a dbt model keeps the parsing next to the rest of your transformations. On Databricks, you can skip the sync entirely: web scraping in Databricks lands pages straight in a Delta table.
The patterns are tied to the target's markup. When a site changes its HTML, you update one query, and the connector and its sync history stay as they are.
Step 9: Contribute to the Airbyte Marketplace
If you want the connector available to every Airbyte user rather than only your workspace:
- Click the Publish chevron and select Contribute to Marketplace.
- Fill in the connector description.
- Provide a GitHub personal access token so Airbyte can open a pull request on your behalf.
- Click Contribute.
Airbyte opens the pull request automatically. Once it is merged, the connector appears in the Airbyte catalog and syncs for anyone who selects it.
How it works under the hood
The manifest is about 180 lines of YAML, and four parts of it do the work.
One request per URL. The Parameterized Requests section maps the user's URL list onto one request per URL:
partition_router:
type: ListPartitionRouter
cursor_field: target_url
values: "{{ config['urls'] }}"
Each address in the list produces its own request and its own record. A user who enters 500 URLs gets 500 rows from a single sync.
A response Airbyte can read. The plain Scrape.do API returns the target page as HTML, and Airbyte's default decoder expects JSON. With returnJSON=true, Scrape.do wraps the rendered page in a JSON object instead:
{
"statusCode": 200,
"content": "<!doctype html><html lang=\"en-us\" ...",
"networkRequests": [ ... ],
"websocketRequests": [],
"actionResults": [],
"screenShots": []
}
returnJSON only works together with render=true, so the manifest sets both on every request instead of offering them as toggles.
Provenance in, browser logs out. Two fields are added to every record after the response comes back, so each row carries its own provenance, and the browser logs are dropped before the row is written:
transformations:
- type: AddFields
fields:
- path: [source_url]
value: "{{ stream_partition.target_url }}"
- path: [scraped_at]
value: "{{ now_utc().strftime('%Y-%m-%dT%H:%M:%SZ') }}"
- type: RemoveFields
field_pointers:
- [networkRequests]
- [websocketRequests]
- [actionResults]
- [screenShots]
networkRequests holds the XHR and fetch calls the page made, response bodies included. On an Amazon product page, that list made the response more than twice the size of the HTML alone. If those API responses are the data you're after, delete the RemoveFields block and they land in the warehouse too.
Failures handled, not swallowed. The manifest handles failures rather than letting them stop a sync:
- A
401or403fails immediately with a message telling the user to check their token. Scrape.do also answers401when the account is out of credits or the subscription is suspended, so the message points there too. - A
429means you hit your plan's concurrency limit. It's treated as a rate limit, so Airbyte slows down instead of erroring. - A
502or503is retried with exponential backoff, since proxy failures are usually transient. Scrape.do doesn't charge for a failed502.
Gotchas (the real ones)
1. Without returnJSON, the sync finishes and the table stays empty. The plain API answers with HTML, and Airbyte's JSON parser can't read it. The stream reads zero records, the sync doesn't fail, and the only sign of trouble is a "Failed to parse JSON data" line in the sync log. If you edit the manifest, keep render and returnJSON together.
2. Every URL is a browser request, and browser requests cost more. With render=true, a successful request costs 5 credits on datacenter proxies and 25 with residential proxies (super=true), against 1 for a plain request. Some domains have their own price: Amazon pages, for example, are routed through Scrape.do's Amazon Scraper API at 1 credit. The request costs table lists them, and the Scrape.do-Request-Cost response header shows what each call actually cost. Size the sync before you schedule it: 500 URLs every hour at 5 credits each is 60,000 credits a day.
3. Pages are big. An Amazon product page comes back as about 2.5 MB of HTML. Keep URL lists to the pages you actually parse, and check your destination's limits on row size before you sync thousands of them.
4. Keep the error handling if you contribute. Airbyte reviewers look closely at error handling. A connector that treats every failure the same way is a common reason for a contribution to be sent back, so leave the three response filters in place.
5. Never hardcode the token in the YAML. The api_token field is marked airbyte_secret, so Airbyte stores it encrypted and the exported manifest stays safe to share. A token typed into the request parameters would ship with every export.
Run order
- Builder > New custom connector > Import a YAML manifest > upload the template.
- Testing values > enter your token and two or three URLs > Test.
- Publish to workspace > Create connection > pick a destination and a frequency > Sync now.
- Parse
contentinto columns in the warehouse.
Downloads:
scrapedo-airbyte-manifest.yaml- the importable Airbyte custom connector (Connector Builder YAML manifest)
Start scraping
If your pipeline runs as orchestrated code rather than a sync, the same API works as a web scraper for Airflow or a web scraper in Windmill.
Scrape.do's free plan gives you 1,000 credits a month, which covers about 200 rendered pages: enough to build this connector, test it, and run your first syncs before you commit to a schedule. Proxy rotation, JavaScript rendering, and geo targeting are included, and you're charged only for successful responses.
Get your free Scrape.do API token and add the open web to the sources your warehouse already syncs.

Growth

