# How to Build a Custom Web Scraping Connector in Airbyte with Scrape.do > Source: https://scrape.do/blog/airbyte-web-scraper/ Published: 2026-10-05 · Updated: 2026-10-05 · Authors: Bugrahan Saka · Categories: Scraping Use Cases, Scraping Tools Airbyte moves data from APIs and databases into your warehouse, but its catalog has no general web scraping source you can point at a list of URLs. So when the data you need lives on web pages (competitor prices, product catalogs, listings), it ends up in a script that runs outside the pipeline, or in a custom connector that sends plain HTTP requests from a single IP address until the target starts blocking it. Scrape.do fills that gap, the same way it plugs into [n8n, Zapier and other workflow tools](/documentation/integrations/). With Airbyte's Connector Builder, you can turn the Scrape.do API into a proper source connector that takes a list of URLs, scrapes each page, and loads one row per page into Snowflake, BigQuery, Postgres, or any other destination in your stack. Everything below happens in the Airbyte UI. No local development environment, no Docker, no Python. ## What you'll build An Airbyte custom connector with a single stream, `scraped_pages`: 1. You give it your Scrape.do token, a list of URLs, a residential proxy toggle and a proxy country. 2. Airbyte sends each address through Scrape.do, one request per URL. Scrape.do renders the page in a headless browser and returns it as JSON. 3. Each URL becomes one row: the rendered HTML in `content`, the target's `statusCode`, and two provenance fields, `source_url` and `scraped_at`. You import it from a ready-made YAML manifest, test it, publish it to your workspace, and sync it on a schedule like any built-in source. The columns you actually want (title, price, rating) come out of `content` with a few lines of SQL in the warehouse. ## Why Scrape.do Scrape.do is a web scraping API: send a target URL, get the page back. Server-side it handles what a plain HTTP request from a connector can't: - **Rotating proxies:** datacenter by default, residential and mobile with `super=true` for protected targets. - **JavaScript rendering** with `render=true`: a headless browser loads the page before it's returned. - **JSON output** with `returnJSON=true`: the rendered page comes back inside a JSON object, which is the shape Airbyte reads. - **Geo targeting** with `geoCode`, for prices and stock that change by country. - **You're charged only for successful responses.** Failed requests cost nothing. - **Free plan: 1,000 credits a month.** [Sign up here](https://dashboard.scrape.do/signup). It's a single GET request with query parameters, which is exactly the kind of HTTP API Connector Builder turns into a source. ## Step 1: Get your Scrape.do API token - Sign up at [dashboard.scrape.do](https://dashboard.scrape.do) if you do not have an account. - Copy your API token from the dashboard. You'll paste it into Airbyte in step 4. ## Step 2: Import the manifest into Connector Builder - From your Airbyte workspace, go to **Builder** in the left sidebar. - Click **New custom connector**. - Choose **Import a YAML manifest**. - Upload **scrapedo-airbyte-manifest.yaml**. It's in the Downloads section at the end of this guide. The connector opens in the visual editor with everything already wired up. You can switch between the form view and the YAML view at any time using the toggle in the top right. > If you prefer to build it yourself, choose **Start from scratch** instead and follow steps 3 and 4 as configuration instructions rather than as a description of what the manifest already contains. ## Step 3: Review the stream configuration The manifest defines a single stream called `scraped_pages`. Open it in the builder and you'll see: - **Base URL** - `https://api.scrape.do` - **Path** - `/` - **HTTP method** - `GET` - **Authentication** - **No Auth**, because Scrape.do takes the token as a query parameter rather than a header - **Query parameters** - `token`, `url`, `render`, `returnJSON`, `super`, `geoCode` `render` and `returnJSON` are fixed to `true`. That pair is what turns a web page into a record Airbyte can store, and the under-the-hood section below explains why the connector needs both. ## Step 4: Fill in the testing values Open the **Testing values** panel on the right and enter: - **API Token** - Paste the token from step 1. It is stored encrypted and never written into the YAML. - **URLs to Scrape** - Add two or three addresses to test with, for example `https://us.amazon.com/dp/B0BLRJ4R8F` - **Use Residential Proxies** - Leave off unless the target is heavily protected - **Proxy Country** - `us` by default; use another two-letter code such as `gb` or `de` for other markets ## Step 5: Test the connector - Click **Test** in the top right. - Open the **scraped_pages** stream in the testing panel. Verify that: - The request returns a 200 status - One record appears per URL in your list - `content` holds the page HTML and `statusCode` is `200` - `source_url` matches the URL that produced each row If the response comes back with a structure you did not expect, adjust the **Record Selector** in the stream configuration. The manifest assumes the response body is the record itself, which is what `returnJSON` sends: one JSON object per page. ## Step 6: Publish to your workspace - Click the **Publish** chevron. - Select **Publish to workspace**. - Give the connector a version number and a short description. The connector now appears in your source list like any built-in Airbyte source. ## Step 7: Create a connection and sync - Go to **Connections** and click **Create connection**. - Select **Scrape.do** as the source and enter your token and URL list. - Select your destination, for example Snowflake, BigQuery, or Postgres. - Choose a sync frequency. Hourly and daily both work well for price monitoring and catalog tracking. - Click **Sync now** and check the destination table. ## Step 8: Turn the HTML into columns Airbyte stores rows, not documents. Each row from this connector carries the whole rendered page in `content`, so the last step is pulling out the fields you want, and the right place for that is the warehouse the rows land in. For the Amazon product pages above, in BigQuery: ``` SELECT source_url, scraped_at, REGEXP_EXTRACT(content, r'id="productTitle"[^>]*>\s*([^<]+?)\s*<') AS title, REGEXP_EXTRACT(content, r'class="a-offscreen">([^<]+)<') AS price, REGEXP_EXTRACT(content, r'data-hook="rating-out-of-text"[^>]*>([^<]+)<') AS rating FROM your_dataset.scraped_pages WHERE statusCode = 200; ``` That gives you `title`, `price`, and `rating` columns alongside `source_url` and `scraped_at`. Snowflake (`REGEXP_SUBSTR`) and Postgres (`substring ... from`) have equivalent functions, and putting the query in a view or a dbt model keeps the parsing next to the rest of your transformations. On Databricks, you can skip the sync entirely: [web scraping in Databricks](/blog/databricks-web-scraping/) lands pages straight in a Delta table. The patterns are tied to the target's markup. When a site changes its HTML, you update one query, and the connector and its sync history stay as they are. ## Step 9: Contribute to the Airbyte Marketplace If you want the connector available to every Airbyte user rather than only your workspace: - Click the **Publish** chevron and select **Contribute to Marketplace**. - Fill in the connector description. - Provide a GitHub personal access token so Airbyte can open a pull request on your behalf. - Click **Contribute**. Airbyte opens the pull request automatically. Once it is merged, the connector appears in the Airbyte catalog and syncs for anyone who selects it. ## How it works under the hood The manifest is about 180 lines of YAML, and four parts of it do the work. **One request per URL.** The **Parameterized Requests** section maps the user's URL list onto one request per URL: ```yaml partition_router: type: ListPartitionRouter cursor_field: target_url values: "{{ config['urls'] }}" ``` Each address in the list produces its own request and its own record. A user who enters 500 URLs gets 500 rows from a single sync. **A response Airbyte can read.** The plain Scrape.do API returns the target page as HTML, and Airbyte's default decoder expects JSON. With [`returnJSON=true`](/documentation/headless-browser/returnjson/), Scrape.do wraps the rendered page in a JSON object instead: ```json { "statusCode": 200, "content": " **New custom connector** > **Import a YAML manifest** > upload the template. 2. **Testing values** > enter your token and two or three URLs > **Test**. 3. **Publish to workspace** > **Create connection** > pick a destination and a frequency > **Sync now**. 4. Parse `content` into columns in the warehouse. **Downloads:** - [`scrapedo-airbyte-manifest.yaml`](/uploads/blog/scrapedo-airbyte-manifest.yaml) - the importable Airbyte custom connector (Connector Builder YAML manifest) ## Start scraping If your pipeline runs as orchestrated code rather than a sync, the same API works as a [web scraper for Airflow](/blog/airflow-web-scraper/) or a [web scraper in Windmill](/blog/windmill-web-scraper/). Scrape.do's free plan gives you **1,000 credits a month**, which covers about 200 rendered pages: enough to build this connector, test it, and run your first syncs before you commit to a schedule. Proxy rotation, JavaScript rendering, and geo targeting are included, and you're charged only for successful responses. [Get your free Scrape.do API token](https://dashboard.scrape.do/signup) and add the open web to the sources your warehouse already syncs.