Category:Anti BotView as Markdown

How to Block Web Scraping: Insights from a Pro Web Scraper

Clock31 Mins Read
calendarCreated Date: October 05, 2026
calendarUpdated Date: October 05, 2026
author

Web Scraping Expert

githublinkedin

Online, you can find plenty of resources on how to block web scraping, but how many of them are actually backed by a professional web scraper with years of hands-on experience and millions of pages scraped over the course of their career?

In this guide, I will present the most important anti-scraping mechanisms and explain how they work, how to implement them, how much they cost, and how experienced scrapers can bypass them.

Why You Want to Protect Against Web Scraping in the First Place

There is a reason why the famous Economist article published in 2017 described data as the new oil: data has become one of the most valuable assets for modern companies. Today, that value has increased even further with the rise of LLMs, which rely heavily on large datasets collected from the web for training and evaluation.

As a result, companies increasingly have an incentive to control who can access and reuse their data rather than giving it away for free. At the same time, the web has become increasingly automated.

According to Cloudflare Radar, bots now account for more than 50% of the global HTML traffic, meaning automated requests have become more common than human visits. Now, not all bots are problematic. Search engines, monitoring tools, and legitimate scrapers can provide useful services.

The problem arises when scraping is performed irresponsibly or at excessive scale. Large numbers of automated requests can consume bandwidth, CPU, database capacity, and other resources, potentially degrading performance for legitimate users.

For these reasons, protecting valuable data and maintaining website availability are two of the main motivations behind modern anti-scraping systems.

How to Block Web Scraping: Top Approaches

The game of web scraping detection and blocking is hard, but it can be addressed with the following mechanisms:

Name Layer Short description Effectiveness Implementation effort Estimated cost
IP-based blocking Network Scores IPs by location, reputation, and proxy or hosting status. Medium Low to medium $0 to $200+/mo
HTTP fingerprinting Network Checks HTTP versions, headers, ordering, and consistency against expected browser patterns. Low Low Mainly development costs
TLS fingerprinting Network Compares TLS handshake traits with known browser fingerprints. Low to medium Low to medium Mainly development costs, with possible costs for managed solutions ($99+/mo)
Rate limiting Backend Caps request volume by IP, account, session, or other client signals. Medium Low to medium Mainly development costs
Request pattern detection Backend Flags systematic URLs, timing, repetition, or navigation sequences over time. Low to medium High Mainly development costs
Custom session validation Backend Requires short-lived, signed session tokens on protected routes. Low Low to medium Mainly development costs
Content obfuscation Backend Encrypts selected API or WebSocket data for client-side decryption. Medium Low to medium Mainly development costs
Content differentiation Backend Varies response markup or schema based on request signals. Low Medium Mainly development costs
JavaScript rendering Frontend Serves valuable data only after JavaScript runs or calls page endpoints. Very low High Mainly development costs
CAPTCHA and JavaScript challenges Frontend Challenges or scores suspicious clients before allowing access. High Medium $0 to $200+/mo
Browser fingerprinting Frontend Combines browser, device, and API signals to identify automation. Medium to high Medium $0 to $200+/mo
Behavioral analysis Frontend Evaluates timing, navigation, typing, scrolling, and pointer activity. High Medium $15 to $5,000+/mo
WAF and anti-bot systems Multi-layer Filters traffic at the edge using rules and, optionally, bot-management signals. Very high Medium to high $0 to $13,000+/mo

If we have not met before, I have 10+ years of experience in the web scraping industry. One of my first applications as a developer was actually a PHP script for retrieving comic data from a website with thousands of pages.

Since then, I have collaborated with some of the most respected companies in the industry, from giants like Bright Data to growing solutions like Scrape.do. I have written more than 600 technical articles about web scraping, and I also write biweekly for The Web Scraping Club, one of the largest and most reliable newsletters in the field.

My experience is not purely theoretical, either. Over the years, I have personally scraped several million pages. With that background in mind, here is how to block web scraping!

Network Layer

Web scraping requests start with a network connection. This is the first layer where you can apply anti-scraping measures.

IP-Based Blocking

IP-based blocking involves looking at characteristics associated with the source IP address to determine whether an incoming request should be allowed, restricted, or given a lower priority.

Analyzed signals include:

  • IP ranges: Checking whether the source IP belongs to a known data center, cloud provider, hosting company, VPN provider, proxy network, or other infrastructure associated with automated traffic.
  • Geographic origin: Determine the approximate country or region from which the request originates. Some countries impose their own nationwide access restrictions. In those cases, additional application-level geolocation controls might not be necessary.
  • IP reputation: Check whether the IP has previously been associated with suspicious or abusive activity.

Based on my experience, this approach is particularly strong against low-cost scraping operations, where the scraping team cannot afford high-quality residential proxies. It is also useful when you want to restrict or deprioritize requests originating from specific regions.

Implementation

A common implementation combines:

  • For IP geolocation, services such as MaxMind GeoIP can be used to determine the approximate geographic origin of an IP address.
  • For IP reputation and threat intelligence, solutions like AbuseIPDB, Spamhaus, IPQS, and GreyNoise.
  • For data center, proxy, and VPN identification, providers such as MaxMind, IPinfo, and IPQualityScore maintain databases that classify IP ranges and networks.

These signals can then be combined into a simple decision process. For example, a request might be treated differently depending on whether the IP:

  1. originates from an allowed geographic region,
  2. belongs to a residential ISP or a known datacenter,
  3. is associated with a VPN, proxy, or Tor exit node, and
  4. has a poor reputation or history of abusive activity.

Note: IP-based geolocation and reputation systems are not exact. Geolocation is estimated, and reputation data is based on observed or inferred activity. IPs can also be reassigned or shared, leading to false positives and false negatives. As a result, these signals should not be used alone to make definitive decisions. Instead, they are better as part of a broader anti web scraping system.

Types of Blocked Scrapers

  • Scrapers relying on datacenter or other low-quality proxy IPs.
  • Scrapers sending requests from blacklisted countries or regions.
  • Scrapers using IPs that are already associated with suspicious or abusive activity.

Estimated Costs

Estimated cost / effort Notes
Implementation Low to medium (1 to 3 engineering days) Involves sending the incoming IP address to a third-party API, receiving the geolocation or reputation data, and applying a rule based on the result.
Third-party services Fraction of a cent per lookup, depending on provider and plan Solutions like MaxMind, AbuseIPDB, and IPinfo offer different pricing models, with plans ranging from ~$20/month to hundreds of dollars per month, depending on usage and features.

How Scrapers Bypass It

The easiest way to bypass IP-based blocking is to route requests through high-quality proxies. Scrapers should choose a provider with a large pool of rotating residential IPs, especially in the country or specific region you want the requests to originate from.

Scrape.do provides access to a proxy pool with 150M+ datacenter, residential, and mobile IPs across 150 countries. IPs are rotated automatically while maintaining consistent session data. Geolocation options include country-level and even ZIP code targeting.

HTTP Fingerprinting

HTTP fingerprinting examines:

  • HTTP protocol version (HTTP/1.1, HTTP/2, or HTTP/3) and related connection characteristics.
  • Header presence and values, such as User-Agent, Referer, Accept-Language, and Accept-Encoding.
  • Header ordering.

The objective is to detect inconsistencies between requests and normal browser traffic. After all, legitimate browsers tend to send recognizable combinations and patterns that scrapers may not reproduce correctly.

In my experience, I have rarely been blocked because of this. It is relatively straightforward to inspect successful requests from a real browser and reproduce the relevant HTTP-level characteristics in a scraper.

Implementation

Most web servers implement HTTP fingerprinting through middleware request-processing rules that inspect incoming traffic. Common approaches include:

  • Header validation: Checking for the presence and format of important headers, such as User-Agent.
  • Header ordering analysis: Comparing the order of headers against patterns commonly produced by different browsers and HTTP clients.
  • HTTP version analysis: Checking which HTTP version is being used and whether it is consistent with the expected client behavior. Some sites may restrict access to specific protocol versions.
  • Header correlation: Checking whether different headers are consistent with one another. For example, a request claiming to come from Firefox based on its User-Agent but containing Chrome-specific client hints in Sec-CH-UA can be flagged.

Types of Blocked Scrapers

  • Scrapers using HTTP clients such as Requests or Axios with default configuration.
  • Scrapers that reuse the same static header configuration across all requests.
  • Scrapers that randomly combine headers without maintaining consistency between related values.

Estimated Costs

Estimated cost / effort Notes
Implementation Low (1 to 2 engineering days) Can be implemented directly on the web server or within the backend application code, without requiring external dependencies.

How Scrapers Bypass It

Virtually all popular HTTP clients support header customization. So, the main challenge is rotating headers appropriately while maintaining consistency.

The idea is to analyze real browser traffic, identify the typical header names, values, and ordering, and then replicate those patterns in your scraper. Scrapers could define an array of valid header sets and randomly select one for each automated HTTP request.

Scrape.do handles HTTP fingerprinting through its Header and User Agent Rotator feature, which dynamically adjusts headers, encoding, connection type, and cookies to better match real browser behavior. It automatically rotates headers and employs domain-aware user-agent selection to ensure header values align with the target site's expected traffic profile.

TLS Fingerprinting

TLS fingerprinting examines characteristics of a client's TLS connection to help identify the software making a request. It operates below the HTTP layer, during the TLS handshake, making it a low-level mechanism to block web scraping.

The most common TLS fingerprinting methods rely on JA3 or JA4, which generate fingerprints from characteristics of the TLS ClientHello, such as supported TLS versions, cipher suites, extensions, supported curves, and their ordering.

Browsers and HTTP clients tend to produce recognizable fingerprints. A server can calculate the JA3/JA4 fingerprint of an incoming connection and compare it against known fingerprints to identify traffic that does not match the expected profile.

From what I have experienced, TLS fingerprinting is an interesting anti-scraping blocking mechanism. That is because very few standard HTTP clients support TLS fingerprint customization.

Plus, most developers I have talked with assume that spoofing HTTP headers is sufficient, but this overlooks the TLS layer that sits underneath HTTP. That is an issue, as scrapers using Requests or Axios can send perfectly crafted browser-like headers, but they will still present a TLS fingerprint that differs from a real browser.

Implementation

Implementation can be accomplished through:

  • Libraries such as read-tls-client-hello can expose JA3/JA4 TLS fingerprint information for incoming requests.
  • A reverse proxy or edge layer can inspect and log TLS fingerprints before forwarding requests to the application.

Once a JA3/JA4 fingerprint is available, you can compare it against an allowlist or known browser fingerprints, or pass it to a third-party service for additional analysis.

Types of Blocked Scrapers

  • Scrapers using HTTP clients without TLS customization.
  • Custom scrapers relying solely on header manipulation without addressing the TLS layer.

Estimated Costs

Estimated cost / effort Notes
Implementation Low to medium (1 to 3 engineering days) Requires integration with TLS fingerprinting capabilities at the reverse proxy, web server, or backend.

How Scrapers Bypass It

To avoid TLS fingerprinting issues, scrapers can develop scripts around browser automation tools such as Playwright, Puppeteer, or Selenium. In this case, requests originate from an actual browser, which naturally produces browser-like TLS characteristics.

Another approach is to use an HTTP client that supports TLS customization. One popular option is curl-impersonate, a specialized version of curl designed to reproduce the TLS and network characteristics of major browsers. It is available in Python through curl_cffi.

Scrape.do addresses this through its Dynamic TLS Fingerprinting feature, which generates authentic, real-browser TLS signatures per request. Rather than using static or spoofed handshakes, Scrape.do rotates between legitimate browser-aligned JA3/JA4 hashes, cipher suite combinations, and ALPN values on every request.

Backend Layer

After going through the connection layer, automated scraping requests reach your web server. This is where you can apply the most relevant protection mechanisms.

Rate Limiting

Rate limiting caps the number of incoming requests a client can make within a defined time period. Limits can be applied by IP address, account, session, or a combination of these signals.

Basically, if a client sends more requests than the configured allowance, the server can react by:

  1. returning a 429 Too Many Requests response, or
  2. delaying or throttling the response, or
  3. temporarily blocking the client.

Speaking from personal experience, rate limiting is one of the most common mechanisms for limiting scraping. That is especially true for scraping at scale or for protecting websites against aggressive crawling.

Types of Blocked Scrapers

  • Scrapers exposing a consistent IP, session, or other identifiable client signals while making excessive requests.
  • Site crawlers that exceed per-IP or per-session thresholds.
  • API-based scrapers that repeatedly call endpoints without respecting configured rate limits.

How to Implement It

Frameworks such as Express provide rate-limiting middleware, including:

One important aspect to define is the blocking behavior. Some common techniques include:

  • Token bucket: Allows controlled bursts while enforcing a defined average request rate.
  • Sliding window: Limits requests over a rolling time period, avoiding fixed-window reset spikes.
  • Leaky bucket: Processes requests at a steady rate, smoothing traffic bursts.

Standard rate-limiting headers can also communicate limits and retry information to clients. These are useful for legitimate developers and ethical scrapers that want to respect your limits. However, they can also reveal information about your rate-limiting policy, potentially helping scrapers understand how your limits are structured.

Estimated Costs

Estimated cost / effort Notes
Implementation Low to medium (1 to 3 engineering days) Requires adding and configuring rate limiting, potentially with different rules for different page types or endpoints.

How Scrapers Bypass It

Solutions to triggering rate limits generally involve:

  1. Reducing request frequency.
  2. Distribute requests across multiple IP addresses by integrating web scraping proxies.

Scrape.do automatically routes requests through a rotating pool of 150M+ residential, mobile, and datacenter IPs. It also changes headers and TLS fingerprints between requests to reduce client tracking.

Request Pattern Detection

Request-pattern detection looks at how a client moves through a site over time, rather than judging an isolated request.

A crawler that visits product IDs in strict sequence, requests every page at a fixed interval, or jumps directly between endpoints without normal navigation can stand out even when its headers look plausible.

The method can work well against aggressive crawlers that maintain a consistent fingerprint over time. Still, it involves considerable effort for limited success, as more sophisticated scrapers usually vary their fingerprints between requests, making it harder to reliably associate those requests with the same client.

Types of Blocked Scrapers

  • Static HTTP clients that enumerate URLs in a predictable order.
  • Browser automation that visits pages quickly, skips normal navigation paths, or repeats the same action sequence.
  • Site crawlers that make broad, repetitive requests without realistic pauses.

Implementation

Request pattern detection requires keeping track of request timestamps, paths, status codes, session or account identifiers, and navigation context. A practical way to store this information is Redis or another temporary data store that can be accessed by all your backend instances.

You then need to calculate metrics such as the number of requests within a rolling interval, time between requests, and common path transitions. You should also monitor patterns such as repeated pagination requests, sequential resource IDs, or unusually systematic navigation.

Finally, based on these signals, define rules to determine when a request should be allowed, throttled, challenged, or blocked.

Estimated Costs

Estimated cost / effort Notes
Infrastructure Low to medium (1 to 3 engineering days) Set up an in-memory data store such as Redis and configure it to store the required request-tracking data.
Implementation High (3 to 5+ engineering days) Implement request monitoring, pattern analysis, and rules to determine whether requests should be allowed, throttled, challenged, or blocked.

How Scrapers Bypass It

Scrapers can randomize delays, vary URL order, follow links naturally, and distribute requests across multiple sessions. They can also change their IP addresses and fingerprints, as explained earlier. All of this makes it harder to reliably track and identify recurring request patterns.

Scrape.do rotates proxies, headers, and TLS characteristics, which greatly reduces the ability to repeatedly associate requests with the same client and apply request-pattern recognition mechanisms.

Custom Session Validation

This mechanism involves issuing short-lived session cookies as part of normal browsing. The server then verifies that the required cookies are present and that the session has not expired.

The goal is to add friction and block web scraping attempts based on basic HTTP clients and crawlers. However, it is relatively weak as a standalone defense because browser traffic can be inspected and replicated. Scrapers can easily automate the needed session flow.

Types of Blocked Scrapers

  • Static HTTP request scripts that do not set the expected headers.
  • Site crawlers that reuse cookies across many workers or sessions.

Implementation

On page initialization, create a session and issue a short-lived, signed token tied to it. Require the token in a header on protected routes, validate its signature and expiry, and rotate session identifiers or tokens after login or privilege changes. Use a shared session store across application instances.

Estimated Costs

Estimated cost / effort Notes
Implementation Low to medium (2 to 4 engineering days) Add/Extend existing session middleware with token issuance, validation, rotation, logging, and route rules.

How Scrapers Bypass It

Playwright or Puppeteer can run the site's JavaScript to obtain the session cookies. Scrapers can then continue via regular browser automation. Otherwise, a persistent HTTP client can carry those values.

Scrape.do supports custom headers and cookies, helping you avoid issues with custom session validation.

Content Obfuscation

One popular example of content obfuscation is encrypting JSON responses for AJAX calls or WebSocket messages and decrypting them in JavaScript after the page loads.

Note how the WebSocket channel sends encrypted binary data

Since the decryption logic is included in JavaScript code loaded by the browser, a scraper can potentially inspect and replicate that logic.

Still, understanding and reproducing the full process can be tough, especially when the JavaScript files are heavily minified. From what I know, not many websites use this technique, but it can be quite effective because many developers might not have advanced reverse-engineering skills. Sure, AI can help you, but reproducing the decryption workflow is not a piece of cake.

Types of Blocked Scrapers

  • Scrapers that target APIs and web sockets directly.

Implementation

Encrypt only the selected API fields on the server and decrypt them in the application when needed. The browser's Web Crypto API and Node.js's built-in node:crypto module provide encryption primitives.

Estimated Costs

Estimated cost / effort Notes
Implementation Low to medium (2 to 4 engineering days) Add response encryption and browser-side decryption.

How Scrapers Bypass It

Scrapers can inspect a page's JavaScript files, identify the decryption routine, and replicate it. Alternatively, and more easily, you can access the website directly through a browser automation tool and extract the data after it has already been decrypted and rendered in the browser.

Scrape.do's headless browser can render JavaScript-driven pages and execute browser actions, allowing scrapers to access data after client-side processing. This tool comes with IP rotation, CAPTCHA solving, and anti-bot bypass capabilities to avoid other blocking mechanisms.

Content Differentiation

Content differentiation is all about serving different markup or API responses based on inputs such as User-Agent, device type, or the geographic location of the incoming request. It can make a scraper's parser flaky and can be useful when a site already supports separate mobile and desktop experiences.

Over the years, I have mainly seen this approach used to differentiate content based on location, but it can also be employed to make scraping more difficult. That is true considering that scrapers frequently change user agents or use device simulation techniques.

Types of Blocked Scrapers

  • Any scraper that assumes a fixed response schema or structure.

Implementation

Use Express middleware or a similar mechanism to return a small set of intentional response variants based on signals such as IP location or HTTP headers. Keep the variants documented and ensure that every legitimate user receives the desired content.

Note: This mechanism can interfere with internal caching if not configured properly, so extra attention should be paid to cache configuration and cache keys.

Estimated Costs

Estimated cost / effort Notes
Implementation Medium (2 to 4 engineering days) Build and document response variants, apply selection rules, and account for cache behavior.

How Scrapers Bypass It

Scrapers can rotate or spoof User-Agent, client-hint headers, and IP addresses through proxies to identify and map the different response variants. Alternatively, AI models can be used to parse, extract, and normalize the returned data into a consistent schema.

Scrape.do's Web Scraping API offers options for obtaining LLM-ready Markdown from web pages, which can be processed and parsed by AI to extract the required data.

Frontend Layer

The web server returns web pages, which are then rendered by the client's browser. JavaScript execution opens up many additional web scraping blocking techniques.

JavaScript Rendering

One of the most basic ways to try to block web scraping is to enforce JavaScript rendering. For example, since early 2025, Google has restricted access to SERP pages only to clients that support JavaScript execution. That stopped many scraping scripts, leading to a SERP data crisis.

In general, scrapers prefer the static HTTP request approach. This should not come as a surprise, as browsers are significantly more resource-intensive, making large-scale scraping with browser automation more challenging than just sending HTTP requests.

Clearly, this is not really an anti-scraping mechanism by itself, but rather an additional requirement that increases the cost and complexity of scraping.

Types of Scrapers It Can Block

  • Web scrapers and crawlers that rely on static HTTP requests instead of browser automation frameworks.

How to Implement It

Avoid embedding valuable data in the initial page or JavaScript bundles. Instead, put valuable data behind endpoints that are called by the page at render time.

Estimated Costs

Estimated cost / effort Notes
Implementation High (5 to 20+ engineering days) Potentially require rebuilding the entire website, or the most relevant sections, to move from static to dynamic pages.

How Scrapers Bypass It

Scrapers can run a browser, wait for the page to load, and read the resulting DOM. Alternatively, they can target the underlying JSON endpoints directly.

Scrape.do's Web Scraping API offers a render=true option for extracting data programmatically from JavaScript-heavy pages and supports page interactions when needed.

CAPTCHAs and JavaScript Verification Challenges

CAPTCHAs and JavaScript challenges are mechanisms aimed at distinguishing real users from bots. CAPTCHAs require you to complete a challenge, such as identifying images or solving a puzzle.

On the other hand, JavaScript challenges run silently in the browser and analyze various browser and client signals to determine whether the request appears to come from a legitimate browser or an automated client.

JavaScript verification produces a bot-likelihood risk score. If the score is high enough, the request can be blocked, or the user is shown a CAPTCHA. If the verification succeeds, the request is allowed to continue.

Cloudflare Turnstile verification based on a JavaScript challenge + a one-click CAPTCHA

JavaScript challenges can be bypassed by using a well-configured browser automation setup that closely resembles normal browser behavior. However, if the client is flagged and a CAPTCHA is presented, it might be game over for scrapers...

Types of Scrapers It Can Block

  • Static scrapers that cannot execute JavaScript.
  • Scrapers using browser automation tools without properly configured and patched browsers.

How to Implement It

Add a third-party CAPTCHA provider, such as Google reCAPTCHA and Cloudflare Turnstile, to the pages you want to protect. You can also introduce mandatory CAPTCHAs from hCaptcha, GeeTest CAPTCHA, or other providers on high-risk actions (e.g., form submissions) to reduce automated activity.

After the user completes the challenge, the provider generates a token. You must validate that token on the backend with the provider's SDK before accepting the request. Tokens are typically short-lived and, depending on the provider, may also be single-use.

Estimated Costs

Estimated cost / effort Notes
Implementation Medium (2 to 4 engineering days) Embed the widget, validate tokens server-side, and define a fallback for verification failures.
Third-party services Fraction of a cent per CAPTCHA, depending on provider and plan Most CAPTCHA providers come with plans ranging from a few dollars a month to hundreds of dollars per month.

How Scrapers Bypass It

JavaScript challenges can often be bypassed by using browser automation tools with patched or stealth-oriented browsers, such as Patchright, Camoufox, or invisible_playwright.

CAPTCHAs are much more difficult to deal with. In some cases, ML-based tools such as Botright can solve certain CAPTCHA types, particularly image-based challenges.

Yet, CAPTCHA bypass is more complex than simply solving the challenge. While you do it, these solutions generally analyze various browser and behavioral signals to determine whether you appear to be a real user or a bot.

For this reason, scrapers prefer to rely on third-party solutions that provide CAPTCHA-handling capabilities, such as Scrape.do. For mandatory CAPTCHAs (e.g., in forms), human-based solving services like 2Captcha and Anti-Captcha might be required.

Browser Fingerprinting

Browser fingerprinting combines client-side characteristics such as Canvas and WebGL output, installed fonts, screen dimensions, and browser API behavior to distinguish between different browser environments. Systems can also look for headless-browser characteristics or apply CDP detection.

Example of a tool to get browser fingerprinting intelligence

Fingerprinting analysis focuses on the overall picture rather than relying on a single unusual signal. After speaking with multiple developers experienced in browser patching, I have consistently heard the same story. You need to keep up with browser API changes, which can introduce new fingerprinting methods. You also need to operate at a low level, usually by patching the browser engine with C++, to avoid introducing inconsistencies that could reveal the spoofing.

In other terms, this requires a level of engineering effort and skill that not every scraper has or can afford. As a result, anti web scraping systems based on browser fingerprinting can be very powerful.

Types of Scrapers It Can Block

  • Stock headless-browser setups that expose common automation characteristics.
  • Browser-based scrapers reusing identical environments and fingerprints across many sessions.
  • All static scrapers.

How to Implement It

Integrate a browser fingerprinting solution into your page to collect client-side signals and determine whether to allow or block the request. You can use libraries such as FingerprintJS BotD, an open-source, MIT-licensed library for detecting common bots and automation frameworks.

For more demanding applications, you might require an API-based bot detection service such as Fingerprint Pro Bot Detection or CreepJS.

Estimated Costs

Estimated cost / effort Notes
Implementation Medium (2 to 4 engineering days) Integrate a library for browser-side checks, send the results to the backend, and tune server-side actions against suspicious traffic.
Third-party services (optional) Free to a fraction of a cent per fingerprint Pricing can range from free to a few hundred dollars per month, depending on usage, volume, and features.

How Scrapers Bypass It

Scrapers may use modified or patched browsers specifically designed to produce more realistic browser fingerprints. These solutions can be open source, such as Patchright, invisible_playwright, Camoufox, or they can be provided as cloud-based anti-detect browser services.

Scrape.do combines headless browser rendering with browser-aligned headers, dynamic TLS fingerprints, and other browser-level techniques to produce more realistic browser characteristics and reduce the likelihood of detection.

Behavioral Analysis

User behavioral analysis evaluates session activity such as request timing, typing speed and patterns, page transitions, scrolling, pointer activity, and repeated actions. Rules or machine-learning models can then analyze those inputs to identify unusual behavior and decide whether to allow, challenge, throttle, or block a session.

This is one of the most advanced mechanisms to block web scraping, and I have rarely seen it applied in practice. It is also broader than traditional anti-scraping because the goal is to detect and stop automation itself, rather than simply identifying web data retrieval activity.

Types of Scrapers It Can Block

  • Browser automation that skips human-like interaction patterns.

How to Implement It

This is probably not something you want to implement in-house. It is just better to trust an all-in-one anti-bot solution that comes with behavioral analysis features, such as Cloudflare Bot Management, DataDome, Imperva Advanced Bot Protection, or Akamai Bot Manager.

Estimated Costs

Estimated cost / effort Notes
Implementation Medium (2 to 5 engineering days) Integrate a third-party anti-bot solution.
Third-party services $15 to $5,000+ per month Pricing varies by provider, traffic volume, and features.

How Scrapers Bypass It

Scrapers can trust sophisticated browser automation solutions that include humanization features. For instance, CloakBrowser provides a humanize=true option to generate more human-like mouse movements, keyboard timing, and scrolling patterns.

Multi-Layer

Some anti-scraping systems do not focus on a specific layer but cover the entire stack for full protection.

Web Application Firewalls (WAFs) and Other Anti-Bot Systems

A web application firewall (WAF) inspects HTTP traffic at the edge or reverse proxy before it reaches the application. Managed rules block known exploits, while custom rules match paths, methods, request rates, and risky networks.

That is a useful baseline, but ordinary WAF rules rarely identify careful scrapers reliably. Bot-management features or dedicated products can add browser and TLS fingerprinting, IP reputation checks, session analysis, and user behavior analysis. In particular, many WAFs come with optional anti-bot features.

Ask anyone experienced in the industry, and they will tell you that multi-layer anti web scraping systems are the most stressful to deal with. Some, such as Cloudflare, may be easier to bypass than others, such as DataDome, but the overall challenge is still big.

Types of Web Scraping It Can Block

  • All types of web scraping scripts and applications

How to Implement It

Route web and API traffic through a WAF or dedicated anti-bot platform. Where available, enable the platform's native bot-management capabilities or integrate a dedicated anti-bot solution such as DataDome or Kasada. The main options are:

Solution WAF Dedicated anti-bot Bot management features
Cloudflare WAF + Bot Management Yes No Yes
AWS WAF + Bot Control Yes No Yes
Akamai Bot Manager Yes (Optional WAF available) Yes Yes
Kasada Bot Defense No Yes Yes
DataDome Bot Protect No Yes Yes

Estimated Costs

Cost item Estimated cost / effort Notes
Implementation Medium to high (2 to 5+ engineering days) Initial integration, configuration, tuning, and monitoring.
Cloudflare Subscription-based WAF starts at Free; Pro from $20/month, Business from $200/month, Enterprise custom-priced. Bot features vary by plan.
AWS WAF Usage-based $5 per web ACL/month + $1 per rule/month + $0.60 per million requests. Bot Control adds $10 per ACL/month + $1 per million requests after the first 10 million.
Akamai Bot Manager Quote-based Public pricing is not listed.
DataDome Bot Protect Subscription-based Essentials $3,830/mo; Advanced $8,670/mo; Premium $10,160/mo; Enterprise from $13,270/mo.
Kasada Bot Defense Quote-based Public pricing is not listed.

How Scrapers Bypass It

To bypass these solutions, scrapers typically need to patch or modify browsers, integrate rotating residential proxies, and keep cookies, headers, and session state consistent.

Even if they manage to combine all of these techniques using open-source tools, the quality of the proxies can have a significant impact on the results against WAFs and native bot-detection systems.

Still, the bigger challenge is maintenance and scalability. Keeping all of these components working reliably is a full-time job or even requires a team of engineers.

This is why many scrapers prefer managed scraping APIs when targeting heavily protected websites. Scrape.do's Web Scraping API provides proxy rotation, dynamic TLS fingerprints, header rotation, CAPTCHA handling, and anti-bot bypass from the same endpoint.

The provider handles the operational complexity of maintaining and scaling these capabilities, so scrapers do not have to manage the underlying infrastructure themselves.

Which Anti-Scraping Techniques Should You Choose?

The right anti-scraping setup depends on the data you need to protect, the traffic your site receives, your engineering capacity, and how much friction you can introduce for legitimate visitors.

As a rule of thumb:

  • Small sites and limited engineering time → Rate limiting
  • Sites facing API harvesting → Rate limiting, HTTP fingerprinting, TLS fingerprinting, custom session validation
  • Sites that need to minimize friction for regular visitors → IP-based blocking, rate limiting, request-pattern detection
  • Sites targeted by systematic crawling → Request-pattern detection, rate limiting, IP-based blocking, JavaScript rendering
  • Valuable API data → Content obfuscation, content differentiation
  • Sites seeing automated browsers → Browser fingerprinting, CAPTCHAs and JavaScript verification challenges
  • High-value sites or persistent attacks → Web application firewalls (WAFs) and other anti-bot systems, behavioral analysis

Note that these techniques work best in combination. For example, a site might apply rate limits to everyone, use IP reputation and TLS fingerprints to assess risk, and show a CAPTCHA only when several signals indicate likely automation. Still, there is no real way to completely prevent web scraping.

For sites facing persistent or sophisticated scraping, and when budget is not a concern, my preferred setup would be a dedicated anti-bot platform such as DataDome, combined with custom, uncommon mechanisms to catch scrapers off guard.

Conclusion

Above, I covered how to block web scraping, based on my multi-year experience in the industry. As you have seen, these mechanisms cover different layers of the stack, including the network, backend, and frontend. The most effective approaches combine multiple techniques across these layers.

No matter which protections a target website has in place, an all-in-one scraping solution like Scrape.do reduces the effort required to bypass them immensely. Instead of building and maintaining the entire scraping infrastructure yourself, scrapers just need to make an API request and get the unlocked HTML or directly extracted JSON data.

FAQ

Are honeypots an anti-scraping technique? And how do they work?

Honeypots are primarily a detection and intelligence technique rather than a standalone blocking mechanism. They expose hidden links, fields, or endpoints that legitimate users normally never access. Requests to these resources can identify automated clients and trigger logging to study the scraping system and its behavior.

What are the most effective ways to block web scraping?

There is no single mechanism that reliably stops all scraping. The strongest anti web scraping approach is layered: combine IP and reputation checks, rate limiting, TLS and HTTP fingerprinting, behavioral analysis, browser fingerprinting, JavaScript challenges, and bot-management systems. Different layers address different scraper capabilities.

Why can't you just protect all pages with mandatory CAPTCHAs?

Mandatory CAPTCHAs add significant friction for legitimate users and can hurt conversion, accessibility, and user experience. They are better used selectively for high-risk requests.

Are IP reputation systems working against residential proxies?

They can be useful, but residential proxies greatly reduce their effectiveness because requests originate from IP addresses associated with legitimate residential users. Reputation systems work better when combined with other signals such as request behavior, fingerprints, TLS characteristics, and session consistency.

Is rate limiting alone sufficient to stop web scraping?

No. Rate limiting can slow or restrict aggressive scrapers, but distributed scrapers can spread requests across many IP addresses, sessions, or accounts. It works best as one layer alongside behavioral detection, fingerprinting, IP analysis, and other bot-management techniques.

What is the difference between detecting scrapers and blocking them?

Detection determines whether traffic appears automated or suspicious. Web scraping blocking is the action taken afterward, such as denying the request, returning a challenge, throttling traffic, or requiring additional verification. Robust systems tend to separate detection from enforcement so responses can be tuned to risk.

Do JavaScript rendering requirements stop modern web scrapers?

Not by themselves. They primarily prevent simple HTTP clients that cannot execute JavaScript. Scrapers can use browser automation to render pages and execute JavaScript, although doing so increases resource requirements, complexity, and maintenance compared with direct HTTP requests.

How often should anti-scraping defenses be updated?

Continuously rather than on a fixed schedule. Scraping tools, browsers, proxies, and automation techniques evolve constantly, so detection rules and fingerprints should be monitored and adjusted as new patterns emerge. Major changes to browser behavior or scraper techniques may require immediate updates.