# Web scraping protection that stops copycats and lets good crawlers in

Source: https://kavralab.com/solutions/web-scraping/

**Web scraping** is the automated copying of a website's prices, listings, content or stock data by bots, usually hidden behind residential proxies and headless browsers so they look like shoppers. Kavra tells scrapers from real visitors and verified search or AI crawlers on every request, so you can block, slow or allow each one.

- **Who it hits:** Retail, travel, marketplaces, publishers, classifieds
- **What it costs:** Undercut prices, copied content, server load
- **Tools used:** Headless browsers, scraping APIs, proxy pools
- **Where to stop it:** Product, search, price and listing endpoints

## What is web scraping?

Web scraping is the use of software to collect data from web pages or APIs at scale. A person browsing a product page reads one price. A scraper requests every product page, every search result and every hidden API behind them, then stores the data in a database someone else controls.

Not all scraping is hostile. Search engines crawl your site so people can find it, and some AI assistants fetch pages to answer a user's question. The problem is the other traffic: competitors pulling your full price list every hour, aggregators republishing your listings, resellers watching your stock for restocks, and AI crawlers that ignore your rules and copy content you paid to create. Good [web scraping protection](https://kavralab.com/glossary/web-scraping/) separates those groups instead of treating every bot the same.

## How a modern scraping operation works

Scraping moved from simple scripts to rented infrastructure that imitates real shoppers.

1. **Map the targets**: The operator finds the pages and, more often, the internal JSON endpoints your own front end calls for prices, availability and search results. APIs return clean data and are cheaper to hit than full pages.
2. **Pick a browser or a client**: Lightweight HTTP clients that imitate a browser handle easy targets. When a site checks for JavaScript, the operator switches to [headless browsers](https://kavralab.com/detect/headless-browsers/) driven by Playwright, Puppeteer or Selenium, often with stealth plugins that hide automation flags.
3. **Hide behind real home IPs**: Requests are spread across a pool of [residential and mobile proxies](https://kavralab.com/detect/residential-proxies/), so each one comes from a different household in the right country and never trips a per-IP rate limit.
4. **Rotate and retry**: Device fingerprints, headers and sessions are rotated between requests. Blocked requests are retried through a new exit. Commercial scraping services sell this whole stack as a single API call.
5. **Use the data**: Prices feed a competitor's repricing engine, listings appear on a copycat site, stock levels trigger reseller bots, and articles end up in someone else's product or training set.

## What scrapers take and what it costs you

The cost depends on what the data is worth to someone else.

| What gets scraped | Who is usually behind it | Business impact |
|---|---|---|
| Prices and promotions | Competitors and repricing tools | Undercut within minutes, margin pressure, lost sales |
| Product and listing content | Copycat stores, aggregators | Duplicate content, lost search traffic, stolen photos |
| Inventory and availability | Resellers and [scalping](https://kavralab.com/solutions/scalping/) bots | Restocks hoarded before real customers see them |
| Fares, rates and seat maps | Metasearch and resale sites | Look-to-book ratios explode, supplier API bills rise |
| Articles, reviews and user posts | Unverified AI crawlers, content farms | Paid content reused without permission or credit |
| Seller, user or job data | Lead brokers, spammers | Privacy exposure, poached sellers, spam to your users |

## Warning signs of scraping bots

Scrapers rarely trip one alarm. They leave patterns that humans do not.

- **API calls without page views**: Requests go straight to price or search endpoints without loading the page, images or scripts a browser would fetch first.
- **Systematic coverage**: Every product ID, every page of results, every date in a calendar, requested in order and at a steady pace no shopper keeps.
- **Home IPs that never repeat**: Thousands of visits from different residential addresses, each seen once, many belonging to commercial proxy networks.
- **Browsers that contradict themselves**: A visitor claims to be Chrome on a laptop but has no real screen, no input history and automation traces under a thin disguise.
- **Fake search engine user agents**: Requests announce themselves as a well-known crawler but come from networks the crawler's operator does not use.
- **Origin load with no revenue**: Traffic and infrastructure costs climb while conversion rates fall, especially on search and availability pages.

## Why common anti-scraping defenses fall short

Scraping tools are built to pass the checks most sites rely on.

| Defense | What it assumes | How scrapers get past it |
|---|---|---|
| robots.txt | Bots follow the rules | It is a request, not a lock; hostile scrapers ignore it |
| Per-IP rate limits | One IP is one visitor | Residential proxy pools give each request a fresh address |
| IP and ASN blocklists | Bad traffic comes from data centers | Proxy traffic exits from real home and mobile networks |
| User-agent filtering | Bots say who they are | Scrapers copy a current browser or a search crawler name |
| JavaScript challenges | Bots cannot run scripts | Headless browsers run them like any other browser |
| CAPTCHA | Only humans can solve it | Solving services clear puzzles cheaply; real shoppers get the friction |

## Scrapers, search crawlers and AI agents: allow the right ones

Blocking all automation would hurt you. Search crawlers bring organic traffic, and AI assistants increasingly visit on behalf of real users who are researching a purchase. The question is not "bot or human" but "who is this, and do I want it here?"

Verified crawlers can be proven. Major search engines and AI operators publish the IP ranges their bots use, and a growing number sign their requests cryptographically. A request that claims to be a known crawler but arrives from a residential proxy is an impostor, and it should be treated like any other scraper. See how Kavra handles [AI agents](https://kavralab.com/detect/ai-agents/) and [good bots](https://kavralab.com/glossary/good-bot/).

- **Allow** verified search crawlers and the AI agents you want to be visible to.
- **Check** unverified declared bots: rate-limit them or serve cached pages.
- **Block** impostors and stealth scrapers on price, inventory and API endpoints.

## How to prevent web scraping

Effective protection looks at each request from several angles and focuses on the data that matters most to you. A practical checklist:

- List the endpoints that expose value: prices, availability, search, listing detail and any internal API your front end calls.
- Assess requests on those endpoints, including API calls, not only page loads. Scrapers skip the page when they can.
- Recognize proxy traffic by what it is, measured against real exit IPs, not by country or ASN alone.
- Check that the browser is what it claims to be, and flag automation even when it runs a full browser.
- Track actors, not IPs: one scraper that rotates fingerprints and addresses is still one actor.
- Verify declared crawlers against their operators' published ranges and signatures before you trust the name.
- Choose the response per endpoint: block the price API, slow the catalog, serve stale data to suspected scrapers, allow verified bots.

## A shopper browsing vs a scraper harvesting

**Real shopper**

- Loads pages, images and scripts in a natural order
- Views a handful of products, compares, goes back
- One home or mobile connection that matches the device
- Consistent browser on every layer

**Scraper bot**

- Calls JSON endpoints directly, skips assets
- Walks the full catalog in ID order
- New residential proxy exit for every request
- Headless browser dressed up as desktop Chrome

> **Key takeaway:** Scrapers win when defenses trust IP addresses and user agents. Stop them by judging each request on network, device and behavior together, protect the endpoints that expose prices and stock, and keep the door open for crawlers that can prove who they are. Related: [API abuse](https://kavralab.com/solutions/api-abuse/).

## How Kavra stops web scraping

Kavra analyzes 3,000+ data points on every request and tells your backend whether it came from a person, a scraper or a verified crawler.

- **Proxy intelligence**: Kavra measures real exit IPs of commercial residential and mobile proxy networks, so a scraper on a home IP is still seen as proxy traffic.
- **Automation exposed**: Headless browsers, stealth plugins and HTTP clients imitating browsers are caught by contradictions between what they claim and how they behave.
- **Seen at the edge**: Kavra's own edge network sees the real connection, not only what the browser reports about itself.
- **Verified crawlers allowed**: Search crawlers and AI agents are verified by signature and published IP ranges. Impostors using their names are flagged as unverified.
- **Protects APIs too**: A server-side request API assesses calls to price, search and inventory endpoints that never load your page script.
- **Your policy per endpoint**: Allow, check or block by bot type and route. Start in observe-only mode to see who is scraping before you act.

## FAQ

### Is web scraping illegal?

It depends on what is scraped, how and where. Collecting public data is not automatically illegal, but scraping can breach terms of service, copyright, database rights or privacy law, especially when it copies personal data or protected content, or bypasses technical barriers. Because legal remedies are slow, most businesses rely on technical protection first and legal action second.

### Can you block scrapers that use residential proxies?

Yes, but not with IP blocklists. Residential proxies route traffic through real household connections, so the address alone looks clean. Kavra measures the exit IPs of commercial proxy networks directly and combines that with device and behavior evidence, so a scraper is recognized even when each request arrives from a new home IP.

### Should I block AI crawlers?

That is a business choice. Some AI crawlers collect training data, others fetch pages live to answer a user's question and can send you visitors. A sensible policy allows verified agents you want to be visible to, rate-limits or serves cached pages to the rest, and blocks anything that fakes an operator's identity.

### Is robots.txt enough to stop scraping?

No. robots.txt tells well-behaved crawlers which pages to skip, and reputable search engines respect it. Hostile scrapers simply ignore it, and some read it to find the paths you consider sensitive. Use robots.txt to guide good bots, and use bot detection to enforce your rules on the rest.

### How do I tell Googlebot from a fake Googlebot?

Check where the request comes from, not what it says. Google and other operators publish the IP ranges their crawlers use and support reverse DNS verification. A request that claims to be a search crawler but arrives from a residential or cloud address outside those ranges is an impostor. Kavra runs this verification automatically on every request.

### Does scraping protection affect SEO?

Not if verified crawlers are allowed. Search engines only need to reach your pages from their own published networks, and Kavra recognizes them there. Problems start when sites block by user agent or show CAPTCHA puzzles to everything automated, which can lock out real crawlers while stealth scrapers slip through.

---
Kavra Lab: bot and fraud detection that explains every decision. Book a demo: https://kavralab.com/contact/
