GET /api/v1/prices?sku=*BlockedHeadless scraper on a rotating proxy pool
- Headless Chrome, stealth plugin
- Residential proxy exit
- 2,400 SKUs, no page views
Solution
Web scraping is the automated copying of a website's prices, listings, content or stock data by bots, usually hidden behind residential proxies and headless browsers so they look like shoppers. Kavra tells scrapers from real visitors and verified search or AI crawlers on every request, so you can block, slow or allow each one.
GET /api/v1/prices?sku=*BlockedHeadless scraper on a rotating proxy pool
Web scraping is the use of software to collect data from web pages or APIs at scale. A person browsing a product page reads one price. A scraper requests every product page, every search result and every hidden API behind them, then stores the data in a database someone else controls.
Not all scraping is hostile. Search engines crawl your site so people can find it, and some AI assistants fetch pages to answer a user's question. The problem is the other traffic: competitors pulling your full price list every hour, aggregators republishing your listings, resellers watching your stock for restocks, and AI crawlers that ignore your rules and copy content you paid to create. Good web scraping protection separates those groups instead of treating every bot the same.
Scraping moved from simple scripts to rented infrastructure that imitates real shoppers.
The operator finds the pages and, more often, the internal JSON endpoints your own front end calls for prices, availability and search results. APIs return clean data and are cheaper to hit than full pages.
Lightweight HTTP clients that imitate a browser handle easy targets. When a site checks for JavaScript, the operator switches to headless browsers driven by Playwright, Puppeteer or Selenium, often with stealth plugins that hide automation flags.
Requests are spread across a pool of residential and mobile proxies, so each one comes from a different household in the right country and never trips a per-IP rate limit.
Device fingerprints, headers and sessions are rotated between requests. Blocked requests are retried through a new exit. Commercial scraping services sell this whole stack as a single API call.
Prices feed a competitor's repricing engine, listings appear on a copycat site, stock levels trigger reseller bots, and articles end up in someone else's product or training set.
The cost depends on what the data is worth to someone else.
| What gets scraped | Who is usually behind it | Business impact |
|---|---|---|
| Prices and promotions | Competitors and repricing tools | Undercut within minutes, margin pressure, lost sales |
| Product and listing content | Copycat stores, aggregators | Duplicate content, lost search traffic, stolen photos |
| Inventory and availability | Resellers and scalping bots | Restocks hoarded before real customers see them |
| Fares, rates and seat maps | Metasearch and resale sites | Look-to-book ratios explode, supplier API bills rise |
| Articles, reviews and user posts | Unverified AI crawlers, content farms | Paid content reused without permission or credit |
| Seller, user or job data | Lead brokers, spammers | Privacy exposure, poached sellers, spam to your users |
Scrapers rarely trip one alarm. They leave patterns that humans do not.
Requests go straight to price or search endpoints without loading the page, images or scripts a browser would fetch first.
Every product ID, every page of results, every date in a calendar, requested in order and at a steady pace no shopper keeps.
Thousands of visits from different residential addresses, each seen once, many belonging to commercial proxy networks.
A visitor claims to be Chrome on a laptop but has no real screen, no input history and automation traces under a thin disguise.
Requests announce themselves as a well-known crawler but come from networks the crawler's operator does not use.
Traffic and infrastructure costs climb while conversion rates fall, especially on search and availability pages.
Scraping tools are built to pass the checks most sites rely on.
| Defense | What it assumes | How scrapers get past it |
|---|---|---|
| robots.txt | Bots follow the rules | It is a request, not a lock; hostile scrapers ignore it |
| Per-IP rate limits | One IP is one visitor | Residential proxy pools give each request a fresh address |
| IP and ASN blocklists | Bad traffic comes from data centers | Proxy traffic exits from real home and mobile networks |
| User-agent filtering | Bots say who they are | Scrapers copy a current browser or a search crawler name |
| JavaScript challenges | Bots cannot run scripts | Headless browsers run them like any other browser |
| CAPTCHA | Only humans can solve it | Solving services clear puzzles cheaply; real shoppers get the friction |
Blocking all automation would hurt you. Search crawlers bring organic traffic, and AI assistants increasingly visit on behalf of real users who are researching a purchase. The question is not "bot or human" but "who is this, and do I want it here?"
Verified crawlers can be proven. Major search engines and AI operators publish the IP ranges their bots use, and a growing number sign their requests cryptographically. A request that claims to be a known crawler but arrives from a residential proxy is an impostor, and it should be treated like any other scraper. See how Kavra handles AI agents and good bots.
Effective protection looks at each request from several angles and focuses on the data that matters most to you. A practical checklist:
How Kavra helps
Kavra analyzes 3,000+ data points on every request and tells your backend whether it came from a person, a scraper or a verified crawler.
Kavra measures real exit IPs of commercial residential and mobile proxy networks, so a scraper on a home IP is still seen as proxy traffic.
Headless browsers, stealth plugins and HTTP clients imitating browsers are caught by contradictions between what they claim and how they behave.
Kavra's own edge network sees the real connection, not only what the browser reports about itself.
Search crawlers and AI agents are verified by signature and published IP ranges. Impostors using their names are flagged as unverified.
A server-side request API assesses calls to price, search and inventory endpoints that never load your page script.
Allow, check or block by bot type and route. Start in observe-only mode to see who is scraping before you act.
FAQ
Something else? Talk to our team.
It depends on what is scraped, how and where. Collecting public data is not automatically illegal, but scraping can breach terms of service, copyright, database rights or privacy law, especially when it copies personal data or protected content, or bypasses technical barriers. Because legal remedies are slow, most businesses rely on technical protection first and legal action second.
Yes, but not with IP blocklists. Residential proxies route traffic through real household connections, so the address alone looks clean. Kavra measures the exit IPs of commercial proxy networks directly and combines that with device and behavior evidence, so a scraper is recognized even when each request arrives from a new home IP.
That is a business choice. Some AI crawlers collect training data, others fetch pages live to answer a user's question and can send you visitors. A sensible policy allows verified agents you want to be visible to, rate-limits or serves cached pages to the rest, and blocks anything that fakes an operator's identity.
No. robots.txt tells well-behaved crawlers which pages to skip, and reputable search engines respect it. Hostile scrapers simply ignore it, and some read it to find the paths you consider sensitive. Use robots.txt to guide good bots, and use bot detection to enforce your rules on the rest.
Check where the request comes from, not what it says. Google and other operators publish the IP ranges their crawlers use and support reverse DNS verification. A request that claims to be a search crawler but arrives from a residential or cloud address outside those ranges is an impostor. Kavra runs this verification automatically on every request.
Not if verified crawlers are allowed. Search engines only need to reach your pages from their own published networks, and Kavra recognizes them there. Problems start when sites block by user agent or show CAPTCHA puzzles to everything automated, which can lock out real crawlers while stealth scrapers slip through.
Run Kavra on your own traffic in observe-only mode. No risk to your customers, and a clear report of the fraud it finds.