Web scraping in plain terms
A person reading a product page sees a photo, a name and a price. A scraper sees the same page as text with a predictable structure, finds the price by its position or label, stores it, and moves to the next page in a fraction of a second. Repeat that across a whole catalog every hour and you have a live copy of someone else's pricing. Nobody had to break in; the data was public, just never meant to be collected in bulk.
Scraping is not a single tool. It ranges from short scripts that fetch raw pages, to headless browsers that run the site's JavaScript, to paid scraping services that bundle rotating residential proxies and CAPTCHA solving. Scraping that goes through a site's private API instead of its pages is common too, because the data arrives already structured.
What gets scraped, and by whom
| Target | Typical scraper | Why it hurts |
|---|---|---|
| Prices and stock levels | Competitors, price trackers | Your prices are undercut within minutes |
| Listings and reviews | Rival marketplaces, lead sellers | Your supply and content appear elsewhere |
| Odds and fares | Arbitrage tools, resellers | Arbitrage, inflated look-to-book costs |
| Articles and media | Content farms, AI training crawlers | Traffic and licensing value lost |
| User profiles | Spammers, data brokers | Privacy risk and phishing lists |
How scraping shows up in real traffic
Scrapers rarely announce themselves. On dashboards they look like a steady rise in page views with no matching rise in sessions that add to cart, sign up or buy. Search and filter pages get hit far more than people use them. Product pages are read in catalog order, around the clock, and prices on a rival's site change minutes after yours.
The cost is not only lost margin. Scrapers eat server capacity, skew conversion and advertising data, and in travel they raise the cost of every fare search you pay a supplier for.
- Many IP addresses from residential proxy pools, each making only a few requests.
- Browsers that never scroll, hover or pause, or pause at perfectly even intervals.
- Direct calls to internal APIs with no page load before them.
- Visitors claiming to be a search crawler from outside that engine's IP ranges.
Web scraping vs web crawling
Crawling
- Discovers and indexes pages across many sites
- Done by search engines that declare themselves
- Follows robots.txt and crawl-rate limits
- Usually sends visitors back to you
Unwanted scraping
- Extracts specific data from one site, repeatedly
- Hides behind proxies and spoofed browsers
- Ignores robots.txt when it gets in the way
- Uses your data to compete with you
How to detect and stop scraping
Rate limits and IP blocklists catch the lazy scrapers. The rest rotate addresses and devices, so detection has to recognize automation by how the visitor behaves and whether its device, browser and network agree with each other. Kavra does this on every request and verifies declared crawlers so real search engines keep access. For the full playbook, including signals, endpoints to protect and response options, see web scraping protection. Scraping of odds and fares is covered for travel and ticketing as well.