Glossary

What is web scraping?

Web scraping is the automated collection of data from websites: a program requests pages, pulls out the parts it wants, such as prices, product details, listings or articles, and saves them in a structured form. It powers search engines and research, and also lets competitors copy a site at scale. Kavra separates welcome crawlers from unwanted scrapers.

Web scraping in plain terms

A person reading a product page sees a photo, a name and a price. A scraper sees the same page as text with a predictable structure, finds the price by its position or label, stores it, and moves to the next page in a fraction of a second. Repeat that across a whole catalog every hour and you have a live copy of someone else's pricing. Nobody had to break in; the data was public, just never meant to be collected in bulk.

Scraping is not a single tool. It ranges from short scripts that fetch raw pages, to headless browsers that run the site's JavaScript, to paid scraping services that bundle rotating residential proxies and CAPTCHA solving. Scraping that goes through a site's private API instead of its pages is common too, because the data arrives already structured.

What gets scraped, and by whom

TargetTypical scraperWhy it hurts
Prices and stock levelsCompetitors, price trackersYour prices are undercut within minutes
Listings and reviewsRival marketplaces, lead sellersYour supply and content appear elsewhere
Odds and faresArbitrage tools, resellersArbitrage, inflated look-to-book costs
Articles and mediaContent farms, AI training crawlersTraffic and licensing value lost
User profilesSpammers, data brokersPrivacy risk and phishing lists

How scraping shows up in real traffic

Scrapers rarely announce themselves. On dashboards they look like a steady rise in page views with no matching rise in sessions that add to cart, sign up or buy. Search and filter pages get hit far more than people use them. Product pages are read in catalog order, around the clock, and prices on a rival's site change minutes after yours.

The cost is not only lost margin. Scrapers eat server capacity, skew conversion and advertising data, and in travel they raise the cost of every fare search you pay a supplier for.

  • Many IP addresses from residential proxy pools, each making only a few requests.
  • Browsers that never scroll, hover or pause, or pause at perfectly even intervals.
  • Direct calls to internal APIs with no page load before them.
  • Visitors claiming to be a search crawler from outside that engine's IP ranges.

Web scraping vs web crawling

Crawling

  • Discovers and indexes pages across many sites
  • Done by search engines that declare themselves
  • Follows robots.txt and crawl-rate limits
  • Usually sends visitors back to you

Unwanted scraping

  • Extracts specific data from one site, repeatedly
  • Hides behind proxies and spoofed browsers
  • Ignores robots.txt when it gets in the way
  • Uses your data to compete with you

How to detect and stop scraping

Rate limits and IP blocklists catch the lazy scrapers. The rest rotate addresses and devices, so detection has to recognize automation by how the visitor behaves and whether its device, browser and network agree with each other. Kavra does this on every request and verifies declared crawlers so real search engines keep access. For the full playbook, including signals, endpoints to protect and response options, see web scraping protection. Scraping of odds and fares is covered for travel and ticketing as well.

FAQ

Frequently asked questions

Something else? Talk to our team.

What is the difference between web scraping and screen scraping?

Web scraping pulls data out of a site's pages or APIs as text and structure. Screen scraping, the older term, reads what is displayed on a screen, sometimes from images, and was used to pull data from legacy systems. Modern scrapers blur the line: some AI-driven tools read screenshots of pages, which is screen scraping applied to the web.

What tools are used for web scraping?

Common categories are HTTP libraries that fetch raw pages, parsing libraries that pull fields out of HTML, headless browser frameworks such as Playwright and Puppeteer for JavaScript-heavy sites, and hosted scraping services that add rotating proxies and CAPTCHA solving. Point-and-click scraping extensions let non-developers do the same on a smaller scale.

Can scraping slow down my website?

Yes. Aggressive scrapers request thousands of pages, including search and filter pages that are expensive to generate, and they skip the cache-friendly paths real users take. The result can be slower pages, higher hosting bills and inflated analytics, even when the scraped data itself is not sensitive.

See who is really on your site.

Run Kavra on your own traffic in observe-only mode. No risk to your customers, and a clear report of the fraud it finds.