# Headless browser detection that sees the script behind the page

Source: https://kavralab.com/detect/headless-browsers/

A **headless browser** is a real browser engine, usually Chrome, run without a visible window and driven by code through frameworks like Playwright, Puppeteer or Selenium. Attackers use it to scrape, stuff credentials and buy out stock while looking like a normal visitor. Kavra catches it by comparing what the browser claims with how it actually runs.

- **What it is:** A real browser driven by code, not a person
- **Attacks it enables:** Scraping, credential stuffing, scalping, fake signups
- **Common tools:** Playwright, Puppeteer, Selenium, stealth plugins
- **Where to check:** Login, signup, checkout, search and pricing

## What is a headless browser?

A headless browser is an ordinary browser engine that runs without drawing a window on a screen. It still loads pages, runs JavaScript, stores cookies and submits forms. The difference is who is at the controls: a program sends it commands such as "open this URL, type this email, click Log in" and reads back the result.

Developers built these tools for testing, and they are excellent at it. The same qualities make them attractive to attackers. A headless browser passes the simple checks that stop basic scripts, because it really does execute your page. Today the term covers a whole family of automation, from headless Chrome to full desktop browsers driven by a framework, with or without a visible window.

## The automation toolkit, from simple to stealthy

Bots sit on a ladder. Each rung costs the operator more and gets past more checks.

| Tool category | What it is | What it gets past |
|---|---|---|
| Plain HTTP scripts | Code that sends raw requests with a copied browser header | Rate limits spread across many IPs; nothing that needs JavaScript |
| HTTP clients that imitate browsers | Libraries tuned to make the connection itself look like Chrome or Safari | Checks that only read headers or the connection's surface |
| Headless Chrome | The real Chrome engine, run without a window | JavaScript checks, cookie checks, simple bot tests |
| Automation frameworks | Playwright, Puppeteer and Selenium driving Chrome, Firefox or WebKit | Multi-step flows: logins, carts, checkouts, forms |
| Stealth plugins and patched drivers | Add-ons and forks that hide the automation flag and fake missing browser features | Checks for the obvious automation markers |
| Automation plus antidetect and proxies | A framework driving [antidetect browser](https://kavralab.com/detect/antidetect-browsers/) profiles over [residential proxies](https://kavralab.com/detect/residential-proxies/) | Device checks and IP reputation, one layer at a time |

## How attackers use headless browsers

The framework is only the engine. The damage depends on which of your pages it is pointed at.

1. **Scraping prices, stock and content**: Scripts crawl product pages, search results and fares, often logged in, to copy catalogs or undercut prices. Headless browsers are chosen when pages render with JavaScript or hide data behind interactions. See [web scraping](https://kavralab.com/solutions/web-scraping/).
2. **Testing stolen passwords**: A framework loads the real login page, types each leaked email and password pair and reads the response, so the attempt looks like a normal form submission rather than a raw API call. See [credential stuffing](https://kavralab.com/solutions/credential-stuffing/).
3. **Buying out limited stock**: Scalping bots keep sessions warm, watch for a drop, then add to cart and check out faster than any person can. See [scalping](https://kavralab.com/solutions/scalping/).
4. **Creating fake accounts**: The same script fills signup forms, confirms emails from disposable inboxes and claims a welcome offer, again and again. See [fake account creation](https://kavralab.com/solutions/fake-accounts/).
5. **Card testing and form spam**: Checkout and donation forms are hit with small payments to validate stolen cards, and contact or review forms are flooded with spam.

## Why stealth plugins are not enough

Stealth plugins and undetected drivers exist because out-of-the-box automation is easy to spot. Headless Chrome used to announce itself in its user agent, set a flag saying it was under remote control, and lacked plugins, languages and graphics features a real desktop browser has. Stealth tools patch each of these: they rename the browser, clear the flag, and fill the gaps with plausible fake values.

The trouble for the attacker is that each patch fixes one question, not the whole story. A browser is thousands of properties that depend on each other and on the machine and network underneath. When a patch says one thing and the engine, the hardware or the connection says another, the disguise shows. The more values are faked, the more places there are for them to disagree.

- The browser claims to be a Mac, but renders fonts and graphics like a Linux server with no graphics card.
- The version string says the latest Chrome, but the features present belong to an older or different engine.
- The patched properties look right when read directly, but the patch itself leaves traces in how the page's code behaves.
- The page says Chrome, but the network connection underneath was made by a scripting library.
- The window has a normal size, yet nothing moved, scrolled or paused the way a person does.

## The layers that give automation away

No single check is decisive. Kavra looks at every layer and at the contradictions between them.

- **Browser integrity**: Missing, extra or altered browser features, signs of injected code and properties that were rewritten after the page loaded.
- **Device and environment**: Software graphics instead of a real graphics card, server-class hardware, default screen sizes and fonts that match no consumer device.
- **Network**: Datacenter and cloud hosting ranges, proxy exits, and a connection whose shape does not match the browser it claims to be.
- **Behavior**: Instant form fills, pasted values, clicks at the exact center of buttons, straight-line or missing pointer movement, no reading time.
- **Velocity and history**: Many logins or checkouts from one environment, identical sessions repeating across accounts, visits at machine intervals.
- **Rotation**: The same automated setup returning with a fresh fingerprint on every run. Changing identity on each visit is a signal in itself.

## Good automation: QA, monitoring and partners

Not every headless browser is an attacker. Your own team runs end-to-end tests against staging and production. Uptime and performance monitors load key pages every few minutes. Accessibility scanners, SEO audit tools, partners with an agreed data feed and verified search crawlers all automate your site with permission. Blocking them breaks releases, triggers false alerts and damages search visibility.

The answer is to identify good automation positively, instead of hoping it slips through. Detection should still recognize these visitors as automated; your policy then decides to let them in.

- Tag your own test and monitoring runs so they can be matched by identity, not by guesswork. Kavra's server-side API lets you mark known sessions and devices as good.
- Allow partner automation by account or API key, and route bulk data through an API instead of your web pages.
- Let verified search crawlers and signed [AI agents](https://kavralab.com/detect/ai-agents/) through, and flag anything that only claims to be one.
- Scope allowlists narrowly: a monitor needs your status page, not your login form.
- Start new rules in observe-only mode and review what would have been blocked before enforcing.

## A person in Chrome vs a stealth-patched headless browser

**Person in a real browser**

- Browser, hardware and fonts tell one consistent story
- Home broadband or mobile carrier network
- Scrolls, pauses, moves the pointer on curved paths
- A handful of logins, spread over days

**Stealth-patched automation**

- Patched values that contradict the engine and hardware
- Cloud, datacenter or proxy exit
- Fields filled in milliseconds, no pointer path
- Hundreds of near-identical sessions in an hour

## Checklist: defending against headless automation

A defense that holds up against modern automation assesses the session, not only the request, and decides at the pages bots care about.

- Assess every login, signup, checkout and search, not a random sample.
- Do not trust the user agent or any single browser property. Compare them with the engine, hardware and network.
- Protect your APIs as well as your pages. Bots skip the page when they can, so use a server-side check on backend traffic too.
- Bind decisions to the action with short-lived, single-use tokens so a token earned by a real browser cannot be replayed by a script.
- Prefer invisible challenges over puzzles. CAPTCHA farms and solving services pass puzzles; consistency checks are harder to outsource.
- Allowlist good automation by verified identity, never by user agent string.

> **Key takeaway:** Stealth tools can make a headless browser answer any single question correctly. They struggle to make every layer agree at once. Detect automation by its **contradictions**, allow the automation you trust by identity, and decide at the pages where bots do damage: [login](https://kavralab.com/solutions/account-takeover/), signup and checkout.

## How Kavra detects headless browsers and automation

Kavra analyzes 3,000+ data points on every visit and explains in plain language why a session looks automated, so your backend can act with confidence.

- **Automation frameworks recognized**: Playwright, Puppeteer, Selenium, headless Chrome, stealth plugins and patched drivers are identified by how they run, not by what they claim.
- **Contradictions across layers**: Browser, device, network and behavior are checked against each other, so a patch that fixes one layer exposes itself in another.
- **Own edge network**: Kavra sees the real connection, which catches HTTP clients that imitate a browser and browsers driven by scripting libraries.
- **Signed, single-use tokens**: Each token is bound to one action and expires quickly. Replays from scripts are refused and reported.
- **Good bots stay welcome**: Verified crawlers and AI agents are recognized by signature and operator IP ranges. Mark your own QA and monitoring as trusted.
- **Your rules decide**: Allow, verify or block per endpoint. Start in observe-only mode and see every automated session before enforcing.

## FAQ

### Can websites detect headless Chrome?

Yes. Older headless Chrome was easy to spot from its user agent and missing features. The newer headless mode is much closer to regular Chrome, so detection now relies on the combination: the environment it runs in, the graphics and hardware it reports, the network it uses, how it interacts with the page and whether patched values contradict the engine underneath.

### Does Playwright or Puppeteer stealth mode avoid bot detection?

It avoids the best-known checks, such as the automation flag and missing plugins. It does not make the whole session consistent. Stealth patches leave traces, fake values disagree with the hardware and network, and scripted interaction still looks scripted. Detection that compares layers catches stealth sessions that single checks miss.

### Is using a headless browser illegal?

No. Headless browsers are standard tools for testing, monitoring and research. What can be illegal or a breach of terms is what they are used for: logging in with stolen credentials, bypassing purchase limits, collecting personal data or overloading a site. Site owners decide which automation they allow on their own pages.

### How do I allow my own test automation without allowing bots?

Identify it positively. Run tests from known accounts or environments, tag those sessions or devices as trusted through your detection provider's API, and scope the exception to the pages the tests need. Do not allowlist by user agent string, because any bot can copy it. Keep detection on so you still see everything else.

### What is the difference between a headless browser and an HTTP client bot?

An HTTP client sends requests directly and never runs your page's code, so it is fast and cheap but fails anything that needs JavaScript. A headless browser runs the full page like a normal visitor. Some HTTP clients now mimic a real browser's connection to get past network checks, but they still cannot produce a genuine browser session.

### Does blocking headless browsers hurt SEO?

Not if good crawlers are verified rather than guessed. Search engines render pages with headless browsers, but they publish IP ranges and identify themselves in ways that can be checked. Allow verified crawlers, and flag visitors that only claim to be one. That keeps your pages indexed while impostor crawlers are treated as the bots they are.

---
Kavra Lab: bot and fraud detection that explains every decision. Book a demo: https://kavralab.com/contact/
