Glossary
bot protection
Bot protection is commercial software - DataDome, Cloudflare challenges - that filters automated traffic, usually stacked on terms-of-use bans. G2 keeps its 2,239 category pages openly readable while its product and review pages sit behind DataDome bot protection, gating depth rather than existence.
What is bot protection?
Bot protection is the commercial filtering tier - DataDome, Cloudflare challenges and similar systems - that decides which automated clients see content at all. Its defining feature in practice is selective depth: some of a site stays open while the valuable part hardens.
G2 is the clean example in this slice. Its 2,239 category pages across 43 top-level groups are openly readable, but per its catalog record "product and review pages sit behind DataDome bot protection". Product Hunt shows the same split differently: public launch pages sit behind Cloudflare while an official GraphQL API serves structured data. Capterra permits AI and search crawlers by name yet "generic agents face Cloudflare challenges" across its 86,267 product profiles. Similarweb closes the loop in writing, with terms section 6 banning "any robot, spider, scraper, or other automated means" outright.
Why does bot protection matter when choosing a dataset?
It gates depth, not existence. The visible surface - category listings, leaderboards - is often deliberately open for search engines, so a pilot crawl succeeds and the production crawl fails at exactly the records you needed. That asymmetry is why scraping estimates built on sampling category pages misprice projects on software directories.
Legal exposure stacks on top. Where protection exists, terms usually ban automation explicitly (Similarweb section 6), so a blocked request is also a documented violation. The durable alternatives are official APIs where offered, sitemap-enumerated polite crawling where tolerated, or licensing.
How do you evaluate bot protection in a source?
- Map which paths are hardened. At G2, categories read openly while product pages do not - test the specific record type you need, not the homepage.
- Check named-agent rules. Capterra allows ClaudeBot, GPTBot and peers while challenging generic agents; your user agent determines your experience.
- Prefer the sanctioned API. Product Hunt's GraphQL API exists precisely because its Cloudflare-fronted pages discourage scraping.
- Read the terms clause numbers. Similarweb's section 6 ban is explicit enough to quote; vaguer sources deserve more caution, not less.
- Assess refresh cost honestly. Daily-updated directories behind challenges mean daily maintenance, which is a salary line, not a one-off fee.
Where to see it in context: application-software data.
Related terms
Entries adjacent to bot protection:
- robots-txt-crawl-rules — the declared permissions that sit above Capterra's named-crawler allowances
- web-catalog-scraping — the extraction practice challenging 86,267 product pages like G2's and Capterra's
Frequently asked questions
What is the difference between DataDome and Cloudflare bot protection?
Both are commercial anti-bot systems; the difference in this catalog is where they sit. G2 runs DataDome on product and review pages while keeping category pages open, and Product Hunt and Capterra surface Cloudflare challenges to generic agents.
Can AI crawlers access sites with bot protection?
Sometimes yes, selectively. Capterra's robots.txt admits ClaudeBot, GPTBot, ChatGPT-User, PerplexityBot and Google-Extended while generic scripted agents meet Cloudflare challenges.
Is bypassing bot protection legal?
It typically breaches the site's terms regardless of technical feasibility - Similarweb's terms, for example, ban any robot, spider or scraper accessing the platform. Licensing or official APIs are the defensible routes.
Datasets containing this field
Datasets containing bot protection
6 datasets carry bot protection in the catalog. Open one, count the fields, judge for yourself.
Apple App Store - 10k Apps Dataset (Kaggle - ramamet4)
Apple iTunes Search API
BIS Data Portal - Bulk Downloads
FREQ · L_MEASURE / L_POSITION / L_INSTR / L_DENOM · L_CURR_TYPE …+8 more
Capterra Software Directory
Chrome Web Store - Extensions & Apps
Data.gov - Software Datasets Catalog
programCode · mediaType (per resource) · views-last-month …+2 more
Every listing shows the field dictionary, sample rows, and coverage before you commit. API, files, or your warehouse. Daily, weekly, or hourly.
Get sample rows