# ============================================================================= # Datadini robots.txt # ============================================================================= # # LEGAL NOTICE — please read before automated access. # # Automated access, scraping, crawling, harvesting, indexing, mirroring, or # systematic downloading of any portion of this website, the Datadini # platform, our APIs, or our underlying datasets is PROHIBITED without prior # written consent from Datadini Ltd, including (without limitation) by: # # * AI / machine-learning training systems, large language model providers, # foundation model operators, and retrieval-augmented generation systems; # * Third-party data aggregators, business-information vendors, lead- # generation platforms, SEO intelligence services, and market-research # firms; # * Search-indexing, archiving, or cataloguing services other than the # major general-purpose web search engines operating within their # published, polite crawl behaviour and respecting these directives. # # Datadini's databases constitute a qualifying database under the Copyright # and Rights in Databases Regulations 1997 (UK). Extraction or re-utilisation # of all or a substantial part of our database without authorisation # infringes our database right and gives rise to a civil claim, including # for damages, costs of mitigation, and injunctive relief, as well as # claims under the Computer Misuse Act 1990 and the Copyright, Designs and # Patents Act 1988. # # Any access by an automated agent constitutes acceptance of our full Terms # of Use, available at: https://datadini.com/docs/legal/terms # # Licensing enquiries: legal@datadini.com # # This notice is in addition to the per-User-agent Disallow rules below; # absence of an explicit Disallow does NOT imply consent to scrape. # ============================================================================= # Verified search crawlers — indexable public pages; block sensitive paths. User-agent: Googlebot Allow: / Allow: /company/ Allow: /people/ Allow: /property/ Allow: /search Allow: /about Allow: /blog/ Allow: /sitemap-static.xml Allow: /sitemaps/ Disallow: /cdn-cgi/ Disallow: /dashboard/ Disallow: /dashboards/ Disallow: /account/ Disallow: /api/ Disallow: /admin/ Disallow: /testing Disallow: /interview Disallow: /accenture Disallow: /accenture-v2 Disallow: /maintenance Disallow: /unsubscribe Disallow: /honeypot Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out User-agent: Google-InspectionTool Allow: / Disallow: /honeypot Disallow: /api/ Disallow: /admin/ User-agent: Bingbot Allow: / Allow: /company/ Allow: /people/ Allow: /property/ Disallow: /honeypot Disallow: /api/ Disallow: /admin/ User-agent: * # Cloudflare internal endpoints (email obfuscation, analytics, etc.) Disallow: /cdn-cgi/ # Auth-gated and internal paths, intentionally blocked Disallow: /dashboard/ Disallow: /dashboards/ Disallow: /account/ Disallow: /api/ Disallow: /admin/ Disallow: /testing Disallow: /interview Disallow: /accenture Disallow: /accenture-v2 Disallow: /maintenance Disallow: /unsubscribe # Honeypot — see app/honeypot/route.ts. Hitting this URL auto-blocks the # requesting IP for 24 h. Real crawlers must respect this Disallow. Disallow: /honeypot # Block specific legal form pages from indexing Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out # Explicitly allow sitemaps and company profiles Allow: /sitemap.xml Allow: /sitemap-static.xml Allow: /sitemaps/ Allow: /company/ Allow: /property/ Allow: /enterprise Allow: /services Allow: /tools/ # Allow all public content Allow: /docs/legal/ Allow: /docs/ Allow: / # --------------------------------------------------------------------------- # Crawl pacing + Disallow mirror # --------------------------------------------------------------------------- # We have ~5.5M /company/ URLs in the sitemap. An aggressive crawl burns # Vercel Function Duration on every ISR cache miss. Googlebot ignores # Crawl-delay (it self-paces), but Bing/Yandex/DuckDuckBot/etc. honour it. # 10s = ~8,640 pages/day per crawler — plenty for new-page discovery while # bounding worst-case daily cost. # # IMPORTANT (robots.txt spec gotcha): when a bot finds a named User-agent # section it uses ONLY that section and IGNORES the `User-agent: *` block # above. So every named bot below must repeat the Disallow rules — most # critically `/honeypot`, otherwise the bot crawls the hidden trap link on # /search, gets auto-blocklisted, and effectively bans itself from the site # forever. Keep the Disallow list here in sync with the wildcard above. User-agent: bingbot Crawl-delay: 10 Disallow: /cdn-cgi/ Disallow: /dashboard/ Disallow: /api/ Disallow: /admin/ Disallow: /maintenance Disallow: /unsubscribe Disallow: /honeypot Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out User-agent: Yandex Crawl-delay: 10 Disallow: /cdn-cgi/ Disallow: /dashboard/ Disallow: /api/ Disallow: /admin/ Disallow: /maintenance Disallow: /unsubscribe Disallow: /honeypot Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out User-agent: DuckDuckBot Crawl-delay: 10 Disallow: /cdn-cgi/ Disallow: /dashboard/ Disallow: /api/ Disallow: /admin/ Disallow: /maintenance Disallow: /unsubscribe Disallow: /honeypot Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out User-agent: BaiduSpider Crawl-delay: 10 Disallow: /cdn-cgi/ Disallow: /dashboard/ Disallow: /api/ Disallow: /admin/ Disallow: /maintenance Disallow: /unsubscribe Disallow: /honeypot Disallow: /docs/legal/data-rights Disallow: /docs/legal/dinibot-opt-out User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: SerpstatBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: PetalBot Disallow: / User-agent: ZoominfoBot Disallow: / # --------------------------------------------------------------------------- # AI TRAINING crawlers — full block. # --------------------------------------------------------------------------- # These crawlers harvest your data to train commercial AI models that compete # with you, and send zero traffic back. No upside, all cost. We also block # them at middleware.ts (defence in depth) — robots.txt alone is voluntary. # # NOTE: We DO allow real-time user-fetch bots (ChatGPT-User, OAI-SearchBot, # Perplexity-User). Those only fire when a real human asks an AI about a # specific company, so they bring traffic and intent. Blocking them means # AI assistants tell users "I don't have info on Datadini" — bad for # discovery as AI search becomes a real acquisition channel. # --------------------------------------------------------------------------- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: cohere-ai Disallow: / User-agent: Google-Extended Disallow: / # Canonical sitemap declarations. # Order: static (homepage/search) → priority companies → full company index. Sitemap: https://datadini.com/sitemap-static.xml Sitemap: https://datadini.com/sitemaps/companies-priority.xml Sitemap: https://datadini.com/sitemaps/companies.xml Sitemap: https://datadini.com/sitemap.xml # Chunk files live on Cloudflare R2: https://sitemaps.datadini.com/