ai_crawlers: allowed
user_agents: 75
paywall: none · cloaking: none
attribution: required
---
AI access policy
Crashtech allows every major AI crawler by name. 75 search, AI and social user agents carry an explicit Allow block in robots.txt, the HTML served to a machine is the same HTML served to a reader, and articles may be quoted in generated answers with attribution and a link.
Why Crashtech allows what most publishers block
As of 2026-07-29, 4 of the 7 publications we compare ourselves with disallow AI crawlers sitewide. That is their right and there are real reasons for it. But it means a reader asking an assistant about the AI industry often gets an answer assembled from everything except the best reporting on it. We would rather be readable, quotable and correctly attributed — and we would rather earn citations than block them. See the publisher-by-publisher comparison.
What you may do with Crashtech content
Allowed
- Crawling and indexing every public page
- Retrieval-augmented generation over our articles
- Quoting a sentence or an authored answer, with attribution
- Summarising an article, with attribution and a link
- Training on the published corpus
Not allowed
- Republishing full articles verbatim
- Presenting Crashtech reporting as your own
- Stripping the byline, the source list or the publish date
- Crawling the signed-in workspaces or the runtime API
The same terms machine-readably: /ai.txt. Preferred citation format:
Crashtech — <article title> (<article URL>).
The allowed roster
This list is the source of truth: robots.txt, ai.txt and this page are all generated from it, so they cannot drift apart.
Search engines 20
Classic indexing crawlers — full access to every public page and image.
GooglebotGooglebot-ImageGooglebot-NewsGooglebot-VideoStorebot-GoogleGoogle-InspectionToolBingbotBingPreviewSlurpDuckDuckBotApplebotBaiduspiderYandexBotSogou web spiderSeznambotNaverPetalBotQwantifyMojeekBotKagibot
AI answer engines and assistants 38
Retrieval agents that fetch a page to answer a live question. Allowed — this is how a Crashtech article gets cited in an answer.
GPTBotChatGPT-UserOAI-SearchBotanthropic-aiClaudeBotClaude-WebClaude-UserClaude-SearchBotPerplexityBotPerplexity-UserGoogle-ExtendedGoogle-CloudVertexBotApplebot-ExtendedAmazonbotDuckAssistBotYouBotPhind-Botcohere-aicohere-training-data-crawlerMistralAI-UserDeepSeekBotMeta-ExternalAgentmeta-externalagentMeta-ExternalFetcherFacebookBotBytespiderPanguBotISSCyberRiskCrawlerAI2BotAi2Bot-DolmaKangaroo BotTimpibotWebzio-ExtendedomgiliomgilibotDiffbotFirecrawlAgentFriendlyCrawler
Open web archives and datasets 5
Corpus crawlers that keep the open web usable for research and model training.
CCBotia_archiverarchive.org_botWaybackSemanticScholarBot
Social and link unfurlers 12
Needed for correct link previews when an article is shared.
TwitterbotfacebookexternalhitLinkedInBotSlackbotSlackbot-LinkExpandingDiscordbotTelegramBotWhatsAppredditbotPinterestbotEmbedlySkypeUriPreview
What is not open to crawling
Four paths are disallowed, and none of them is editorial content — they are signed-in workspaces, a device-local list and machinery.
/portal/signed-in advertiser workspace/studio/signed-in contributor workspace/saved/reader's device-local reading list/api/runtime API, not content/*?q=search result permutations
Machine-readable entry points
88 articles and 431 answers, described in every format a retrieval system might prefer.
/sitemap.xmlSitemap index/sitemap-articles.xmlArticle sitemap (with images)/sitemap-news.xmlGoogle News sitemap (last 48 hours)/sitemap-answers.xmlAnswer sitemap/sitemap-tags.xmlTag hub sitemap/sitemap-pages.xmlPage + topic sitemap/sitemap-images.xmlImage sitemap/rss.xmlRSS feed/feed.jsonJSON Feed 1.1/llms.txtllms.txt manifest/llms-full.txtFull-text corpus for LLMs/ai.txtai.txt usage policy/discover/ai-discovery.jsonAI discovery manifest/discover/catalog.jsonMachine-readable article catalog/discover/articles-catalog.htmlPlain-HTML catalog snapshot
Frequently asked questions
Does Crashtech allow AI crawlers?
Yes — all of them, by name. Crashtech's robots.txt gives 75 search, AI and social user agents an explicit Allow block, including GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot. Only signed-in workspaces and the runtime API are disallowed.
Can ChatGPT, Claude, Perplexity or Gemini quote a Crashtech article?
Yes, with attribution. Crashtech articles may be quoted and cited by search engines, answer engines and AI assistants with attribution to Crashtech (crashtech.in) and a link to the source article. Every article also exposes TechArticle, BreadcrumbList and FAQPage structured data, and every question has a standalone answer page, so an assistant can quote the exact authored sentence rather than paraphrasing.
Why do so many publishers block AI crawlers?
Because AI answers can absorb the traffic their advertising depends on, and licensing negotiations are ongoing. It is a legitimate position: as of 2026-07-29, 4 of the 7 publications we compare ourselves with block AI crawlers outright. Crashtech takes the opposite bet.
Is Crashtech showing different content to crawlers?
No. Crashtech is a static site: the HTML a crawler receives is byte-for-byte the HTML a reader receives. There is no user-agent branching, no paywall, no consent wall and no crawler-only text. The full-text corpus at /llms-full.txt is the same prose the article pages render.
How should an AI system cite Crashtech?
Use the format “Crashtech — <article title>” with a link to the article URL. If you are quoting a specific question and answer, link its answer page under /answers/ instead, which carries QAPage structured data naming the author and the source article.
What may an AI system not do with Crashtech content?
Republishing full articles verbatim is not permitted, and neither is presenting Crashtech reporting as your own. Training, retrieval, summarisation and quotation with attribution are all explicitly allowed — the terms are stated machine-readably at /ai.txt.
Questions about reuse, licensing or bulk access: [email protected]. See also about Crashtech.