User-Agent bots
Bots
Crawlers, monitors, and HTTP clients we've seen in real traffic. Who runs them, what they do, and example User-Agent strings.
AI crawlers
Training and retrieval bots from AI labs, the ones people actually search for.
- ccBot crawlerAICommon Crawl open web corpus crawler.8.4K hits
- DiffbotAIDiffbot knowledge-graph crawler that extracts structured data from pages.1.6K hits
- The Knowledge AIAIThe Knowledge AI — automated agent (classification unknown).1.3K hits
- IAS CrawlerAIIntegral Ad Science brand-safety crawler.324 hits
- ccBot crawlerAICommon Crawl open web crawl — the public corpus many AI labs train on.
- BytespiderAIByteDance Bytespider crawler.20 hits
- SentiBotAISentiment / brand-monitoring crawler.14 hits
- DevinAIDevin — automated agent (classification unknown).10 hits
- OAI-SearchBotAIOpenAI crawler for SearchGPT / ChatGPT search results — separate from training crawls.6 hits
- YouBotAIYou.com AI search crawler.6 hits
- ClaudeBotAIAnthropic crawler that collects public web data for Claude model training.5 hits
- AmazonbotAIAmazon crawler used for Alexa and Amazon services that need public web content.4 hits
- GPTBotAIOpenAI crawler used to gather training data for GPT models. Respects robots.txt via GPTBot user-agent.4 hits
- ChatGPT-UserAIOpenAI user-triggered fetcher when ChatGPT browses the live web on a user's behalf.3 hits
- Meta-ExternalAgentAIMeta AI crawler (external agent) that fetches public pages for AI features.3 hits
- PerplexityBotAIPerplexity AI crawler that indexes pages for answer citations.3 hits
- Claude-UserAIAnthropic user-initiated fetcher for Claude's web browsing features.2 hits
- Perplexity-UserAIPerplexity user-triggered fetch when a query needs a live page.2 hits
- Claude-SearchBotAIAnthropic crawler for Claude's search / citation features.1 hits
- cohere-training-data-crawlerAICohere training-data crawler — explicit training corpus collection.1 hits
- DeepseekBotAIDeepSeek AI web crawler.1 hits
- DuckAssistBotAIDuckDuckGo DuckAssist crawler for AI-assisted instant answers.1 hits
- Gemini-Deep-ResearchAIGoogle Gemini Deep Research agent fetching pages for multi-step answers.1 hits
- Google-CloudVertexBotAIGoogle Cloud Vertex AI crawler.1 hits
- ImageSiftAIImageSift reverse-image / visual search crawler.1 hits
- AI2BotAIAllen Institute for AI (AI2) research crawler.
- Applebot-ExtendedAIApple robots.txt token controlling use of crawled content for Apple Intelligence training.
- Claude-WebAIAnthropic web-fetch agent used when Claude retrieves a page during a conversation.
- cohere-aiAICohere crawler for model training data collection.
- Google-ExtendedAIControls whether Google may use your content for Gemini / Vertex AI training (robots.txt).
- Meta-ExternalFetcherAIMeta fetcher for on-demand retrieval of public URLs.
- MistralAI-UserAIMistral AI user-triggered web fetch.
Search & crawlers
- aHrefs BotAhrefs SEO crawler for backlink and site research.217K hits
- YandexBotYandex main web search crawler.201K hits
- OmgilibotWeb crawler associated with content / news aggregation.80K hits
- Facebook CrawlerMeta / Facebook link-preview crawler (facebookexternalhit and related).80K hits
- NuzzelNuzzel news-link crawler (historical).56K hits
- Hatena BookmarkHatena Bookmark social-bookmarking crawler.53K hits
- Google FaviconGooglebot variant that fetches site favicons for search results.49K hits
- SemrushBotSemrush SEO crawler used for site audits and competitive research.40K hits
- MeltwaterNewsMeltwater media-monitoring crawler.36K hits
- Trendiction BotTrendiction / Talkwalker social listening crawler.28K hits
- HubSpotHubSpot crawler for marketing / CMS link and content features.27K hits
- MuckRackMuck Rack journalist / media database crawler.27K hits
- GrapeshotGrapeshot (Oracle) contextual advertising crawler.19K hits
- VagabondoWiseNut / Looksmart legacy search crawler.17K hits
- QwantbotQwant privacy-oriented search crawler.16K hits
- Sogou SpiderSogou (Tencent) Chinese search crawler.304K hits
- GooglebotGoogle's main web search crawler.495K hits
- Mediatoolkit BotMediatoolkit media-monitoring crawler.13K hits
- DuckDuckBotDuckDuckGo search crawler.13K hits
- YandexMobileBotYandex mobile-oriented search crawler.12K hits
- BingBotMicrosoft Bing web search crawler.331K hits
- Yeti/NaverbotNaver Yeti search crawler (Korea).12K hits
- MJ12 BotMajestic SEO crawler (MJ12bot).262K hits
- SMTBotSimilarTech crawler for technology-detection research.9.9K hits
- aHrefs BotAhrefs SEO crawler / brand user-agent.
- YandexImageResizerYandex image thumbnail / resizer bot.9.7K hits
- TrendsmapTrendsmap social / trend crawler.9.2K hits
- evc-batchEasyBib / Chegg citation crawler (evc-batch).8.5K hits
- FlipboardFlipboard content crawler for magazine-style feeds.8.0K hits
- BLEXBot CrawlerWebMeUp / BLEXBot backlink crawler.7.5K hits
- SurdotlyBotSurdotly SEO / research crawler.6.9K hits
- DuckDuckGoDuckDuckGo search crawler / brand UA.
- archive.org botInternet Archive Wayback Machine crawler.6.4K hits
- Domain Re-Animator BotDomain re-registration / monitoring crawler.6.2K hits
- AdbeatAdbeat competitive ad-intelligence crawler.5.9K hits
- YandexImagesYandex Images crawler.5.7K hits
- BarkrowlerBabbar (Barkrowler) SEO crawler.4.4K hits
- ZoominfoBotZoomInfo B2B data crawler.4.3K hits
- PanscientPanscient web research crawler.4.2K hits
- Scraping RobotGeneric commercial scraping service user-agent.4.2K hits
- DatanyzeDatanyze technographic research crawler.3.6K hits
- Yandex BotYandex main web search crawler.
- Baidu SpiderBaidu web search crawler.116K hits
- ApplebotApplebot — Apple's web crawler for Spotlight, Siri, and related features.14K hits
- DotBotMoz / Open Site Explorer DotBot crawler.26K hits
- BLEXBot CrawlerWebMeUp / BLEXBot backlink crawler.
- SEOkicksSEOkicks backlink crawler.1.2K hits
- Googlebot NewsGooglebot variant for Google News.113 hits
- ExaBotExalead / Dassault Systèmes search crawler.9.4K hits
- SeobilitySeobility SEO audit crawler.57 hits
- Yahoo! SlurpYahoo Search crawler (Slurp).40K hits
- MegaIndexMegaIndex SEO crawler.24K hits
- HeritrixHeritrix archival crawler (Internet Archive and others).10K hits
- TwitterbotX (Twitter) card / link-preview crawler.40K hits
- Yeti/NaverbotNaver search crawler.
- RogerbotMoz Rogerbot SEO crawler.26K hits
- serpstatbotSerpstat SEO crawler.3 hits
- AhrefsBotAhrefs SEO crawler for backlink and site research.
- LinkedInBotLinkedIn link-preview and content crawler.
- OutbrainOutbrain content-recommendation crawler.104 hits
- PetalBotHuawei Petal Search crawler.
- PinterestPinterest crawler for pin previews and discovery.1.9K hits
Monitors & scanners
- Amazon Route53 Health CheckAWS Route 53 health-check probes from Amazon's monitoring fleet.2.3M hits
- Nagios check_httpNagios / Icinga HTTP check plugin used for uptime monitoring.349K hits
- masscanHigh-speed Internet port scanner (often seen as noisy reconnaissance).192K hits
- UptimeRobotUptime monitoring service that periodically hits configured URLs.116K hits
- WebPageTestWebPageTest synthetic performance measurement agent.26K hits
- RuxitSyntheticDynatrace (Ruxit) synthetic monitoring agent.24K hits
- zgrabZMap application-layer scanner used in Internet-wide surveys.21K hits
- YandexMetrikaYandex Metrica analytics / availability checker.20K hits
- Notify NinjaUptime / change-notification monitoring bot.14K hits
- SiteimproveSiteimprove accessibility and quality-assurance crawler.6.6K hits
- ComscorecomScore audience-measurement crawler.5.2K hits
- WebMoney AdvisorWebMoney Advisor security / reputation checker.4.3K hits
- Google-Site-VerificationGoogle Search Console site-verification fetch.4.1K hits
- IPS AgentVerisign / IPS agent used in Internet surveys.3.8K hits
- NmapNmap network scanner HTTP probes.3.4K hits
- LighthouseGoogle Lighthouse performance audit agent.3.3K hits
- Screaming Frog SEO SpiderScreaming Frog desktop SEO spider (site audits).3.2K hits
- GTmetrixGTmetrix performance testing agent.400 hits
- StatusCakeStatusCake uptime monitoring agent.78 hits
- SitebulbSitebulb desktop SEO auditor.54 hits
- Site24x7 Website MonitoringSite24x7 monitoring agent.
HTTP clients
- httplib2Python httplib2 HTTP client library.
- Windows CryptoAPIWindows CryptoAPI HTTP client (often certificate / update traffic).
- Windows Push Notification ServicesWindows Push Notification Services (WNS) client traffic.
- Windows Delivery OptimizationWindows Delivery Optimization peer-to-peer update client.
- Node Fetchnode-fetch HTTP client for Node.js.
- WordPressWordPress core / plugin HTTP client strings (updates, pingbacks, oEmbed).63K hits
- OkHttpAn HTTP & HTTP/2 client for Android and Java applications.
- Python RequestsPython Requests HTTP library default user-agent.
- Go httpPackage http provides HTTP client and server implementations. Get, Head, Post, and PostForm make HTTP (or HTTPS) requests.
- EmbedlyEmbedly oEmbed / link-unfurl service.8.8K hits
- SynapseMatrix Synapse federation / HTTP client.6.0K hits
- curlcurl command-line HTTP client.
- Quora Link PreviewQuora link-preview fetcher.1.9K hits
- JavaGeneric Java HTTP client user-agent.
- BitlyBotBitly link-preview and metadata fetcher.1.2K hits
- Python urllibPython urllib HTTP client.
- IframelyIframely link-preview / oEmbed service.801 hits
- TelegramBotTelegram link-preview crawler.743 hits
- ScrapyScrapy Python crawling framework default UA.6.5K hits
- RubyRuby HTTP client user-agent.
- Guzzle (PHP HTTP Client)Guzzle PHP HTTP client.
- PerlPerl HTTP client user-agent.
- WgetGNU Wget file-retrieval client.
- SlackbotSlack link-unfurl and integration HTTP client.41K hits
- inoreaderInoreader feed fetcher.24 hits
- FeedlyFeedly RSS / Atom fetcher.10K hits
- httpxPython HTTPX client.7 hits
- HTTPieHTTPie command-line HTTP client.
- Microsoft BITSBackground Intelligent Transfer Service (BITS) to transfer files asynchronously between a client and a server.
- DiscordbotDiscord link-embed crawler.
- Google Sitemap GenOpen-source XML creation in Python. Script can be freely downloaded.
- NewsBlurNewsBlur feed fetcher.918 hits
- WhatsAppWhatsApp link-preview fetcher.
Other
149 bots with real write-ups from our archive. Useragent API