FacebookExternalHit
OthersVerify FacebookExternalHit IP Address
Verify if an IP address truly belongs to Meta / Facebook, using official verification methods. Enter both IP address and User-Agent from your logs for the most accurate bot verification.
FacebookExternalHit is Facebook’s (Meta’s) crawler used to fetch webpage content for link previews across Facebook, Messenger, Instagram, and other Meta surfaces. It retrieves metadata such as Open Graph tags, titles, descriptions, images, and structured data. These requests are user-triggered, occurring when someone shares or pastes a URL on a Meta platform. The bot does not index or rank websites and has no connection to search algorithms. Blocking it may prevent accurate link previews. Crawl activity is lightweight and focused on fetching just enough content to generate rich social previews. It ignores the global user agent (*) rule. RobotSense.io verifies FacebookExternalHit using Meta’s official validation methods, ensuring only genuine FacebookExternalHit traffic is identified.
User Agent Examples
Contains: facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)
Contains: facebookexternalhit/1.1
Contains: facebookcatalog/1.0Robots.txt Configuration for FacebookExternalHit
FacebookExternalHitUse this identifier in your robots.txt User-agent directive to target FacebookExternalHit.
Recommended Configuration
Our recommended robots.txt configuration for FacebookExternalHit:
User-agent: FacebookExternalHit
Allow: /Completely Block FacebookExternalHit
Prevent this bot from crawling your entire site:
User-agent: FacebookExternalHit
Disallow: /Completely Allow FacebookExternalHit
Allow this bot to crawl your entire site:
User-agent: FacebookExternalHit
Allow: /Block Specific Paths
Block this bot from specific directories or pages:
User-agent: FacebookExternalHit
Disallow: /private/
Disallow: /admin/
Disallow: /api/Allow Only Specific Paths
Block everything but allow specific directories:
User-agent: FacebookExternalHit
Disallow: /
Allow: /public/
Allow: /blog/Set Crawl Delay
Limit how frequently FacebookExternalHit can request pages (in seconds):
User-agent: FacebookExternalHit
Allow: /
Crawl-delay: 10Note: This bot does not officially mention about honoring Crawl-Delay rule.
Put these rules to work
Frequently Asked Questions
- What is FacebookExternalHit, and why is it visiting my website?
- FacebookExternalHit is Meta's crawler for generating link previews across Facebook, Messenger, Instagram, and other Meta products. It retrieves webpage metadata such as Open Graph tags, titles, descriptions, images, and structured data so that shared links can display rich previews. Requests are typically triggered when a user shares, posts, or pastes a URL on a Meta platform, or when Meta refreshes preview information. For publicly accessible websites that are shared on Meta services, this bot traffic is expected and usually low in volume.
- Is FacebookExternalHit a legitimate bot, or is it commonly spoofed?
- FacebookExternalHit is a legitimate crawler operated by Meta. Like many well-known bots, its User-Agent string can be spoofed by scrapers, scanners, or malicious actors attempting to disguise automated traffic. Attackers may impersonate FacebookExternalHit because websites sometimes allow social preview crawlers broader access than unknown bots. User-Agent strings alone cannot verify authenticity and should not be relied upon as the sole verification method. You can use Meta's recommended methods mentioned below to verify a legitimate visit, or use RobotSense.io API to easily verify FacebookExternalHit bot visits.
- How can I verify that a request is really coming from FacebookExternalHit?
- You can use Meta's recommended official methods to verify FacebookExternalHit bot visits, these include: - IP range checks Do not use User-Agent based detection as that can be easily spoofed. Alternatively, you can use RobotSense.io API to easily verify FacebookExternalHit bot and all other bots from Meta.
- Should I allow or block FacebookExternalHit on my website?
- For most public websites, allowing FacebookExternalHit is beneficial because it enables accurate link previews when content is shared on Meta platforms. Rich previews often improve how shared content appears to users by displaying the correct title, description, and image. Blocking may be appropriate when: - Content is private or restricted. - Internal applications or APIs should not be previewed. - Sensitive resources should not be fetched by external services. - Access policies require strict crawler restrictions. For websites that depend on social sharing, allowing the crawler is generally recommended.
- How can I control or block FacebookExternalHit using robots.txt or other methods?
- You can add a rule in your robots.txt, as given above to control (crawl-delay) or disallow FacebookExternalHit bot. The FacebookExternalHit bot honors it's own specific robots.txt directives, but does not honor global directives. Also, you can use further controls in your WAF, or in RobotSense enforcement settings to manage the bot behavior.
- How often does FacebookExternalHit crawl websites, and can it impact server performance?
- FacebookExternalHit primarily performs event-driven fetching. Requests are commonly generated when users share URLs, when previews need to be generated, or when metadata is refreshed for previously shared content. For most websites, the impact is minimal because: - Requests are targeted to specific URLs. - Bandwidth usage is generally low. - Crawl activity is not site-wide. High-traffic websites with frequently shared content may see more requests, but the overall load is typically much lower than that of search engine crawlers.
- What happens if I block FacebookExternalHit? SEO, visibility, and feature impact explained.
- Blocking FacebookExternalHit does not affect search engine rankings because it is not a search indexing crawler. Potential impacts include: - Missing or incomplete link previews on Facebook, Messenger, Instagram, and related Meta products. - Missing preview images, titles, or descriptions. - Reduced quality of shared content presentation. - Metadata refreshes may fail or become outdated. Blocking does not directly affect: - Google indexing. - Bing indexing. - Organic search rankings. - Traditional SEO databases.
- Does FacebookExternalHit collect, scrape, or use my content for training or reuse?
- FacebookExternalHit retrieves webpage content and metadata needed to generate and maintain link previews. This commonly includes page titles, descriptions, Open Graph tags, structured data, and preview images. Documented uses include: - Link preview generation. - Metadata extraction. - Preview validation and refreshes. - Content enrichment for shared URLs. There is no public documentation indicating that FacebookExternalHit is used for general web indexing, SEO datasets, or AI model training. Its documented purpose is focused on social sharing and preview generation rather than large-scale content collection.
Other Meta / Facebook Bots
Meta / Facebook operates other crawlers you may also need to configure.
Meta-ExternalAds
AdsMeta-ExternalAds is Meta’s crawler used to evaluate landing pages associated with ads running on Facebook, Instagram, and other Meta platforms. It performs targeted checks to assess page load behavior, policy compliance, redirects, content quality, and overall ad safety. These fetches are ad-driven, not general web crawling, and help Meta determine whether landing pages meet advertising standards. Blocking it may affect ad review accuracy or eligibility. Crawl activity is focused, low-volume, and typically triggered when advertisers submit new ads, update creatives, or undergo automated policy reviews. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-ExternalAds using Meta’s official validation methods, ensuring only genuine Meta-ExternalAds traffic is identified.
Meta-ExternalAgent
AI TrainingMeta-ExternalAgent is a Meta crawler used to fetch webpage content for AI, integrity, and content understanding systems that operate outside classic social preview or ads workflows. It performs broader content retrieval to support tasks like classification, safety analysis, and model training. This traffic is not user-triggered and is separate from Meta’s ad review or link preview bots. Crawl activity is moderate and targeted toward pages relevant to Meta’s internal systems. It does not affect search rankings, as Meta has no public web search engine. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-ExternalAgent using Meta’s official validation methods, ensuring only genuine Meta-ExternalAgent traffic is identified.
Meta-ExternalFetcher
OthersMeta-ExternalFetcher is a Meta crawler that retrieves webpage content to support link previews, metadata extraction, and other external content processing tasks across Facebook, Instagram, and related Meta products. It fetches titles, descriptions, images, and structured data required for rendering shared links or enriching user interactions. These requests are typically user-driven but may also support automated metadata refreshes. Crawl volume is lightweight and focused, targeting only the URLs needed for previews or content enrichment within Meta’s ecosystem. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-ExternalFetcher using Meta’s official validation methods, ensuring only genuine Meta-ExternalFetcher traffic is identified.
Meta-WebIndexer
AI TrainingMeta-WebIndexer is Meta’s web crawler used to discover and fetch publicly available webpage content for internal indexing, AI research, and content understanding tasks. It performs broader, more systematic crawling than Facebook’s preview-focused bots. The crawler analyzes text, metadata, and structured elements to improve Meta’s machine learning models and content classification systems. Crawl activity ranges from moderate to wide-reaching depending on Meta’s data needs. Meta-WebIndexer does not influence external search rankings, as Meta does not operate a web search engine. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-WebIndexer using Meta’s official validation methods, ensuring only genuine Meta-WebIndexer traffic is identified.
Similar Bots
Other Others bots from different operators.
Amazonbot
Othersby Amazon
[Amazon Bots can take upto 30 days to read your Robots.txt updates.] Amazonbot is Amazon’s official web crawler, used to discover and fetch webpage content for applications such as Alexa, product-related features, and Amazon’s AI and search systems. Crawl activity varies based on Amazon services that rely on external web content, but it is generally moderate and focused on structured data, text content, and page metadata. Its purpose is to enhance Amazon’s search, AI models, and user-facing features. It ignores the global user agent (*) rule. RobotSense.io verifies Amazonbot using Amazon’s official validation methods, ensuring only genuine Amazonbot traffic is identified.
Amzn-User
Othersby Amazon
[Amazon Bots can take upto 30 days to read your Robots.txt updates.] Amzn-User is a bot associated with Amazon services that fetch webpage content on behalf of end users or Amazon applications rather than acting as a general-purpose crawler. It typically appears when Amazon apps, devices, or internal systems request metadata, previews, or content needed for features like link expansion, in-app browsing, or contextual analysis. The traffic is user-driven, not designed for large-scale indexing or scraping. Amzn-User usually performs lightweight, targeted fetches limited to specific URLs users interact with. Its purpose is to support Amazon product experiences by retrieving just enough page data to power user-facing functionality. It ignores the global user agent (*) rule. RobotSense.io verifies Amzn-User using Amazon’s official validation methods, ensuring only genuine Amzn-User traffic is identified.
BingPreview
Othersby Microsoft
BingPreview is Microsoft’s rendering and compatibility crawler used to evaluate how webpages appear in browsers and Bing search features. It fetches pages to test layout, mobile responsiveness, JavaScript rendering, and visual elements. These checks help Bing understand how content will display in search results and improve snippet generation and ranking signals tied to user experience. Crawl activity is moderate and often concentrated on pages important to Bing’s index. Its purpose is to simulate real-browser behavior and refine Bing’s presentation quality. RobotSense.io verifies BingPreview using Microsoft’s official validation methods, ensuring only genuine BingPreview traffic is identified.
DuplexWeb-Google
Othersby Google
[This crawler is officially retired as per Google] DuplexWeb-Google is a Google crawler associated with Duplex and Assistant-related technologies that fetch web content to help generate conversational responses and perform task-oriented actions. It retrieves page information needed to understand structured data, business details, menus, appointment flows, and other interactive elements. Crawl activity is selective and generally tied to user-initiated tasks or systems that prepare content for automated assistance. Its purpose is to support natural-language interactions by ensuring Google’s assistant technologies can interpret and use real-time webpage information accurately. It ignores the global user agent (*) rule. RobotSense.io verifies DuplexWeb-Google using Google’s official validation methods, ensuring only genuine DuplexWeb-Google traffic is identified.
Google Messages
Othersby Google
Google Messages uses a fetcher to retrieve webpage data for generating link previews when users send URLs in chat conversations. It fetches metadata such as titles, descriptions, images, and structured tags (e.g., Open Graph) to render rich previews inside messages. This traffic is strictly user-triggered, occurring only when a link is shared. It is not a crawler for indexing or discovery and has no impact on Google Search rankings. Blocking it may prevent previews from displaying correctly. Activity is lightweight and targeted, focused solely on enhancing the messaging experience with accurate link previews. Google Messages bot does not respect robots.txt rules. RobotSense.io verifies Google Messages using Google’s official validation methods, ensuring only genuine Google Messages traffic is identified.
GoogleOther
Othersby Google
GoogleOther is a general-purpose crawler used by Google for internal research, large-scale data analysis, and non–Search-related fetching. It is part of Google’s secondary crawling infrastructure, designed to offload tasks that don’t require the full capabilities or strict policies of Googlebot. GoogleOther typically performs broad but lower-priority fetches, such as machine learning dataset generation or internal experiments. Its activity is generally lightweight compared to Googlebot and is separate from indexing operations that directly influence Google Search results. RobotSense.io verifies GoogleOther using Google’s official validation methods, ensuring only genuine GoogleOther traffic is identified.