G

Google-Extended

Visit Bot Homepage

Verify Google-Extended IP Address

Verify if an IP address truly belongs to Google, using official verification methods. Enter both IP address and User-Agent from your logs for the most accurate bot verification.

Google-Extended is a special user-agent that allows website owners to control whether their publicly accessible content can be used to train and improve Google’s AI models, including products like Gemini. It does not crawl the web itself; instead, it serves as a policy signal interpreted by Google’s AI systems. Site owners can allow or block AI training access by configuring robots.txt rules for Google-Extended. Blocking this agent does not affect Google Search ranking, crawling, or indexing. Its purpose is purely governance-giving publishers a transparent way to manage how their content contributes to Google’s AI research and model development.

This bot does not honor Crawl-Delay rule.

User Agent Examples

Google does not identify any specific user-agent pattern in HTTP requests for Google-Extended bot!
Example user agent strings for Google-Extended

Robots.txt Configuration for Google-Extended

Robots.txt User-Agent:Google-Extended

Use this identifier in your robots.txt User-agent directive to target Google-Extended.

Recommended Configuration

Our recommended robots.txt configuration for Google-Extended:

User-agent: Google-Extended
Allow: /

Completely Block Google-Extended

Prevent this bot from crawling your entire site:

User-agent: Google-Extended
Disallow: /

Completely Allow Google-Extended

Allow this bot to crawl your entire site:

User-agent: Google-Extended
Allow: /

Block Specific Paths

Block this bot from specific directories or pages:

User-agent: Google-Extended
Disallow: /private/
Disallow: /admin/
Disallow: /api/

Allow Only Specific Paths

Block everything but allow specific directories:

User-agent: Google-Extended
Disallow: /
Allow: /public/
Allow: /blog/

Set Crawl Delay

Limit how frequently Google-Extended can request pages (in seconds):

User-agent: Google-Extended
Allow: /
Crawl-delay: 10

Note: This bot does not officially mention about honoring Crawl-Delay rule.

Frequently Asked Questions

What is Google-Extended?
Google-Extended is a control-oriented user-agent defined by Google that allows website owners to manage whether their content can be used for AI model training. It does not function as a traditional crawler and does not actively send requests to websites. Instead, it acts as a policy signal interpreted when Google systems process robots.txt rules. You will not typically see Google-Extended generating traffic in website logs.
Is Google-Extended a legitimate bot, or is it commonly spoofed?
Google-Extended is an official Google-defined user-agent, but it is not a crawling bot and therefore not typically seen in server requests. As a result, spoofing is less relevant compared to active crawlers, though malicious actors could still misuse the name in headers. Since it does not generate traffic, any requests claiming to be Google-Extended should be treated with skepticism. As always, user-agent strings alone cannot verify authenticity.
How can I verify that a request is really coming from Google-Extended?
There is generally nothing to verify because Google-Extended does not make HTTP requests to websites. If you see traffic using this user-agent in logs, it is likely not legitimate. Standard verification methods (reverse DNS, forward DNS, IP range checks) apply only to actual Google crawlers. In this case, user-agent-based detection is not applicable because the agent functions as a robots.txt policy identifier, not a network client.
Should I allow or block Google-Extended on my website?
Allowing or blocking Google-Extended is a policy decision rather than a traffic control decision. It determines whether your publicly accessible content may be used for Google’s AI model training. Blocking may be appropriate if: - You do not want your content used in AI training datasets - Your content is proprietary or sensitive - You require stricter content usage controls Allowing it enables participation in Google’s AI ecosystem but is entirely optional.
How can I control or block Google-Extended using robots.txt or other methods?
Google-Extended is controlled exclusively via robots.txt directives. You can add a rule in your robots.txt, as given above to signal no AI training usage of your content. No WAF, rate limiting, or IP blocking is needed since it does not generate server requests.
How often does Google-Extended crawl websites, and can it impact server performance?
Google-Extended does not crawl websites and does not generate HTTP requests. It has no crawl frequency, bandwidth usage, or request rate. As a result, it has zero impact on server performance. Any perceived traffic under this name is not from the legitimate Google-Extended agent.
What happens if I block Google-Extended? SEO, visibility, and feature impact explained.
Blocking Google-Extended does not affect search indexing, rankings, or standard crawling by Google Search. Its impact is limited to AI-related usage i.e., your content will not be used for training or improving Google AI models. This is strictly a content usage control, not an SEO setting.
Does Google-Extended collect, scrape, or use my content for training or reuse?
Google-Extended itself does not collect or scrape content. Instead, it defines whether Google systems are permitted to use publicly accessible content for AI training and model improvement. If allowed, content may be included in training datasets; if blocked, it is excluded. It does not store or process content independently, and it does not function as a crawler or indexing system.

Other Google Bots

Google operates other crawlers you may also need to configure.

View all 29 Google bots →

AdsBot

Ads

AdsBot-Google is Google’s crawler responsible for evaluating landing pages used in Google Ads campaigns. It performs desktop-focused checks on page quality, load speed, relevance, and policy compliance. These assessments directly influence ad quality scores, cost efficiency, and overall eligibility. Blocking AdsBot prevents Google from reviewing landing pages, which can degrade or disable ad performance. Crawl activity is selective and tied to active or recently modified ad campaigns rather than broad indexing. Its purpose is to ensure that advertisers maintain fast, trustworthy, and policy-compliant landing pages. It ignores the global user agent (*) rule. RobotSense.io verifies AdsBot/AdsBot-Google using Google’s official validation methods, ensuring only genuine AdsBot/AdsBot-Google traffic is identified.

AdsBot Mobile Web

Ads

[This crawler is officially retired as per Google] AdsBot-Google-Mobile bot is Google’s mobile-focused crawler used to evaluate the landing page experience for Google Ads. It simulates mobile device conditions to assess page quality, load performance, mobile usability, and policy compliance. These evaluations directly influence Google Ads quality scores and ad eligibility. If you run ads, blocking it may negatively affect ad performance because Google cannot verify the mobile landing page experience. Crawl activity is targeted and low-volume, triggered when ads are created, updated, or actively running. Its purpose is ensuring advertisers provide fast, compliant, and user-friendly mobile pages. It ignores the global user agent (*) rule. RobotSense.io verifies AdsBot Mobile Web/AdsBot-Google-Mobile using Google’s official validation methods, ensuring only genuine AdsBot Mobile Web/AdsBot-Google-Mobile traffic is identified.

AdSense / Mediapartners-Google

Ads

Mediapartners-Google is Google’s crawler dedicated to evaluating webpages for Google AdSense. It scans pages to understand content, layout, and context so Google can deliver relevant ads and optimize revenue for publishers. Unlike Googlebot, this crawler does not index content for Search - its role is purely advertising-related. Blocking it may prevent AdSense from analyzing pages and serving targeted ads effectively. Crawl activity is generally light and focused on pages where AdSense code is present, helping Google match ad inventory with page themes and user interests. It ignores the global user agent (*) rule. RobotSense.io verifies AdSense / Mediapartners-Google using Google’s official validation methods, ensuring only genuine AdSense / Mediapartners-Google traffic is identified.

APIs-Google

Developer Tools

APIs-Google is a service crawler used by Google to verify and interact with endpoints tied to various Google APIs. It is typically triggered when applications, scripts, or integrations using Google services need to fetch or validate external resources. Common use cases include OAuth flows, link previews, data validation, push notification messages, and API-driven checks performed on behalf of Google products. Crawl activity is usually low-volume and event-driven, reflecting specific API operations rather than broad crawling or indexing associated with Google Search. It ignores the global user agent (*) rule. RobotSense.io verifies APIs-Google using Google’s official validation methods, ensuring only genuine APIs-Google traffic is identified.

Chrome Web Store / Google-CWS

Developer Tools

Chrome Web Store fetcher or Google-CWS is a Google user-agent associated with Chrome Web Services, typically used for link preview generation, safe browsing checks, and content fetching triggered by Chrome features. It performs lightweight requests to retrieve metadata, page titles, favicons, and safety signals from URLs that developers provide in the metadata of their Chrome extensions and themes. This bot is not a search crawler and does not influence Google Search indexing or rankings. Activity occurs when Chrome or Google services need to quickly inspect a URL for previews, safety evaluation, or rendering behavior. It ignores robots.txt rules. RobotSense.io verifies Chrome Web Store fetcher using Google’s official validation methods, ensuring only genuine Chrome Web Store fetcher traffic is identified.

DuplexWeb-Google

Others

[This crawler is officially retired as per Google] DuplexWeb-Google is a Google crawler associated with Duplex and Assistant-related technologies that fetch web content to help generate conversational responses and perform task-oriented actions. It retrieves page information needed to understand structured data, business details, menus, appointment flows, and other interactive elements. Crawl activity is selective and generally tied to user-initiated tasks or systems that prepare content for automated assistance. Its purpose is to support natural-language interactions by ensuring Google’s assistant technologies can interpret and use real-time webpage information accurately. It ignores the global user agent (*) rule. RobotSense.io verifies DuplexWeb-Google using Google’s official validation methods, ensuring only genuine DuplexWeb-Google traffic is identified.

Similar Bots

Other AI Training bots from different operators.

GPTBot

AI Training

by OpenAI

GPTBot is an AI Training bot operated by OpenAI. GPTBot is used by OpenAI to make their generative AI foundation models more useful and safe. It is used to crawl content that may be used in training their generative AI foundation models. Disallowing GPTBot indicates that a site’s content should not be used in training generative AI foundation models. RobotSense.io verifies OpenAI GPTBot using OpenAI’s official validation methods, ensuring only genuine GPTBot traffic is identified.

Meta-ExternalAgent

AI Training

by Meta / Facebook

Meta-ExternalAgent is a Meta crawler used to fetch webpage content for AI, integrity, and content understanding systems that operate outside classic social preview or ads workflows. It performs broader content retrieval to support tasks like classification, safety analysis, and model training. This traffic is not user-triggered and is separate from Meta’s ad review or link preview bots. Crawl activity is moderate and targeted toward pages relevant to Meta’s internal systems. It does not affect search rankings, as Meta has no public web search engine. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-ExternalAgent using Meta’s official validation methods, ensuring only genuine Meta-ExternalAgent traffic is identified.

Meta-WebIndexer

AI Training

by Meta / Facebook

Meta-WebIndexer is Meta’s web crawler used to discover and fetch publicly available webpage content for internal indexing, AI research, and content understanding tasks. It performs broader, more systematic crawling than Facebook’s preview-focused bots. The crawler analyzes text, metadata, and structured elements to improve Meta’s machine learning models and content classification systems. Crawl activity ranges from moderate to wide-reaching depending on Meta’s data needs. Meta-WebIndexer does not influence external search rankings, as Meta does not operate a web search engine. It ignores the global user agent (*) rule. RobotSense.io verifies Meta-WebIndexer using Meta’s official validation methods, ensuring only genuine Meta-WebIndexer traffic is identified.