Web Light / googleweblight
Developer ToolsVerify Web Light / googleweblight IP Address
Verify if an IP address truly belongs to Google, using official verification methods. Enter both IP address and User-Agent from your logs for the most accurate bot verification.
[This crawler is officially retired as per Google] googleweblight is a Google fetcher used by the now-deprecated Google Web Light service, which provided simplified, faster-loading versions of webpages for slow mobile networks. It requested pages to generate lightweight, transcoded versions optimized for low-bandwidth conditions. Site owners allowed it to ensure better accessibility for users on slow connections. Since Web Light has been discontinued, activity from this user-agent is now rare or legacy in nature. Any remaining traffic is typically minimal and related to leftover systems or outdated client requests rather than active Google services. It ignored the global user agent (*) rule. RobotSense.io verifies Web Light / googleweblight using Google’s official validation methods, ensuring only genuine Web Light / googleweblight traffic is identified.
User Agent Examples
Mozilla/5.0 (Linux; Android 4.2.1; en-us; Nexus 5 Build/JOP40D) AppleWebKit/535.19 (KHTML, like Gecko; googleweblight) Chrome/38.0.1025.166 Mobile Safari/535.19Robots.txt Configuration for Web Light / googleweblight
googleweblightUse this identifier in your robots.txt User-agent directive to target Web Light / googleweblight.
Recommended Configuration
Our recommended robots.txt configuration for Web Light / googleweblight:
# This bot is officially retired by Google
User-agent: googleweblight
Allow: /Completely Block Web Light / googleweblight
Prevent this bot from crawling your entire site:
User-agent: googleweblight
Disallow: /Completely Allow Web Light / googleweblight
Allow this bot to crawl your entire site:
User-agent: googleweblight
Allow: /Block Specific Paths
Block this bot from specific directories or pages:
User-agent: googleweblight
Disallow: /private/
Disallow: /admin/
Disallow: /api/Allow Only Specific Paths
Block everything but allow specific directories:
User-agent: googleweblight
Disallow: /
Allow: /public/
Allow: /blog/Set Crawl Delay
Limit how frequently Web Light / googleweblight can request pages (in seconds):
User-agent: googleweblight
Allow: /
Crawl-delay: 10Note: This bot does not officially mention about honoring Crawl-Delay rule.
Put these rules to work
Frequently Asked Questions
- What is googleweblight, and why is it visiting my website?
- googleweblight was a crawler operated by Google for the now-retired Google Web Light service. The service fetched webpages and generated lightweight, optimized versions designed for users on slow mobile networks or low-bandwidth connections. Historically, visits were triggered when users accessed webpages through Web Light-enabled experiences on mobile devices. Since the service has been discontinued, legitimate crawl activity from this bot is now rare and typically associated with legacy systems, cached requests, or outdated client behavior. Public websites may still occasionally see this User-Agent in server logs, but sustained activity is uncommon.
- Is googleweblight a legitimate bot, or is it commonly spoofed?
- googleweblight was an official Google fetcher when the Web Light service was active. However, because the service has been retired, any modern traffic using this User-Agent deserves closer scrutiny. User-Agent spoofing is common across well-known Google crawler names because attackers and scrapers sometimes imitate trusted bots to bypass rate limits, security filters, or bot protection systems. User-Agent strings alone cannot verify authenticity, since any automated client can send the same header. You can use Google's recommended methods mentioned below to verify a legitimate visit, or use RobotSense.io API to easily verify googleweblight visits.
- How can I verify that a request is really coming from googleweblight?
- You can use Google's recommended official methods to verify googleweblight visits, these include: - IP range checks - Reverse DNS → forward DNS Do not use User-Agent based detection as that can be easily spoofed. Alternatively, you can use RobotSense.io API to easily verify googleweblight and all other bots from Google.
- Should I allow or block googleweblight on my website?
- Because Google Web Light has been discontinued, allowing googleweblight is generally optional for most websites today. Blocking it typically has little or no impact on modern search visibility or user experience. Blocking may make sense when: - The traffic appears suspicious or excessive - Requests target sensitive endpoints or APIs - You want to reduce unnecessary bot traffic - Server resources are limited If verified Google-origin traffic is minimal and non-disruptive, allowing it is usually harmless, but most sites no longer depend on this crawler.
- How can I control or block googleweblight using robots.txt or other methods?
- You can add a rule in your robots.txt, as given above to control (crawl-delay) or disallow googleweblight. googleweblight honors robots.txt directives. Also, you can use further controls in your WAF, or in RobotSense enforcement settings to manage the bot behavior.
- How often does googleweblight crawl websites, and can it impact server performance?
- When the Web Light service was active, crawling was primarily request-driven and tied to mobile users accessing optimized pages. The bot fetched content as needed to generate lightweight versions for low-bandwidth environments. Today, crawl activity is typically very limited. Any remaining traffic usually has minimal impact on bandwidth, request rates, or dynamic page generation. On most websites, modern server performance effects from legitimate googleweblight traffic are negligible.
- What happens if I block googleweblight? SEO, visibility, and feature impact explained.
- Blocking googleweblight generally has little to no effect on modern SEO or Google Search rankings because the underlying Web Light service has been retired. Most websites no longer rely on this crawler for accessibility or content delivery features. Potential impacts are limited and mostly historical: - No significant impact on Google Search indexing - No effect on standard Googlebot crawling - No meaningful impact on modern mobile search visibility - Legacy low-bandwidth page transcoding would no longer function - Minimal or no effect on analytics or developer tools For most sites, blocking this crawler has no noticeable operational consequence today.
- Does googleweblight collect, scrape, or use my content for training or reuse?
- googleweblight fetched webpage content in order to generate simplified, compressed versions of pages for low-bandwidth mobile users. This included retrieving HTML, images, CSS, and other resources necessary to create optimized page renderings. Its documented purpose was content transcoding and mobile delivery optimization rather than AI training or large-scale content reuse. The crawler could temporarily process or cache page content and metadata to serve lightweight versions, but there is no documented indication that googleweblight was specifically used for machine learning training pipelines. The bot historically processed: - Full page HTML - Metadata and structured content - Images and page resources - Mobile rendering information - Performance optimization data Because the service has been retired, active collection activity from legitimate googleweblight infrastructure is now extremely limited.
Other Google Bots
Google operates other crawlers you may also need to configure.
AdsBot
AdsAdsBot-Google is Google’s crawler responsible for evaluating landing pages used in Google Ads campaigns. It performs desktop-focused checks on page quality, load speed, relevance, and policy compliance. These assessments directly influence ad quality scores, cost efficiency, and overall eligibility. Blocking AdsBot prevents Google from reviewing landing pages, which can degrade or disable ad performance. Crawl activity is selective and tied to active or recently modified ad campaigns rather than broad indexing. Its purpose is to ensure that advertisers maintain fast, trustworthy, and policy-compliant landing pages. It ignores the global user agent (*) rule. RobotSense.io verifies AdsBot/AdsBot-Google using Google’s official validation methods, ensuring only genuine AdsBot/AdsBot-Google traffic is identified.
AdsBot Mobile Web
Ads[This crawler is officially retired as per Google] AdsBot-Google-Mobile bot is Google’s mobile-focused crawler used to evaluate the landing page experience for Google Ads. It simulates mobile device conditions to assess page quality, load performance, mobile usability, and policy compliance. These evaluations directly influence Google Ads quality scores and ad eligibility. If you run ads, blocking it may negatively affect ad performance because Google cannot verify the mobile landing page experience. Crawl activity is targeted and low-volume, triggered when ads are created, updated, or actively running. Its purpose is ensuring advertisers provide fast, compliant, and user-friendly mobile pages. It ignores the global user agent (*) rule. RobotSense.io verifies AdsBot Mobile Web/AdsBot-Google-Mobile using Google’s official validation methods, ensuring only genuine AdsBot Mobile Web/AdsBot-Google-Mobile traffic is identified.
AdSense / Mediapartners-Google
AdsMediapartners-Google is Google’s crawler dedicated to evaluating webpages for Google AdSense. It scans pages to understand content, layout, and context so Google can deliver relevant ads and optimize revenue for publishers. Unlike Googlebot, this crawler does not index content for Search - its role is purely advertising-related. Blocking it may prevent AdSense from analyzing pages and serving targeted ads effectively. Crawl activity is generally light and focused on pages where AdSense code is present, helping Google match ad inventory with page themes and user interests. It ignores the global user agent (*) rule. RobotSense.io verifies AdSense / Mediapartners-Google using Google’s official validation methods, ensuring only genuine AdSense / Mediapartners-Google traffic is identified.
APIs-Google
Developer ToolsAPIs-Google is a service crawler used by Google to verify and interact with endpoints tied to various Google APIs. It is typically triggered when applications, scripts, or integrations using Google services need to fetch or validate external resources. Common use cases include OAuth flows, link previews, data validation, push notification messages, and API-driven checks performed on behalf of Google products. Crawl activity is usually low-volume and event-driven, reflecting specific API operations rather than broad crawling or indexing associated with Google Search. It ignores the global user agent (*) rule. RobotSense.io verifies APIs-Google using Google’s official validation methods, ensuring only genuine APIs-Google traffic is identified.
Chrome Web Store / Google-CWS
Developer ToolsChrome Web Store fetcher or Google-CWS is a Google user-agent associated with Chrome Web Services, typically used for link preview generation, safe browsing checks, and content fetching triggered by Chrome features. It performs lightweight requests to retrieve metadata, page titles, favicons, and safety signals from URLs that developers provide in the metadata of their Chrome extensions and themes. This bot is not a search crawler and does not influence Google Search indexing or rankings. Activity occurs when Chrome or Google services need to quickly inspect a URL for previews, safety evaluation, or rendering behavior. It ignores robots.txt rules. RobotSense.io verifies Chrome Web Store fetcher using Google’s official validation methods, ensuring only genuine Chrome Web Store fetcher traffic is identified.
DuplexWeb-Google
Others[This crawler is officially retired as per Google] DuplexWeb-Google is a Google crawler associated with Duplex and Assistant-related technologies that fetch web content to help generate conversational responses and perform task-oriented actions. It retrieves page information needed to understand structured data, business details, menus, appointment flows, and other interactive elements. Crawl activity is selective and generally tied to user-initiated tasks or systems that prepare content for automated assistance. Its purpose is to support natural-language interactions by ensuring Google’s assistant technologies can interpret and use real-time webpage information accurately. It ignores the global user agent (*) rule. RobotSense.io verifies DuplexWeb-Google using Google’s official validation methods, ensuring only genuine DuplexWeb-Google traffic is identified.