HomeGuidesE-commerce SEO

    What Is Crawl Budget?

    Crawl budget is the number of pages Googlebot will crawl on your site within a given period, and it is a finite resource that gets consumed by every URL the crawler encounters. For large e-commerce sites with thousands of products and faceted filters, uncontrolled URL generation can exhaust the crawl budget before Googlebot reaches your most important pages, effectively removing them from the index.

    Tharindu Gunawardana
    Tharindu Gunawardana
    March 22, 2026
    10 min read read
    E-commerce SEO
    What Is Crawl Budget?

    What Is Crawl Budget?

    Crawl budget is the number of URLs Googlebot will crawl on your website within a given period of time. Google does not crawl every URL it discovers on every visit. Instead, it allocates a finite crawling capacity to each site based on the site's authority and its server's ability to handle crawler requests. That capacity is divided across all URLs Googlebot finds on the site.

    For small websites with a few hundred pages, crawl budget is rarely a constraint. Googlebot will visit all pages frequently. For large e-commerce stores with tens of thousands of product pages plus uncontrolled faceted navigation generating millions of filter URLs, crawl budget becomes a critical bottleneck. If Googlebot exhausts its allocated budget on thin filter pages, your newest products may go unindexed for days or weeks.

    Google published guidance on crawl budget in 2017 and has updated it since. The core principle is unchanged: sites should make it easy for Googlebot to find and crawl their most important content by reducing the number of low-value URLs Googlebot has to wade through.

    The Two Factors That Set Your Crawl Budget

    Crawl budget is the product of two separate signals: crawl rate limit and crawl demand. These operate independently, and improving one without addressing the other produces limited gains.

    What Determines Your Crawl BudgetCrawl Rate LimitHow fast Googlebot can crawlwithout overloading your servers.Controlled by:Server response timeTime between requests (crawl delay)Server capacity and stability+Crawl DemandHow much Google wants to crawlyour site based on its signals.Influenced by:Site authority and backlink profileURL popularity (internal + external links)Content freshness and change frequency= Your Allocated Crawl BudgetShared across ALL URLs Googlebot can discover on your site

    Crawl rate limit

    Crawl rate limit is how fast Googlebot crawls without overloading your server. If your server responds slowly to requests, Googlebot backs off to avoid causing errors. This means server performance is a crawl budget input. A hosting environment that responds to bot requests in under 200 milliseconds will receive significantly more crawls per day than one that takes 2 seconds per response. Google allows site owners to reduce the crawl rate manually in Search Console if server load is a concern, though this is rarely beneficial.

    Crawl demand

    Crawl demand is how much Google wants to crawl your site. Sites with strong backlink profiles, high authority, and frequently updated content receive higher crawl demand. When you publish new products or update prices, Google's systems detect that your site changes often and allocate more crawl capacity to keep its index current. Newly launched or low-authority sites receive lower crawl demand, so fewer URLs are crawled even if the server is fast.

    What Wastes Crawl Budget on E-commerce Sites

    Most e-commerce crawl budget problems are caused by URL proliferation, not server speed. The six most common sources of wasted crawl budget are below.

    Common Crawl Budget Wasters on E-commerce SitesFaceted Filter URLs?size=M&colour=blue type URLswithout canonical or noindex controlOften 80%+ of all site URLsOut-of-Stock ProductsDiscontinued SKUs returning 200status with thin contentGrows silently over timePaginated Listing PagesPage 50+ of category listingswith little unique contentEspecially past page 3-4Duplicate Title PagesMultiple URLs with identicaltitle tags and meta descriptionsSignals thin content at scaleRedirect Chains301 → 302 → 200 chainsthat slow crawl and dilute equityCommon after migrationsSoft 404 PagesEmpty search results orcategories returning 200 statusInvisible to basic auditsEvery URL Googlebot visits consumes crawl budget that could have gone to your most valuable pages.

    Faceted navigation URLs

    As covered in the Faceted Navigation guide, uncontrolled filter combinations are the single largest source of crawl budget waste on e-commerce sites. A store with 200 categories and five filter types can generate over 200,000 unique URLs, the vast majority of which are near-duplicates of root category pages.

    Out-of-stock and discontinued products

    Product detail pages for discontinued SKUs remain crawlable long after the product is removed from sale. If these pages return 200 status codes with thin "product unavailable" content, Googlebot continues to visit them on each crawl cycle. For large catalogues that cycle through seasonal inventory, this accumulates into hundreds of thin pages consuming regular crawl visits.

    Soft 404 pages

    A soft 404 is a page that returns a 200 HTTP status code but contains no useful content. Empty search result pages, empty category filters, and deleted product pages that display a generic template without a proper 404 or 410 status are common examples. Google's Search Console identifies these separately from hard 404s, and they are often invisible to basic auditing tools.

    How to Prioritise Your Crawl Budget

    Crawl budget optimisation is fundamentally about signal concentration. The more low-value URLs you remove from Google's discoverable URL set, the higher the proportion of your crawl budget that goes to pages that can actually rank and drive revenue.

    Crawl Priority Hierarchy for E-commerce SitesHomepage + Core Service PagesTop Category Pages (root, no filters)Product Detail Pages (PDPs) + Blog/GuidesFilter URLs, Pagination, Out-of-Stock SKUsHighestPriorityLowestCanonicalise or block lower-tier URLs so crawl budget concentrates at the top two tiers.

    The hierarchy is straightforward: your homepage and key category pages should receive the most frequent crawl visits, followed by product detail pages, then supporting content like blog posts and guides. Filter URLs, paginated pages beyond page 2-3, and out-of-stock product pages should either be blocked from crawl or given signals that reduce their crawl priority.

    Practical steps to implement this hierarchy include: adding canonical tags to filter URLs that point to root category pages, applying noindex to out-of-stock product pages or returning 410 (Gone) status for discontinued SKUs, blocking pure parameter URLs in robots.txt, and configuring Google Search Console's URL Parameters tool for sorting and display parameters.

    Sitemaps, Robots.txt and Crawl Control

    An XML sitemap tells Google which URLs you consider most important. Googlebot prioritises sitemap-listed URLs in its crawl queue. For e-commerce sites, this means your sitemap should list root category pages and active product pages, nothing else. Including filter URLs in your sitemap actively directs crawl budget to pages you probably do not want indexed.

    The robots.txt file controls which paths Googlebot is allowed to access. It is an effective tool for blocking entire URL patterns, such as all URLs containing ?sort= or ?page= beyond page 1. Be careful with robots.txt: blocking a URL prevents crawling but does not prevent indexing if the URL is linked from an external site. For index control, noindex meta tags are more reliable.

    A common misstep is to disallow crawl of a URL in robots.txt while leaving it indexable. Google cannot see the noindex tag if it cannot crawl the page, so the page may remain indexed even with a robots.txt disallow directive. If you want a URL excluded from the index, either serve a noindex tag (and allow crawling) or return a 404/410 status.

    How Crawl Budget Affects AI Indexation

    For smaller e-commerce sites (under 10,000 pages), crawl budget is rarely limiting. Google can efficiently crawl the entire site regardless of a moderate number of filter URLs. The effort of managing crawl budget becomes worthwhile once your site exceeds roughly 10,000 indexable URLs or when you notice in Search Console that new pages are taking more than a week to appear in the index.

    The most reliable diagnostic is the Coverage report in Google Search Console. If you see large numbers of pages in the "Crawled - currently not indexed" or "Discovered - currently not indexed" categories, Googlebot is finding and queuing more pages than it is processing. Reducing the low-value URL count is the primary lever to move those pages out of the queue and into the index.

    For AI search engines, crawl budget connects to citation probability. Systems like Google's AI Overviews and Perplexity index content from pages they can access. A site with many thin, near-duplicate pages signals low content quality to both traditional crawlers and AI retrieval systems. Concentrating crawl on high-quality pages improves the signal quality of the entire site's content graph, which benefits AI citability in addition to traditional rankings.

    The SEO Audit service includes a full crawl budget analysis that maps every URL on your site, identifies the sources of crawl waste, and delivers a prioritised implementation plan. For checking individual product and category pages against technical SEO best practices, the Category Page SEO Checker and Product Page SEO Checker audit pages at the component level.

    Frequently Asked Questions

    How do I know if crawl budget is a problem for my site?

    Check Google Search Console's Coverage report for large numbers of pages in "Crawled - currently not indexed" or "Discovered - currently not indexed". Also check if new products take more than 1-2 weeks to appear in Google's index. If your site has more than 10,000 indexable URLs or uses faceted navigation without canonical controls, crawl budget is likely being wasted.

    Does crawl budget affect rankings?

    Crawl budget does not directly affect rankings. A page's ranking depends on its content quality, authority, and relevance. However, if a page is not crawled, it cannot be indexed. If it is not indexed, it cannot rank. Crawl budget constraints can therefore indirectly suppress rankings by delaying or preventing indexation of important pages.

    Should I include all my product pages in my XML sitemap?

    Yes, include all active, in-stock product pages in your sitemap. Exclude out-of-stock or discontinued products, filter URL variants, paginated pages beyond page 1, and any URL with a noindex tag. A well-maintained sitemap directly signals to Googlebot which pages deserve crawl priority.

    How quickly does reducing low-value URLs improve crawl efficiency?

    Improvements are typically visible in Google Search Console within 4-8 weeks of implementing canonical tags and parameter blocking. The Coverage report should show a reduction in "Crawled - currently not indexed" URLs and an increase in indexed pages. Large sites with hundreds of thousands of filtered URLs may see gradual improvement over several months as Google recrawls and processes the signals.

    Can I increase my crawl budget by improving site speed?

    Yes, but with limits. Faster server response times reduce crawl rate constraints, allowing Googlebot to visit more pages in the same timeframe. However, crawl demand — Google's desire to crawl your site — is primarily driven by authority and content quality, not server speed. Improving page speed alone will not significantly increase crawl budget for a low-authority site.

    What should I do with out-of-stock product pages?

    For temporarily out-of-stock products, keep the page live with in-stock notification functionality and maintain its internal links. For permanently discontinued products, redirect to the closest in-stock alternative or the parent category using a 301, or return a 410 (Gone) status to signal permanent removal. Avoid leaving discontinued product pages returning 200 with thin "unavailable" content.

    Is Your Crawl Budget Being Wasted?

    A technical audit reveals the true scale of your crawlable URL set and identifies exactly which pages are consuming budget that should go to your product and category pages. Most e-commerce sites are surprised by what the data shows.

    Tharindu Gunawardana

    Tharindu Gunawardana

    Founder and Director of SearchMinistry

    Tharindu Gunawardana is the Founder of SearchMinistry Media and a search strategist with 17 years of experience across Sri Lanka, Singapore, and Australia. A former Agency SEO Director, he specialises in helping brands transition from traditional SEO to AI-driven discovery.

    Leave a Reply