Free SEO audit for new clients — claim yours
Home/Blog/Technical SEO/Crawl Budget: What It Is and When It Matters
Technical SEO

Crawl Budget: What It Is and When It Matters

Crawl Budget: What It Is and When It Matters
Table of contents
  1. Key Takeaways
  2. What Is Crawl Budget?
    1. Crawl capacity limit
    2. Crawl demand
  3. Does Crawl Budget Matter for Your Site?
  4. What Wastes Crawl Budget
  5. How to Diagnose Crawl Budget Issues: Step by Step
  6. How to Optimise Crawl Budget
  7. What Doesn't Help (and Common Mistakes)
  8. Related Guides
  9. Frequently Asked Questions
    1. Does my small website need to worry about crawl budget?
    2. Is crawl budget a ranking factor?
    3. Does noindex save crawl budget?
    4. Can I ask Google to crawl my site more?
    5. Where can I see how Google crawls my site?
    6. Do redirects waste crawl budget?
  10. Conclusion
  11. References

Crawl budget is one of those SEO terms that gets more attention than it deserves on small sites and far too little on large ones. A 50-page local business site will not see a difference if it “optimises crawl budget”. A store with two million product and filter URLs, on the other hand, can have new products waiting weeks to be discovered because Googlebot is busy crawling parameter combinations nobody searches for.

This guide explains what crawl budget actually is according to Google, how to tell whether it matters for your site, how to diagnose crawl waste with Search Console and log files, and the practical fixes that free up crawling for the pages you care about.

Key Takeaways

  • Crawl budget is the time and resources Google allocates to crawling a site, shaped by crawl capacity limit and crawl demand.
  • It mainly matters for large or fast-changing sites; most small sites can ignore it.
  • Faster, more reliable servers raise the capacity limit; popular and fresh content raises demand.
  • The biggest crawl waste comes from faceted navigation, parameters, duplicates, soft errors and redirect chains.
  • Use robots.txt to stop crawling of worthless URLs; noindex does not save crawl budget.

What Is Crawl Budget?

The web is effectively infinite, so Google cannot crawl every URL it knows about as often as it would like. It has to decide how much time and how many requests to spend on each site.

“A site's crawl budget is determined by two main elements: crawl capacity limit and crawl demand.”

— Google, Large site owner's guide to managing your crawl budget, Crawl budget management

Crawl capacity limit

This is how many simultaneous connections Googlebot can use, and how long it waits between fetches, without overloading your server. Google's guide explains that if your site responds consistently and response times remain stable or improve, the limit goes up. If the site slows down or returns server errors or rate-limiting signals such as 429, the limit goes down and Google crawls less.

Crawl demand

This is how much Google wants to crawl your site. The guide lists factors including the URLs Google perceives as worth crawling, popularity (more popular URLs tend to be crawled more often) and staleness (Google wants to recrawl documents often enough to pick up changes). Large events such as site moves can also increase demand temporarily.

Does Crawl Budget Matter for Your Site?

Google's guide is explicit about its audience. It is written for:

  • Large sites with 1 million+ unique pages whose content changes moderately often (weekly).
  • Medium or larger sites with 10,000+ unique pages whose content changes very rapidly (daily).
  • Sites with a large share of URLs classified in Search Console as “Discovered – currently not indexed”.

The guide also says that if your site does not have a large number of rapidly changing pages, or your pages seem to be crawled the same day they are published, you don't need to read it. For most small business sites, indexing problems are about quality and duplication, not crawl budget. Our guide to “Crawled – currently not indexed” covers those cases.

Site profileCrawl budget priorityFocus instead on
Local business site, under 500 pagesVery lowContent quality, internal links, local signals
Blog with a few thousand postsLowPruning thin content, sitemaps, internal linking
Ecommerce store with faceted navigationHighParameter control, duplicates, server speed
Marketplace, classifieds or news site with daily changesHighFreshness signals, sitemaps with lastmod, removals
Enterprise site with 1M+ URLsVery highLog analysis, URL inventory, infrastructure

What Wastes Crawl Budget

  • Faceted navigation and filters generating near-endless combinations such as ?color=red&size=9&sort=price.
  • Session IDs and tracking parameters that create unique URLs for identical content.
  • Duplicate content across protocol, host, trailing-slash and case variants. See our guide to duplicate content.
  • Soft 404s that return 200 for empty or error pages; see how to fix 404 errors.
  • Infinite spaces such as calendars with “next month” links forever.
  • Redirect chains, where each hop is an extra request.
  • Slow server responses, which lower the capacity limit.
  • Hacked or spam pages injected into the site.

How to Diagnose Crawl Budget Issues: Step by Step

  1. Open the Crawl Stats report in Search Console (Settings → Crawl stats). Review total crawl requests, average response time, and breakdowns by response code, file type and Googlebot type.
  2. Check host status in the same report for robots.txt fetch, DNS or server connectivity problems.
  3. Review the Page indexing report for large counts of “Discovered – currently not indexed”, duplicates and soft 404s.
  4. Analyse server logs. Filter for verified Googlebot requests and group by URL pattern. Work out what share of crawling goes to parameters, redirects or errors versus your key templates.
  5. Crawl the site yourself with Screaming Frog, Sitebulb or Semrush Site Audit to map parameter URLs, chains and orphan pages.
  6. Compare crawled URLs against your list of pages you want indexed. The gap is your crawl waste.

How to Optimise Crawl Budget

Google's guide groups its advice around managing your URL inventory, keeping the site efficient and monitoring crawling. Here is how that translates into actions:

ProblemFixWhy it works
Worthless parameter and filter URLsDisallow patterns in robots.txt; avoid linking to themStops Googlebot requesting URLs you never want crawled
Duplicate URL variants301 redirects and rel=canonical; consistent internal linksHelps Google consolidate and crawl duplicates less often
Removed pages returning 200Return 404 or 410Google crawls 4xx URLs less over time
Redirect chainsRedirect straight to the final URLRemoves extra hops
Slow responsesCaching, CDN, faster TTFB, efficient databasesRaises the crawl capacity limit
New or updated pages discovered slowlyAccurate XML sitemaps with lastmod; strong internal linksRaises discovery and recrawl demand
Heavy pages and resourcesSmaller pages, fewer render-blocking resourcesFewer bytes and requests per page

Two examples of robots.txt rules for faceted navigation:

User-agent: *
# Block sort and session parameters anywhere in the URL
Disallow: /*?*sort=
Disallow: /*?*sessionid=

Sitemap: https://www.example.com/sitemap_index.xml

Test any pattern carefully before deploying; an overly broad rule can block important pages. Our guide to robots.txt vs meta robots explains which tool to use when, and our XML sitemap guide shows how to keep sitemaps lean.

Server speed matters too. Improvements to time to first byte help users as well as crawlers; see our guide on how to improve page speed, or let our page speed optimization team handle it.

What Doesn't Help (and Common Mistakes)

“Don't use noindex, as Google will still request, but then drop the page.”

— Google, Large site owner's guide to managing your crawl budget, Crawl budget management
  • Using noindex to save crawl budget. Google still requests the page.
  • Toggling robots.txt to shift budget. Google advises against using robots.txt to temporarily reallocate crawl budget to other pages; use it only for URLs you never want crawled.
  • Blocking pages you want de-indexed. If a URL is blocked, Google cannot see a noindex on it, and it may still be indexed without content.
  • Relying on nofollow internal links. They are not a dependable way to control crawling; remove the links or block the URLs instead.
  • Returning 5xx or 429 to slow Googlebot for long periods. This is an emergency measure only; prolonged errors can lead to URLs being dropped.
  • Expecting crawl rate to boost rankings. More crawling does not mean better rankings; it only ensures your content can be indexed and refreshed.

Frequently Asked Questions

Does my small website need to worry about crawl budget?

Almost certainly not. Google's guide is aimed at very large sites (1 million+ pages changing weekly) or medium-to-large sites (10,000+ pages changing daily), plus sites with many URLs stuck in 'Discovered – currently not indexed'.

Is crawl budget a ranking factor?

Being crawled more often does not by itself improve rankings. Crawl budget matters because pages that are not crawled cannot be indexed or updated in search results.

Does noindex save crawl budget?

No. Google still has to request the page to see the noindex tag, then drops it. Use robots.txt for URLs you never want crawled, and 404/410 for removed pages.

Can I ask Google to crawl my site more?

Not directly. Google increases crawling when your server responds quickly and reliably and when your content is popular and frequently updated. You can, however, temporarily reduce crawling in emergencies.

Where can I see how Google crawls my site?

In Search Console under Settings, the Crawl Stats report shows total requests, download size, response times, status codes and file types. Server logs give an even more detailed view.

Do redirects waste crawl budget?

Each redirect hop is an extra request. Google's guide advises avoiding long redirect chains because they have a negative effect on crawling.

Conclusion

Crawl budget is a real constraint, but only for sites big or dynamic enough to hit it. If that is you, the playbook is straightforward: shrink your URL inventory to pages worth crawling, make your server fast and reliable, keep sitemaps and internal links pointing at canonical URLs, and monitor Crawl Stats and logs. For everyone else, effort is better spent on content quality and a regular SEO audit, as our technical SEO guide explains.

Running a large ecommerce or enterprise site? Our enterprise SEO team specialises in log file analysis and crawl optimisation at scale. Get a free quote to find out where your crawl budget is going.

References

  1. Google Crawling Infrastructure: Large site owner's guide to managing your crawl budget
  2. Google Search Central: How HTTP status codes and network errors affect Google Search
  3. Google Search Central: Site moves with URL changes
  4. Google Search Central: Introduction to robots.txt
  5. Google Search Central: Build and submit a sitemap
  6. Semrush: Crawl Budget — What Is It and Why Is It Important for SEO?
My SEO Experts Team

Our team of SEO and digital marketing specialists has been helping businesses grow organically since 2018. More about us.

Keep reading

Related articles

Ready to grow your business online?

Get a free, no-obligation proposal tailored to your goals and budget.

CallFree Quote