We address an important matter in the world of SEO positioning that applies to any website, from the smallest to the largest. It is about the crawl budget, an aspect that can affect the visibility of your website in search engines and how these index their web pages.
What is the crawl budget?
The crawl budget is simplythe rate of pages of a site that a search engine is willing to crawl over a period of time. In Spanish, the term would mean ‘presupuesto de rastreo’ in Spanish. This budget is directly proportional to thedomain authorityof the website, being higher if the page is more important, but it can also be affected by the frequency of updates to its content.
Importance of crawl budget in SEO
In the world of SEO it is essential to understand the crawl budget because much depends on it in terms of how search engines will interact with the website. Let’s look at the most important points to understand its relevance.
How crawl budget influences the visibility of your website
The crawl budget directly influences the visibility of websites because the search engine needs to crawl the pages in order to add them to itsdatabase.Greater relevance of a website implies a higher crawl frequency and also more chances of appearing in searches.
Impact of crawl budget on content indexing
If you have a low crawl budget it is possible that not all pages of your site will be crawled, which will affect indexing. As the crawl budget increases the search engine can spend more time crawling the site, so there will be a greater likelihood that it will index its content or review page updates.
How does the crawl budget work?
The algorithms used by search engines are not published, so it is not possible to provide data about their operation with certainty. However, from talks given by the Google team and articles they have been publishing, we can generally deduce how it works. Basically, as we have said, it is an indicator of the time a search engine can spend traversing a website. Therefore,a high crawl budget is always more desirable than a low one, since it is ideal that the search engine can take the necessary time to crawl the site and publish its content in the form of links on its results pages.
Factors that influence the crawl budget
The factors that carry weight in the crawl budget are the following.
- Website authoritySites with greater authority will generally be assigned a higher crawl budget.
- Frequency of content updatesIn addition, sites that update their content regularly can receive more attention from search engines, in the form of a higher crawl budget.
How to optimise the crawl budget
As we gain authority we can begin to receive more crawl budget. Authority is obtained as the site gains relevance. We can increase relevance throughlink building, thanks to inbound links to our site.
Relevance is therefore very important to achieve an increase in our crawl budget. But it is not just about being better regarded by the search engine, but also about having a properly optimised site so that our current crawl budget, whatever it may be, is used more efficiently. So let’s see how we can optimise the crawl budget assigned to us,
Audit of your website’s structure
For our site to be more indexable it is important that it has a clear structure and is easy for search engines to navigate. If this is the case we will be optimising the time the crawler spends traversing the site.
A good website structure is one that has a well-defined hierarchy, with links that allow navigation through the pages in a simple way, without requiring too many clicks and without orphan pages or pages that are too far from the homepage.
Optimisation of the robots.txt file
We can also improve the way the search engine uses our crawl budget through proper configuration of therobots.txt file. For example, we can define which parts of the site we do not want to be indexed or crawled, to prevent the assigned crawl budget from being wasted crawling irrelevant or duplicate content.
Managing page load speed
If we look after the site’s load speed we will also achieve a higher crawl speed, as the search engine will be able to visit more pages in less time. Therefore, sites that load quickly potentially offer higher rates of optimisation of the crawl budget.
Prioritising the most relevant content on your website
Always try to ensure that the search engine is able to understand which content on your website is most important, so that it is crawled more regularly. You can achieve this with a linking structure, making more pages link to the important content. You can also define it via the priority of links expressed in thesitemap.
Also, regarding content we should always try to ensure it is of quality. Content copied from other websites, low-quality content or spam will cause the crawl budget to be wasted and will prevent Google or other engines from reaching the content that really matters.
Removal of duplicate and non-indexable content
Duplicate content can consume your crawl budget, reducing the chances that the search engine will index the most relevant pages. Therefore, it is important to eliminate it as much as possible or to prevent it from being crawled via robots.txt. Also remember to use techniques like canonical to declare which pages are primary, if you cannot completely remove your duplicate content.
Recommended tools to monitor your crawl budget
Now let’s look at some of the tools we can recommend to monitor how crawlers interact with your content.
Google Search Console
Google Search Consoleis the tool provided by Google to understand the degree of indexing and possible problems your site may have when it is crawled by the spider. It replaces the former ‘Webmaster Tools’ and is completely free for any website. To be able to use it you simply have to verify ownership of the site by following the instructions presented when you first enter.
Screaming Frog SEO Spider
This application is a crawler in itself. Screaming Frog SEO Spider, therefore, crawls websites and detects possible problems they may present in terms of SEO. It will be very useful in helping us to solve problems with our website and to optimise the time Google spends crawling it.
Deepcrawl or Lumar
Deepcrawl, currently known as Lumar, is an advanced web crawling platform that offers detailed reports on your site’s structure. It can detect technical problems that prevent crawlers from interacting with your content. It is also able to monitor your website and send alerts if problems are detected at any time.
Sitebulb
Sitebulb is anSEO auditing toolthat works both as a web application and as a desktop application. Once launched it crawls the website and provides numerous notable details for its optimisation in search engines.
Botify
Botify is a tool aimed at brands that want to monitor their visibility in search engines. It provides a technical SEO audit with real-time data analysis, which is also combined with reports on content indexing in Google.
How Google manages the crawl budget
As we have said before, it is not possible to know with certainty how Google manages a site’s crawl budget, since its specific mode of operation is not published on the Internet. However, from the information published by the search engine itself, we know thatGoogle manages the crawl budget dynamically, adjusting the frequency and depth of crawling for each site according to its relevance, authority, the updates the site has and its performance.
Googlebot is Google’s crawler and assigns more resources to sites it considers important, also trying to spend the crawl budget on the most relevant pages, so that these are more likely to be crawled and indexed.
Mechanisms for managing the crawl budget
Google sets some parameters for managing the crawl budget on websites.
TheCrawl Rate Limitsets the crawl rate limit, which is the maximum number of requests that Googlebot makes to your server in a given period of time. We can generally adjust how Googlebot can use our site in the robots.txt file by configuring a delay between each crawled page with a configuration like this:
crawl-delay: 5
However, it is important to know that this configuration could cause your crawl budget to be wasted, because in the time assigned to your site it will be able to crawl fewer pages.
On the other hand we have theCrawl Demand, which defines the demand for crawling that can be adjusted depending on how interesting the site is (deduced based on the people who access the content), or the frequency with which it is updated.
Finally, we have the parameterCrawl Schedulingthat defines the scheduling of crawling, something defined internally by Google.
Advanced tips to optimise the crawl budget
To optimise your website’s crawl budget you can take into account some advanced tips:
Continuously monitor your website, especially via the Search Console tool, as it can inform you of your site’s crawling problems.
Use sitemaps: this will make it easier for Google to locate your content and will increase the chances that it indexes it.
Avoid soft errors: they occur when the server responds with the HTTP 200 OK status code when it is actually a page not found. Make sure that when a page does not exist on your server the crawler receives an HTTP 404 code.
Redirect URLs that have changed address correctly, with codes of HTTP 301 for permanent redirects, so that Google does not waste time with the old address ever again.
Use canonical tags to mark which page the search engine should take into accountwhen there is duplicate or very similar content.