As SEO experts, we use many tools every day, and each one gives us different outputs. Crawling tools are among the most important because they let us crawl selected pages or an entire website and quickly identify technical problems or gaps in the results. Some crawling tools are desktop applications, while others are cloud-based. Desktop tools must be downloaded and use your computer's hardware during a crawl. Cloud-based tools such as Deepcrawl do not require a download. You can go to deepcrawl.com and sign in with your username and password. The crawl runs entirely on Deepcrawl's servers rather than using your computer's hardware.
In this guide, we will cover Deepcrawl in detail. Let's begin with how to start a crawl for your website.
How to Start a Crawl Using Deepcrawl
After signing in to deepcrawl.com with your username or e-mail address and password, you will see the screen below. To start a crawl, click the "New Project" button in the upper-right corner of the page.
The first step is the "Domain" screen. Enter the domain name in the first field and the project name in the second. If necessary, you can also enable the Javascript rendering setting. When you have completed these steps, click the "Save and Continue" button.
Next, you will see the second step, the "Sources" screen. I will explain each setting and what it does.
1) Website: Here, you can include all sub-domains in the crawl, along with HTTP and HTTPS pages.
2) Sitemaps: After you enter the domain on the first screen, Deepcrawl finds your website's sitemaps. Make sure they are selected so that no URL is overlooked. You can also upload a different sitemap or include any other sitemap you want to crawl.
3) Backlinks: Here, you can use the Majestic tool to include URLs that link to your site, or upload a backlink list manually. When you include this source in the crawl, some reports may also show backlink data.
4) Google Search Console: Select a Search Console property to include its data in the crawl. This helps reduce the chance of overlooking pages that are not linked internally or that Deepcrawl bots cannot access for other reasons. If your properties do not appear here, you can add the Google account connected to your Search Console property in the upper-right section of the page.
5) Analytics: You can select your Analytics account here. This allows Deepcrawl bots to discover pages that are not linked from your site. Some crawl outputs may also include your Analytics data.
6) Log Summary: In this section, you can use summary data from log file analysis tools such as splunk and Logz.io, or upload your log files manually and include the data in the crawl.
7) URL Lists: Finally, you can upload a manual list of the URLs you want to crawl.
We have now covered all the settings that determine which sources are included in the crawl. Once you have configured this section, click the "Save and Continue" button at the bottom of the page.
The third step is the “Limits” screen. Here, you can set the number of URLs crawled per second and the maximum total number of URLs to crawl. When the specified limit is reached, you can choose to receive a notification or complete the crawl regardless. After configuring these settings, click the "Save and Continue" button again.
The fourth and final step is the "Settings" screen. The main settings are now in place. To view more detailed options, click the "Advanced Setting" button. Here, you can select additional domains or subdomains to include in the crawl, define URL paths to include or exclude, and choose the user agent that will run the crawl. Once you have finished configuring the settings, click the "Start Crawl" button.
After starting the crawl, you will see the screen below. It shows the number of URLs crawled in real time and lets you pause, stop, or delete the crawl. If you do not need to take one of those actions, you can leave the screen. Deepcrawl will send you a notification e-mail when the crawl is complete.
Deepcrawl Dashboard
When the crawl is complete, you will see the Dashboard screen below. Let's look briefly at its sections.
You can see your site's main problems, such as broken links, pages that are not linked from anywhere, and pages with a 4xx status code. Click any of these links to view the affected URLs and details of the problem.
This section shows basic details such as primary pages, duplicate pages, pages with or without the status code 200, and pages that cannot be indexed for any reason, along with the related pie chart.
If some time has passed since your last crawl and you want to see what has changed on your site, click the "Run Crawl" button in the "Changes" section to start a new one.
Pages that we want indexed must have the status code 200 (200 pages). This section shows pages that do not have the status code 200 (non-200 pages).
This section lists URLs that cannot be crawled and explains why.
We may choose not to index some pages to optimize the crawl budget or for another reason. This section lists non-indexable pages and explains why they cannot be indexed. I recommend reviewing these pages and acting quickly if an important page is not indexed because of a problem.
This section shows pages that are not linked from any other page on your site, known as "orphaned" pages. I recommend reviewing them carefully. If any important pages appear here, make sure other pages link to them. If an overlooked page is unimportant, you can consider options such as redirection.
This section lets you check whether your pages appear in the specified sources. In the example below, you can access Search Console, Analytics, sitemap, and pages that are on or not on the web. To view Search Console, Analytics, or Backlink data here, you need to establish those connections when starting the crawl.
This section shows duplicate pages, non-200 pages, pages that cannot be indexed, and related trend graphs.
This section contains a graph of your pages by depth level. It helps you quickly see where most pages sit and find pages deep in the structure that may have gone unnoticed.
This section shows the number of HTTP and HTTPS pages. If you have HTTP pages, make sure they all open securely over HTTPS.
Here, Search Console data shows which indexable pages receive impressions and which do not. To see this data, you need to connect Search Console when starting the crawl.
This section shows the click-through rate of indexable and non-indexable pages by device. Again, you need to connect Search Console when starting the crawl to see this data.
We have briefly covered the Dashboard sections you see when a crawl finishes. Now, let's look at the important sections in the left-hand menu that you should review.
- Issues
The "Summary" section lists the main problems found after your pages are crawled.
- Changes
If you have run a crawl before, this section shows how problems on your web pages have changed.
- All Pages
Here, you can access the two basic tables and the full list of pages from the Dashboard. Use the options in these tables to filter pages by relevant features or problems. You can also export all the pages found during the crawl.
- Indexable Pages
This section lists all your indexable pages. The pie chart on the right also shows which of those pages are unique or duplicate.
- Non-Indexable Pages
This section lists pages that cannot be indexed and shows the reasons in the graph on the left. I recommend reviewing these pages carefully because an overlooked problem may prevent an important page or group of pages from being indexed. In the sample crawl below, all the non-indexable pages point to different pages with the canonical tag. Pages that you want indexed need to refer to themselves with the canonical tag.
- 200 Pages
This section lists pages with the status code 200. Our important pages need to be served without problems and return the status code 200. The graph on the right shows the number of 200 pages found in previous crawls.
- Non-200 Pages
Here, you can access pages that do not have the status code 200. Pages that users and search engine bots need to reach must be served with the status code 200, so review every page listed here that does not have the status code 200. If pages return status codes 404 or 500 rather than a redirect, resolve the underlying problems and make sure the pages open without errors with the status code 200.
- Uncrawled URLs
This section lists URLs that cannot be crawled and explains why. In the example below, every URL is blocked by the disallow command in the robots.txt file. Review both the URLs and the commands in the robots.txt file. Otherwise, overlooked pages or page groups may not be indexed.
- Primary Pages
This section lists your indexable, unique pages. The pages shown here can be considered important primary pages for your website.
- Duplicate Pages
This section lists pages with duplicate titles, description tags, or identical or substantially similar content. I recommend reviewing these pages and giving the pages you expect to receive organic traffic unique titles, descriptions, and content.
- Self Canonicalized Pages
This section lists pages that refer to themselves with the canonical tag. Each page that should be indexed needs to refer to itself with the canonical tag. Even so, review the pages and page groups in this section. If some pages should not be duplicated or indexed, you can adjust their canonical tags as needed.
- Noindex Pages
This section lists pages with the "noindex" tag. I recommend reviewing them and removing the "noindex" tag from any overlooked pages that you want indexed.
- Canonicalized Pages
Here, you can access pages whose canonical tag points to a different URL. Review these pages and make any page you want indexed refer to itself with the canonical tag.
- 301 Redirects
This section lists pages redirected with the 301-status code and the pages they redirect to.
- Non-301 Redirects
This section lists pages with 302 status codes. If any of them need to be redirected permanently, you can use a 301 redirect.
- 5xx Errors
This section lists pages with 5xx status codes. Examine the server-side problems causing these responses, resolve them, and make sure the pages open with a 200-status code.
- Broken Pages (4xx Errors)
This section lists pages with 4xx status codes. Pages that users and search engine bots should visit need to open without problems with the status code 200. Resolve the problems on these 4xx pages, or add the necessary redirects if the pages will not become active again.
- Failed URLs
This section lists failed pages. A page may appear here because it does not open or loads too slowly. In the example below, the pages take too long to load and return a "0" status code. Fix the problems on these pages so they open without errors with the status code 200.
- Content Overview
Here, you can access graphs and a list of general content problems. If you have run another crawl before, the side tab also shows how these content-related problems have changed.
- Missing Titles
This section lists pages that are missing a title tag. I recommend adding an appropriate, unique title tag to each one.
- Short Titles
This section lists pages with title tags that are short in character length. You do not need to lengthen every title listed here. However, reviewing and improving the title tags of certain pages or page groups may help you use this report more effectively.
- Max Title Length
This section lists pages with title tags that are longer than required. Title tags above a certain pixel length will appear truncated on results pages, so I recommend optimizing the length of the titles listed here.
- Pages with Duplicate Titles
This section lists pages with duplicate title tags. I recommend adding a unique title tag to every page shown here, especially pages that are indexed and expected to receive organic traffic.
- Missing Descriptions
This section lists pages without a description tag. Although description tags are not a direct ranking factor, they may affect click-through rates. It is therefore useful to add a unique description tag to each indexed page that is expected to receive organic traffic. Otherwise, Google will display random text from the page as its description tag.
- Short Descriptions
Here, you can access pages with description tags that are short in character length. Google usually does not show very short description tags on results pages and instead displays random text from the page. It is therefore useful to optimize description tags for indexed pages that are expected to receive organic traffic. Deepcrawl considers description tags short when they contain fewer than 50 characters.
- Max Description Length
This section lists pages with description tags that are longer than necessary. Description tags above a certain pixel length will appear truncated on results pages, so I recommend optimizing the length of the descriptions listed here.
- Pages with Duplicate Descriptions
This section lists pages with duplicate description tags. It is useful to add a unique description tag to each important page shown here.
- Empty Pages
This section lists empty pages. If a page was created by mistake, users and search engine bots should not encounter a blank page. You can redirect it or fix the underlying problem and add rich content.
- Thin Pages
This section lists pages with thin content. Improving these pages can provide a better experience for users and search engine bots. If you expect them to receive organic traffic, make the necessary content improvements.
- Missing H1 Tags
This section lists pages without an H1 tag. In particular, I recommend adding a unique H1 that summarizes the content of every indexed page expected to receive organic traffic.
- Multiple H1 Tag Pages
This section lists pages with more than one H1 tag. The H1 acts as the page's main title and needs to be unique on each page. Adjust the pages listed here so that each one has only one H1 tag.
- Canonical to Non-200
This section shows pages whose canonical tag refers to a URL that does not have the status code 200. Make sure the canonical URL opens without problems with the status code 200, or update the canonical tag.
- Redirect Chains
This section lists pages with a redirect chain. A redirect chain occurs when one page redirects to another page that then redirects somewhere else. Search engine bots send a new request at every step. To avoid this problem, remove the redirect chains and redirect each page directly to the final page.
- All Redirects
This section lists all redirects on the crawled pages, including their types and details.
- All Broken Redirects
This section lists broken redirects whose destination does not open with the status code 200. Fix these problems by removing the redirect or pointing it to a page that opens without errors with the status code 200.
- HTTP Pages
This section lists the HTTP pages on your site. Make sure every page shown here opens securely over HTTPS.
- Broken Sitemap Links
This section lists sitemap pages that do not open with the status code 200. Your sitemap needs to include pages that you want indexed and that open with the status code 200. Make sure the URLs shown here return the status code 200, or remove them from the sitemap.
- Non-Indexable URLs in Sitemaps
This section lists URLs that are not indexed but still appear in the sitemap. If these URLs should not be indexed, remove them from the sitemap as well.
- Indexable Pages without Search Impressions
This section lists indexable pages that never receive impressions. To see this data, you need to connect your Search Console account at the beginning of the crawl. A general review of these pages may be useful. If a page or group of pages has no potential to receive impressions, you can prevent it from being indexed to optimize the crawl budget. Otherwise, make the necessary improvements so the pages can receive impressions and clicks.
Conclusion
We have covered the important issues you may encounter after running a crawl with Deepcrawl. Some are minor technical problems, while others are significant enough to directly affect your website's SEO performance. We therefore recommend crawling your website with Deepcrawl periodically and resolving technical problems or gaps in order of priority.
Penned by Metehan Urhan - SEO Executive, Zeo Agency

