A free, working cannabis & CBD SEO library, 500 guides. Browse the library ›
Home » Blog » Analytics & Measurement » Using Log Files to Improve Cannabis Crawl Efficiency
Analytics & Measurement

Using Log Files to Improve Cannabis Crawl Efficiency

In short

Your server logs are the only record of what Googlebot actually did on your site, as opposed to what a crawler simulation guesses it might do. On a big cannabis catalog, they show where crawl budget gets burned, on filtered menu URLs and dead parameters, so you can steer it back to the pages you need indexed.

This is an advanced move, and it doesn't earn its keep on a ten-page brochure site. But for a dispensary or brand running a large menu with faceted filters, logs answer a question no other tool can: is Google spending its limited crawling on your product and content pages, or drowning in near-infinite filter combinations. Screaming Frog's crawl reports estimate; the log file shows the truth.

What a log file actually contains

Every request to your server leaves a line: the URL, the timestamp, the response code, and the user agent. Filter those lines to the Googlebot user agent and you have a literal diary of Google's visits, which URLs it hit, how often, and what status code it got back. No sampling, no modeling, just the record. That's why log analysis sits above tool-based crawl estimates for this specific question.

Get the logs and load them

Pull the access logs from your host, usually through cPanel's raw access logs or from whoever manages the server. Load them into the Screaming Frog Log File Analyser, which parses the format and verifies that the Googlebot hits are real rather than spoofed bots pretending to be Google. Once imported, you can group hits by directory, by response code, and by frequency.

Find where crawl budget leaks

On cannabis sites the leak is almost always faceted navigation. A menu that generates a unique URL for every combination of category, potency, price, and brand can spin up thousands of low-value crawlable pages, and the logs will show Googlebot dutifully hammering them while your actual product pages get visited rarely. Other classic leaks: long redirect chains eating hits, and 404s that Google keeps recrawling. Each line spent on junk is a line not spent on a page you care about.

Steer crawling back to what matters

Once the logs show the waste, you act on it: block worthless parameter URLs in robots.txt, add noindex or canonical signals to filter pages, fix redirect chains, and clean up the 404s Google keeps hitting. Then pull logs again a few weeks later and confirm the crawl shifted toward your important directories. The goal is simple, get Google spending its finite crawl on pages that can actually rank.

A worked example

A multi-location brand had thousands of products indexing slowly. Their logs showed 70 percent of Googlebot hits landing on filtered menu URLs like "?brand=x&potency=y" combinations, and their core category pages getting crawled maybe once a week. We disallowed the filter parameters in robots.txt and canonicalized the filtered views to their clean category pages. A month later the logs showed category and product hits up sharply, and new products started getting indexed in days instead of weeks. Nothing else changed, just where the crawl went.

Key takeaways

  • Logs are the only true record of Googlebot's behavior; crawl tools estimate, logs show what actually happened.
  • This pays off on large faceted cannabis catalogs, not small sites, where crawl budget is a real constraint.
  • Load host access logs into Screaming Frog's Log File Analyser and verify the Googlebot hits are real.
  • The usual leak is filter-generated URLs; block or canonicalize them, fix redirect chains and 404s, then re-pull logs to confirm the shift.

Frequently asked questions

Why use log files instead of a crawl tool?

Because logs record what Googlebot actually requested, while a crawl tool only simulates what it might. For questions about crawl budget, which pages Google visits and how often, only the server log gives you the real answer, with no sampling or modeling in between.

Is log file analysis worth it for a small cannabis site?

Usually not. Crawl budget only becomes a constraint on large catalogs where faceted filters generate thousands of URLs, so a small brochure site sees little benefit. The technique earns its keep on big menus where Google may waste its crawl on filter combinations instead of your product pages.

Want this done for your brand?

We build cannabis & CBD search visibility with full transparency, for dispensaries, brands, and MSOs.

Get a free cannabis SEO audit ›