- Which pages they fetch.
- When they fetch them.
- If these pages agree with the pages that they cite.
Why CDN logs show how answer engines use your site
The citation tracking tells you which of your pages appear in the answers of answer engines. Your logs tell you the other half: what answer engines do on your site first. The logs show two different behaviors. The main purpose of the analysis is to keep these two behaviors separate.1. Training crawls
Answer engines crawl the web to make and update the models behind their answers. Your logs show which of your pages they visit for training. Thus, you see which parts of your site the model knows and which parts it ignored.2. Real-time answer enrichment
This signal has more value. Answer engines also fetch pages in real time, while they answer a specific question. They use these pages to add current data to the answer. When a page has this type of access, an answer engine adds your content to a live answer at that time. Thus, you see which pages are important for the answers of answer engines now. This data is very different from “a crawler visited this page at some time”, and you can act on it more easily.Questions that the analysis answers
- Do answer engines crawl my most important content, or do they ignore it?
- Which pages do answer engines fetch in real time to support live answers?
- Which high-value pages do AI bots never visit?
- Which pages do crawlers visit frequently, but answer engines never cite?
Use the CDN logs with GA4 data
The log data is most useful when you read it together with your Google Analytics (GA4) data. Your CDN logs show what the answer engines did: the crawls and the real-time fetches. GA4 shows what persons did. When you compare the two, you get a much fuller view of these items:- How answer engines use the content that you publish
- What the persons who come to your site from these answers do
What log data Genezio ingests
Genezio ingests your CDN or web server access logs. These are the standard logs that your CDN or your origin server already records for each request:- The URLs and the paths of the requests
- The user agents that identify the source of each request, which include AI crawlers
- The time of each request, which makes the real-time answer enrichment visible
What Genezio analyzes in the logs
Raw logs are long lists of separate requests. Alone, they tell you very little. Genezio analyzes them and gives you reports on three items.Clustering by keyword
Genezio clusters your log entries by keyword. It puts the requested paths into themes. Thus, Genezio reports the activity at the level of topics, not of single URLs. When a different view is useful, Genezio can cluster the same logs in more than one way.Crawler breakdown
Genezio makes a crawler breakdown. It identifies the bots and the crawlers that access your site. It gives special attention to AI crawlers, which are the crawlers that answer engines use to fetch content. This keeps the traffic of answer engines separate from search engines, scrapers, and usual visitors.AI-crawler traffic compared with citations
The analysis with the most value connects the two halves. It compares the AI-crawler traffic with your citations. It examines if the pages that answer engines fetch agree with the pages that they really cite in their answers. This analysis shows the gaps. Examples are important pages that answer engines ignore, or pages that they fetch frequently but never cite.Request the log ingestion
The CDN log ingestion is not self-serve. Genezio does it on request for each customer.1
Contact Genezio
Contact Genezio and tell us that you want us to ingest your CDN logs or server logs.
2
Agree on the delivery
Genezio works with you and your infrastructure team to agree on how you send the logs.
3
Get the reports
Genezio analyzes the data and gives you custom reports from your CDN data and your GA4 data. The reports show what answer engines crawl, what they fetch in real time, and how this compares with the pages that they cite.