7 min
GPTBot: what it is and how to configure it in your robots.txt
GPTBot is OpenAI’s crawler that trains its generative AI models. We explain how to tell it apart from OAI-SearchBot and configure your robots.txt.
Read
What stops an AI crawler from ever reading your site: firewalls, inherited robots.txt rules and anti-bot protection. It is the knock-out pillar: while the way in is closed, no other fix counts.
6 articles
7 min
GPTBot is OpenAI’s crawler that trains its generative AI models. We explain how to tell it apart from OAI-SearchBot and configure your robots.txt.
Read
7 min
Google Search Console is Google’s most reliable source on how it sees your site. What it is, what it is for and what each of its main reports measures.
Read
9 min
What a site’s crawl budget is, how Google calculates it and which practices help optimise it. With common myths and how to check it in Search Console.
Read
8 min
What a sitemap is, when you really need one, and how to create it and submit it to Google Search Console.
Read
9 min
Structured data, schema markup and JSON-LD are not synonyms. Here is what each one is, why they are no longer enough to show FAQs on Google, and how to check yours.
Read
7 min
Permission in robots.txt is no guarantee the crawler gets through. How to spot a firewall block in two minutes, and why the most dangerous case returns HTTP 200.
Read