Publisher information

SemanticHubBot

SemanticHubBot is the crawler behind SemanticHub. It fetches articles from sites chosen by our users, groups them by topic and shows them to newsrooms as sources, with a link to the original.

User-Agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; SemanticHubBot/1.0; +https://semantichub.app/bot) Chrome/140.0 Safari/537.36
robots.txt token
SemanticHubBot
IP addresses
185.193.113.24783.243.37.22
robots.txt
Honoured per RFC 9309
Content usage signals
Content-Signal, AIPREF, TDMRep
Signal checks
Daily

1. What SemanticHubBot does

The bot only visits sites that SemanticHub users have added as sources. It fetches new articles, and newsrooms see them grouped by topic, always with a link to the original.

Every request identifies itself with the header Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; SemanticHubBot/1.0; +https://semantichub.app/bot) Chrome/140.0 Safari/537.36. The header follows the Googlebot format: it starts like a browser header, because some WAFs drop every other client. It always contains the SemanticHubBot token and the address of this page, so the bot never hides who it is.

In robots.txt and in firewall rules, match us by the SemanticHubBot token, not by the whole header.

2. What we don’t do

  • We do not train language models on the texts we fetch.
  • We do not pass the fetched texts on as a dataset.
  • We do not bypass paywalls or logins.

3. Why allow SemanticHubBot

We show every article to newsrooms with a link to the original. Material our users publish on their own sites may include links to its sources. That means referral traffic and backlinks that support your SEO.

It also supports GEO, your visibility in answers from AI tools. Content cited with its source can reach those answers together with a link to your site.

That is why we encourage you to allow SemanticHubBot in your WAF or bot management system. The addresses for your allowlist are listed under How to verify SemanticHubBot.

4. How to verify SemanticHubBot

SemanticHubBot fetches content only from 185.193.113.247, 83.243.37.22. The current address list in machine-readable form, the same format as Googlebot's lists, is at /semantichubbot.json.

A request with our User-Agent from any other address does not come from us, and you can block it.

5. How to block us

We honour robots.txt as defined in RFC 9309. To block our access completely, add this to your /robots.txt:

User-agent: SemanticHubBot
Disallow: /

6. Content usage signals

You can allow crawling and still reserve how your content is used. For this we read three kinds of signals:

SignalWhere to put it
Content-SignalA declaration in robots.txt, e.g. ai-input=no, ai-train=no
IETF AIPREFA Content-Usage declaration in robots.txt, e.g. train-ai=n
W3C TDMRepThe /.well-known/tdmrep.json file, the TDM-Reservation header or the <meta name="tdm-reservation" content="1"> tag

Example in robots.txt:

User-agent: SemanticHubBot
Content-Signal: ai-input=no, ai-train=no
Allow: /

7. How often we check

We check consent signals daily for sites in active use and record every change with its date. A change to your robots.txt reaches us by the next day at the latest.

8. Contact

Questions and requests: use our contact form.