Thursday, 27 August 2026

Aube.

News of progress
Crawler

AubeBot

Aube is a techno-optimist online journal: it covers what is moving forward and what works, across technology at large. It follows that news through the feeds publishers make public. This page states who is knocking at your door, what it reads, how often, and how to close that door — it is written for the administrator who has just found “AubeBot” in their logs.

Identity

User-Agent
AubeBot/0.1 (+https://aube.news/robot; collecte RSS polie)
Operator
Aube.https://aube.news
Contact
[email protected]
Request origin
A single virtual machine on a fixed IP address. No address rotation, no proxy pool, no headless browser.

What it reads

  • Your public RSS or Atom feed, and nothing else for as long as that suffices: at most forty items per read.
  • Conditionally (If-None-Match, If-Modified-Since). If nothing has changed since the previous read, your server answers 304 and sends no body at all.
  • An article page, only when the feed carries nothing but a summary under six hundred characters — and only for an address never read before. A page read once is never requested again.
  • Your robots.txt before every request, the feed included, interpreted per RFC 9309: most specific group, Allow directives, wildcards, and Crawl-delay honoured up to ten seconds.

What it does not do

  • It does not crawl your site: it only opens addresses your own feed published, and follows no internal links.
  • It runs no JavaScript, signs into no account, accepts no cookie and circumvents no paywall. A paid article is never extracted, even when the page carries its text in the clear.
  • It never republishes your article. What appears on Aube is an original French text that cites the source and links back to it.
  • It reuses none of your images: Aube’s illustrations are generated from the article’s key figure.
  • Aube trains no model. The texts it reads are used to write an article, never to build a training corpus.

Rate

  • One feed read per source every fifteen minutes at most — and no data transferred at all when nothing has changed.
  • At least fifteen hundred milliseconds between two requests to the same host, counted from the end of one to the start of the next. Your Crawl-delay wins if it is longer.
  • At most five article pages per source per quarter-hour, sixty across all sources.
  • Requests go out one at a time. AubeBot never opens two connections at once, to you or to anyone.

Blocking it

Three lines in your robots.txt are enough, and they are honoured without argument:

User-agent: AubeBot
Disallow: /

The file is re-read once per host per cycle: the block takes effect within the quarter-hour. To slow it down rather than shut it out, a Crawl-delay is enough; to close a single section, a Disallow on that path.

If you would rather write than configure — removal of a published article, a question about how your content is used, an exclusion request — the address above receives real messages and they are answered.

What gets published from your content

An Aube article groups several sources covering the same fact, then rewrites them in French. Each source’s title, date and address appear at the foot of the article, with a clickable link to the original page. A short quotation between quotation marks is possible; lifting whole sentences is not.

Legal notice and intellectual property

AubeBot — Aube’s crawler · Aube.