archive.org_bot
Cratus neverforgetus
A joke about the drawing A storage box that remembers everything put in it.
The line above plays on the name. It says nothing about how the bot behaves: everything below is read from the request log, for today and the 29 days before it, by UTC date.
Alignment
Insufficient data. A judgement needs at least 200 requests and 5 sessions in the trailing 30 days. This one would rest on unverified sessions, so it is a statement about the species and not about any operator: 1 request in 1 session over the trailing 30 days.
How the score is worked out, and the count behind every input, is on the Census page.
Who runs it, and how that is checked
- Who runs it
- the Internet Archive
- A user-agent string it sends
Mozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)Usually followed by the name and version of the crawling software.- What it says it is for
- archive.org_bot collects pages for the Internet Archive’s Wayback Machine, so that old versions of the web can still be read later.
Everything above is a claim: the operator’s own account where it has published one, and a third party’s where it has not. This site has not tested it. What the bot has done is in the sections below, read from the request log.
Customs has no documented check for it. A session that says it is archive.org_bot can be neither confirmed nor refuted, so every one is shown as unverified and none is put down to the Internet Archive.
The name it is matched on
A request is filed under this species when its user-agent string contains the product name
archive.org_bot as a whole word, in any mix of capitals. Where a string holds more than one product name, the
longest wins, and a named bot wins over a library it is built on.
How to address this bot in robots.txt
A group in robots.txt addresses this bot by the line:
User-agent: archive.org_bot
An example group. The path is a placeholder.
User-agent: archive.org_bot Disallow: /a-path/ Allow: /a-path/one-page
This is reference, not advice. Systema Machinae is not a bot-blocking product.
This site’s own robots.txt names it, in the group that keeps the guest list off the Private Beach.
Landings
1 landing in the trailing 30 days. A landing is the first request of a session.
Verified sessions passed Customs. Unverified sessions said they were archive.org_bot and could not be checked, or have not been answered yet: they are shown for the species and are not put down to any operator. Sessions that failed are not on this page at all.
By day
VerifiedUnverified
The same figures as a table
| Day | Verified | Unverified | Total |
|---|---|---|---|
| 9 Sep | 0 | 0 | 0 |
| 10 Sep | 0 | 0 | 0 |
| 11 Sep | 0 | 0 | 0 |
| 12 Sep | 0 | 0 | 0 |
| 13 Sep | 0 | 0 | 0 |
| 14 Sep | 0 | 0 | 0 |
| 15 Sep | 0 | 0 | 0 |
| 16 Sep | 0 | 0 | 0 |
| 17 Sep | 0 | 0 | 0 |
| 18 Sep | 0 | 0 | 0 |
| 19 Sep | 0 | 0 | 0 |
| 20 Sep | 0 | 0 | 0 |
| 21 Sep | 0 | 0 | 0 |
| 22 Sep | 0 | 0 | 0 |
| 23 Sep | 0 | 0 | 0 |
| 24 Sep | 0 | 0 | 0 |
| 25 Sep | 0 | 0 | 0 |
| 26 Sep | 0 | 0 | 0 |
| 27 Sep | 0 | 0 | 0 |
| 28 Sep | 0 | 0 | 0 |
| 29 Sep | 0 | 0 | 0 |
| 30 Sep | 0 | 0 | 0 |
| 1 Oct | 0 | 0 | 0 |
| 2 Oct | 0 | 0 | 0 |
| 3 Oct | 0 | 0 | 0 |
| 4 Oct | 0 | 0 | 0 |
| 5 Oct | 0 | 0 | 0 |
| 6 Oct | 0 | 0 | 0 |
| 7 Oct | 0 | 0 | 0 |
| 8 Oct | 0 | 1 | 1 |
By hour of the day, UTC
VerifiedUnverified
The same figures as a table
| Hour | Verified | Unverified | Total |
|---|---|---|---|
| 00:00 | 0 | 0 | 0 |
| 01:00 | 0 | 0 | 0 |
| 02:00 | 0 | 0 | 0 |
| 03:00 | 0 | 0 | 0 |
| 04:00 | 0 | 0 | 0 |
| 05:00 | 0 | 0 | 0 |
| 06:00 | 0 | 0 | 0 |
| 07:00 | 0 | 0 | 0 |
| 08:00 | 0 | 0 | 0 |
| 09:00 | 0 | 0 | 0 |
| 10:00 | 0 | 0 | 0 |
| 11:00 | 0 | 0 | 0 |
| 12:00 | 0 | 0 | 0 |
| 13:00 | 0 | 0 | 0 |
| 14:00 | 0 | 0 | 0 |
| 15:00 | 0 | 0 | 0 |
| 16:00 | 0 | 0 | 0 |
| 17:00 | 0 | 0 | 0 |
| 18:00 | 0 | 0 | 0 |
| 19:00 | 0 | 0 | 0 |
| 20:00 | 0 | 0 | 0 |
| 21:00 | 0 | 0 | 0 |
| 22:00 | 0 | 1 | 1 |
| 23:00 | 0 | 0 | 0 |
Busiest hour: 22:00 UTC, with 1 landing.
What it asks for
| Part of the site | Verified requests | Unverified requests |
|---|---|---|
| The island, on the home page | 0 | 1 |
Parts of the site, from a fixed list. The address of a request is never shown.
Its record with each rule
| Trailing 30 days | Verified | Unverified |
|---|---|---|
| Sessions | 0 | 1 |
| Requests | 0 | 1 |
| Sessions that asked for anything besides robots.txt | 0 | 1 |
| read robots.txt first | 0 | 0 |
| had read it in the 24 hours before | 0 | 0 |
| read it only afterwards | 0 | 0 |
| did not read it | 0 | 1 |
| Sessions that entered the Pitfall | 0 | 0 |
| having read robots.txt | 0 | 0 |
| Requests turned back at the vault, with a 403 | 0 | 0 |
| Highest Climb level reached | Has not climbed | Has not climbed |
| Private Beach visits while on the guest list | 0 | 0 |
| Turnstile waits kept | 0 | 0 |
| Turnstile waits broken | 0 | 0 |
| Requests over the rate limit | 0 | 0 |
| sessions that went over it | 0 | 0 |
These are counts, not a verdict. A session is judged on what it did inside the 30 days.
Customs
| Sessions that said they were this bot | Trailing 30 days |
|---|---|
| Passed: verified | 0 |
| Could not be checked: unverified | 1 |
| Not yet answered: unverified | 0 |
| Failed, and filed under Googlebot (unverified) | 0 |
Nothing a failed session did is counted anywhere else on this page.
Where it comes from
| Network owner | Country | Sessions | Requests |
|---|---|---|---|
| Internet Archive Canada | Canada | 1 | 1 |
Network owner and country only. An address is never shown.