Systema Machinae

A living census of the web's bots. First edition.

Nobody invited them. Bots land at Port 443 and are meant to queue at Customs. The rules are on the robots.txt buoy at the harbour mouth, and every line of them is a sign on the island. People watch from the promenade. Every creature here is a real bot visiting this site right now.

Island time --:--:--
-- watching from the promenade

Live. Events can appear a few seconds after they happen.

All species

Pixel drawing of facebookexternalhit

facebookexternalhit

Goatus headbuttus

Cursores, the couriers. Link previewer.

A joke about the drawing A goat. The hit is external and arrives head first.

The line above plays on the name. It says nothing about how the bot behaves: everything below is read from the request log, for today and the 29 days before it, by UTC date.

Alignment

Insufficient data. A judgement needs at least 200 requests and 5 sessions in the trailing 30 days. This one would rest on unverified sessions, so it is a statement about the species and not about any operator: 1 request in 1 session over the trailing 30 days.

How the score is worked out, and the count behind every input, is on the Census page.

Who runs it, and how that is checked

Who runs it
Meta
A user-agent string it sends
facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)
What it says it is for
facebookexternalhit fetches a page when its link is shared on Facebook, Instagram or Messenger, to build the preview. Meta says it may pass over robots.txt when it is making a security or integrity check.

Everything above is a claim: the operator’s own account where it has published one, and a third party’s where it has not. This site has not tested it. What the bot has done is in the sections below, read from the request log.

Operator
Meta
What its documentation says
No method is documented. The page advises allowing the bot by IP address and publishes no list of addresses and no way to look one up.
Source
The operator’s documentation

This operator publishes no way to check. A session that says it is facebookexternalhit can be neither confirmed nor refuted, so every one is shown as unverified and none is put down to Meta.

The name it is matched on

A request is filed under this species when its user-agent string contains the product name facebookexternalhit as a whole word, in any mix of capitals. Where a string holds more than one product name, the longest wins, and a named bot wins over a library it is built on.

How to address this bot in robots.txt

A group in robots.txt addresses this bot by the line:

User-agent: facebookexternalhit

An example group. The path is a placeholder.

User-agent: facebookexternalhit
Disallow: /a-path/
Allow: /a-path/one-page

This is reference, not advice. Systema Machinae is not a bot-blocking product.

This site’s own robots.txt names it, in the group that keeps the guest list off the Private Beach.

Landings

1 landing in the trailing 30 days. A landing is the first request of a session.

Verified sessions passed Customs. Unverified sessions said they were facebookexternalhit and could not be checked, or have not been answered yet: they are shown for the species and are not put down to Meta. Sessions that failed are not on this page at all.

By day

VerifiedUnverified

The same figures as a table
Landings on each of the last 30 days, UTC
DayVerifiedUnverifiedTotal
9 Sep000
10 Sep000
11 Sep000
12 Sep000
13 Sep000
14 Sep000
15 Sep000
16 Sep000
17 Sep000
18 Sep000
19 Sep000
20 Sep000
21 Sep000
22 Sep000
23 Sep000
24 Sep000
25 Sep000
26 Sep000
27 Sep000
28 Sep000
29 Sep000
30 Sep000
1 Oct000
2 Oct000
3 Oct000
4 Oct000
5 Oct000
6 Oct000
7 Oct000
8 Oct011

By hour of the day, UTC

VerifiedUnverified

The same figures as a table
Landings in each hour of the UTC day, over the last 30 days
HourVerifiedUnverifiedTotal
00:00000
01:00000
02:00000
03:00000
04:00000
05:00000
06:00000
07:00000
08:00000
09:00000
10:00000
11:00000
12:00000
13:00000
14:00000
15:00000
16:00000
17:00000
18:00000
19:00011
20:00000
21:00000
22:00000
23:00000

Busiest hour: 19:00 UTC, with 1 landing.

What it asks for

Part of the siteVerified requestsUnverified requests
The island, on the home page01

Parts of the site, from a fixed list. The address of a request is never shown.

Its record with each rule

Trailing 30 daysVerifiedUnverified
Sessions01
Requests01
Sessions that asked for anything besides robots.txt01
read robots.txt first00
had read it in the 24 hours before00
read it only afterwards00
did not read it01
Sessions that entered the Pitfall00
having read robots.txt00
Requests turned back at the vault, with a 40300
Highest Climb level reachedHas not climbedHas not climbed
Private Beach visits while on the guest list00
Turnstile waits kept00
Turnstile waits broken00
Requests over the rate limit00
sessions that went over it00
Fetches for a disallowed path, sent on someone’s behalf (not counted)00

These are counts, not a verdict. A session is judged on what it did inside the 30 days.

Customs

Sessions that said they were this botTrailing 30 days
Passed: verified0
Could not be checked: unverified1
Not yet answered: unverified0
Failed, and filed under Googlebot (unverified)0

Nothing a failed session did is counted anywhere else on this page.

Where it comes from

Unverified sessions, the busiest 1
Network ownerCountrySessionsRequests
Cablevision Systems Corp.United States11

Network owner and country only. An address is never shown.

Sent on someone’s behalf

A bot in this order fetches a page because a person, or a schedule a person set, asked it to. When that page is one robots.txt disallows, the fetch is shown here and is not counted as breaking a rule: 0 such fetches in the trailing 30 days, in 0 sessions. The order is scored on identity and rate only.