Pull, do not log in

Step 160 minutes

By the end of this step

You finish this step with one scheduled task that pulls data out of one system, into a directory you own, with nobody opening a browser.

Words used in this step Pull · API · Credential · Timer · Pagination · Directory
Pull
A small job that asks a system for its data and saves a copy somewhere you own. Nobody logs in and nobody copies anything by hand.
API
The way one piece of software lets another ask it for data. Your pull uses it with a key, instead of a person reading a screen.
Credential
The key your pull shows to prove it is allowed in. Give this job its own, able to read and nothing more, so you know what stops when you change it.
Timer
A setting that starts a job at fixed times with nobody pressing anything. The line in this step that starts with 23 runs the pull at 23 minutes past every hour.
Pagination
Systems hand data over a page at a time. A pull that only asks for the first page gets part of the data, and the part looks complete.
Directory
A folder on a computer. Here it is one folder that only this job writes into, so every file in it came from the pull.

You open six tabs every Monday and copy the numbers into a sheet. It takes half an hour, it works, and for a year nobody has a reason to touch it. Then you take a fortnight off. The sheet has a gap in it now. Nobody spots the gap. The sheet still opens and still looks like a sheet. A missing fortnight is shaped exactly like a flat fortnight once it has been typed into a row.

Those numbers were never collected. They were remembered, and memory takes holidays.

Logging in is not collection. It is a person doing a machine's job, weekly, from memory, with no record of what they skipped.

Concede the obvious first. Opening a tab today is faster than writing a pull that works. It stays faster right up until it does not. By then the sheet has habit behind it, and somebody will defend it in a meeting.

Logging in every Monday

  • Six tabs opened and copied into a sheet
  • A fortnight off leaves a gap in the rows
  • The gap looks exactly like a flat fortnight
  • No record of what got skipped

A pull on a timer

  • A script fetches the data every hour
  • Files land in a folder only the job writes to
  • Its own read-only key, named after the job
  • Anything not pulled yet is marked unverified
Logging in only works while somebody remembers. A pull on a timer keeps collecting when everybody is away.

The rules

Start with the system you would miss first, not the one you find most interesting. Money in, usually. You will run this pull every day for years, so pick the feed whose absence you would feel by Wednesday. Interesting systems make good demonstrations and bad foundations.

Pull on a schedule you set once, not when you remember. A schedule set in advance is a decision. A pull you trigger when you want a number is the six tabs again, wearing a script.

An export on a timer beats an API you never finish. Not every system has a usable API, and not every API is worth the week it costs. The pure version of this step is the version that never ships. A nightly CSV landing in a directory is a feed. Ugly is allowed.

Give the pull its own credential, read-only, named after the job. Not your login. Not the shared admin account, not a key pasted into a script three people edit, not a token nobody can find the owner of. When you rotate it you want to know exactly what stops.

Write down what you cannot pull today and mark it unverified. Unverified is not zero. A system you have not connected is a hole you can see. A hole you can see beats a summary that quietly leaves it out.

Do it now

Pick one system. Read its auth page and nothing else that site has to offer.

Make a directory that only this job writes into.

mkdir -p /srv/hub/landing/billing

Now write the smallest thing that fetches one page and saves it under a timestamp.

#!/usr/bin/env bash
set -euo pipefail
stamp=$(date -u +%Y%m%dT%H%M%SZ)
out=/srv/hub/landing/billing/$stamp.json
curl -sS --fail \
  -H "Authorization: Bearer $BILLING_TOKEN" \
  "$BILLING_API/charges?limit=100" > "$out.part"
mv "$out.part" "$out"

Write to a part file and rename it at the end. A half-written file that looks finished is worse than no file.

Put it on a timer.

23 * * * * /srv/hub/bin/pull-billing.sh >> /var/log/hub/pull-billing.log 2>&1

Now deal with the second page. A pull that silently stops after page one still returns valid data. Every time you look at it. That is what makes it so hard to spot. Read the source's own pagination rules. Loop until it tells you there is nothing more to fetch.

A worked example

Say you run three cafés under one name. Sales go through the till system. The staff rota sits in a scheduling app. Supplier invoices land in the accounts package. Every Monday somebody logs in to all three and types the numbers into a sheet.

Money in is the feed you would miss by Wednesday, so the till goes first. The other two wait, written down as unverified.

Worked example

A three-site café group

The till pull, set up on a machine the group already runs, with nobody logging in.

Feed:        till sales, all three sites, one system
Directory:   /srv/hub/landing/till
Script:      /srv/hub/bin/pull-till.sh
Timer:       23 * * * * /srv/hub/bin/pull-till.sh >> /var/log/hub/pull-till.log 2>&1
Credential:  read-only key called pull-till, used by this job and nothing else
Unverified:  staff rota, not pulled yet
Unverified:  supplier invoices, not pulled yet
$ ls -l /srv/hub/landing/till/
-rw-r--r-- 1 hub hub 61204 Sep 21 10:23 20260921T102301Z.json
-rw-r--r-- 1 hub hub 58817 Sep 21 11:23 20260921T112302Z.json
Records in the newer file: 124. Page limit: 100.

What to notice

Two files, two timestamps, both non-zero. The newer one holds more than one page, so the loop past page one works. The rota and the invoices stay on the list as unverified until they get pulls of their own.

Nobody opened a browser to get any of it.

The failure you will hit

The credential expires and the job keeps running. The script exits and the log fills with a 401 nobody reads. The directory stops gaining files while every downstream number carries on looking reasonable.

Fix it at the source. Keep the non-zero exit the script already has on a bad status. Treat an empty pull as a failure rather than as a quiet day. Step 5 builds the check that shouts. Until then, check the directory by hand, every Monday.

Not for you if

You are not willing to own something that runs while you sleep. Stay with the six tabs, and accept that your numbers stop the week you do.


A pull you can run twice beats a login you have to remember. The Vault is the paid library that holds the whole of it, and the price is on the page.

Done means

Leave it for two scheduled runs, then run ls -l /srv/hub/landing/billing/ and see two files, different timestamps, both non-zero. Count the records in the newer one. If that count is exactly your page limit, pagination is not working and you have one page pretending to be a feed.

I keep this tick in this browser and send it nowhere. It goes when you close the browser.

Untick it if the check stops passing.

This step files under Data & Analytics in the Vault.

Open Data & Analytics
Three people talking with drinks in hand at a daylit networking event

Networking

Included in Membership

The next step is included in Membership.

This step is open to everyone. The rest of the track is not.

Step 2, Land it raw before you tidy it

What this step installs

You finish this step with source data sitting unchanged in one table, and every tidy number you read built downstream from it.

75 minutes

It hands off to Data & Analytics in the Vault, which the same membership opens.The track list stays open to anyone.

Included in Membership — £79/month.

You sign in first, so the membership lands on the right account.

Your ticks in this browser stay. They come with you when you join.