A Job Board That Remembers
I’ve put jobctl.net online. It’s a job board that keeps the history other job boards throw away. Postings are polled, reconciled against the last look, and every difference is written down: appeared, changed, went missing, closed, came back. It’s small and it’s opinionated. It exists because of two grudges and one curiosity.
Grudge one: they took RSS away
Browsers used to understand feeds. There was an icon in the address bar, you clicked it, and from then on the site came to you. Firefox removed live bookmarks in 2018. Chrome never cared. Safari still has an RSS button, though I don’t think anyone at Apple remembers why. The web didn’t stop publishing feeds. Most of it still does. It just stopped telling you.
What replaced the feed was the tab. I had fifty of them open, each one a site that would have pushed its changes to me if anyone had left the door open. So the first thing jobctl is, under everything else, is a feed reader that remembers. It polls, it keeps what it saw, and it hands the result back as a feed. Every filter on /jobs travels in the URL, and the URL is an Atom feed. Paste it into any reader and it stays current. It needs no account, no app and no notification permission dialog.
Grudge two: hunting for a job in 2026 is a tab marathon
This started when a couple of friends asked me to help them look. They’re mid-level and junior, which is a separate post about who the industry has decided it no longer hires. I said yes and found out what the procedure actually is. You open Greenhouse for company A, Lever for company B, a Rippling board for C, four aggregators that scrape each other (we’re no different, I know), and two national boards behind a login wall. Then you do it again on Thursday, because the one from Tuesday is gone and nobody tells you when it went.
The postings are public and the APIs are documented. Greenhouse alone publishes a JSON board for every company that uses it. There’s no reason a human should be the loop that walks them, so jobctl walks them: 91 sources today, 57 Greenhouse boards, 28 feeds, plus Manfred, Rippling and JobFluent, listed with their health on /sources. The number goes up whenever I have an evening, and there’s a suggest a source page for that. If it’s a public feed or a documented API and its terms allow it, it goes in. Anything that has to be tricked into responding stays out. Conditional requests, a real User-Agent with a contact URL, per-source rate limits and backoff. A failing source gets recorded as failing and left alone until its backoff expires.
The curiosity: what do companies do with their postings?
This is the part I care about most and the part that’s least finished.
A job board shows you a snapshot: what’s open right now. It can’t tell you that a posting has been reopened five times in eight months, that its salary range vanished after week one, or that it’s been “urgently hiring” since March. That information exists. It’s just never kept. Once you keep it, a few questions that are currently folklore become answerable:
- Which postings flap, close, reopen, close, reopen, and how often?
- Which companies keep an evergreen “Senior Engineer” open for a year with no visible change? That’s what CV harvesting looks like from outside.
- Do salary ranges get narrowed after publication? Removed?
- How long does a role at company X stay open, on median?
jobctl is built to make those queries possible later without lying now.
The core is an append-only event log: ItemObserved, ItemSeen,
ItemChanged, ItemMissing, ItemClosed, ItemReopened, plus
per-source fetch results. The vocabulary is about what I observed, not
about what a publisher did. A feed never says a job was deleted; it just
stops showing it. So “closed” is a conclusion reached after N consecutive
misses, N per source, never a fact anyone reported. I tried closing on
the first miss and got a timeline full of jobs dying and coming back that
had never gone anywhere, which is worse than no timeline at all. The
reconciler that decides this is a pure function,
(previous, snapshot) -> []Event, with no I/O, and that’s what lets me
property-test it: replay is deterministic, an item present in every
snapshot never closes, an item missing for exactly the threshold closes
once and only once, a reopening is always visible in the timeline.
Every observation also keeps the source’s original bytes. My normalisation rules are wrong today in ways I haven’t found yet, and the raw payload is the only thing that lets me re-derive them when I do.
Boring on purpose
Dan McKinley’s Choose Boring Technology gives every project a handful of innovation tokens. jobctl spends none of them on the stack, so it can spend all of them on the one idea above.
It’s one Go binary and one Postgres. The event log is a table. The current view and the timeline are projections replayed from it, so a bad mapping gets fixed by fixing the code and replaying. There is no migration to write against data I no longer trust. Pages are HTML rendered on the server. The JavaScript budget is 41 lines, there to make the feed URL select itself on focus. There’s no queue, no cache tier, no Kubernetes: a compose file on a machine in my homelab, nginx in front, and deploy is build-and-ship. Subscriptions go out over Atom, plus a Telegram bot for people who’d rather be pinged than poll. I made every one of those choices so that when a source breaks at 03:00, and sources have opinions about JSON, the failure is in the one place I haven’t seen before and never in the plumbing.
What it does today
- /jobs: one list across every source, filterable by text, company, source, status, remote policy, “states a salary”, minimum completeness and posted-after. Titles are shown in full. Salary is rendered as the source stated it (currency code, period, “up to X” when only a ceiling was given) and never guessed.
- A completeness score on every posting: what fraction of the fields a candidate needs (salary, location, remote policy, employment type, description) the source filled in. Right now 9,429 of 17,011 open postings state a salary. That ratio is itself a finding.
- A timeline per posting: first observed, every field change with before and after, every miss, closure and reopening.
- /sources: every source with its health, last success, consecutive failures and item count, in the open. If something’s broken you can see it before I do.
- Feeds: every filtered list is an Atom URL.
- A Telegram bot for the same subscriptions.
- Suggest a source: a form. I read them.
What it doesn’t do yet is the analysis: the flapping report, the “posted N times” badge, the salary-drift chart. The recorded history is five days old as I write this, and that’s on purpose. I threw the previous history away once, before launch, because it had been parsed by mappings I was in the middle of rewriting, and I’d rather start the clock on data I trust. The honesty signals come when there’s enough of it to say something true.
Where it goes
More sources, one evening at a time. A proper place filter: remote, or a continent, or a country or list of countries. Right now “where” is whatever string the source wrote, and that’s not a filter, it’s a search box in a costume. Cross-source deduplication, the same job on four boards, which I kept out of the identity model on purpose so it could be designed on its own terms. And then the analysis above, once the log has a few months in it.
If you’re looking, or helping someone who is, I hope it saves you the tab marathon. If you run a public feed or a board on a documented API and want it in, the suggest page is there. And if you find a bug, and you will, the /sources page shows the same thing I see. I’d rather hear about it than not.