The Art of Not Buying Disk
I have released rpkv v0.1.0: key-value reads over Redpanda topics without duplicating values. That is the whole pitch. “Without duplicating values” is not a feature. It is a grudge.
The economics of 2026 hardware
RAM is currently priced like it is mined by hand. NVMe went up with it, out of solidarity. The industry’s advice has not changed: buy more anyway. Storage is cheap, say the articles, all of them written when it was.
My homelab didn’t get the memo. It has the disk it has. Everything I care about lives in Redpanda topics; if it isn’t in a topic, it doesn’t exist. And I’m not buying disk to store the same bytes twice.
The problem
A compacted topic already holds the latest value for every key. What it can’t do is hand you the value for key X. The log knows. You just have to replay the whole thing to ask.
Every respectable answer is a copy. Kafka Streams copies every value into RocksDB. A cache copies it into RAM, see above. Fine engineering, all of it, billed in gigabytes I have already spent.
The grudge, implemented
rpkv is a pure-Go sidecar that keeps only an index: key -> (partition, offset), in embedded Pebble.
That’s 14.2 bytes per key, measured, whether the value is a heartbeat or
a novel. A read fetches that one record from the log over the plain Kafka
protocol. The values stay where they were. The index is a projection:
delete it and it rebuilds from the topic at about 996k keys per second,
which is faster than I can regret it.
The trade-off is printed on the tin. Every read pays a broker round-trip: 11.1 ms p50, measured. If you need microseconds, buy the RAM. Big values, modest read rates, no disk budget: that’s the sweet spot, and it’s the shape of my rack.
Compaction and retention are handled by verification. A fetch is never trusted blindly. A superseded pointer waits for the index to catch up. When retention deletes a record, rpkv answers 410 and moves on. The topic is the system of record; if the log decided those bytes were not worth keeping, that’s the log’s call. Details are in the repo, which for once has more tests than excuses.
On the icon
ChatGPT drew the logo: a little panda, which is what the Chinese call the red one, sitting on a stack of red boxes you have seen before. Infringing one trademark is carelessness. Infringing two in one icon is curation. Any resemblance is a matter between my lawyer and theirs. Only one of them exists.
None of this is new. Rick Houlihan called it data pressure in his re:Invent 2018 talk: data presses against the technology holding it until something gives, and what gives is usually somebody too stubborn to buy more hardware. rpkv is my small entry in that tradition. I hope it’s useful to somebody else’s stubborn rack.
So: go install, point it at a topic, and stop paying for your data
twice.