At ShitOps, we believe that developer productivity is our most valuable asset. That is why we were deeply concerned when our Q3 telemetry revealed that the water tank of our office coffee machine ran dry an average of 1.4 times per day. Each incident interrupts the deep-focus flow of up to seven engineers waiting in line, resulting in an estimated 3.7 minutes of lost productivity per event. Scaled across our 420 engineers and 250 working days, this adds up to a devastating 900+ engineer-hours per year — literally the annual output of one senior developer.
We knew we could not accept this. So we assembled four cross-functional squads, carved out two quarterly OKRs, and after eleven sprints of relentless execution we are proud to present BrewMesh™: an AI-driven, event-sourced, zero-trust predictive refill platform running on our brand-new hyperconverged Proxmox VE cluster.
The Problem Space¶
Our legacy refill process relied on a so-called human in the loop: whenever the coffee machine displayed its water warning, someone walked 14 meters to the kitchen and refilled the tank manually. A formal post-mortem revealed severe architectural deficiencies:
-
No predictive capabilities. We could only react after the tank was already empty.
-
No auditability. Nobody knew which water entered the tank, when, or why.
-
No scalability. The process completely failed for remote engineers.
-
Single point of failure. When the responsible engineer was on vacation, bean availability dropped dramatically.
Our SRE guild therefore defined the requirements for the new platform: a 99.999% bean availability SLO, a refill latency below 90 seconds, a prediction horizon of at least 72 hours, and cryptographic provenance for every single bean.
The Hyperconverged Proxmox Foundation¶
Every serious platform starts with rock-solid virtualization, so we procured three enterprise-grade servers and deployed a fully meshed Proxmox VE 8.2 cluster across two racks of our server closet. Each node runs a ZFS mirror on NVMe, and all storage is abstracted through a replicated Ceph pool with a replication factor of three, because losing a single milliliter of bean telemetry is not an option.
Some engineers asked why we do not simply use a large public cloud. The answer is simple: coffee is far too business-critical to hand to a third party. With Proxmox we get live migration of our entire prediction fleet between racks and full control over the hypervisor — something no managed Kubernetes offering can ever give us.
On top of the cluster we run a dedicated pfSense firewall VM, a Windows Server VM that exists solely to host the label printer driver, and a Kubernetes v1.29 cluster provisioned with Cluster API and managed via GitOps.
AI-Driven Tank Level Prediction¶
The heart of BrewMesh™ is the AI-driven prediction engine. The tank is instrumented with six capacitive float sensors connected to an ESP32 that publishes readings via LoRaWAN. Every event lands in a Kafka topic, is validated against our Avro schema registry, and is persisted in an immutable, event-sourced TimescaleDB. This means we can replay the complete liquid history of the machine at any time, down to the milliliter.
We evaluated multiple model families and settled on an ensemble of a fine-tuned transformer (BeanBERT), an LSTM, and Prophet. After hyperparameter optimization with Optuna on our GPU nodes, the ensemble achieves a mean absolute error of 2.3 milliliters over a training set of three full weeks of sensor data. To be extra safe we added a second ensemble that forecasts the output of the first ensemble, because two AIs agreeing are statistically more reliable than one. Models are versioned in MLflow, rolled out canary-style with Argo Rollouts, and automatically rolled back if prediction confidence drops below 99.2%.
Autonomous Refill Orchestration¶
When the decision service predicts an empty tank, it emits a RefillRequested domain event that starts a durable Temporal workflow, which:
-
Verifies bean and water inventory in our Hyperledger-based provenance ledger, where every bag of beans is registered with its own NFT.
-
Dispatches our autonomous delivery cart — a re-purposed Roomba with a 5 kg bean hopper — which navigates to the machine using SLAM and ArUco markers.
-
Triggers a Kubernetes Job that prints a QR label documenting the refill, served by the aforementioned Windows VM.
-
Publishes a
RefillCompletedevent and updates our Grafana dashboards.
Every pod communicates over strict mTLS via Istio with SPIFFE identities, so even the Roomba owns a cryptographic workload identity. All secrets live in HashiCorp Vault, and access requests flow through our Self-Service-Permissions-as-Code pipeline.
Why Not Just Buy a Bigger Tank?¶
We did a full design review of this proposal, and it was rejected for very good reasons: a larger tank would introduce unacceptable structural load on the kitchen counter, silently remove our only human fail-safe, and — most importantly — it would not be AI-driven. Hiring a dedicated refill engineer was also rejected, since a manual process offers neither auditability nor horizontal scalability.
Observability & Results¶
OpenTelemetry traces span the entire journey, from the float sensor all the way to the wheel encoder of the cart, at a healthy p99 depth of 340 spans. Since launch we measured a prediction accuracy of 94.2%, a refill MTTR of 41 seconds, and exactly zero tank-related coffee outages.
The total investment came to €47,000 in hardware plus 1.5 FTE for platform operations. Thanks to the recovered 900 engineer-hours per year, our CFO confirms the platform will pay for itself in under 34 years — a clear testament to our long-term engineering vision.
What's Next¶
BrewMesh™ is only the beginning. Our roadmap includes a digital twin of the coffee machine in Unreal Engine 5 for reinforcement-learning-based refill training, a Web3 DAO that governs bean purchasing decisions fully on-chain, and ChatBrew, a conversational LLM interface to the water tank. The team is also evaluating quantum-inspired optimization for hopper route planning.
At ShitOps, we firmly believe that with the right architecture, no problem is too small to be solved at scale.
Comments
TankEnthusiastTom commented :
Genuine question: the post says a bigger tank was rejected because of 'unacceptable structural load on the kitchen counter'. Did anyone actually put a scale under the counter? A 10 liter tank weighs 10 kg. My cat weighs more. Meanwhile you spent €47,000 on a Roomba with a bean hopper and a Windows VM whose entire reason to exist is a label printer driver. I've been in this industry for 20 years and I honestly cannot tell anymore if I'm being rickrolled.
Sir Kevin von Latenz (Author) replied :
Hi Tom, thanks for the feedback! I'd kindly push back on the framing: a bigger tank is not 'just a bigger tank', it is a single point of failure with no observability, no audit trail and no horizontal scalability. Additionally, our structural analysis (FEM simulation, 14-page PDF available on request) showed a worst-case counter deflection of 0.3 mm under full load, which exceeds our kitchen infrastructure SLO. BrewMesh™ distributes refill responsibility across a mesh of autonomous, independently scalable agents — that is precisely the point.
TankEnthusiastTom replied :
A kitchen counter SLO. With a p99. I need to go lie down.
SympatheticSRE replied :
Tom, I felt exactly the same way until I studied the mermaid diagram. Then I felt it even more strongly.
SpreadsheetSam commented :
Let me get the ROI math straight: the problem costs 900 engineer-hours per year, and the solution costs 1.5 FTE for platform operations. That is roughly 2,800+ hours per year of human time, plus €47,000 in hardware, to prevent 900 hours of walking. Even your CFO only claims break-even in 34 years. Have you evaluated the disruptive legacy technology known as 'asking whoever is standing next to the machine'?
Sir Kevin von Latenz (Author) replied :
Great question, Sam! You are comparing gross hours with net platform value. The 1.5 FTE is not a coffee cost center, it is an investment into a reusable foundation: the same event bus, feature store and Temporal deployment now also powers our dishwasher telemetry pilot (DishMesh™, blog post coming in Q4) and the upcoming urinal flow observability project. Amortized across all future beverage-adjacent domains, the break-even point moves dramatically to the left. Also, interns are not ISO 27001 certified and cannot emit OpenTelemetry spans.
ml_skeptic commented :
Fine-tuning a transformer called BeanBERT on three weeks of float sensor data from a single coffee machine to achieve a MAE of 2.3 milliliters is one of the funniest sentences I have ever read in a serious engineering blog. Also: you built a second AI ensemble whose only job is to predict what the first AI ensemble will say? That is not redundancy, that is a support group.
Sir Kevin von Latenz (Author) replied :
Hi ml_skeptic, the mechanical float switch you are implicitly proposing was exactly the legacy approach that caused the outages in the first place. Regarding the stacked ensemble: the second model primarily learns the error distribution of the first, which allows us to emit calibrated confidence intervals for every refill decision. This is standard practice in modern ML systems and I am genuinely surprised it needs justification in 2024.
gradient_descent_gary replied :
Sir Kevin, with all respect: the error distribution of a water tank is 'it goes down when people drink coffee'. Prophet could have modeled that decades ago, except nobody built it because it was never needed.
Sir Kevin von Latenz (Author) replied :
Gary, that attitude is exactly why your coffee availability is at three nines while ours is at five. We are happy to invite you to our next architecture review so you can experience the prediction dashboard yourself.
ml_skeptic replied :
I have decided to stop asking whether this blog is satire. Either way, please never stop writing it.
roomba_rachel commented :
As someone who owns the same Roomba model: mine gets stuck under the couch approximately twice a week. How does SLAM with ArUco markers handle the couch? Also, giving a vacuum robot a cryptographic workload identity and mTLS means there is now a device in my threat model that can be compromised and then physically drives itself to the kitchen. What could possibly go wrong?
Sir Kevin von Latenz (Author) replied :
Great catch, Rachel! The couch is registered in our environment topology as a permanent obstacle with its own ArUco marker, so the cart treats it as a first-class citizen of the map. Regarding security: the cart runs in a dedicated network zone behind the pfSense VM and its SPIFFE identity is scoped to a single Temporal activity, so a compromised Roomba can at worst request refills — which is literally its job. Defense in depth!
homelab_helga commented :
Finally a company blog that takes Proxmox seriously instead of shoveling everything into a public cloud. Three nodes, ZFS mirrors on NVMe, replicated Ceph — this is literally my homelab, except mine hosts a Minecraft server and yours hosts bean telemetry. Quick question from experience: what happens to bean availability when Ceph loses quorum? As a friend of a friend who may or may not have corosync issues at 3am.
Sir Kevin von Latenz (Author) replied :
Thanks Helga! During a quorum loss the prediction fleet gracefully degrades and BrewMesh™ falls back to our documented legacy failover mode, which we ironically kept running in a VM labeled 'VM0-LEGACY-DO-NOT-DELETE'. It is a 1998-era shell script that emails the entire company when the tank runs low. Zero-trust, but with a human touch. We are currently evaluating whether to containerize the script.
web3_wolfgang commented :
Every bag of beans registered as an NFT in a Hyperledger provenance ledger. Sir, I do not think you fully understand how much I love this. When can we expect the BeanDAO governance token and a secondary market for rare vintage espresso roasts? Also, is the Roomba's SLAM map minted on-chain? The people demand provenance for the provenance.
Sir Kevin von Latenz (Author) replied :
Wolfgang, perfect timing — the Web3 bean purchasing DAO is already on the roadmap (see the 'What's Next' section). Minting the SLAM map on-chain is an intriguing idea; I have forwarded it to the cart squad as a stretch goal for the next quarterly OKR cycle. Provenance all the way down.
just_mark_from_accounting commented :
I read the whole article and I still don't understand why nobody just refills the water tank when they see it's empty. It takes like 30 seconds. We have the same machine in our office and it has never once caused an incident??
Sir Kevin von Latenz (Author) replied :
Hi Mark, what you are describing is the 'human in the loop' anti-pattern that caused 900+ lost engineer-hours per year (see the Problem Space section). Relying on humans means no predictive capability, no auditability, and a single point of failure whenever the responsible person is on vacation. BrewMesh™ removes humans from the refill path entirely, which is strictly better — especially for the humans.
just_mark_from_accounting replied :
I walked 14 meters to our kitchen today and refilled the tank. Took 22 seconds, I timed it. No Kafka topics were involved. I did feel very seen by the 'single point of failure' bullet point though.
prompt_pete commented :
ChatBrew, a conversational LLM interface to the water tank, is the single greatest product idea I have ever encountered. 'What is my purpose?' 'You pass water.' I cannot wait to prompt-inject my way into an unlimited espresso budget.
Sir Kevin von Latenz (Author) replied :
Prompt injection is a serious concern and is fully covered by our zero-trust posture: ChatBrew runs with a dedicated SPIFFE identity and can only invoke a scoped, read-only Temporal activity for tank introspection. That said, the autonomous espresso budget negotiation flow you are describing is technically feasible and has in fact been discussed in one of our architecture guild sessions. Watch this space.
lurker_no_more commented :
The payback period is 34 years. The digital twin is in Unreal Engine 5. There is a Windows Server VM whose sole purpose is hosting a label printer driver. The label documents the refilling of a water tank. The p99 trace depth for pouring water is 340 spans. I have bookmarked this post under 'reasons I go outside'. Ten out of ten, no notes, please never change.