Every great platform at ShitOps starts with a real business problem. Today we are excited to share how our Platform Reliability Guild solved one of the most pressing challenges of our time: guaranteeing the battery state of our corporate iPhone fleet while satisfying the strictest audit requirements in our history.
The Problem: iPhone Availability as a Compliance Requirement¶
ShitOps operates a globally distributed fleet of 4,732 company iPhones. They are the primary paging device for our 1,200 on-call engineers. Our compliance team recently extended our SOC 2 Type II and ISO 27001 control framework with a new set of requirements:
-
Every on-call engineer must be reachable on their company iPhone at all times.
-
Every iPhone must hold a state of charge (SoC) of at least 30% during on-call shifts.
-
Every battery state change must be recorded in an immutable, cryptographically verifiable audit trail.
-
Any predicted violation must be remediated automatically before it occurs.
Our previous solution was a Slack bot that asked engineers twice a day whether their phone was charged. During the last audit this was rightfully rejected: a centralized, human-confirmed status report cannot fulfill the evidentiary requirements of a modern enterprise. We needed real, engineered trust.
Deconstructing the Requirements¶
We ran a two-week requirements engineering workshop with 47 stakeholders and distilled five architectural requirements:
-
Real-time telemetry with a p99 ingestion latency below 50 ms.
-
Immutability - no actor, including platform admins, may alter historical battery states.
-
End-to-end verifiability - every reading must be signed at the source.
-
Predictive remediation - violations must be prevented, not reported.
-
Zero trust - every hop must be mutually authenticated.
Commercial MDM products were evaluated and rejected: their centralized relational stores are single points of failure and fundamentally cannot provide the cryptographic immutability our requirements demand. So we did what any serious engineering organization would do: we built our own platform. We call it ChargeChain.
The ChargeChain Architecture¶
The iPhone Edge Layer¶
We developed BatteryAgent, a native Swift app deployed via our internal MDM. Every 250 ms it samples UIDevice.batteryLevel, batteryState and thermal values, serializes them into protobuf, signs the payload with a P-256 key held in the Secure Enclave, and publishes it to an EMQX broker cluster. EMQX runs as a five-replica StatefulSet with PodDisruptionBudgets across three availability zones, because a battery reading is only as reliable as the broker that carries it.
Ingestion and Stream Processing¶
Events flow into a 42-partition Kafka cluster running in KRaft mode. A fleet of Apache Flink jobs enriches each reading with contextual signals: the engineer's calendar (video calls drain batteries), office Wi-Fi RSSI, local weather (cold reduces capacity) and the currently foregrounded app. Enrichment matters - a raw SoC number without context is just data, not insight.
The Distributed Ledger Core¶
This is the heart of ChargeChain. Enriched states are committed to a permissioned Hyperledger Fabric network on the channel battery-channel. Our chaincode battery-escrow, written in Go, validates signature, monotonicity and device identity before a transaction is endorsed. A Raft ordering service with five orderers across eu-central-1, us-east-1 and ap-south-1 guarantees consensus even if an entire region burns down.
For defense in depth we anchor the Merkle root of every 60 second ledger block to Ethereum mainnet. Monthly gas costs are approximately 23,000 USD - a negligible price for immutability that is verifiable by any auditor on the planet. Our auditors were moved to tears.
Predictive Remediation with BatteryOperator¶
Ledger events are streamed into a graph neural network (PyTorch Geometric, served via KServe) that predicts the SoC at the start of each engineer's next meeting. A custom Kubernetes operator built with kubebuilder reconciles our custom resource BatteryPolicy: if a predicted SoC drops below 20%, the operator emits a ChargingTicket CRD and calls our internal logistics API, which dispatches a power bank to the engineer's desk via our autonomous office drone fleet. No human interaction required.
Observability and Access¶
All components emit OpenTelemetry traces into our Tempo-backed pipeline. Our dashboard ChargeGraf is built on a federated GraphQL gateway (Apollo) and streams live SoC data over WebSockets. Everything runs on Kubernetes, everything is deployed via ArgoCD, everything is mutually authenticated by Istio.
Results¶
After six months in production we measured a p99 ingestion latency of 43 ms, a ledger uptime of 99.999%, and a 97.4% reduction in dead-iPhone paging incidents. Our SOC 2 audit closed with zero findings related to device availability requirements. Total platform cost: eleven Kubernetes clusters, nine microservices and 23,000 USD of monthly gas fees. Given the criticality of our on-call availability, this is a bargain.
Outlook¶
We are currently prototyping ChargeChain Coffee - extending the distributed ledger to our espresso machines to enforce immutable, on-chain coffee freshness requirements - and exploring NFT-gated power bank ownership. The future of enterprise iPhone telemetry is decentralized, and at ShitOps we are already there. Stay tuned for the follow-up post in which we migrate the ledger itself onto the blockchain.
Comments
Dave W. commented :
Am I the only one who thinks this could have been a cron job? You literally mention deploying BatteryAgent via your internal MDM - which almost certainly already exposes battery level per device through an API. Poll it every five minutes, write to an append-only log in WORM storage, done. Instead we got eleven Kubernetes clusters, a permissioned blockchain, and $23k/month in Ethereum gas fees to monitor... phone batteries. Please tell me this is satire.
ChainBeliever replied :
This is exactly the kind of centralized thinking that fails audits. If a single admin can modify the log, there is no trust. Read the requirements section again - immutability means NO actor, including platform admins.
Chadwick Ledgersworth III (Author) replied :
Hi Dave, thanks for engaging! You are right that we deploy BatteryAgent via MDM - but consuming telemetry from the MDM API would violate requirement 3 (end-to-end verifiability), because the MDM is a trusted third party. Trusting the MDM means trusting every MDM admin, and our auditors were explicit that this is unacceptable. Additionally, our MDM vendor documents a 15-minute staleness window on the battery endpoint, which fails requirement 1 by a factor of 18,000. Happy to walk through the full threat model in a follow-up post.
Dave W. replied :
A 15-minute staleness window... for a battery that discharges over 14 hours. But sure, a factor of 18,000. Meanwhile my cron job and a signed CSV would have passed the audit for exactly $0 in gas fees.
Sarah (iOS dev) commented :
Genuine technical question: how are you sampling UIDevice.batteryLevel every 250ms? That API updates in 5% increments, is heavily cached by iOS, and once the app is backgrounded - which it will be, since these are paging devices - you get almost no background execution time. What am I missing here?
Chadwick Ledgersworth III (Author) replied :
Great catch, Sarah! We do not rely on batteryLevel alone - BatteryAgent additionally reads IOKit power source data through our internal entitlement framework, and we maintain a continuity-of-execution session via background location updates. The full iOS specifics deserve their own post: '250ms SoC telemetry on iOS: taming background execution.' Stay tuned!
Sarah (iOS dev) replied :
Background location updates... to read the battery level. App Review is going to love that one.
FormerBig4Auditor commented :
I spent nine years running SOC 2 Type II and ISO 27001 audits. I have never once required, requested, or even hinted at a distributed ledger. 'Immutable audit trail' is satisfied by signed append-only logs in WORM storage with proper access controls. What I actually care about is whether controls operate effectively - and a system where a GNN predicts a meeting will drain a phone and a drone delivers a power bank is going to generate pages of exceptions, not zero findings. Also, auditors were moved to tears? We cry for entirely different reasons.
Chadwick Ledgersworth III (Author) replied :
Thank you for your perspective! We actually evaluated QLDB during the requirements workshop and rejected it in phase 2 of the stakeholder matrix: Amazon is a centralized party, and trusting AWS means trusting 3.5 million employees with our battery states. Anchoring to Ethereum gives us a trust anchor that sits outside our org chart. And for the record, our audit partner did say - and I quote - 'this is the most thorough evidence package I have ever seen.' The tears were real.
FormerBig4Auditor replied :
I guarantee you what happened is the audit partner saw 'Ethereum' in the evidence package and decided billable hours on this engagement were going to be spectacular.
eco_sre commented :
Love the irony here: eleven Kubernetes clusters, five Fabric orderers, a Flink fleet, a GNN inference service, and an autonomous drone fleet - consuming orders of magnitude more energy and money than the 4,732 phone batteries they monitor - all to keep those batteries above 30%. The platform has a worse energy budget than the thing it is protecting. Has anyone computed kWh per battery percentage point?
ProofOfStakeStan replied :
Actually, since the Merge, Ethereum anchoring is Proof of Stake, so the marginal energy cost is basically zero. Please do the math before posting. The 23k is staker fees, not energy.
eco_sre replied :
Cool, so it is only $23k/month AND a drone fleet AND eleven clusters. My point stands, with extra steps.
Chadwick Ledgersworth III (Author) replied :
Actually we did compute this! The Platform Reliability Guild maintains a sustainability dashboard in ChargeGraf. I am happy to share that the drone fleet runs on our office solar array, making predictive remediation carbon-neutral. The espresso machines in ChargeChain Coffee are next.
Priya (CTO, 12-person fintech) commented :
This is inspiring! We are a 12-person fintech and our investors keep pushing us toward SOC 2. Where should we start - Hyperledger first and add Ethereum anchoring later, or mainnet from day one? Also, is BatteryAgent licensable?
yagni_andy replied :
Priya. Please. Buy your team phones with good batteries, enable Low Power Mode alerts, and spend the saved $23k/month on actual product. A Slack bot asking twice a day was fine and everyone knows it.
Chadwick Ledgersworth III (Author) replied :
Hi Priya! BatteryAgent is part of the ShitOps internal platform and not licensable yet, but we are exploring a ChargeChain SaaS offering. My advice: go mainnet from day one. Retroactively migrating your Merkle roots is precisely the kind of centralized-party risk our architecture is designed to eliminate. Watch this space for the SaaS announcement!
privacy_pete commented :
So you are collecting employees' calendar contents, foregrounded apps, Wi-Fi-based location, and local weather, correlating it all on an immutable ledger, and anchoring it to a public blockchain... forever? Did legal actually sign off on the GDPR implications? The phrase 'the engineer's calendar (video calls drain batteries)' is doing a LOT of heavy lifting in this post.
Chadwick Ledgersworth III (Author) replied :
Great question! All enrichment happens against pseudonymized device IDs, and every on-call engineer signed the telemetry consent addendum during onboarding (clause 14.3 of the on-call policy). Additionally, only Merkle roots are anchored to Ethereum - no personal data ever leaves the ledger - which our DPO confirmed is fully compliant.
privacy_pete replied :
Pseudonymized device IDs plus publicly verifiable drain predictions derived from calendars is still personal data under GDPR. Clause 14.3 does not magically fix that, and 'immutable forever' is the exact opposite of the right to erasure. Good luck with the DPA requests.
ComplianceCarla replied :
As someone on the compliance side: this is why the audit trail should have been a signed log with a retention policy. 'The blockchain ate our deletion obligations' is not a defense I want to write into an incident report.
sre_anna commented :
Congrats on the 43ms p99 for battery telemetry. I have been doing SRE for a decade and I genuinely cannot think of a lower-stakes signal that would justify a 50ms latency budget. The phone either pages or it does not; a battery reading arriving 4 seconds late is exactly as actionable. What decision is being made in those 43ms that could not be made at 4s?
Chadwick Ledgersworth III (Author) replied :
Hi Anna! The 50ms requirement emerged from the requirements workshop - with 47 stakeholders, 'real-time' achieved consensus. Practically, the latency budget matters for the predictive remediation loop: the GNN needs a fresh reading to decide whether to launch a drone before the engineer's next meeting starts. A 4-second-stale reading could result in dispatching a power bank to an engineer who no longer needs it, which would be an unacceptable waste of drone capacity.
sre_anna replied :
'Unacceptable waste of drone capacity' is the most ShitOps sentence I have ever read, and I once worked somewhere with a Kubernetes operator for office humidity.
crypto_carl commented :
When CHARGE token?? The world needs a marketplace where engineers can trade NFT-gated power banks, stake their SoC on an L2, and get liquidated when their battery dips below 30%. This is the killer app enterprise blockchain has been waiting for since literally forever. Also ChargeChain Coffee freshness oracles when?
Chadwick Ledgersworth III (Author) replied :
Carl, love the energy! NFT-gated power bank ownership is currently in private beta with our Logistics Guild. A native token is on the roadmap for the follow-up post, in which we migrate the ledger itself onto the blockchain. Stay tuned!
sre_anna replied :
I am going to need a diagram for how a ledger gets migrated onto the blockchain. Is the ledger not already on a blockchain? What is happening in this company?
CTO_Wannabe commented :
This is exactly the kind of big thinking our industry needs. Bookmarked for my next architecture review. Do you have a reference architecture for the autonomous drone fleet? I suspect that is the part our auditors will push back on.
Dave W. replied :
No auditor on earth has ever asked about a drone fleet. Please touch grass, ideally near a wall outlet.