Splitting Light: Season 3 - Episode 13


Splitting light

Season 3 Episode 13

Bills reconciliation

If you are no longer interested in the newsletter, please unsubscribe

January to March 2019

When you have a complex system there are always going to be discrepancies between them. In our case the discrepancies were in the billing systems. We had five different data sources. None agreed on the numbers.

Let me introduce our five systems. Observability of the disk usage as well as the incoming and outgoing bandwidth. The client observability, client consumption and bandwidth. Then our internal billing system, which probed the cluster for usage and received bandwidth telemetry. Lastly it was the end bills, which were generated by the billing team.

All these systems conversed in different units. Some of them were point in time usage. Bytes used at a specific time. Some were accumulating bytes over the course of a month. Some were in gigabytes per hour and lastly some in euros per month.

I remember spending a lot of time on these systems with Gaspard (a). The values varied widely. Part of the issue was that we did not segregate internal and external clients. I created the right dashboards and filters to segregate them. Then it was data which was not completely garbage collected. We investigated and fixed the issues. Then it was data in transit between the STANDARD and GLACIER storage class. Step by step we aligned the data.

There was only so much we could do without implementing the billing ourselves. I did dashboards which approximated the calculations from the billing team. They converted units using crude methods. It half worked but it did give me and Gaspard a good idea where to look for discrepancies.

A few issues were interesting.

One was that the software that received bandwidth telemetry did not follow the request curve. I found out its performance was capped at 4,000 rps (request per second). This happened because we maxed out the single thread performance. We changed the software to parallel execution and the problem was settled.

Another one was that the reported usage did not align precisely. It turns out that increasingly complex SQL queries are not a solution. The SQL was getting increasingly arcane to handle edge cases. It was directly impacted by any blip or incident in production. A state machine was the solution. I wrote the code of that machine to make the system more accurate. I baked in rewind and replay mechanisms that would come in very handy a few months later.

But! Why does this matter? The first half was that we needed to predict hardware consumption. We needed to predict how many servers we would need at a specific date. Without clean data this was impossible to do. We needed clean data also to make sure the clusters and systems were behaving correctly.

You cannot optimize blindly. I like this quote from Donald Knuth: “We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%.” To find that 3% where optimization is worth it, you need to understand and map the systems together. Otherwise, you are just whacking the system with a stick in hopes of making it better.

The second half is that in a complex interdependent service offering, the revenue attribution is important. I was about to try to make that point clear.

  1. Gaspard Plantrou, VP of storage and database back then, now CPO at Numspot

If you have missed it, you can read the previous episode here

To pair with :

  • Seven churches for st Jude - GAIKA
  • EBADSLT

Vincent Auclair

Connect with me on your favorite network!

Payson, Payson, AZ 85541
Unsubscribe · Preferences

Symbol Sled

Business, tech, and life by a nerd. New every Tuesday: Splitting Light: The Prism of Growth and Discovery.

Read more from Symbol Sled

Splitting light Season 3 Episode 12 Warsaw If you are no longer interested in the newsletter, please unsubscribe End of January 2020 to beginning of February 2020. In January 2020, we had a new challenge. A new opponent to ring in and control. Let me introduce him. It was the Warsaw region! We had to deploy object storage in Warsaw. Back in September 2019 we had selected the hardware setup and had done a brief discovery of it. Then the whole region’s hardware had been shipped by truck. Once...

Splitting light Season 3 Episode 11 OpenIO reaching out If you are no longer interested in the newsletter, please unsubscribe End of 2019 While we were dealing with the 500k object problem in object storage, Jean-Francois Smigielski (a), OpenIO’s CTO reached out to me on their community Slack. They were looking to hire more engineers and several people that they trusted had told them I was what they needed. Most likely I had also done a good impression when we had worked together to bootstrap...

Splitting light Season 3 Episode 10 Team of brothers If you are no longer interested in the newsletter, please unsubscribe Around September 2019 When you are very close to a team, it can feel like the team is a band of brothers. It felt to me that the team storage was cohesive. I felt like their older brother guiding them. Maybe it was because I was the eldest in my family and the second among my cousins on my dad’s side. My parents tell me I led my younger cousins for many years even though...