Business, tech, and life by a nerd. New every Tuesday: Splitting Light: The Prism of Growth and Discovery.
Share
Splitting Light: Season 3 - Episode 12
Published about 13 hours ago • 3 min read
Splitting light
Season 3 Episode 12
Warsaw
If you are no longer interested in the newsletter, please unsubscribe
End of January 2020 to beginning of February 2020.
In January 2020, we had a new challenge. A new opponent to ring in and control. Let me introduce him. It was the Warsaw region! We had to deploy object storage in Warsaw.
Back in September 2019 we had selected the hardware setup and had done a brief discovery of it. Then the whole region’s hardware had been shipped by truck. Once rolled in, the network team had worked on its setup in December 2019. Afterwards it was then sent for the platform team’s basic setup. Then, it was our turn. We now had to bring object storage up.
Initial sequence of bringing a new region up. Object storage is one of the first requirements.
I was the one to take up this foe. I jumped in the ring and started to work on it. Little did I know, it would prove much harder than expected.
The region was set up completely differently than all the previous regions Scaleway had done. It was a new, better design. But that implied there were a lot of new constraints to handle and implement.
The first one was the boot system. It was different. We could not use our own system. We had to add Scaleway’s one into our stack. After I painstakingly did this, I realized we had forgotten a very important thing. While we had discovered the servers. We had missed the IPMI discovery. We couldn’t reboot the servers remotely. We had to rely on datacenter technicians. After some fiddling I was able to finally start one server. Then I tweaked some knobs, essentially magic, on the IPMI and I was able to get access to the others servers via bridging the networks.
The network for operations and IPMI are isolated. You don't have the same tool freedom on both sides.
Then we had to port the deployment code. We relied on mechanisms that worked in Paris and Amsterdam, but in Warsaw we couldn’t. There were a few added complexities to handle. After fighting with these new constraints, hitting down all the right implementations, I was able to get the storage servers up and running. But, the load balancers proved to be even more difficult. It was a very long fight.
I started by probing the unknown. I was able to discover the storage server’s IPMI but not the load balancer’s one. I flooded the network with packers designed to trigger responses. I dumped all the traffic going in and out and used filtering to see very specific information. The network team helped with the router configuration. Finally, just as I was about to resign myself to physically sending someone, on a Monday morning, I saw a packet trace. A single packet had gone through during the weekend.
From that single packet, I was able to understand the missing link and find the invisible servers. I hit them with several connection tunnels and was able to access the graphical boot interface. To understand what was wrong, I needed to see what was happening when it was trying to boot. After setting the right boot sequence, I could see the server get an IP, request a boot and then stop. The fight continued.
If any of these steps fail, the system does not work. Most of these steps are blind. Meaning the only way to know if it worked is because it advanced to the next step.
I expanded the ring’s perimeter. Finally, an uppercut. One of the critical elements in the boot sequence was a different version than the one we had to boot our systems in Paris and Amsterdam. After uploading that version, adding the right hooks, it booted. The load balancers spun up. I had landed the final blow. It was a critical hit. KO. Wasted.
Then the merging of chances and final configuration. In February 2020, the object storage in Warsaw was live. Not publicly, but that would come.
It was a successful fight. We amended our procedure so this one wouldn't happen again. But, it never ends. The next one would bring me in the realm of accounting.
If you have missed it, you can read the previous episode here
Splitting light Season 3 Episode 11 OpenIO reaching out If you are no longer interested in the newsletter, please unsubscribe End of 2019 While we were dealing with the 500k object problem in object storage, Jean-Francois Smigielski (a), OpenIO’s CTO reached out to me on their community Slack. They were looking to hire more engineers and several people that they trusted had told them I was what they needed. Most likely I had also done a good impression when we had worked together to bootstrap...
Splitting light Season 3 Episode 10 Team of brothers If you are no longer interested in the newsletter, please unsubscribe Around September 2019 When you are very close to a team, it can feel like the team is a band of brothers. It felt to me that the team storage was cohesive. I felt like their older brother guiding them. Maybe it was because I was the eldest in my family and the second among my cousins on my dad’s side. My parents tell me I led my younger cousins for many years even though...
Splitting light Season 3 Episode 09 First recall on our tech loans If you are no longer interested in the newsletter, please unsubscribe Autumn 2019 We had our first tech loan recalled in November 2019. One of the limitations of OpenIO was the number of objects in a bucket. They had recommended not putting more than 500 000 objects in a single bucket. We had documented this publicly. It was a recommendation but, of course, we didn’t limit the number of objects for customers. The limitation...