← All writing
Moving

September 20th Move Log

A live updated moving log while we are transporting systems.

As mentioned in our status page and in our last journal on the September 5th, we will be moving SEC1 and RIC1 production systems to our new datacenter: ROK1!

This exciting adventure begins tomorrow at 9am, at which time we will begin logging the journey here on the journal.

All times are in ET

  • [9/20 9am] Arrived at SEC1, took 40 minutes to get someone to get us into the cage.
  • [9/20 9:15am] RIC1 team departs
  • [9/20 9:45am] Cage is locked, needs to get the key
  • [9/20 10:20am] Informed that the person who runs this cage with my servers in it had a security issue and is locked out of his key room, called the DC to unlock it
  • [9/20 10:40am] Told it'll be another 20 minutes, we're now 2 hours behind our scheduled departure from SEC1
  • (Author's Note: I understand now why these folks took so long to do anything with our servers, and I'm glad we're leaving)
  • [9/20 11:20am] Finally got access to the SEC1 cabinet, took out servers and departed
  • [9/20 1:00pm] RIC1 Team arrives and delivers equipment
  • [9/20 4:00pm] SEC1 Team arrives in region and begins picking up equipment shipments
  • [9/20 5:00pm] SEC1 Team arrives at ROK1 with all equipment, begins racking
  • [9/20 6:30pm] Racking is complete with some issues. M18 is physically too big for the rack, so a compromise has to be made.
  • [9/20 7:00pm] Networking Begins, this involves reestablishing BGP networks and sessions with our new upstream peers
  • [9/20 8:00pm] Sessions are stabilized using the old upstream peer as the new peer is determined to not have set things up properly, we now move to reconfiguring all Servers to the temporary configuration
  • [9/20 10:00pm] All servers are reconfigured, but a latent Proxmox bug detaches all networking interfaces from VMs that bridge them to physical networking (TAPs)
  • [9/20 11:00pm] All TAPs are reconnected, infra is handed back to customers

What we learned.

  1. SEC1 was a garbage fire. We had to wait for 2-3 hours just to get access to the cage because of security lapses in their Datacenter.
  2. New DC BGP sessions were not configured appropriately. This caused a lot of headache in troubleshooting an issue we couldn't solve that night.
  3. A lot of manual configuration was needed for the BGP workaround which needed to be automated. This has been setup in Mikrotik with a netwatch protocol.
  4. M18 was physically too big for the rack. This was. Incredibly unfortunate. We ordered a new Node which will be installed during maintenance this weekend.

What's changed?

  1. We're now officially in one stack in ROK1
  2. Networking in the Edge is now properly routed instead of needing to use a Firewall Mangle Hack, which makes expansion and general operation much smoother.
  3. Our /24 network is now better split to serve customers and networks.
  4. We have a proper 10G/1G network split to allow for high speed server operations ensuring that 1G devices don't bring the entire network speed down.

The maintenance window was scheduled for 6 hours, with the delays and troubleshooting it took well over 13 hours to go back to fully operational. While we don't plan on moving datacenters every other day, we do apologize that it took this long.

Thanks for sticking with us.

Back to the journal →