September 5th, 2026 Outage Report
Going over the outage that occurred on September 5th and what we learned.
No one likes downtime, espically small hosts like us. Downtime degrades the public's trust, hurts customer workload, and in general is just a stressful time for operators, customers, and clients alike. When we have downtime like we had on September 5th, I use it as a moment to step back and review how we got here and what we can do better in the future.
This is the report of that day as we strive to learn from what went wrong, and what we're doing to improve right now, in the near future, and in the long term future.
The takeaway is this: We got lucky. this could have been much worse, and it almost was. The response time was abysmal, and we cannot take the good fortune we got for granted, and we will not.
(All times are in Eastern, GMT-4)
The Timeline
- At about 3am on 9/5/2026, PHYS-M18 went offline unexpectedly. Status pages and monitoring alarms were on this machine, so no alerts were sent out. One alarm was sent from a personal cron via telegram, but did not break through Sleep Focus.
- At 6:45am, the personal cron alarm was acknowledged, and an investigation began.
- At 6:54am, Jouleworks Customers were alerted to the outage.
- At 6:57am, the Colocation Provider was alerted to a network failure or equipment malfunction which was acknowledged by them at 7am.
- At 7:16am, a public announcement was released on my Bluesky account, after realizing Damocles (our Status Site) was unavailable.
- At around 7:30am, it was discovered that IPMI access to PHYS-M18 was not available, causing Remote Hands to be activated for the colocation provider.
- Status Updates were requested from the Colocation Provder and Acknowledged at: 10:18am / 10:33am, 11:35am / Not Acknowledged
- The colocation provider informed us the equipment had been power cycled at 12:40pm. The server did not come back online after a power cycle, and we informed the colocation provider of this at 12:52pm, requesting IPMI to be reset. This was not acknowledged.
- A followup was requested from us at 2:39pm. The Colocation provider acknowledged this at 3:33pm, saying there was no update yet.
- At 4:07pm we were granted access to a server. However, the Colocation provider gave us access to the wrong server. We let them know there was a mistake and elevated.
- From 4:07pm to 7pm, exchanges were made back and forth to ensure we are on the right hardware. We were. We then performed a login recovery on the PVE Environment, and found it was entirely wiped.
- At 7:24pm, we ran IPMI tools for SuperMicro where we discovered that we had indeed been pwned. All data had been erased and a new system had been put in its place.
- At 10:00pm as we wrapped up the investigation and began the process of reimaging, we noticed the attacker only formatted one of the four ZFS drives, so we took a closer look and realized the data was all still in tact but degraded. So, we began to resilver it.
- At 6:22am on September 6th, Resilvering completed and we rebooted into PHYS-M18
- At 8:00am Disk settling finished and the incident was marked closed.
We have regained full functionality and have confirmed all vectors from this attack are CLOSED. No data was lost during this event.
The attack progression
- Attacker gained control of IPMI at 2:55:41 AM and changed the Admin Password, deleting all users and configuration
- Attacker cold halts PHYS-M18 at 3:10:02 AM via IPMI, this coincides with disconnect logs on m1, m2, and m4. M18 never reconnects.
- Attacker installs a new Proxmox VE environment at 3:20:14 AM. This was confirmed via creation logs on the new environment after we regained control of the box.
- The Attacker then misconfigures our public IP address for M18, trying various networking configuration commands via IPMI, then disconnects for the last time.
- Jouleworks regains control of M18 at 7:08PM. Locks down IPMI and networking from M18.
This attack chain is simplified, and we are certain no customer data ever came in contact with the attacker. We have, however, saved an image snapshot of the OS installed and the access logs and have submitted them to the FBI.
Jouleworks' biggest failures
- We left an IPMI exposed to the public internet. There's no excuse for that. We fucked up. It was the prime vector of attack, and it was the reason this all happened.
- Our Status Monitors were Dogfooding from the very services they were meant to observe. For those who don't know, "dogfooding" is the process of using ones own services to operate your company or network. We take dogfooding at Jouleworks as a badge of honor, because it shows that we believe in the same things we offer to our customers. In this case however, all of our monitoring and watchdog processes were on the same hosts that went down, leading to us not knowing our services were even offline until I woke up that morning.
- Our colocation provider delayed remote-hands work. It's not a good look to blame someone else when your services go down, but in this case I have no choice but to count this as one of our mistakes. Our colocation provider was alerted at 6:57am, they acknowledged at 7:00am, the first reset that they issued on our equipment was at 12:40pm (6 hours later). The intermediary responses were apologies and intermediary prodding. This colocation provider is fitted well for homelabs that have outgrown the home, not for ISPs like we have become.
- Customer VMs that required high availability were not configured to fail over gracefully. The impact of this outage could have been mitigated if we had configured VMs for high-traffic customers to fail over to alive nodes gracefully. There is a complication with this due to our network configuration which doesn't allow this to be possible while maintaining reliable network connectivity.
What we're doing to fix this
- We are leaving this colocation provider. We have a one strike rule when it comes to providers, because we believe our customers and clients should only give us one strike for failures of this magnitude. We have signed an agreement with a new colocation provider that is much closer geographically to my team so we can control hardware failures directly and not have to wait on 3rd parties, because that is a shit excuse no matter who you talk to. Full stop. That is not to say this provider is bad, in fact, we're deliberately not naming them because we believe they offer a good value to the colocation ecosystem. It is clear to me; however, that we have outgrown them in space and in need of support. And that means we need to change and take on additional costs as it comes to that. Moving to a new colocation provider also lets us revisit the feasibility of providing High Availability solutions for customers that need it.
- IPMI Configuration Testing & Monitoring will become part of our maintenance schedule. All colocated systems undergo a bi-yearly examination of their hardware components physically. We ensure that everything is blinking green and nothing appears or sounds out of the ordinary. As part of this maintenance we check RAID arrays, component health, and OS Security Patches. We will be adding IPMI testing to this regimen to ensure this component which we rely on for disasters like this are not compromised at game time, or worse: Becomes the vector for the attack.
- We are installing ZFS on every new Hypervisor moving forward. ZFS was at first a martyr as we were diagnosing potential problems. Later as the investigation progressed, we found it was actually in fact our savior. Software RAID saved the day here and made something that would have been a major disaster into a unusually long outage.
- FluxFur will be moved to High Availability configuration ASAP. One of our biggest clients, FluxFur, is in discussion with us to be one of the first workloads to be moved to a High Availability, Hyperconverged state to ensure their service stays up even if the node they are on goes down. We are also looking into creating staging environments and an isolated subnet for them to conduct their activities without being interrupted by the R&D subnets which also occupy Jouleworks' IP Space.
- We are re-reviewing our Disaster Recovery Policies. This scenario was understood to be possible, but every failsafe we had to mitigate it fell through at game time. We are re-writing and re-examining our risk posture when it comes to these scenarios and ensuring we have a more stable DR plan in the future.
- We are going to continue dogfooding, but on seperate infra. We believe in the value of dogfooding our software; however, we will be offloading the status, monitoring, and update systems to seperate cloud infra. As we should have at the start of all of this.
My sincerest apologies.
Though I am sure people will forgive this incident and we will move past it, I refuse to treat it as an inevitability. People put their faith in my systems, and I genuinely believe it is my responsibility to ensure they are operated, maintained, and acted on professionally. This was a low point for us as an organization; however, I want this to be an experience we take with us moving forward.
For the rest of us.
I want to take this last few sentences to explain our motto: "Development for the Rest of Us." Jouleworks is a company that aims to be for those who can't or won't fit into the molds of big corporations. We offer services, solutions, and software that is bound for the small or medium sized business that can't operate Salesforce at their scale, or can't pay that premium for software they only need a fraction of. We also support people who otherwise would be kicked off other services because they are "too risque" or "too abnormal". We think that's bullshit. Everything we do, we want to be done For the Rest of Us. Period.
I hope you continue with us as we accomplish that mission with this experience in mind. See you around.
All the best,
Kaisa
Jouleworks Owner