Feb 062026
 

Delta Electronics DPS-420CV A REV.0 logoWell, last night I got up to take a leak and heard my VFI UPS unit beep again (the batteries are pretty dead, I just didn’t take care about it by now), so I walk past the ancient IBM server that hosts this place to switch the alarm off. And as I walk back, I notice something. This machine – as [previously described] – features a maximum of three power supplies for redundancy and load balancing. Each unit has two LEDs to indicate whether everything is working fine (green light) or whether something is wrong (orange light). One is for indicating whether the AC input is fine and the other indicates whether the DC outputs are okay. I’ve had two units die on me over the decades that this machine has been running, and last night, in the 30th year of operation – there’ll be an anniversary soon – it seems another one bit the dust. In this case, it was just completely off, no lights lit at all.

I unplugged it and plugged it back in just to see what would happen. Both LEDs went on one after another – like it should be – but the power supply then immediately went offline after less than a second. Well, that didn’t look too good, so I removed it while the server was running (as it’s the hotplug version) and replaced it with another unit, given I have several replacement units in stock.

Broken Delta Electronics DPS-420CV A REV.0

Another broken Delta Electronics DPS-420CV A REV.0 (I didn’t bother to clean it, as it’s going to enter the recycling chain soon enough, click to enlarge)

So while it was in the middle of the night, I still fetched a replacement unit and swapped them, it’s just four screws you have to remove after all, and the “new” unit appears to work just fine:

Back online

Back online (The top one was broken, click to enlarge)

After decades of experience with computer hardware I do consider power supplies consumables, so it’s kinda fine I guess. Those things have been running for almost 30 years after all, and I put them under pretty high load for at least 20 years.

Thank god I’ve collected all kinds of replacement parts for almost 20 years from all over the planet. Wouldn’t be easy these days anymore…

May 172022
 
Fan failure reporting on the IBM PC Server 704, logo

Two days ago I got woken up in the middle of the night by some noise that sounded like somebody was trying to break into my home. Or so I thought in my still half-asleep state. Left my bedroom rather cautiously to check things out, only to find no problem anywhere. Then, while I was standing in my hall, I heard it again. And nope, that wasn’t people trying to break in, but actually a noise coming from the stone-age server that runs this web site. And it sounded pretty nasty to boot. And all that at 02:00 am in the middle of the week… Hah. So what’s happening here? One of the Delta Electronics WFB1212ME-R00 fans running in the machine is experiencing bearing problems. Something like this has happened not too long ago for the unit cooling the CPU & RAM thermal zone compartment. Back then I managed to repair the fan, as I [described in detail here].

Another ancient Delta WFB1212ME-R00 fan with bearing failure

Another ancient Delta WFB1212ME-R00 fan with bearing failure as indicated by one of the front LEDs (click to enlarge)

This time, the affected fan is one of two intakes feeding the PCI & system board thermal zone and the additional internal fan for said CPU & RAM compartment. Since there are two intakes right on top of each other, there is actually some redundancy there. So this time it isn’t critical, like the CPU & RAM one was, resulting in a thermal trip and lockup of the server two years ago.

So I took a pair of scissors (first thing I could find), opened the server up and rammed it right into the fan after getting it to spin down and stop fully by carefully pressing against the underside of its shaft. This isn’t actually a problem as these special types of fans won’t constantly attempt to spin up, so there’s no overheating issues. The electronics simply briefly attempt a spin up every few seconds and stop immediately if it’s stuck.

 

I did this (not pictured :roll: ) because the bearing had started to wobble somehow, and every 20-30 seconds the blades would violently hit some metal parts of the server’s case as the whole rotor lifted off for several millimeters. This also resulted in the fan failure reporting LED constantly going on and off again. Might be the circlip locking the shaft’s axle into the casing that’s broken, not sure yet.

Next weekend or the weekend after that there will be a brief downtime for XIN.at, in which I will attempt to repair this newly failed fan as well. If it’s not possible, it will be replaced with one of the spares I managed to collect in the meantime.

This one has so far made it to ≈30 billion revolutions over the course of about 26 years, exceeding its rated continuous runtime of 8 years by a factor of 3.25×. Not too bad I’d say, but I’m guessing since two of them have now hit their limits within the past 2 years, I might see more failures from now on. We’ll see.

PS.: In case you’re wondering why I’m not replacing them with newer fans: It’s not possible. First, they have physically different 3-pin plugs and second, they use the third pin for a special failure signal instead of for fan speed.

Nov 082016
 
G.SHDSL extender failure (logo)

…and it wasn’t even my fault! Can you believe it?! Probably not if you know me, but it’s true nonetheless… Almost 4 days of downtime and we’re back up since just about 2½ hours or so. Given that I already had to do maintenance on the server once this year (replacing a bad hard drive and doing a thorough cleaning as well as dust filter installation), this has crushed the yearly 99%+ availability that I was so proud of. So for the first time since 2006, XIN.at failed to satisfy my personal requirement in that regard. Including the maintenance done on the server and several regular ISP maintenances on the G.SHDSL line, the full downtime should now amount to roughly 90 hours in 2016. If we assume a sum of 8760 hours per year, I’m now down to an availability of ~98.97%.

That value might get a bit worse though if my ISP decides to do another few rounds of maintenance on the DSLAMs in the automatic exchange hub.

So, how did this happen?

It all began when my RAID-6 started acting up, the one in my workstation though, not in the server. Ok, I know, that’s entirely unrelated, but still. It died no pretty death right there last Friday. And once again (this happened before!) it was not the disks to blame, neither the controller, nor the FBM, not even the hotplug bay that I suspected because all disk failures where happening in the same bay. It was the power cable extensions. Again. Even though they’re brand new! I mean, what the hell. At least I know now, that an Areca controller can force RAID-6 arrays to come back to life even if already completely failed with 3+ disks down. Nice one, Areca, I’ll have a cold one in your honor!

And when that RAID was back up, I wanted to pull up my rolling shutters a bit, just because. Which is when the belt ripped in half and the shutters went crashing down, damning me to darkness. Ok, after that I had a beer and just went to bed. Not my day. Next day I did some makeshift repairs on the shutters so they would at least be rolled all the way up and stay there. Having 0% daylight at 09:00am is pretty depressing after all. Ok, after that was done (it was Saturday now), I sat back down in my chair and thought: “Ok, let’s just read my emails…”.

And then my G.SHDSL extender burned up, sending me, my email client, my server and the rest of my digital existence offline…

And that’s when I just knew I had to get up, drive to the supermarket and get a TON of beer!

Seriously… There is bad luck and then there is…

Bad luck never comes alone!

When it rains, it pours, they say

So, the thing just went dark from one moment to the next! No fan, no LEDs, no nothing. At first I thought it might be its external power supply, some standard 12V DC unit. But I measured the voltage and it was perfectly fine. So the extender itself was obviously dead. Never seen such a thing happen with Paradyne/Zhone hardware, but what can you do. So here’s the new one (or maybe it’s refurbished, you never know with this stuff):

Paradyne/Zhone SNE2040G G.SHDSL network extender

Paradyne/Zhone SNE2040G G.SHDSL network extender (click to enlarge)

Now all that’s left is to send the defective unit back and that’s that. I hope I won’t see anything like that happen again… :( At least I got them on the phone on Saturday (business level support), but I only have the small service level agreement with my current contract, so I couldn’t get a technician on weekends. And I wasn’t available “on-site” (at home) on Monday, so the replacement unit had to be shipped via parcel service.

Oh, and neither the 3G fallback solution nor the large SLA (full 24/7 on-site support) will ever be agreed upon for XIN.at – too expensive at ~40€ a month. :( There is just so much money I can pour into a free server after all.

At least everything is back up now, so cheers! Prost!

Feb 062015
 

Network[1] Everybody hates servers going offline. Especially email servers. Or web servers. Or MY SERVER! Now I prepared for a lot of things with my home server, I prepared for power failures, storage failures, operating system kernel crashes, everything. I thought I can recover from almost any possible breakdown even remotely, all but one: My four bonded G.SHDSL lines all failing at once. Which is what just happened. After lots of calls and even a replacement Paradyne/Zhone SNE2040G-S network extender having been brought to me within the time allowed by my SLA, all four lines still remained dark.

Now, today the telecommunication company which is responsible for the national network fixed the issue in the local automatic exchange. I tried to find out what had happened exactly, but ran into walls there. My Internet provider UPC got no information feedback from the telecommunication company A1 either, or at least nothing besides “it’s been fixed at the digital exchange”. Plus, as I am not an A1 customer exactly, so they won’t answer me directly. The stack is: UPC (internet provider) <=> Kapsch (field technicians handling UPC branded Internet access hardware, via outsourcing by UPC) <=> A1 (field technicians regarding the whole telecommunications infrastructure), while UPC may also communicate with A1 directly to handle outages. Communication seems to be kept to a minimum though. :(

Bad thing is, for a “business class” line, an outage of almost two days or 47 hours is a bit extreme. In such a case, more efficient communication could easily fix it faster. But it is what it is, I guess. And now I have to send one of the two Paradyne/Zhone G.SHDSL extenders back to UPC, this little bugger here:

The actively cooled Zhone 2040 G.SHDSL extender

The actively cooled Zhone SNE2040G-S G.SHDSL extender (click to enlarge)

There is actually a HSDPA (3G) fallback option, which works by implementing an OSI layer 2 coupling between the G.SHDSL line and the 3G access, keeping all IP addresses and domains the same and the services reachable during times of complete DSL failure. But I won’t order that upgrade, because it’s a steep 39€ before tax per month, or 46.80€ after tax. That’s just too expensive on top of what that connection’s already draining from my wallet.

All in all, this greatly endangers my usual, self-imposed yearly service availability of >=99%. 47 hours is a lot after all. So to maintain 99%, the server cannot go offline for more than 3 days, 15 hours and 36 minutes per regular year, and now I already have 1 day and 23 hours on the clock, and it’s just the beginning of the year! Let’s hope it runs more smoothly for the rest of 2015.

[1] Logo image is © Kyle Wickert, Do You Really Understand The Applications Flowing Through Your Network?