NAS & RAID · case file · BHD-2025-8642
The Rebuild Cost More Than the Power Cut.
A press-tool subcontractor in Coventry came back from the bank holiday to a six-bay shelf running degraded. One drive had been amber since Friday
, and nobody had thought much of it. By Tuesday a second had left the set. The contractor who looks after their machines put a spare in and set it rebuilding
, and on the Wednesday it froze at about sixty per cent
. The power cut cost them nothing at all. The rebuild cost them a fortnight.
Sounds like yours? Give us a ring.
0800 6890668
What that means.
The failures were survivable. The rebuild was not. RAID 5 keeps one disk in hand and no more, and that allowance had been spent the moment the first bay went amber; once a second member left, the controller was being asked to solve for two things it could no longer see. Told to rebuild in that state it restores nothing, and it does real harm, laying down new parity across stripes a recovery needs left exactly as the failure found them. Enterprise shelves add a problem of their own. Their disks are often formatted with sector sizes desktop equipment cannot address, and vendor structures sit where consumer software either mangles them or walks straight past. Underneath all of it is that amber flag, weeks old, while parity carried the set single-handed.
What did the work here.
The order of the work →| Equipment | The job it did | What it adds |
|---|---|---|
| PC-3000 SAS/SCSI | Assessed all six server disks over the interface they were designed to speak | Pull a SAS or SCSI disk from a server and ordinary desktop kit cannot read it |
| Atola TaskForce 2 | Copied all six at the same time rather than one after the next, saving days | Every disk in the array goes through the imager together, not one after another |
| UFS Explorer RAID Recovery | Worked out the array layout and mounted the volume off the six images | Reads the volume layers a NAS puts over its array, not only the RAID underneath |
What we did.
Copy all six, including the ones still working
The shelf was not switched on again. All six disks came out and were numbered against their bays, then went onto imagers. The sound ones were taken first, leaving the pair that had failed to be assessed before anything was asked of them. Interface and sector format have to be right at this point; get either wrong and every copy taken afterwards is quietly spoiled.
Getting the sectors out of the two failed disks
One of them was weak on a single head; the other was picking up bad sectors by the hour. Neither disk was close to empty. Both were read in short, slow passes, and most of each surface made it into an image. Holding a copy of both mattered: where the two disagreed about a sector, the cleaner read settled it, since no parity was left to arbitrate.
Read the layout off the array, not off habit
Member order, stripe size, the direction of parity rotation and its delay all came out of the array's own structures rather than from what a given controller usually does. With that fixed, the six images were built into a virtual volume and the filesystem read out of that, while the disks themselves stayed powered down.
How it closed.
The assembled volume mounted and was verified against the directory tree it carried, then went back on new media, postage home at our cost. One qualification goes on the record: within the region Wednesday's rebuild had already covered, a live project folder returned short. The engineer working on it kept a fortnight-old copy on his laptop.
Other pages worth a look here.
The RAID & NAS cases.
Is yours behaving the same way?
Switch the device off, post it in, and decide nothing until the diagnosis comes back to tell you where you stand.