Case file · NAS & RAID · MHD-2025-6642
RAID 5, Two Disks Down, Then a Rebuild.
A bank holiday power cut left a Windsor engineering consultancy's six-bay shelf running degraded, with one bay flagged before the weekend
and no great alarm about it. On the Tuesday a second disk left the set. Their IT contractor fitted a spare from the drawer and pressed rebuild
, and by Wednesday that rebuild had stopped at sixty per cent
. The rebuild, not the power cut, was the expensive part.
Much the same fault? Call us.
0800 6890668
In everyday terms.
Parity in a RAID 5 buys you exactly one missing disk, and that allowance was spent the moment the first bay was flagged. With a second member gone the controller holds more unknowns than it has equations to solve them with, and a rebuild begun in that state restores nothing at all; what it does instead is write fresh parity across stripes that a recovery needs left as the failure found them. The shelf itself compounded matters, because enterprise disks are often formatted with sector sizes ordinary desktop equipment cannot address, and vendor metadata sits where consumer software either misreads it or passes over it. The quietest part of the account is that first flag, which had been showing for weeks while parity did two jobs at once.
What the bench used here.
How a job is handled →| Tool | Why we used it | What it brings |
|---|---|---|
| PC-3000 SAS/SCSI | Read and assessed each of the six enterprise disks on the interface they are built for | SAS and SCSI disks out of a server will not talk to standard desktop hardware |
| Atola TaskForce 2 | Imaged the whole set concurrently instead of disk by disk, which saved days | Parallel imaging across the whole set, so the copying stage runs in days, not weeks |
| UFS Explorer RAID Recovery | Derived the array geometry and mounted the volume from the six copies | Understands the layers a NAS stacks on top of its disks, not just the array |
The lab work.
Every bay imaged, the healthy ones included
The shelf was never powered up again. All six disks came out, numbered against their bays, and went onto imagers, the sound ones first so that the two casualties could be assessed before being asked to work. Getting the interface and the sector format right at this stage is what the rest of the case rests on: an error here spoils every copy taken afterwards without announcing itself.
Coaxing the two casualties into giving up their sectors
One casualty was reading weakly on one head; the other was accruing bad sectors by the hour. Neither was anywhere near empty. Both were read slowly, in short passes, and most of each surface went into images; holding a copy of both meant that a sector they disagreed about could be settled by whichever read it more cleanly, rather than by parity that no longer existed.
Read the geometry off the disks, never assume it
Member order, stripe size, the direction parity rotates and its delay were all taken from the array's own structures rather than assumed from what most controllers do. That settled, the six images were assembled into a virtual volume and the filesystem walked out of that, the original disks staying switched off.
How the job closed.
The assembled volume mounted, verified against its own directory tree, and went back on fresh media with the return carriage at our cost. One qualification belongs on the record: inside the area Wednesday's rebuild had overwritten, one live project folder came back incomplete. A copy of it, a fortnight old, sat on an engineer's laptop.
More of the jobs we've closed.
Other RAID & NAS jobs.
Does any of that match your case?
Nothing needs deciding until the diagnosis is back: switch the device off, send it in, and let the findings settle it.