Nobody had looked at the alert panel for months, so the first failure went unnoticed, and the second one arrived twenty minutes into the rebuild. Or the controller card died with the layout held in its memory and nowhere else. RAID 5 data recovery is a weekly job on this bench rather than an unusual event, and so are levels 0, 1, 6 and 10 pulled out of servers, workstations and NAS units. Since everything travels by tracked post, the RAID recovery London firms need runs alongside work from Reading, Thames Valley Park and the units around Bracknell, with disks turning up from Croydon, Richmond and Hammersmith just as readily as from Wokingham, and RAID array recovery UK enquiries arriving from well beyond either. No theory is tested until every disk has been copied. Anybody who has stopped trading goes to the top of the list.
Anything sent in for RAID recovery is diagnosed free. The written figure that follows is settled before a screwdriver leaves the drawer: from £500 + VAT for an array, rising with the member count.
No fix, no fee covers logical recoveries. Outside it sit electronic and mechanical failures, chip-level work, DVR jobs and forensic jobs, and physical work is 50% up front. Every band is published on the data recovery cost page.
Reading a symptom back to the failure underneath it is where the work genuinely starts, and this set of thirty accounts for all but a handful of the parcels that reach the Guildford bench. If yours is not on the list, it will still be recognised — describe it on the phone and you will get a straight answer about the odds before you post anything.
More arrays are destroyed by this than by any component in them. A rebuild asks every surviving member to read from end to end for hours while it writes to the replacement, which is the heaviest possible work to demand of disks that have shared a warm cupboard for six years. One of them gives out around the halfway mark, and the set now holds two failures and a half-written newcomer. Copy everything first and none of that can happen.
Somebody lifts all the disks out to read the serial numbers off them and cannot say afterwards which slot each one came from. Member order, stripe width and the offset where data actually begins all have to agree before a volume appears, and the wrong order produces something that mounts and reads as nonsense. The order is recoverable from the contents themselves. It is also an afternoon of arithmetic that a strip of masking tape would have saved.
The cache is holding writes, the backup battery has been flat since last year, and the power goes. What comes back has parity blocks describing stripes that were never committed. That gap is the write hole, and it does exactly what the name suggests. Both versions can be set side by side on copies and the right one chosen. Live, the controller settles the argument by overwriting one of them without asking.
Hardware controllers keep their array definition in memory on the card, so when the card fails the definition stops being readable. Nobody here goes hunting for the identical model on an auction site. Stripe boundaries, parity rotation and the start of the file system are all legible in what the member disks contain, and reading them back off the platters is quicker and more reliable than sourcing a replacement.
Adding a disk to a live RAID 5 rewrites every stripe on the set at a new width, in place, over a period of days. Interrupt that and the array is carrying two geometries at once: the old one behind the point it reached, the new one in front of it. The boundary shows up plainly on the images, and each region is then read with whichever geometry belongs to it.
A disk that dropped out in April, pushed back into the set in June, is offering the controller a version of the file system that is two months stale. The two accounts tear against each other along the line where they stop agreeing. Pulling that seam apart is done on copies, where a wrong assumption costs an hour instead of a directory tree.
chkdsk, fsck and the vendor repair utilities all assume the volume beneath them is fundamentally sound. Run one over an array that has come back together in the wrong order and it will dutifully rewrite metadata to match a layout that never existed. That is a second job created on top of the first, and it is the reason nothing gets repaired here until the geometry has been proved.
Somebody opens the controller utility to look at the state of things, works through the wizard, and creates a fresh array on hardware that already had one. Initialisation writes from the front, so if it was stopped in the first minutes most of the previous volume is still lying there behind it, untouched and perfectly readable once the old geometry is worked out.
A dying PSU or a spike on the rail can take out several members in the same instant, and no level of redundancy is designed for that. Each casualty is treated as a drive recovery in its own right on the mechanical bench before any of them is asked to contribute to a reconstruction. Arithmetic across the group comes last, not first.
When an entire column of bays drops out together, six disks failing simultaneously is the least likely explanation on the table. Chassis wiring, a tired expander or a cracked solder joint on the backplane accounts for most of these. Establishing which it is costs nothing but an hour, and it is a great deal cheaper than treating six healthy disks as casualties.
The spare sits powered and idle for four years, the array loses a member at two in the morning, and the automatic rebuild begins onto a disk that has not been asked to do anything since it was installed. Idle is not the same as healthy. Nothing here joins a reconstruction before it has been imaged and read clean, and a spare gets no exemption from that.
The controller offers to import metadata it has found, somebody clicks yes because the alternative box looked worse, and the member order changes underneath a working file system. The genuine layout is still written across the drives, which keep a more honest account of themselves than a card that has just been handed stale information.
Occasionally the geometry falls out cleanly, every parity check agrees, and the NTFS or ext4 riding on top is still wrecked. That is two jobs stacked: an array recovery underneath a file-system recovery. The second half is finished by reading the structures directly rather than by pointing a downloaded scanner at the result.
A disk with failed heads goes to the clean bench for a head transplant, has its firmware brought back into line, and only then joins the imaging queue with the rest of the set. The other members wait, because an array is only as recoverable as the worst disk in it and there is nothing to be gained by reconstructing early.
Small-business arrays are almost always built from a single order of identical drives, which then run the same hours at the same temperature in the same box. They reach the end of their working lives within weeks of each other, and the failure that gets noticed is rarely the first one. Every member is imaged in parallel here for exactly that reason.
mdadm superblocks vanish after a distribution upgrade. Storage Spaces stops recognising a pool it created itself. Neither takes much provoking, and neither is as well documented as the wiki page suggests. Both are reassembled by hand from the raw members, which is slower than any automatic route and considerably more likely to work.
A disk remapping sectors on a schedule is not failing suddenly; it is failing predictably. Imaged while there is still slack in that schedule, the whole set comes home intact. Left in service until the count runs out, part of it will not, and the part that goes is chosen at random rather than by importance.
On the design document the set is finished, because losing a complete mirror pair puts a hole straight through the stripe. In practice one of those two disks will usually image far enough on lab hardware to fill the gap, and the recovered volume is complete or very nearly so. This is where the theory and the bench tend to part company.
Consumer drives will spend a minute or more trying to read a marginal sector, because on a desktop that is the right behaviour. A RAID controller waits seven seconds and then throws the disk out of the set for insubordination. On imaging hardware that is happy to wait all night, those same drives read perfectly well and the set reassembles.
Buy the cheap high-capacity drive, put it in a degraded array, and the rebuild slows to a crawl as the drive shuffles overlapping tracks about behind the controller. It eventually stops answering, the controller drops it, and the array is worse off than before it was touched. Which drives were shingled is worth knowing before anything is ordered.
Virtualisation hosts running ZFS arrive after a power cut with pool metadata halfway through an update. Full images of every member come first, without exception, and the transaction history is then walked backwards on those copies until a consistent point turns up. That point is usually only seconds behind the failure.
A patrol read or a monthly scrub means hours of uninterrupted reading across every member, which is precisely the work a marginal disk cannot survive. The failure lands part way through the scan, at a weekend, when nobody sees the email. If a set is going to be imaged, image it between the checks rather than after one.
An array short of a member has no way to correct a medium error, so it records the failure into parity as a permanent hole. Those punctures are charted from the images and routed around at file level, which usually means a handful of files are affected rather than the volume. Nobody finds out until the rebuild reaches them.
Sector formats do not mix, the controller rejects the newcomer, and forcing the point creates two more problems on the way to solving nothing. On images the sector size is something that can be adjusted at will, which it never is on hardware that is trying to be helpful.
HBA and expander updates that fail to complete shed disks in ones and twos across the following fortnight. It reads on the console as a run of drive failures and it is nothing of the sort. Identifying that the fault is on the controller side, not the media side, comes before anybody starts replacing perfectly good disks.
A stripe with no parity anywhere in it, chosen because the numbers looked good in a review. One module stops answering and the volume goes with it in the same second. Each module is handled as an NVMe job on its own terms, and the stripe is reassembled from the images afterwards.
Warnings have been arriving in the mailbox of a member of staff who left before Easter, so nobody has known for months that the set was running without protection. By the time the phone rings the second disk has gone too. Everything in the chassis gets copied here, including the members the controller still calls healthy, on the reasoning that a controller which missed the first failure is no judge of the rest.
Nominally identical capacities differ between production runs, sometimes by thousands of sectors, so the controller refuses the new disk and somebody clears the configuration to make it accept one. The original fault now has a second fault standing next to it. Copies can be made whatever size the work requires, so on images the capacity argument simply does not arise.
Drives pulled out of a decommissioned server and put into a new one arrive with the previous array metadata still written on them. A controller that can see both sets of signatures starts making decisions nobody asked it to make. The layout that matters is settled from the data on the platters, not from whichever signature was loudest.
JBOD spans hold no redundancy of any kind. Files are laid down across the members one after another, so a missing disk removes a slice from the middle of the file system rather than truncating the end of it. Everything that happened to sit wholly on the surviving members comes back intact, and on a large span that is a far bigger proportion than most owners expect to hear.
Questions about the layout can wait. Before any of them are asked, every member is duplicated in full, top to bottom, and once that is done your original hardware plays no further part in the recovery. All the analysis happens against the duplicates, which removes the time pressure entirely: the position each disk occupied, the width of the stripe, the direction parity rotated in, and the offset at which data actually begins can all be deduced from the contents, unhurried and repeatedly if necessary. The volume is then assembled in software over the images and the file system read out of that assembly. Typical arrivals are a member that dropped silently out of a RAID 5 half a year before anybody looked at the front panel, a RAID 6 that has lost two, nested 10s, and controllers with no memory left of how they were configured. Where the card is the component that died, there is no reason to go hunting for an identical one, since everything required to rebuild the set is written across the disks rather than stored in the hardware that lost track of it.
Databases produce the difficult conversations. Power goes during a write on a SQL server, the transaction logs sit across the precise instant everything stopped, and a company is standing idle while this gets explained over the phone. Work of that kind is a discipline in itself. Reassembling the array in software comes first, and only then are the database files repaired until they are internally consistent, taking whichever tables the business genuinely depends on ahead of the rest, so people are working again long before the last terabyte has finished copying. Urgent jobs in Berkshire tend to look much the same each time: a firm off the A329M or on the Slough Trading Estate with thirty staff sitting on their hands, and the safest quick plan agreed before the call ends. Three habits account for more destroyed arrays than failing hardware does. Forcing a disk that has dropped out back into the set. Rebuilding onto a member that is visibly failing. And allowing one more person one more attempt. The correct response when a set goes down is short: cut the power, write the bay number on every disk, and pick up the phone before anything further is done to it.
One rule governs array work on this bench: every member gets copied before anybody offers an opinion about the layout. It has never been waived, for anybody, at any hour:
Members are cloned side by side on dedicated hardware, and that finishes before the first theory about disk order is put forward. All the risk in a reconstruction then lands on copies, and your disks sit on a shelf doing nothing until they go back in the post.
Disk order, stripe size, parity rotation and start offset are solved against the images and the volume is stood up in software. The geometry has to be demonstrated against real file content before it is accepted, not merely proposed.
Whichever disk stopped the array becomes a full drive recovery of its own: heads, motor, firmware modules, board work, whatever it takes to make it read consistently enough to image alongside the others.
A candidate layout is checked by running parity arithmetic across the images and by measuring how ordered the result looks. Both have to agree before extraction starts. Nothing here is released on the strength of a plausible-looking directory listing.
NTFS and ReFS, ext4, XFS, Btrfs and VMFS are parsed at structure level, along with any VMDK or VHDX files stored inside them.
No member can be written to while it is being read, at any stage of the job. Whatever you posted comes back bit for bit as it arrived, whether the reconstruction went well or badly.
Card, chipset or software layer, the geometry lives on the member disks and that is where it is read back from, so a dead controller never sends anyone shopping for a matching one. Three roads lead here and they all end at the same bench: batch-bought drives reaching the end of their lives in the same fortnight, desktop-grade members thrown out of the set under load, and a rebuild started before anybody stopped to think. Every one of them is handled the same way. All members imaged in parallel on day one, the reconstruction carried out on those copies, your original disks returned untouched. Array work opens at £500 + VAT and rises with the disk count, because each member is a recovery in its own right. The free diagnostic closes 2 working days after the drives are booked in at Guildford, and the figure that follows it is fixed in writing. Most of this work comes from the office parks along the A329(M) and Western Road, from firms on the Slough Trading Estate, and from IT providers looking after small networks across Bracknell Forest. Say on the call to 0800 689 0668 that the company has stopped trading and the job is dealt with first.
Shut the server down cleanly first, because nothing should be pulled out of a running chassis. Write the slot number on each disk as it comes out, one to eight or however many there are, and photograph the front of the unit before you start. The reconstruction depends on knowing which disk sat where, and thirty seconds with a marker pen saves an afternoon of arithmetic later. Only the drives travel: chassis, rails, cables and the controller card are of no use here and stay in your rack. Send the whole set rather than the member that failed, since no layout can be solved from part of one. Use a signed-for, insured service or a courier of your own, addressed to Guildford Data Recovery, Building 2, Ground Floor, Guildford Business Park, Guildford GU2 8XH. Nothing is collected and there is no Bracknell address to bring drives to. If you would rather hand them over, it is roughly 40 minutes on the A322 and the A3, and the Guildford reception takes drives in person on weekdays, 9:00am to 5:30pm.
Nearly every job here arrived as a parcel. Tracked, insured post is the calmest way to move a drive that is already struggling, and something handed over in Berkshire, Surrey or London is normally on the Guildford bench the next working day.
Is the storage still bolted into a machine — laptop, tower, iMac, MacBook, rack server, a DVR under the till? Free it first and send the bare unit. Stripping hardware is not something this lab does, though it is ten minutes' work for any repair shop on your high street. There is a single case with no way round it: memory chips soldered flat onto the mainboard, which is how Apple Silicon machines and certain ultrabooks are built. Where the storage cannot be unbolted, there is no parcel to make up.
↓ Print the shipping & booking-in form (PDF)
The name on the parcel wants to be Guildford Data Recovery. Driving it over from Bracknell is roughly forty minutes on the A322 then the A3; posting it costs you a stamp and a day. Either way, a message goes out to you as soon as it is logged onto the system, and two working days later the diagnostic is finished.
Unsure whether something should go in the box? Ring 0800 689 0668 while the lid is still open, or work through the free online diagnostic and let it tell you.
No member holds a whole file, so block order decides everything
Both halves imaged, then judged on contents rather than a status flag
One missing piece per stripe, and a rebuild that finishes it off
What the backup plan missed, written from the bench it reaches
Every postcode, worked down the A3 rather than over a counter
Nothing to pay for the diagnosis, one written figure before any work begins, and the band on this page is from £500 + VAT for an array, rising with the member count.