The awkward thing about shared storage is that when it fails, everything sitting on top of it fails in the same instant, and the phone call that follows has a particular sound to it. Broken pool metadata, LUNs that will not come online, VMFS datastores that refuse to mount: all of that is rebuilt here for companies and IT providers across Bracknell, Reading, Slough and the London side of the M25, whether the job is a single shelf or a full rack. Enterprise storage is ordinary weekly work on this bench, not a favour fitted in around something else.
Anything sent in for SAN recovery is diagnosed free. The written figure that follows is settled before a screwdriver leaves the drawer: from £500 + VAT for a SAN, rising with the member count.
No fix, no fee covers logical recoveries. Outside it sit electronic and mechanical failures, chip-level work, DVR jobs and forensic jobs, and physical work is 50% up front. Every band is published on the data recovery cost page.
Reading a symptom back to the failure underneath it is where the work genuinely starts, and this set of thirty accounts for all but a handful of the parcels that reach the Guildford bench. If yours is not on the list, it will still be recognised — describe it on the phone and you will get a straight answer about the odds before you post anything.
Every host that was built on that block device has gone blind at once, which is why the phone call has the tone it does. Rebuilding runs from the floor upwards: RAID groups first, then the pool that sits on them, then the LUN, then whatever file system was riding on top. Start anywhere else in that ladder and you get a result that looks convincing and proves nothing.
The structure mapping physical extents to volumes has broken. Everything it was describing is still sitting untouched on the disks, so the map is rebuilt from the very structures it was supposed to be describing, which is slower than reading it but rather more dependable.
A fully populated shelf does not lose members one at a time in the polite order the design assumed. Each failed disk is repaired and imaged individually on the mechanical bench, and only when all of them read consistently does anyone begin doing arithmetic across the group.
Firmware updates fail from time to time, and one that fails badly enough leaves the management plane dead for good. Not a byte of the data behind it has moved. Recovery ignores the controller altogether and reads the drives directly, which is usually quicker than waiting on a vendor to ship a replacement head for hardware they stopped making years ago.
Two heads sharing a chassis, a production batch and a firmware revision have a habit of going within hours of one another, which is not the scenario the redundancy was sold against. Once both are down the work goes at the disks directly. Nothing in a shelf of drives knows or cares what has happened to the controllers in front of it.
Disks fine, groups fine, datastore refusing to mount. The virtualisation layer has collapsed on storage that is in good order underneath it. The datastore is parsed by hand from the LUN image and each virtual disk lifted out and repaired on its own, in whatever order the business needs them.
A provisioning task aimed at the wrong volume, and a LUN disappears in the time it takes to confirm a dialogue box. How much of it survives depends on what the array has written over the top since. Taking the platform out of production is therefore the one useful thing to do while everybody is still deciding what happens next.
Snapshots depend on their parents, so a single damaged link takes down every volume hanging below it, often dozens of them. The chain is rebuilt from the images link by link, in the order it was created, and the volumes come back in the same order.
Promising more capacity than the shelf physically holds works beautifully right up until the morning it does not, and then every thin volume drawing on that pool is damaged in the same minute. Rebuilt from images, the allocation tables can be put back into a state that survives inspection instead of one that merely mounts.
Writes acknowledged to the hosts and never committed to disk leave an entire RAID group with parity that contradicts its own data. Settled on copies, that contradiction can be examined carefully. Settled on the live array by a controller in a hurry, it tends to be settled destructively.
NetApp aggregates and Dell EMC pools stop knowing where the volumes inside them live. Reconstruction begins down at the RAID groups and climbs one layer at a time until the volumes are visible again, with each layer proved before the next one is attempted.
Enterprise disks are formatted with extra bytes per sector for integrity data, which defeats ordinary tooling by design rather than by accident. The bench here reads those formats natively, without an improvised adapter and a hopeful attitude.
HPE's small-business array turns up with its vDisk metadata damaged after a shelf reboot or a controller swap. It is a rebuild this workshop has done often enough to find dull, and dull is the correct state for a rebuild to be in.
V3700 and V5000 shelves lose extents after a power failure and the volume then refuses to assemble. Extents are remapped from the drive images, patiently and in order, until the volume is whole and its file system parses.
The box has forgotten the target definition, so nothing can log in. The volume behind that definition very rarely goes with it. It is extracted from the images and handed back over plumbing that can be trusted, rather than reconstructed on the array that just lost it.
A dedupe index is a small structure carrying an enormous amount of responsibility, since every reference to every shared block runs through it. Damage it and volumes that were healthy a moment earlier stop resolving to anything at all. Rebuilding it until those references land on real blocks again is what brings the volumes back with them.
A failover that went wrong leaves both heads writing independently to the same back end, so the array is holding two accounts of the same afternoon. The two are reconciled from images, timestamped, with the reasoning written down so your own engineers can check it afterwards.
End-of-support arrays fail with no official route left open, and the quote for an emergency support contract usually arrives before the diagnosis does. Lab work carries no renewal date. The disks were never told the contract had lapsed and they read exactly as they did the week before.
Drop a disk running a newer firmware level into an elderly group and it behaves oddly in ways that look exactly like a drive on its way out. Stabilise the group, image it, reconstruct afterwards. In that order, because the other order removes the evidence needed to tell the two apart.
Edit the fabric or the LUN masking and every host loses sight of its storage at once, even though the volumes underneath are in perfect condition. That mapping is rebuilt and handed back to the people who own the switches, and no disk needs to be touched to do it.
Where a batch of drives shares a power-on-hours defect, the members reach the fatal figure within days of one another. This has happened in public more than once and it is a manufacturing fault rather than misfortune. Surviving members are copied straight away, at whatever speed the shelf will sustain, before the rest of them arrive at the same hour on the clock.
Nearline SATA disks sitting behind SAS interposers fail at the adapter far more often than at the media. The small board gets taken out of the path and the disk imaged on its own, and the majority of them read perfectly once it is gone.
Automatic tiering keeps the most active blocks on flash and leaves the rest on spinning disk, so the slow tier on its own is full of holes. The tiering map and whatever tiers survived are used together to put the hot blocks back where they belong, and anything genuinely lost with the flash is listed by name rather than quietly left out.
Without a clustered file system between them, two hosts given the same volume will each behave as its sole owner and write accordingly. Separating the two histories happens on the images, one operation at a time, and every step of it is checked against timestamps rather than against optimism.
A dead service processor, or credentials that left with a contractor whose number nobody has any more. The array will not talk to anybody at all. The disks inside it have never asked for a password and do not start now.
The space set aside for snapshots is fixed, and when it fills most platforms respond by throwing away the oldest ones without asking. The restore point somebody was relying on went weeks before anyone looked for it, and no message was sent. What remains is carved out of unallocated space on the images, and that is frequently a good deal more than expected.
A failover test, a site labelled incorrectly, or a relationship restarted in reverse, and the stale copy lands on top of the good one at wire speed. That side gets imaged and the earlier structures lifted back out from underneath. Stopping the replication in the first ten minutes is worth more than anything a laboratory can do afterwards.
An array encrypting at rest holds its keys in the controller or on a key manager standing next to it. If both have gone, the disks read perfectly and the contents mean nothing at all. That verdict comes out of the free assessment and it is given straight, because letting anybody spend money on it would be dishonest.
Enclosures are daisy-chained, so one dead expander, one failed cable or one shelf powered up in the wrong order can take an entire tier of disks off the map at once. On the console it reads as total loss. In the workshop it turns out to be copper more often than media, and working out which of the two you have is the first hour of the job.
Present a fresh server with several volumes and sooner or later somebody will initialise whichever one looked unused. A quick format writes very little, so what lies beneath survives largely intact — provided that host is disconnected before it starts filling space it believes is empty.
Nothing in a SAN sits directly on a disk. Drives are grouped into RAID sets, the sets go into a pool, the pool is carved into LUNs, and only at the top of all that do the file systems and hypervisor datastores anybody uses appear. Recovery follows the same ladder in reverse and then back again. First every drive is imaged, without exception. The RAID sets are reconstructed from those images. The pool metadata is repaired until it parses. The LUNs are extracted. Nobody opens a virtual machine or a database until each of those steps has finished. Rarely is the hardware itself unusual: most of what arrives carries a Dell EMC, NetApp, IBM, Fujitsu or HPE MSA and EVA badge, and whatever confidentiality agreements, change control and internal sign-off your organisation runs on can be threaded through that sequence without costing it a day.
Delivering a customer several terabytes of raw blocks is no help to anybody. What is actually needed is the small number of systems without which the business cannot trade, working again by the opening of the next business day. So datastores are mounted on top of the reconstructed LUNs, the VMDK and VHDX files are pulled out and repaired where repair is needed, and any SQL or Exchange workload named during the first phone call is dealt with ahead of the rest. Staff get back to work while the remaining volume copies away quietly behind them. That is triage, not heroics, and it is carried out at a deliberate pace, because hurrying an enterprise job is precisely how the second mistake follows the first.
Shared storage takes specialist connections as well as specialist knowledge, and there is no point holding one without the other:
Bare drives, populated shelves and expansion enclosures all connect to fabric kept here for the purpose, and the non-standard sector formats enterprise disks ship with are read as they are rather than converted into something more convenient first.
Entire RAID groups are captured alongside one another, after any casualty has been round the mechanical and firmware benches first. Nothing is reconstructed until the last member has finished copying.
Vendor metadata is rebuilt in the order the array originally built it — groups, pools, LUNs, then file systems — with each layer verified against real content before the next is attempted.
Once a LUN is back, the datastore on it is mounted and read out, and the guest disks are extracted and repaired one at a time in whatever order the business needs its systems returned.
Say on the first call which systems the company cannot trade without — SQL Server, Exchange, a line-of-business database — and those are verified and released first, while the long tail carries on copying quietly behind them.
Originals sit behind write blockers from booking in to despatch, whether that is one bare drive or a shelf with twenty-four in it. There is no stage of the job where that rule is set aside to save time.
An unfamiliar badge rarely means unfamiliar hardware, because the enterprise vendors buy from a small handful of suppliers and most of what arrives has been apart on this bench before under a different label. The method does not change much between them. Image every drive, rebuild the RAID groups, repair the pool until it parses, extract the LUNs, then mount the datastore and lift the VMDK or VHDX files out one guest at a time. Dead controllers are simply left out of it. Sector sizes of 520, 524 and 528 bytes are read as they are. Thin provisioning, tiering and deduplication all come apart the same way, provided nobody has been persuaded to run a vendor repair tool first. Enterprise recovery opens at £500 + VAT and rises with the number of disks, with the figure confirmed in writing after a free assessment that closes 2 working days from booking in. The work arrives from data rooms along the M4 corridor between Reading and Slough, from managed service providers around Camberley and Frimley, and from firms either side of the A329(M) whose entire business sits on one shelf.
Position matters more in an enterprise layout than almost anything else that comes through the door, so mark each disk with its enclosure and its slot before it leaves the rack, and photograph the front of every shelf while the drives are still in it. Then send the marked drives on their own. Enclosures, controllers, rails and cables are needed where they are and are no use in Guildford. Where a volume spanned more than one shelf, every shelf has to be represented, because a pool cannot be rebuilt from half its members. Packaging should stop the drives moving against each other; the original foam trays are ideal if you still have them. Send by an insured, signed-for service or a courier your own company books, to Guildford Data Recovery, Building 2, Ground Floor, Guildford Business Park, Guildford GU2 8XH. There is no collection service and nowhere in Bracknell to deliver to. An engineer who is passing can hand the parcel in at the Guildford reception during office hours on weekdays, 9:00am to 5:30pm, and the journey is around 40 minutes by way of the A322 and the A3.
Nearly every job here arrived as a parcel. Tracked, insured post is the calmest way to move a drive that is already struggling, and something handed over in Berkshire, Surrey or London is normally on the Guildford bench the next working day.
Is the storage still bolted into a machine — laptop, tower, iMac, MacBook, rack server, a DVR under the till? Free it first and send the bare unit. Stripping hardware is not something this lab does, though it is ten minutes' work for any repair shop on your high street. There is a single case with no way round it: memory chips soldered flat onto the mainboard, which is how Apple Silicon machines and certain ultrabooks are built. Where the storage cannot be unbolted, there is no parcel to make up.
↓ Print the shipping & booking-in form (PDF)
The name on the parcel wants to be Guildford Data Recovery. Driving it over from Bracknell is roughly forty minutes on the A322 then the A3; posting it costs you a stamp and a day. Either way, a message goes out to you as soon as it is logged onto the system, and two working days later the diagnostic is finished.
Unsure whether something should go in the box? Ring 0800 689 0668 while the lid is still open, or work through the free online diagnostic and let it tell you.
Nothing to pay for the diagnosis, one written figure before any work begins, and the band on this page is from £500 + VAT for a SAN, rising with the member count.