A drive or server has failed — what happens now

A drive or a server has failed: what happens now

A drive has failed, or a server will not start. This is one of the few genuinely bad days in small business IT, and it is worth knowing in advance how it is handled, because knowing the shape of it makes the day much less frightening.

First, the calm part. Hardware fails. It is expected, it is planned for, and in most cases the data is recoverable. Many servers and NAS devices are built so that one drive can fail without losing anything at all — the device keeps running on the remaining drives while the failed one is replaced. If that is your situation, you may not even notice a difference in your day.

What we do first

  1. Stop making it worse. The most damaging thing that can happen to failing storage is repeated attempts to power it on and off. Each attempt can turn a recoverable fault into an unrecoverable one. So the first instruction is usually to leave it alone.
  2. Work out what has actually failed. A server that will not start is not necessarily a server with dead data. Power supplies, memory, controller cards and operating systems all fail in ways that look identical from the outside and have very different consequences.
  3. Find out what is still protected. We check the backup: what it covers, when it last ran successfully, and what point in time we can restore to. This is often the moment the whole picture changes from alarming to manageable.
  4. Agree the priority order with you. Rarely does everything need to come back at once. Usually there are one or two systems the business genuinely cannot trade without, and everything else can follow. You decide that order, not us.
  5. Recover, in that order. Either by repairing the hardware, by restoring from backup to different hardware, or by standing the critical service up somewhere temporary while the permanent fix is arranged.

Why we work on a copy wherever we can

Where the storage is readable at all, our strong preference is to take a copy of it first and then do the recovery work on the copy.

The reason is simple. A drive that is partly failing may only give you one good read before it deteriorates further. If we spend that one good read making a complete image of it, we can then attempt recovery on that image as many times as we like, and a failed attempt costs nothing. If instead we experiment directly on the original and the drive dies mid-attempt, there is no second try.

Taking the copy takes time, and it can feel like nothing is happening. It is the most valuable part of the process.

Why we cannot tell you how long in the first hour

You will want a time. Everyone does, and it is a reasonable thing to want. The honest answer in the first hour is usually that we do not know yet, and anyone who gives you a confident figure at that point is guessing.

It depends on what failed, whether the data is readable, whether a replacement part is on the shelf or has to be ordered, and how much data has to be copied back — and copying takes as long as the volume of data dictates, no matter how urgent it is.

What we will do instead of guessing is give you a first assessment as soon as we have one, tell you what we know and what we do not, and update you as each unknown is resolved. Once the picture is clear, you will get a real estimate. We would rather give you an honest "not yet known" than a number that quietly falls apart at four o'clock.

What you can usefully do meanwhile

  • Do not power-cycle the equipment and ask everyone else not to either. This genuinely matters.
  • Work out what your business can do without today. If quoting can wait but invoicing cannot, tell us. That directly changes the order we work in.
  • Fall back to what still works. Email, mobile phones, and any cloud systems are usually unaffected by a local server failure. Keep taking orders and keep talking to customers on paper if needed.
  • Write down anything created while systems are down — orders taken, jobs booked, payments received. After the recovery, that list is what fills the gap between the last backup and the failure.
  • Nominate one person to talk to us. Several people ringing separately with slightly different versions slows everything down. One point of contact who can make decisions is worth an hour.
  • Tell your team something. Even "the server has failed, we are working on it, use your phone and write things down" stops the speculation.

Afterwards

Once things are running, we will write up what failed, what was recovered, anything that was lost, and what would reduce the impact next time. Sometimes that is a hardware change, sometimes it is a change to the backup, sometimes it is nothing — the arrangement worked as designed and it simply took the time it took.

Still stuck?

Log a ticket at portal.yougrowit.com.au or email support@yougrowit.com.au. If it is stopping you working right now, call 03 9028 4358.

Support hours are Monday to Friday, 8:30am to 5:30pm Melbourne time, excluding Victorian public holidays.

    • Related Articles

    • A mapped network drive keeps disconnecting

      You click the shared drive, the one that might be S: or Z: or P:, and Windows tells you the network path cannot be found. Or the drive shows a small red cross beside it and nothing opens. This happens with every brand of storage, whether your files ...
    • The shared drive has disappeared

      You open File Explorer and the drive you use every day is gone, or it is still listed with a red cross next to it and will not open. Almost always this means your computer has lost its connection to the server holding those files. The files ...
    • Storage is nearly full: what to do and what not to delete

      A warning appears saying the shared drive or the NAS is nearly full, or people start getting errors when they try to save. This is one of the most common tickets we see, and it is not a sign anyone has done anything wrong. Business data grows, and ...
    • Testing a restore before you need it

      Every business we look after has a backup. Far fewer have ever seen one come back. That gap is worth closing, because the two are not the same thing. An untested backup is a hope, not a plan A backup job that reports success every night is ...
    • What we back up, how often, and how a restore works

      Most people only think about backups on the day they need one. This article explains, in plain terms, what a backup actually is, what it is not, and what you can reasonably expect when you ask us to get something back. A sync is not a backup This is ...