Backup & continuity
Backup monitoring for MSPs: why a green tick is not a restore
Every MSP has a story about the backup that was green every night for six months and then would not restore. Backup monitoring for MSPs is usually treated as a solved problem: the job ran, the tick is green, move on. But "the job ran" and "the client's data is recoverable" are two different claims, and the distance between them is where businesses get hurt.
The claim a green tick actually makes
When Veeam, Acronis or any other backup product reports success, it is telling you one thing: the job completed without a fatal error. That is worth knowing, but notice everything it does not say. It does not say the backup covered the data that matters. It does not say the backup chain is intact or that the storage it landed on is healthy. It does not say the copy is free of the ransomware that has been sitting quietly in the environment for a fortnight. And it says nothing at all about whether anyone could actually restore it under pressure.
The failure modes that end up in incident reports are rarely loud. They are the quiet ones:
- The job that stopped running. A schedule gets disabled during maintenance, a service account password expires, and no failure alert fires, because nothing failed. Nothing happened at all.
- The scope that drifted. The job still protects the file server it was pointed at in 2023. The line-of-business database that moved to a new VM last year is not in any job.
- The chain that broke. Incrementals keep succeeding against a corrupt full. Every night is green, and every night the estate is one bad block away from an unrecoverable chain.
- The retention that silently shrank. The repository filled up, old restore points were pruned, and the "30 days of history" in the contract quietly became six.
None of these show up as a red job. All of them show up on the day you need a restore.
3-2-1 was never the whole answer, and it has grown two digits
The classic 3-2-1 rule still holds: three copies of the data, on two different media, with one copy off site. It is a good test of architecture. But it describes where copies live, not whether they work, which is why the industry has largely moved to 3-2-1-1-0: the extra 1 is one copy that is offline or immutable, so ransomware cannot encrypt the backups along with the production data, and the 0 is zero errors on verification, meaning backups are tested and confirmed restorable, not assumed.
That final zero is the part most MSP backup setups are missing. Architecture is a design decision you make once. Verification is an operational habit you keep every week, and habits are harder to sell, harder to schedule and easier to drop when the queue is full. Which is exactly why it should be systematised rather than left to good intentions.
What backup monitoring for MSPs should actually check
A monitoring setup you can stand behind checks four layers, in order of how cheap they are to automate:
- Job outcomes, including silence. Alert on failures, obviously. But also alert when an expected job simply does not report. "No result in 26 hours" must be treated as seriously as "failed", because the stopped job is the more dangerous of the two.
- Coverage against inventory. Reconcile the list of protected machines against the list of machines that exist. Every device in the estate should either be in a backup job or explicitly marked as not requiring one. New servers and rebuilt VMs are how scope drift starts.
- Restore point age and retention. Track the age of the newest restore point per protected system, and the depth of history, against what the contract promises. A green job with a three-week-old restore point is a red finding.
- Verification results. Use the vendor's built-in verification where it exists, health checks and scheduled verify jobs in Veeam, validation in Acronis, and record the results centrally. A verification job that never runs is a policy, not a control.
Restore tests that fit in a real week
Full disaster recovery rehearsals matter, but if the bar for testing is "spin up the whole client from bare metal", testing will happen once a year at best. The practical answer is a tiered cadence that trades depth for frequency:
- Weekly, automated: vendor verification jobs and a scripted file-level restore of a handful of files from a random protected machine, checksummed against the source.
- Monthly, lightweight: boot one server backup as an isolated VM and confirm the operating system and key application actually start. Fifteen minutes per client, rotated so every critical system is exercised over a quarter.
- Twice a year, per critical client: a timed restore of the most important workload, with the result written down: how long it took, what was missing, what surprised you. That number is your real recovery time, whatever the proposal document says.
Make the evidence do double duty. Every restore test produces proof: a screenshot of the booted VM, a checksum match, a timed run. Save it and put it in the client's next review. "We restored your finance server in 41 minutes last month" is the most persuasive sentence in any QBR, and it costs nothing extra once testing is routine.
Turn backup state into daily ops, not a monthly report
The last piece is organisational. Backup health should be visible in the same place your technicians already work, not in a separate console per vendor that someone checks when they remember. A failed or silent job should become a ticket with an owner and an SLA, the same as a down server. The counterweight is discipline about noise: one warning per chain of related failures, not forty emails from four consoles. The techniques in our guide to cutting alert noise without missing real incidents apply to backup alerts more than anything else, because backup noise is precisely the kind that trains people to stop reading.
Get those pieces in place, coverage reconciliation, silence detection, retention tracking, tiered restore tests and single-pane visibility, and backup stops being a leap of faith. It becomes a control with evidence behind it, which is what clients think they are already paying for.
Where this fits with Helios
Helios monitors Veeam and Acronis backups from the same agent that already watches patching, antivirus and system health, so backup state sits in the one dashboard your team actually looks at. Missed and silent jobs surface as alerts rather than as absences, and Helio, the AI layer, folds backup findings into device investigations, so "when did this machine last have a good restore point" is a question with an instant answer rather than a console safari. The restore testing habit is still yours to build, but the watching, reconciling and chasing is work software should be doing.
Know your backups will restore before you need them
Helios is an AI-native platform for MSPs: backup monitoring for Veeam and Acronis alongside patching, security and service desk, with a 14-day trial and no feature gating.
Start free