Website hacked again after cleaning, and the scan says clean
A site was cleaned several times and kept being reinfected while the scanner reported clean. Comparing every file with the official release found 211 files.
Your website is cleaned, and a few weeks later it is hacked again. You pay to have it cleaned again and it happens again. Each time, the security scan your host runs comes back clean. So the scan is telling you everything is fine while customers, or Google, are telling you otherwise.
The cost is the obvious one, repeated: the site serving something you did not put there, the trust that goes with it, and the clean-up bill each time. It also gets worse in a quieter way, because every clean report teaches you to stop looking.
We met exactly this. A site kept getting reinfected after cleaning. We ran the usual scanners, the kind that look for code that resembles known malware. Zero hits, a clean report. Then we compared every file in the installation against the official release of that version, and the comparison returned 211 files that should not have been there. Zero and 211, on the same folders, on the same afternoon.
What was actually going on
A scanner asks whether a file looks like malware it has seen before. That is a question about appearance, and appearance is the one thing an attacker controls completely. Anyone writing a backdoor for a widely used website platform knows which scanners the owner will run, because they are free and public, and they test against them until the code passes. The owner is always scanning for last month’s attack.
A file comparison asks something different: is this the file the software maker shipped? That has a plain answer that does not depend on what the file contains. Nobody can write code that is both a working backdoor and identical, byte for byte, to a file in the official release.
The files were also hidden in ways aimed at checks people trust. The platform loads certain specially named files on every single request without anyone switching them on, and they never appear in the list of add-ons. One had been replaced by a restorer, code whose only job was to check whether the malicious files were still there and put them back if not. That is why the site kept being reinfected: each clean-up removed the symptoms and left the thing that reproduced them.
One add-on was registered as active in the database but did not appear in the admin screen, because it had removed itself from the list it was shown in. And every backdoor carried a date more than a year old. The standard move on a hacked site is to sort by recently changed, and with backdated files that puts the malicious ones at the bottom of a list of thousands and quietly returns the wrong answer. Files were also sitting in places they should never be: stray program files next to the real entry points, folders with random names, and runnable code in the folder meant only for uploaded pictures.
What we changed
The order of checks changed. The file comparison against the official release now goes first on any site we suspect, and a clean scanner report is treated as information about the attacker, not about the site. A scan with zero hits tells you the attacker was competent. It does not tell you the site is clean.
After that we check the things no list shows: the files loaded on every request, the add-on records in the database compared with what the screen shows, and anything runnable in the upload folders.
What it did not fix
The comparison only covers files that came from an official release with published fingerprints. Custom themes, bespoke add-ons, uploaded media and settings files have nothing to compare with, and need another baseline, such as your own version history or a record taken when the site was known to be good.
It says nothing about what is stored in the database: spam pages saved as posts, redirects hidden in settings, administrator accounts added by the attacker. And it cannot tell a forbidden change from a permitted one, so on a site where people edit files by hand it produces noise until they stop. That is arguably a benefit, since it forces the discipline that makes the check meaningful.
This checking tells you what is on the site today. It does not reduce how much runs there.
The pattern, for anyone whose site keeps being reinfected
Ask who or what is vouching for the site being clean, and whether that witness sits inside the thing being checked. A scanner run by your host, the plugin list, the admin dashboard and the dates on files are all the site describing itself, using code inside the boundary the attacker has already crossed. A hacked site is under no obligation to describe itself truthfully.
The test worth applying anywhere is where the reference copy lives. If it lives outside what you are inspecting, such as the maker’s published release, the check means something. If it lives inside, you are interviewing the system, not assessing it. Never let a system be the only witness to its own integrity.
Where this ends up
Checking what is on the host today does not reduce how much executes there. That is a decision about the platform, and moving off one that nobody wants to touch and nobody can switch off is the work described under enterprise modernisation.