Ghost Access: When 'Locked' Doesn't Mean 'Blocked'
I spent most of this afternoon chasing a phantom login on a staging database server. The logs showed a successful pubkey authentication for a developer who transitioned out of the engineering team back in November. We had followed the standard offboarding checklist: the Jira ticket was closed, his Active Directory account was disabled, and his laptop was wiped. Yet, there he was, or at least his key was, sitting in the authorized_keys file of a shared deploy user.
This is the "Ghost Access" problem. It is a specific, annoying flavor of technical debt that accumulates every time someone joins or leaves a project. We often focus on the front door—the primary user account—while leaving side doors wide open. In many environments, especially those relying on legacy automation or shared service accounts, SSH key offboarding is the step that everyone assumes someone else is doing.
The core of the issue is how OpenSSH handles authentication. I saw a junior admin try to fix this by locking the user account last week:
# The "fix" that didn't actually work
sudo usermod -L developer_alias sudo passwd -l developer_alias
We hit a wall with this approach almost immediately. In our environment, and most modern Linux distros, usermod -L or passwd -l simply prepends a '!' or '*' to the password hash in /etc/shadow. It stops password-based logins dead. However, because our sshd_config is set up for PubkeyAuthentication yes, the SSH daemon often bypasses the PAM password stack entirely. If the public key is still in ~/.ssh/authorized_keys, the user gets in. The account is "locked," but the session is established anyway. It’s a false sense of security that leads directly to "Key Debt" during ISO27001 or SOC2 audits.
The Manual Grind: Locating and Revoking Keys
When you're dealing with a handful of servers, you might think you can just grep your way out of this. I started by trying to identify which keys belonged to whom. If your team doesn't enforce comments in the public key files, you're looking at a string of base64 nonsense with no context.
I ran this to at least get the fingerprints and see what we were dealing with:
ssh-keygen -lf ~/.ssh/authorized_keys
If you're lucky, the output looks like this:
2048 SHA256:nS... user@laptop-01 (RSA)
4096 SHA256:mP... dev-id-123 (RSA)
Once I identified the offending key for 'employee-id-123', I tried a quick sed command to strip it across the fleet. This is where things usually go sideways. If you have an immutable flag set on the file for "security," your automation will fail silently, and you'll spend an hour wondering why the key keeps reappearing.
# Attempting a quick removal
sudo sed -i '/employee-id-123/d' /home/deploy/.ssh/authorized_keys
If that command returns exit code 0 but the file doesn't change, check lsattr. I've seen plenty of "hardened" systems where chattr +i was applied to authorized_keys to prevent unauthorized changes, which ironically prevents authorized removals during offboarding. You have to chattr -i, run the sed, and then chattr +i again. It's a clunky, error-prone workflow.
Auditing the Aftermath
After you think you've cleared the keys, you need to verify if that ghost was actually active. I've been using a combination of auditctl and journalctl to track who is touching what. If you aren't monitoring changes to your SSH files, you're flying blind.
# Monitor the authorized_keys file for any writes or attribute changes
sudo auditctl -w /home/ubuntu/.ssh/authorized_keys -p wa -k ssh_key_change
To see who has actually been logging in recently via pubkey, I usually pipe the journal logs into something readable. It's a quick way to spot-check if an offboarded user's key is still hitting your infrastructure.
journalctl -t sshd | grep 'Accepted publickey' | awk '{print $9, $11}'
This gives you a list of the user and the fingerprint. If you see a fingerprint that should have been revoked weeks ago, you have a persistence problem. This persistent access is a prime candidate for lateral movement. An attacker—or a disgruntled former employee—only needs that one stale key on a jump box to start probing the rest of the VPC.
Scaling Challenges and the MSP Context
In larger environments, especially in the context of Managed Service Providers (MSPs) where teams rotate frequently, this becomes a nightmare. I’ve seen cases where "Jump Servers" are shared among dozens of contractors. The ec2-user or a generic deploy account becomes a dumping ground for keys. When a contractor leaves the MSP, the AD account is disabled, but the authorized_keys file on that jump box remains a graveyard of stale credentials.
This is where CVE-2015-5600 comes back to haunt you. While it’s an old MaxAuthTries bypass, having a massive authorized_keys file increases the attack surface for anyone trying to brute-force their way through multiple keys. More keys simply mean more chances for an exploit to find a path of least resistance.
The SSH Agent Forwarding Gotcha
Here is something that tripped us up last Friday: even if you remove the key from the server, if the user has an active session through a jump box with SSH Agent Forwarding enabled, they might still have a functional socket. Removing the key stops new connections, but it doesn't necessarily kill the current access if they've already tunneled in.
I had to start hunting down specific processes for the offboarded user to ensure the cleanup was total:
# Check for active sessions for a specific user even after key removalwho | grep 'developer_alias'
If found, you might need to kill those specific PIDs
ps -u developer_alias
Simply locking the account doesn't terminate existing SSH multiplexing (ControlMaster) sessions either. If they have an active socket, they can keep opening new channels without re-authenticating against the (now deleted) key.
Moving Toward Centralized Management
We've realized that managing local authorized_keys files is a losing battle. The drift is inevitable. The smarter move is to pull the authorization logic out of the local filesystem and into a managed access solution.
This tells the SSH daemon: "Don't just look at the local file; run this script to fetch the keys."
# /etc/ssh/sshd_configFetching keys dynamically to prevent local file drift
AuthorizedKeysCommand /usr/local/bin/get-keys-from-ldap.sh %u AuthorizedKeysCommandUser nobody
Ensure we don't allow loose permissions to bypass security
StrictModes yes
The script get-keys-from-ldap.sh would query your LDAP or IdP (like Okta or Azure AD) for the public key associated with the username %u. When the user is offboarded in LDAP, the script returns nothing, and the SSH login fails instantly across every server in the fleet. No sed, no chattr, no ghost access.
Short-Lived Certificates: The Real Fix
If I had my way, we’d kill static keys entirely. Static keys are effectively infinite-life passwords. As noted in this guide on automating rotation, the industry is moving toward SSH Certificates, and for good reason. Instead of a key that lives forever, the user authenticates to an identity provider, gets a certificate signed by a trusted CA, and that certificate expires in 8 hours.
When the employee leaves, you don't have to clean up authorized_keys because their ability to get a new certificate is revoked at the IdP level. Their old certificate simply expires by the time they get home. It turns the offboarding process into a single point of failure (in a good way) rather than a scavenger hunt across 500 nodes.
The Offboarding Checklist (For the Rest of Us)
Since we aren't fully on certificates yet, I’ve had to standardize a "Manual-ish" workflow that’s a bit more robust than just disabling the AD account. If you’re stuck managing static keys, this is the bare minimum:
- Expire the account: Use
usermod --expiredate 1 <username>. Unlike-L, this actually prevents the account from being used for any service, including those that bypass PAM password checks. - Grep by key string, not comment: Users change their email addresses or comments. Target the actual public key string when automating cleanup.
- Kill the sessions: Don't just remove the key; check for active PIDs and open SSH sockets.
- Audit the shared accounts: The
deploy,jenkins, androotaccounts are where ghost keys go to hide. These need a monthly audit, not just an exit-interview trigger.
I'm still seeing some drift between our IaC (Terraform) state and what’s actually on the boxes. I think the next experiment will be a small Go binary that runs as a cron job, compares the local authorized_keys against a "Source of Truth" S3 bucket, and hard-resets the file if it finds an unmanaged key. It’s aggressive, but it’s the only way to stop the "quick fix" keys that devs keep adding manually when they're in a rush.
Wait, I just saw another login from a 'temp-contractor' who left in July. Back to the logs.
Next experiment: Testing AuthorizedKeysCommand with a local cache for when the LDAP server inevitably goes down during a network blip.
