The Mumbai Bastion Incident and the Problem with Agent Forwarding
I spent most of this morning tracing a weird lateral movement alert in the Mumbai VPC. We have a shared jump box that two different offshore MSP teams use to reach their respective client environments. One of the junior devs on the vendor side had ForwardAgent yes set in their global ~/.ssh/config because they were tired of typing passphrases for every downstream microservice. It works, it's seamless, and it's also a massive security hole that I'm still cleaning up.
When you use SSH Agent Forwarding, you aren't sending your private key to the remote server. Instead, the SSH process creates a Unix domain socket on the remote machine and tunnels requests back to your local ssh-agent. The local agent does the signing, and the remote server just passes the challenge-response along. On the surface, it feels safe because the private key never leaves your laptop.
The reality is messier. I hit this realization when I saw how easy it is for anyone with root access on that intermediate jump box to impersonate a connected user. If I'm root on the bastion, I don't need your key. I just need your socket.
Understanding the SSH Agent Forwarding Security Risk
The core risk is socket hijacking. When you run ssh -A user@bastion, the SSH daemon on the bastion creates a socket file, usually tucked away in /tmp. If I'm an attacker who has gained root on that bastion, I can just look for those sockets.
# I ran this on the bastion as root just to prove the point
ls -al /tmp/ssh-XXXXXX/agent.XXXX
srw------- 1 victim_user victim_group 0 Oct 27 14:30 /tmp/ssh-tY4kLp/agent.1234
Even though the permissions say only the user can read it, root doesn't care about permissions. An attacker can simply point their own SSH_AUTH_SOCK environment variable to that path and start using the victim's authenticated identities to hop into production databases or internal git repositories.
# Attacker perspective on the bastion
export SSH_AUTH_SOCK=/tmp/ssh-tY4kLp/agent.1234 ssh prod-db-server.internal
The victim is still logged in, their agent is still active, and now the attacker is essentially "in" as the victim. This is exactly how lateral movement happens in MSP environments where one vendor's compromise leads to another vendor's infrastructure getting poked. It's a "trust" vulnerability—you are trusting every administrator of the intermediate server with the ability to act as you for the duration of your session.
Why Agent Forwarding is Usually "Broken" or Unnecessary
I often get tickets asking, "Why is SSH Agent Forwarding not working?" Usually, it's a mix of environment variable mismatches or sshd_config restrictions. But my answer is becoming: "Good, it shouldn't be working. We need to move you to ProxyJump."
Common failures I see include:
- Environment Variables: The
$SSH_AUTH_SOCKisn't being passed into asudosession or atmuxpane, leaving the user with "Permission Denied (publickey)". - Server-side Restrictions:
AllowAgentForwarding nois set in the global/etc/ssh/sshd_config, which is a sensible default we've started pushing. - Identity Management: The user has too many keys loaded (
ssh-add -l), and the remote server hitsMaxAuthTriesbefore the right key is offered.
Instead of fixing these, we should be addressing the underlying architecture. We also have to consider CVE-2023-38408. This was a nasty one where an attacker could achieve Remote Code Execution (RCE) via the ssh-agent if forwarding was enabled, specifically by leveraging how PKCS#11 providers are loaded. If that doesn't convince you to turn off -A, I don't know what will.
The Better Way: ProxyJump and ProxyCommand
Since OpenSSH 7.3, we've had ProxyJump (the -J flag). It's the "correct" way to handle multi-hop access because it doesn't expose your agent socket to the intermediate server. Instead, it uses the bastion as a simple TCP relay. The end-to-end encryption happens between your local machine and the final destination server.
Here is the failed approach I saw in a dev's config last week:
# DANGEROUS - DO NOT DO THIS
Host * ForwardAgent yes
Host bastion HostName 10.0.1.50 User jumpuser
Instead, I pushed this update to our internal wiki. It forces the use of ProxyJump and explicitly disables agent forwarding to minimize the attack surface.
# The safer way
Host bastion HostName 10.0.1.50 User jumpuser
Host 10.0.2.* ProxyJump bastion User produser ForwardAgent no AddKeysToAgent confirm
If you're on the command line, it's just as simple:
ssh -J [email protected]:22 [email protected]
Dealing with Legacy Systems
I hit a snag yesterday with some old RHEL 7 boxes that were still running OpenSSH 6.x. ProxyJump isn't available there. In those cases, you have to fall back to ProxyCommand using nc (netcat). It's uglier, but it achieves the same goal of not exposing the agent socket.
Host legacy-app
HostName 10.0.2.55 ProxyCommand ssh [email protected] nc %h %p
Restricting the Agent: Confirmation and Constraints
Sometimes you actually do need the agent (e.g., performing a git clone on a remote server using your local keys). If you absolutely must use it, don't just leave it wide open. I've been experimenting with the confirm flag.
ssh-add -c ~/.ssh/id_ed25519
When you add a key with -c, the agent will require a manual confirmation for every signature request. This means if an attacker tries to use your hijacked socket, a dialog box will pop up on your local machine asking for permission. If you aren't currently trying to log in somewhere and a prompt appears, you know something is wrong.
Gotcha: For this to work in a headless or CLI-heavy environment, you need an SSH_ASKPASS helper configured. If you don't have a GUI-based askpass binary set up, ssh-add -c might just fail silently or hang when it tries to prompt you.
Destination Constraints (OpenSSH 8.9+)
If you're running a modern OpenSSH version, you can actually restrict keys to specific destinations. This is a game changer for minimizing the impact of a hijacked socket. You can tell the agent, "This key is only valid for reaching these specific IPs."
# Restrict key to a specific destination
ssh-add -h 10.0.0.5 ~/.ssh/id_ed25519
This way, even if someone steals the socket, they can't use your identity to pivot to the entire subnet—only the host you specifically authorized.
Auditing and Monitoring
We can't just set these configs and hope for the best. I've been building out a small dashboard to monitor socket usage on our jump hosts. While you can't easily see inside the encrypted SSH traffic, you can audit the login attempts and the socket creation logs.
# Check who is logging in via the jump host
journalctl -u ssh | grep 'Accepted' | grep 'port'
In the Mumbai incident, I was able to correlate the logs from the jump host with the access logs on the internal database. We saw a mismatch: the user logged into the bastion was dev_user_a, but the downstream database saw a connection from dev_user_a at a time when that dev was supposed to be at lunch. The hijacked socket was the only explanation.
I've also started using hardware security keys (FIDO2/U2F) for our most sensitive production environments. When the private key is physically backed by a YubiKey (e.g., [email protected]), the "hijacking" risk changes. Even if the socket is hijacked, the attacker still needs a physical touch on the device to sign the request, assuming you haven't cached the touch requirement.
Moving Forward
I'm still cleaning up the leftovers from the Mumbai MSP breach. The immediate fix was easy: we updated the global sshd_config on all bastions to AllowAgentForwarding no. The pushback from the devs was loud for about two hours until they realized ProxyJump actually made their lives easier by simplifying their ~/.ssh/config files.
The next experiment is to see if we can completely eliminate long-lived SSH keys in favor of short-lived certificates issued by a central CA. If the certificate only lasts for 15 minutes, the window for socket hijacking becomes so narrow it's almost not worth the attacker's effort. But that's a project for next quarter.
For now, if you see ssh -A in your command history, maybe ask yourself if you really trust the person who has root on the other end of that connection.
Open question for the team: Has anyone had luck getting ssh-add -c to work reliably across WSL2 and Windows native OpenSSH without the askpass dialog getting buried under other windows?
