The Jump Logic Disaster in /etc/pam.d/sshd
I spent two hours tracking down why an Ubuntu 22.04 box let me log in with just an SSH key and an account password, despite our team ostensibly enabling TOTP across all jump hosts. The prompt for an authenticator token simply never appeared. No error in /var/log/auth.log, no pam_google_authenticator failure, just a straight shell drop.
The host was supposed to be strictly hardened. We had compliance folks asking for proof of multi-factor enforcement on all administrative bastions—part of our baseline alignment with standard regulatory mandates like the CERT-In guidelines for remote privileged access. On paper, the config was done. In practice, the PAM stack had an invisible bypass.
Here was our initial "quick fix" appended to the bottom of /etc/pam.d/sshd:
# /etc/pam.d/sshd
@include common-auth account required pam_nologin.so @include common-account @include common-session @include common-password auth required pam_google_authenticator.so nullok
It looked reasonable at a glance. Run system auth, check account state, apply sessions, and enforce MFA. Except it failed completely. Entering a valid password bypassed the token prompt entirely.
To see why, you have to look inside the generated /etc/pam.d/common-auth file that Debian and Ubuntu maintain via pam-auth-update:
# /etc/pam.d/common-auth
auth [success=1 default=ignore] pam_unix.so nullok auth requisite pam_deny.so auth required pam_permit.so
PAM evaluation is procedural, and the bracketed syntax [success=1 default=ignore] is explicit jump logic. It tells the PAM engine: "If pam_unix.so returns PAM_SUCCESS, skip the next 1 module in the stack."
When we appended pam_google_authenticator.so to the end of /etc/pam.d/sshd, the actual evaluation chain flattened into this sequence:
pam_unix.soevaluates the user password.- Password matches: PAM returns success and executes the jump: skip 1 module.
- The skipped module is
pam_deny.so. - The engine executes
pam_permit.so. - Because
pam_permit.soreturns success and satisfies the stack criteria defined up to that point, the control flow exits early or fails to bind downstreamrequiredblocks properly depending on the PAM implementation release.
Even worse: if you put the module directly after common-auth, the jump target lands directly past it if the stack only has one terminal guard. We tested this explicitly using pamtester from the local console without restarting sshd:
$ pamtester -v sshd testuser authenticate
pamtester: invoking pam_start(sshd, testuser, ...) pamtester: performing authentication Password: pamtester: successfully authenticated
No TOTP prompt. Exit code 0. A clean, silent bypass caused by bad stack positioning.
Control Flag Traps: sufficient vs requisite vs [jump]
The root of most SSH PAM stacking vulnerabilities is a fundamental misunderstanding of how control flags terminate evaluation. Most admins learn the legacy flags (required, requisite, sufficient, optional) and assume they behave like standard boolean logic. They do not.
| Control Flag | On Success | On Failure |
|---|---|---|
required |
Continues stack; records success. | Continues stack; remembers failure to return at the end. |
requisite |
Continues stack. | Terminates stack immediately; returns failure. |
sufficient |
Terminates stack immediately if no prior required failed. |
Ignored; continues stack. |
[action=N] |
Skips the next N modules on matching action. | Evaluates default action rule. |
The danger of sufficient is immediate short-circuiting. If an upstream module like pam_unix.so or a custom SSO handler (e.g., pam_sss.so) is marked sufficient, and it succeeds, PAM halts execution of that management group right there. Any MFA module listed below it is never invoked.
The danger of required is information leakage and user confusion. If an early module fails, required forces the user through the MFA prompt anyway before failing them at the very end. That wastes time and leaks the fact that a second factor is active on the account.
Then there is the nullok parameter. Look at line 7 of our broken config:
auth required pam_google_authenticator.so nullok
People add nullok during migrations so users without a ~/.google_authenticator file do not get locked out. What this actually creates is an unmonitored backdoor. Any compromised account that hasn't manually run the enrollment tool can be accessed with just a stolen private key or password, bypassing MFA entirely. If MFA is mandatory, nullok has no business being in production configs.
The OpenSSH Layer: AuthenticationMethods and the v8.7 Deprecation
PAM does not live in a vacuum. It interacts directly with OpenSSH authentication logic defined in /etc/ssh/sshd_config. You can write a completely sound PAM configuration and still introduce a bypass by botching the sshd directives.
First, check which syntax your running binary actually parses:
# sshd -T | grep -E '(authenticationmethods|kbdinteractiveauthentication|challengeresponseauthentication|usepam)'
challengeresponseauthentication no kbdinteractiveauthentication yes usepam yes authenticationmethods publickey,keyboard-interactive
Notice KbdInteractiveAuthentication. Starting in OpenSSH 8.7, ChallengeResponseAuthentication was deprecated and mapped internally to KbdInteractiveAuthentication. If you are migrating older configuration templates to newer distros (Debian 12, RHEL 9, Ubuntu 22.04+), retaining only the legacy flag can result in sshd silently failing to advertise the keyboard-interactive method to clients.
Second, check the separator in AuthenticationMethods. This trips people up constantly:
AuthenticationMethods publickey,keyboard-interactive(Comma = AND logic). The client must satisfy publickey authentication and then complete keyboard-interactive (PAM).AuthenticationMethods publickey keyboard-interactive(Space = OR logic). The client can satisfy publickey or keyboard-interactive.
A single missing comma turns your forced MFA policy into an optional bypass route where a compromised SSH key avoids PAM altogether.
Here is what debugging that failure looks like over the wire when testing a target:
$ ssh -vvv -o PreferredAuthentications=keyboard-interactive,publickey -o PubkeyAuthentication=no user@target-host
debug1: Authentications that can continue: publickey,keyboard-interactive debug1: Next authentication method: keyboard-interactive debug2: userauth_kbdint debug1: Authentications that can continue: publickey,keyboard-interactive debug2: we sent a keyboard-interactive packet, wait for reply debug1: Server accepts key: /home/user/.ssh/id_ed25519 ED25519 SHA256:abcd... Authenticated to target-host ([192.0.2.10]:22) using "publickey".
If you see OpenSSH accept the key and grant access without ever asking for kbdint prompts, your AuthenticationMethods configuration is broken or allowing OR-based fallbacks.
Building a Predictable, Hardened Stack
To eliminate jump logic traps and prevent short-circuiting, place the MFA requirement ahead of distribution-managed include files, or write a dedicated PAM file that does not rely on dynamic jumps. The MFA module must be evaluated as a requisite condition before password processing or session creation begins.
Here is the functional setup we ended up standardizing on:
# /etc/ssh/sshd_config
UsePAM yes PasswordAuthentication no PubkeyAuthentication yes KbdInteractiveAuthentication yes AuthenticationMethods publickey,keyboard-interactive
And the corresponding /etc/pam.d/sshd:
# /etc/pam.d/sshdEnforce MFA prior to common-auth jump logic
auth requisite pam_google_authenticator.so
Standard system auth processing
@include common-auth account required pam_nologin.so @include common-account @include common-session @include common-password
By making pam_google_authenticator.so a requisite check at the very top of the stack, several things happen:
- If the user fails or lacks a token, PAM fails immediately. It does not waste cycles checking UNIX passwords or querying LDAP.
- The internal jump logic inside
common-auth(e.g.,[success=1 default=ignore]) can only skip modules inside its own sub-stack. It cannot jump past our MFA check because the MFA check has already passed. - Because
PasswordAuthentication nois set insshd_config, the password step handled viakeyboard-interactivecan either be retained as a third factor (Key + MFA + Password) or dropped ifcommon-authis customized to only evaluate the token.
Verify this using pamtester again to ensure the prompt order makes sense:
$ pamtester -v sshd testuser authenticate
pamtester: invoking pam_start(sshd, testuser, ...) pamtester: performing authentication Enter Google Authenticator code: Password: pamtester: successfully authenticated
If you test an invalid TOTP token, the requisite flag terminates evaluation without ever prompting for the UNIX password:
$ pamtester -v sshd testuser authenticate
pamtester: invoking pam_start(sshd, testuser, ...) pamtester: performing authentication Enter Google Authenticator code: 123456 pamtester: returned 7: Authentication failure
This prevents brute-force password guessing against the secondary factor endpoint.
Remaining Blind Spots
While this ordering fixes the common-auth bypass, it introduces a separate operational question: what happens to automation and non-interactive sessions? Service accounts running via SSH key will fail if AuthenticationMethods globally demands keyboard-interactive. Solving that usually requires carving out Match Group or Match User blocks in sshd_config to exempt specific service principals, which introduces its own risk of scope creep.
I still need to run some traces on whether pam_exec.so can be cleanly injected into this stack to trigger an out-of-band alert via auditd on repeated PAM_AUTH_ERR (return code 7) results without adding measurable latency to our interactive logins.
