Lab Notes: Eliminating Static CI/CD SSH Keys Using OpenSSH Match Exec and OIDC Tokens
We had twenty-eight static id_ed25519 private keys sitting in GitLab CI variables and GitHub repository secrets, all granting unrestricted root or deployer shell access to staging and production target hosts. Rotating them meant updating secrets across forty different pipelines and hoping nobody missed an application repository—a problem that prompted our effort to replace shared SSH keys with automated, identity-driven workflows. If a worker node running an untrusted build got popped, those keys were gone before anyone checked audit logs.
Instead of deploying a full HashiCorp Vault or Teleport cluster just to push container updates and run systemctl restarts on raw compute, we wired OpenSSH directly to ambient CI/CD OIDC identity tokens. The core mechanism handling this logic on the wire is openssh match exec authentication and conditional block evaluation.
Things broke immediately during our initial implementation. These notes document the mechanics, the failed attempt, and the working setup.
The Execution Mechanics: How Match Exec Actually Operates
OpenSSH added Match exec to allow runtime shell evaluation before applying a configuration block. While most engineers know Match User or Match Host, Match exec executes an arbitrary command via /bin/sh -c. If the command exits with return code 0, the block matches and its directives apply. If it returns any non-zero code, OpenSSH skips the block entirely.
sshd_config vs. ssh_config
The execution context differs completely depending on which side of the socket you are configuring:
- In
sshd_config(Server): The test command executes on the server host. Crucially, it runs as root during the pre-authentication phase, before privilege separation drops to the authenticated user. This gives the command raw access to system state, but also means any bug or injection vulnerability runs with full system privileges. - In
ssh_config(Client): The command executes locally as the invoking user before the TCP connection opens or before authentication payloads are dispatched. This makes it ideal for fetching short-lived credentials or inspecting the local runtime environment.
Token Expansions and Evaluation Order
OpenSSH interpolates parameters into the command string before handing it to /bin/sh. The most useful expansion tokens include:
%u: The username supplied by the client.%h: Target host (in client config) or user's home directory (in server config—a frequent source of confusion).%C: Connection addresses and ports formatted aslocal_ip,local_port,remote_ip,remote_port(server-side).%k: Host key or client public key depending on the block context.%Lor%l: Local hostname.
OpenSSH evaluates configuration directives in first-match-wins order for general settings, but Match blocks are parsed sequentially and accumulate overrides. If multiple Match exec blocks return 0, every setting inside those blocks merges into the session configuration, with later blocks overriding earlier ones if they define the same parameter.
The Failed Approach: Passing OIDC Tokens via SSH Environment Variables
Our initial design attempted to do the verification entirely inside a server-side Match exec block by reading the CI runner's OIDC JWT directly off an environment variable.
We configured the client to pass the token:
# Client pipeline step
export CI_JOB_JWT=$(cat /var/run/secrets/oidc/token) ssh -o SendEnv=CI_JOB_JWT [email protected] "systemctl restart api"
On the server, we permitted the variable and wrote an evaluation check:
# /etc/ssh/sshd_config (Broken attempt)AcceptEnv CI_JOB_JWT
Match User deployer Exec "/usr/local/bin/verify-jwt.sh" AuthorizedKeysFile /dev/null AuthenticationMethods publickey ForceCommand /usr/local/bin/deploy-restricted.sh
The verification script looked like this:
#!/usr/bin/env bash/usr/local/bin/verify-jwt.sh
set -euo pipefail
Check if the environment variable exists and validate claims
if [[ -z "${CI_JOB_JWT:-}" ]]; then exit 1 fi
Call public keyset endpoint and parse claims...
/usr/bin/jwt-validator --token "$CI_JOB_JWT" --issuer "https://gitlab.example.com"
It failed every time with exit code 255:
$ ssh -vvv -o SendEnv=CI_JOB_JWT [email protected]
debug1: Authentications that can continue: publickey debug1: Next authentication method: publickey debug1: Trying private key: /root/.ssh/id_ed25519 debug1: Authentications that can continue: publickey debug2: we did not send a packet, disable method debug1: No more authentication methods to try. [email protected]: Permission denied (publickey).
Why it failed
This failure comes down to the internals of the SSH transport layer. Match exec executes during the initial handshake and pre-auth phase to construct the active authorization rules. Environment variables requested via SendEnv are not processed until the session channel request phase (SSH_MSG_CHANNEL_REQUEST with string env)—which happens after authentication succeeds.
When sshd executed /usr/local/bin/verify-jwt.sh, CI_JOB_JWT was completely empty in the process environment. The script exited with 1, the Match block was skipped, and the default server configuration rejected the login.
The Working Setup: Client-Side OIDC Exchange and Server-Side Validation
To eliminate static keys cleanly, the client must exchange the ambient OIDC token for an ephemeral, signed credential before the wire handshake completes, while the server uses Match exec to dynamically gate source criteria and enforce strict boundaries.
1. Client-Side ssh_config with Match Exec
Instead of hardcoding a long-lived key in CI secrets, we let the runner generate an ephemeral keypair in memory. The client uses Match exec in ~/.ssh/config to inspect the environment, call our internal certificate authority or token signer using the workload identity, and write a 5-minute ephemeral certificate to disk before the connection initiates.
# ~/.ssh/config on the CI Runner
Match host staging-*.internal exec "test -n \"$CI_JOB_JWT\" && /usr/local/bin/fetch-ephemeral-cert.sh" IdentityFile ~/.ssh/ephemeral_ed25519 CertificateFile ~/.ssh/ephemeral_ed25519-cert.pub IdentitiesOnly yes StrictHostKeyChecking accept-new
The helper script runs in under 300 milliseconds:
#!/usr/bin/env bash/usr/local/bin/fetch-ephemeral-cert.sh
set -euo pipefail
KEY_PATH="$HOME/.ssh/ephemeral_ed25519"
if [[ ! -f "$KEY_PATH" ]]; then ssh-keygen -t ed25519 -N "" -f "$KEY_PATH" -q fi
Exchange OIDC token for an ephemeral signed certificate
curl -s -f -X POST https://ca.internal.net/sign \ -H "Authorization: Bearer ${CI_JOB_JWT}" \ -F "pubkey=@${KEY_PATH}.pub" \ -o "${KEY_PATH}-cert.pub"
exit 0
If the machine is not running inside a CI job (i.e., $CI_JOB_JWT is empty), Match exec returns a non-zero exit code, and SSH falls back to standard developer profiles.
2. Server-Side sshd_config Dynamic Authorization
On the destination hosts, we configure sshd_config to use Match exec for real-time validation of deployment state. For instance, we only permit deployments if the runner belongs to an active, registered pipeline IP and the staging deploy window is open in our control plane.
# /etc/ssh/sshd_configTrustedUserCAKeys /etc/ssh/trusted_oidc_ca.pub
Base configuration for deployer user
Match User deployer AuthenticationMethods publickey AuthorizedKeysFile /dev/null X11Forwarding no AllowTcpForwarding no AllowAgentForwarding no
Dynamic policy check via Match Exec
Match User deployer Exec "/usr/local/bin/validate-pipeline-context.sh %u %C" ForceCommand /usr/local/bin/apply-deployment.sh MaxSessions 2
The server-side validation script parses the source IP from %C and verifies that the deployment registry actually expects an incoming CI sync:
#!/usr/bin/env bash/usr/local/bin/validate-pipeline-context.sh
set -euo pipefail
TARGET_USER="$1" CONN_INFO="$2"
Parse client IP from connection token: local_ip,local_port,remote_ip,remote_port
CLIENT_IP=$(echo "$CONN_INFO" | cut -d',' -f3)
Query internal state API over localhost
STATUS=$(curl -s -o /dev/null -w "%{http_code}" \ --connect-timeout 1 \ "http://127.0.0.1:9099/leases/check?ip=${CLIENT_IP}&user=${TARGET_USER}")
if [[ "$STATUS" -eq 200 ]]; then exit 0 else # Non-zero exit code tells sshd this Match block does NOT apply exit 1 fi
If the control plane has an active deploy lease for that IP, the Match block applies, locking the session into ForceCommand /usr/local/bin/apply-deployment.sh. If the lease expired or never existed, the block evaluates to false. Because the base Match User deployer block lacks a default command and requires tight restrictions, unauthorized attempts terminate cleanly.
Security and Latency Traps with Match Exec
While this architecture cleanly eliminates static keys, running arbitrary shell processes inside the SSH handshake introduces concrete operational risks.
Command Injection in % Expansions
Because Match exec evaluates commands using /bin/sh -c, arguments expanded into the command string are vulnerable to injection vulnerabilities if not handled defensively. Consider this dangerous directive:
# DANGEROUS: Remote user can pass malicious usernames
Match Exec "/usr/local/bin/check-user.sh %u"
A client running ssh 'deployer;reboot;'@target-host causes %u to expand literally into the shell string. Even if the user cannot log in, the subshell executed by root evaluates reboot. Always quote the tokens explicitly:
# Safer: Tokens wrapped in double quotes
Match Exec "/usr/local/bin/check-user.sh \"%u\" \"%C\""
Inside the script itself, validate that $1 contains only expected alphanumeric characters before doing anything else.
Handshake Timeouts and Blocking
OpenSSH handles Match exec synchronously. The SSH daemon blocks the connection thread while the external process runs. If your script calls an external network endpoint that hangs or runs a slow DNS check, the entire SSH connection pauses.
We hit intermittent kex_exchange_identification: Connection closed by remote host errors when an internal metadata service degraded. If the script takes longer than LoginGraceTime (defaults to 120 seconds, but often tuned down to 15 or 30), sshd kills the pre-auth child process.
| Parameter | Recommended Value | Reason |
|---|---|---|
| External script timeout | < 2 seconds |
Prevent connection stalls; use timeout 2 cmd or native curl flags. |
LoginGraceTime |
15s |
Stops stale pre-auth processes from piling up and exhausting worker limits. |
| Script permissions | 0750 root:root |
Prevents unprivileged processes from altering the root-executed check. |
Troubleshooting Match Exec Rules
When an execution script fails silently, OpenSSH gives you almost nothing on the client terminal by default. To understand why a block did or did not apply, inspect both ends using deep debug flags.
Testing Server Rules with sshd -d
Stop the running service or spin up a test daemon on a non-standard port to see real-time Match exec parsing:
# Run sshd in debug mode on port 2222
/usr/sbin/sshd -d -p 2222
Watch the standard error output when the client attempts to connect:
debug1: user deployer matched 'User deployer' at line 45
debug1: Running exec command: "/usr/local/bin/validate-pipeline-context.sh \"deployer\" \"10.0.1.5,2222,10.240.0.10,48124\"" debug1: Exec command /usr/local/bin/validate-pipeline-context.sh returned 0 debug1: user deployer matched 'Exec "/usr/local/bin/validate-pipeline-context.sh \"deployer\" \"10.0.1.5,2222,10.240.0.10,48124\""' at line 52 debug1: Setting ForceCommand: /usr/local/bin/apply-deployment.sh
If the script exits with anything other than 0, you will see:
debug1: Exec command /usr/local/bin/validate-pipeline-context.sh returned 1
debug1: match not found at line 52
Validating Client Evaluation with ssh -vvv
On the client side, running verbose output shows the local subshell execution:
ssh -vvv -F ~/.ssh/config staging-api-01.internal
Look specifically for the configuration processing lines:
debug2: checking match for "host staging-*.internal exec \"test -n \\\"$CI_JOB_JWT\\\" && /usr/local/bin/fetch-ephemeral-cert.sh\"" host staging-api-01.internal
debug3: running command "test -n \"$CI_JOB_JWT\" && /usr/local/bin/fetch-ephemeral-cert.sh" debug2: match found debug1: /home/runner/.ssh/config line 1: Applying options for staging-*.internal
If you see debug2: match not found, check your escaping. Shell quoting inside the SSH config string requires escaped quotes around inner variables, or the subshell will error out before your script even runs.
Where This Leaves Us
We cleared all 28 static private keys out of the pipeline secret stores. CI runners now authenticate using short-lived certificates minted on the fly by validating their platform OIDC signatures, while target servers use Match exec to confirm that incoming requests match real-time deployment windows.
The next problem we are looking at is connection multiplexing (ControlMaster). When an SSH control socket is reused across pipeline jobs, OpenSSH skips configuration evaluation entirely for subsequent sessions, which bypasses the Match exec check on both client and server. For now, we have explicitly forced ControlMaster no inside all automated deployment configurations.
