The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".
Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table
Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them
Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.
Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
Found by testing a real enable/disable cycle (art, fjb, gift-card):
1. compose_gen always swapped build->image, so a never-built challenge
produced 'pull access denied for services-<name>'. Now it only reuses
the image when it exists locally, otherwise keeps build: so
'docker compose up --build' builds it.
2. Canonical templates use 'build: context: .' (written for the shared
services/ tree). In the per-team compose that resolves to the team dir
which has no Dockerfile -> 'failed to read dockerfile'. The renderer
now rewrites the main service's context to ./<name>.
3. Teams created before the XVI/XVII import had no xvi/xvii subpackages
under their local challenges/ dir, so the regenerated receiver main.py
crash-looped on import. gen_receiver_main now mirrors ALL shared
checkers (native + xvi + xvii) into every team receiver on each sync.
4. systemd Environment= keys can't contain hyphens, so
CHALLENGE_PORT_GIFT-CARD was silently dropped. Keys are now
normalized to underscores on both the writer and reader side.
5. Several checkers called 'docker exec' with no timeout; against a
container with accumulated chall.py zombies that blocks forever and
stalls the whole SLA loop. Added mandatory timeouts (Phew, Sheesh,
Carbeat, Poke, Warmup).
Also: enabling a challenge now copies its source tree into each team's
services/ dir (team dirs only held challenges enabled at create_team
time), and the XVII checkers were rewritten to be protocol-aware
(gift-card/gift-voucher are socat TCP, not HTTP) with strict timeouts.
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.