Created a scratch team 5, booted all 16 of its containers, snapshotted its
footprint, deleted it via the API, and re-snapshotted:
before: 16 containers, 1 sidecar, team5_default network, teams/team5 dir
after : 0 containers, 0 sidecars, network gone, dir GONE, 34 UFW ports
closed, Traefik domain removed, 0 score rows
took 102s; the other 4 teams (64 containers, receivers active) were
untouched and still at 64/64 SLA.
Adds panel/team_footprint.sh (before/after proof of a delete),
team_health.sh, watch_load.sh.
Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
Root causes found by prebuilding every challenge image in parallel:
- fjb: ghcr.io base is not anonymously pullable here -> official httpd:2.4.
pnpm 12 (via corepack on node:20) fails the install with
ERR_PNPM_IGNORED_BUILDS unless build scripts are approved; neither
onlyBuiltDependencies in pnpm-workspace.yaml nor --no-ignore-scripts
suppresses it. The working sequence is:
pnpm install --ignore-scripts && pnpm approve-builds --all && pnpm rebuild
- xl + kode-viewer: node:20-slim-bookworm is not a real tag -> node:20-bookworm-slim.
- burvesigner: python-dev no longer exists in bookworm -> dropped (python3-dev
was already there and the source has no py2 syntax).
- burvesigner/hirnfick/s3: apt update and install were separate RUN layers;
with the bundled apt-insecure.conf the second invocation re-resolved against
the EOL bullseye-security mirror and 404'd every package. Merged into one
'update && install' layer (fix_apt_layers.py, idempotent).
- consolidate_images.sh: teams used to build a private image per team
(team1-x ... team4-x) because no shared image existed. Since the password is
applied at runtime via chpasswd, one shared services-<name> build is enough;
this reclaims ~1.5 GB, which matters on a 79 GB disk.
- reconcile_team_state(): a challenge enabled while a team was down left
state.json without ports/flag/password, so the next compose render died with
KeyError. Now both the API and the CLI tools reconcile first.
- fix_dup_volumes.py: 4 canonical templates had TWO volumes: keys inside one
service (invalid YAML -> 'mapping key volumes already defined'), which broke
every enable for anti-alchemy/burvesigner/gemas-notes/kode-viewer.
- fjb: ghcr.io base is not anonymously pullable on this host; swapped to the
official httpd:2.4 (its httpd.conf only uses stock modules). Added
onlyBuiltDependencies to package.json (pnpm >=10 blocks esbuild's postinstall).
- xl + kode-viewer: node:20-slim-bookworm is not a real tag; use
node:20-bookworm-slim. gift-voucher: buster -> bookworm.
- prebuild_images.py: build each challenge's shared services-<name> image once
in parallel (passes a placeholder PASSWORD build-arg, since several Dockerfiles
run chpasswd and fail on an empty arg).
- set_enabled.py / sync_all_challenges.py: batch registry flip + runtime apply
that survives panel restarts and reports per-team results.
- Challenge toggle is now async: PATCH returns a job id, the client polls
/api/challenges/jobs/<id> so a multi-minute build no longer blocks the panel.
Added _SYNC_LOCK to serialize concurrent compose rewrites.
Found by testing a real enable/disable cycle (art, fjb, gift-card):
1. compose_gen always swapped build->image, so a never-built challenge
produced 'pull access denied for services-<name>'. Now it only reuses
the image when it exists locally, otherwise keeps build: so
'docker compose up --build' builds it.
2. Canonical templates use 'build: context: .' (written for the shared
services/ tree). In the per-team compose that resolves to the team dir
which has no Dockerfile -> 'failed to read dockerfile'. The renderer
now rewrites the main service's context to ./<name>.
3. Teams created before the XVI/XVII import had no xvi/xvii subpackages
under their local challenges/ dir, so the regenerated receiver main.py
crash-looped on import. gen_receiver_main now mirrors ALL shared
checkers (native + xvi + xvii) into every team receiver on each sync.
4. systemd Environment= keys can't contain hyphens, so
CHALLENGE_PORT_GIFT-CARD was silently dropped. Keys are now
normalized to underscores on both the writer and reader side.
5. Several checkers called 'docker exec' with no timeout; against a
container with accumulated chall.py zombies that blocks forever and
stalls the whole SLA loop. Added mandatory timeouts (Phew, Sheesh,
Carbeat, Poke, Warmup).
Also: enabling a challenge now copies its source tree into each team's
services/ dir (team dirs only held challenges enabled at create_team
time), and the XVII checkers were rewritten to be protocol-aware
(gift-card/gift-voucher are socat TCP, not HTTP) with strict timeouts.
- all 6 Dockerfiles: vim curl wget netcat git python3-pip now installed
- apt-insecure.conf (AllowInsecureRepositories) copied into images so
participants can apt-get install despite expired Ubuntu/Debian GPG keys
- warmup base ubuntu:20.04 (EOL, GPG expired) -> ubuntu:24.04
- installed vim+git live into all 18 running team containers
- team portal target dropdown reloads after login (was empty pre-auth)
- attack log endpoint + A/D submit (attacker vs target) verified e2e
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.
- guide link now server-side replaced to /team/<idx>/guide (no /team/0 403)
- _check_team_host() applied to ALL team endpoints (login, info, targets,
status, guide, portal, ssh-ws): host must match team domain; panel/gemastik
host only with admin session. Cross-domain session reuse -> 403.
- host check BEFORE auth on info/targets (no team-existence oracle)
- loading overlay (spinner + text) on start/stop all-team/set; JS util
showLoading/hideLoading
- PORT_BASE 20000->30000: team1=31xxx team2=32xxx; syncthing owns 22000
- create_team replaces build: with image: services-<name> so teams reuse base images (was rebuilding 6 images per team, disk 100%)
- recovery: receiver/main.py was corrupted by bad patch (write_file with read_file format); restored from team1 copy + original GitHub
- docker compose -p teamN: project isolation so team compose doesn't overlap (was showing team1 containers for team2)
- FastAPI app at panel/ proxying receiver API server-side (admin creds stay server-side)
- Login-protected dashboard: SLA status, rotate flag, restart/rollback/activate/deactivate, SSH creds, command history
- Runs as systemd service gemastik-panel.service on :18081
- Published at https://panel.gemastik.imrnes.team via Traefik
- blogpost: python:3.11-slim-bullseye is EOL (apt 404s), switch to bookworm
- utils/bashrc: host port 80 is taken by Traefik/Coolify, preexec posts to :18080
- ignore receiver .venv