Commit Graph
17 Commits
Author SHA1 Message Date
root 50cb782ded fix(portal): per-challenge SSH user in web terminal + credential API
The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".

Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
  defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
  6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
  the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table

Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
  holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
  receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
  challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them

Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.

Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
2026-09-26 16:37:40 +08:00
Cyrene aa0bb45633 fix(sla): 100% fleet SLA (64/64) - per-challenge SSH login, phew checker, sidecar detection
Three independent root causes, all found by measuring instead of assuming:

1. SSH failed on 10/16 challenges while state.json looked perfect.
   Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
   XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
   set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
   used -> 'Permission denied' on every team. Registry gains a per-challenge
   'ssh_user'; chpasswd targets the real login and reports failures loudly.

2. phew SLA timed out on a healthy service, four bugs stacked:
   - chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
     set) and does a fresh Pailier keygen (~12 s) before printing its menu;
   - _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
     TextIOWrapper buffer and select() then blocked on data already in memory;
   - its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
     the '> ' the next call waited for -> a race that failed intermittently;
   - reaping killed chall.py it did not own: a blanket pkill -f, a
     snapshot-diff (concurrent sessions diff against the same pre-spawn set),
     and a class-level _children shared across uvicorn's thread pool. The child
     now prints its own pid so exactly one session is reaped.
   Also: ONE interactive session per check instead of five spawns (Paillier is
   randomized per ciphertext, not per process) - 5 keygens were the CPU load
   that starved the checks. And the 6 orphan single-node containers from the
   original deploy were removed; one held 58 leaked chall.py and drove load
   average 76 on 2 CPUs.

3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
   every service sets an explicit container_name, so it reported all 16 running
   challenges as missing and hid the one real gap (anti-alchemy-db, which has
   no container_name). Now reads container_name when present and falls back to
   the compose default otherwise.

Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.

Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
2026-09-26 15:37:31 +08:00
Cyrene ae50acfe40 fix(ssh): per-challenge SSH login + phew buffering/leak/timeout
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
  XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
  set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
  account nobody uses -> 'Permission denied' everywhere.
  Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
  login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
  (PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
  timeout (26 orphans, container saturated) -> reaps the whole exec process
  group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
  _CRYPTO_TIMEOUT, not the 5 s prompt default.

Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
2026-09-26 14:40:10 +08:00
MythEclipse ef385c3397 fix: get all imported challenges building and running
Root causes found by prebuilding every challenge image in parallel:
- fjb: ghcr.io base is not anonymously pullable here -> official httpd:2.4.
  pnpm 12 (via corepack on node:20) fails the install with
  ERR_PNPM_IGNORED_BUILDS unless build scripts are approved; neither
  onlyBuiltDependencies in pnpm-workspace.yaml nor --no-ignore-scripts
  suppresses it. The working sequence is:
    pnpm install --ignore-scripts && pnpm approve-builds --all && pnpm rebuild
- xl + kode-viewer: node:20-slim-bookworm is not a real tag -> node:20-bookworm-slim.
- burvesigner: python-dev no longer exists in bookworm -> dropped (python3-dev
  was already there and the source has no py2 syntax).
- burvesigner/hirnfick/s3: apt update and install were separate RUN layers;
  with the bundled apt-insecure.conf the second invocation re-resolved against
  the EOL bullseye-security mirror and 404'd every package. Merged into one
  'update && install' layer (fix_apt_layers.py, idempotent).
- consolidate_images.sh: teams used to build a private image per team
  (team1-x ... team4-x) because no shared image existed. Since the password is
  applied at runtime via chpasswd, one shared services-<name> build is enough;
  this reclaims ~1.5 GB, which matters on a 79 GB disk.
- reconcile_team_state(): a challenge enabled while a team was down left
  state.json without ports/flag/password, so the next compose render died with
  KeyError. Now both the API and the CLI tools reconcile first.
2026-09-25 21:03:29 +08:00
MythEclipse 6d1ede8c2b fix: bulk challenge enable path (compose YAML, base images, async toggle, tooling)
- fix_dup_volumes.py: 4 canonical templates had TWO volumes: keys inside one
  service (invalid YAML -> 'mapping key volumes already defined'), which broke
  every enable for anti-alchemy/burvesigner/gemas-notes/kode-viewer.
- fjb: ghcr.io base is not anonymously pullable on this host; swapped to the
  official httpd:2.4 (its httpd.conf only uses stock modules). Added
  onlyBuiltDependencies to package.json (pnpm >=10 blocks esbuild's postinstall).
- xl + kode-viewer: node:20-slim-bookworm is not a real tag; use
  node:20-bookworm-slim. gift-voucher: buster -> bookworm.
- prebuild_images.py: build each challenge's shared services-<name> image once
  in parallel (passes a placeholder PASSWORD build-arg, since several Dockerfiles
  run chpasswd and fail on an empty arg).
- set_enabled.py / sync_all_challenges.py: batch registry flip + runtime apply
  that survives panel restarts and reports per-team results.
- Challenge toggle is now async: PATCH returns a job id, the client polls
  /api/challenges/jobs/<id> so a multi-minute build no longer blocks the panel.
  Added _SYNC_LOCK to serialize concurrent compose rewrites.
2026-09-25 17:01:42 +08:00
MythEclipse d4c741926d fix: make challenge toggle actually work end-to-end (5 root-cause bugs)
Found by testing a real enable/disable cycle (art, fjb, gift-card):

1. compose_gen always swapped build->image, so a never-built challenge
   produced 'pull access denied for services-<name>'. Now it only reuses
   the image when it exists locally, otherwise keeps build: so
   'docker compose up --build' builds it.
2. Canonical templates use 'build: context: .' (written for the shared
   services/ tree). In the per-team compose that resolves to the team dir
   which has no Dockerfile -> 'failed to read dockerfile'. The renderer
   now rewrites the main service's context to ./<name>.
3. Teams created before the XVI/XVII import had no xvi/xvii subpackages
   under their local challenges/ dir, so the regenerated receiver main.py
   crash-looped on import. gen_receiver_main now mirrors ALL shared
   checkers (native + xvi + xvii) into every team receiver on each sync.
4. systemd Environment= keys can't contain hyphens, so
   CHALLENGE_PORT_GIFT-CARD was silently dropped. Keys are now
   normalized to underscores on both the writer and reader side.
5. Several checkers called 'docker exec' with no timeout; against a
   container with accumulated chall.py zombies that blocks forever and
   stalls the whole SLA loop. Added mandatory timeouts (Phew, Sheesh,
   Carbeat, Poke, Warmup).

Also: enabling a challenge now copies its source tree into each team's
services/ dir (team dirs only held challenges enabled at create_team
time), and the XVII checkers were rewritten to be protocol-aware
(gift-card/gift-voucher are socat TCP, not HTTP) with strict timeouts.
2026-09-25 15:36:09 +08:00
MythEclipse c6fd9ec268 feat: challenge registry-driven platform + XVI/XVII imports + admin toggle + domain rename
- Rename repo/domain: attack-defense-platform / attackdefense.imrnes.team (all refs replaced)
- challenge_registry.json: single source of truth (28 challs across gemastik18/xvi/xvii)
- teams.py: registry-driven CHALLENGES, set_challenge_enabled, sync_challenge_runtime
  (apply enable/disable to live teams: build/up or stop/remove + receiver restart)
- compose_gen.py: render per-team compose from canonical per-challenge templates
  (image reuse, per-team ports 30xxx, flag mounts, passwords)
- gen_canonical_composes.py: canonical docker-compose.yml for all services
- import_new_challenges.py: import XVI/XVII services + EOL base image fixes
  (debian:buster→bookworm, node:14→20, python:3.7-slim→3.11)
- receiver: xvi package (10 checkers) + xvii package (12 generic checkers),
  Challenge base reads PASSWORD_<team_port> from env; gen_receiver_main.py
  generates per-team main.py from registry
- main.py: /api/challenges returns full registry; PATCH /api/challenges/<name>
  toggles enabled + applies to live teams
- index.html: 🏗️ Challenge Manager tab (toggle per challenge, grouped by set)
- SLA bonus now dynamic (all enabled challenges, not hardcoded 6)
2026-09-25 14:04:33 +08:00
MythEclipse c35b23a37f feat: SLA dashboard per team + points system (100/flag, +50 SLA bonus) + badges juara/runner-up/3rd + scoreboard API public+admin
Panel: /api/scoreboard + /api/public/scoreboard; teams.py: points.json ledger, probe_team_sla_fast, sla_status_all w/ background refresher; team.html: SLA & Skor tab; index.html: admin SLA tab; leaderboard shows points; Phew checker timeouts raised for slow Paillier keygen
2026-09-24 00:49:05 +08:00
root 877f14ecf3 fix: editTeam endpoint + leaderboard live name resolution + domains in team set + apt-insecure.conf in all service dirs 2026-09-23 21:50:16 +08:00
root 24dbf0a662 draggable topology + attack visualizer + tools in containers
- topology nodes draggable (pointer events, SVG transform), layout hint shown
- attack visualizer: /api/attacks logs attacker->target events; red pulsing
  dashed arcs on recent attacks (60s hot), ⚔ counts ok/fail
- submit_flag now takes attacker_idx vs target_idx (A/D semantics); UI has
  target dropdown (enemy teams), leaderboard records target
- containers get vim+curl+wget+netcat+git+pip3 (Dockerfiles blogpost/cdn/
  phew/sheesh/warmup); warmup base ubuntu:20.04 EOL -> 24.04
- team portal: target dropdown refreshed after login (was empty pre-auth)
2026-09-23 18:05:44 +08:00
root 44dc1ac852 feat: target matrix + reset scores/environment buttons
- /api/targets (admin): matrix of all teams' domain:port targets
- /api/reset/scores: wipe leaderboard (admin, confirm dialog)
- /api/reset/environment: stop all teams, remove containers+receivers+
  team dirs+systemd units, wipe scores, drop domains (admin, confirm)
- portal targets now only domain+port (no ssh/labels)
- UI: tab Target Matrix, header buttons Reset Skor / Reset Environment
  with confirm() alerts 'apakah anda yakin ingin mereset...'
2026-09-23 17:14:42 +08:00
root b66530e7fd fix: per-team receivers as independent systemd services
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.
2026-09-23 17:06:24 +08:00
root 43ed3df557 feat: team portal with own login, SSH web terminal, targets & guide
- Portal tim punya login sendiri (password = ssh_pass, session team_token)
- SSH Web GUI: /api/team/{idx}/ssh/ws (WebSocket+paramiko) → xterm.js terminal
  (fix: ssh_to_ws non-blocking poll, chall_passwords per-container auth)
- Root <slug>.gemastik.imrnes.team → portal tim (bukan login admin)
- Host validation: /team/{idx} & /team/{idx}/guide 403 kalau host != team domain
- /api/team/{idx}/targets: daftar service tim musuh (attack target)
- guide.html: panduan SSH/attack/defense untuk peserta
- fix esc() String(s) (bug: (s||'').replace is not a function saat port number)
- set_ssh_passwords retry loop (container boot race)
2026-09-23 16:37:58 +08:00
root 5e44a70049 feat: team-specific subdomains + portal tim
- create_team now takes label -> slug -> <slug>.gemastik.imrnes.team
- ensure_team_domains() writes Traefik dynamic config (gemastik-teams.yaml)
- team portal at /team/<idx> (public, shows chall ports, SSH, submit form)
- /api/team/<idx>/info + /api/team/<idx>/status (server-side receiver auth)
- UI: name inputs per team (set count -> labels), domain link in card
- delete team endpoint drops its domain
2026-09-23 16:18:11 +08:00
root 265159e9d8 feat: set per-team SSH passwords at runtime (chpasswd in shared images) 2026-09-23 15:47:22 +08:00
root 8c061ba1b7 fix: port scheme 30000 (avoid syncthing 22000), reuse base images (no per-team rebuild), recover corrupted receiver/main.py, compose -p project isolation
- PORT_BASE 20000->30000: team1=31xxx team2=32xxx; syncthing owns 22000
- create_team replaces build: with image: services-<name> so teams reuse base images (was rebuilding 6 images per team, disk 100%)
- recovery: receiver/main.py was corrupted by bad patch (write_file with read_file format); restored from team1 copy + original GitHub
- docker compose -p teamN: project isolation so team compose doesn't overlap (was showing team1 containers for team2)
2026-09-23 15:44:53 +08:00
root 1ca963b2b7 feat: multi-team orchestrator - topology UI, team management, flag randomizer, public submit + leaderboard
- panel/teams.py: create/start/stop teams with isolated ports+creds, per-team receivers with CHALLENGE_PORT/CONTAINER env, flag randomization, submit validation + leaderboard
- panel/main.py: /api/teams, /api/teams/{idx}/randomize, /api/flag/submit, /api/leaderboard, /api/public/teams, /submit page
- panel/static/submit.html: public flag submission UI for teams
- receiver: _ch_port/_ch_container read .env per team, container name overrides
- .gitignore: exclude teams/, .venv, __pycache__
2026-09-23 15:14:52 +08:00