Commit Graph
18 Commits
Author SHA1 Message Date
root cf47770798 Fix the topology freezing and make team drags carry their challenges
User report: the graph "disappeared". Reproduced, and the cause was not the
graph at all — the panel's main thread was blocked hard enough that the browser
stopped responding to clicks.

Root cause, found by measuring rather than guessing:
  rAF 2 FPS, setTimeout(0) lag 1353ms, with the graph rendering correctly the
  whole time. The scene was repainting continuously at 60fps even though it is
  static between the 10s data refreshes. Each repaint cost ~305ms on this
  host's software GL, so the main thread never got a free slot.

Fixes:
- Render on demand. The ticker is registered but not started; it runs only while
  an attack pulse is in flight or the user is dragging, and stops once the scene
  is quiet. applyView() (the single funnel for pan/zoom/drag/refit) repaints
  synchronously. Result: 2 FPS -> 73 FPS, 1353ms -> 2ms lag. That is faster than
  a blank page in the same harness (29 FPS), which confirms the ticker was the
  cost, not the scene complexity.
- Stop rebuilding Pixi objects every refresh. _redrawEdges destroyed and
  recreated ~70 Graphics + ~36 Text each cycle; every new Text allocates a
  canvas, rasterises glyphs and uploads a texture. Edges are now kept per
  from->to key and only their geometry is redrawn; a label's text is re-set only
  when the string actually changes. Same for node titles/subtitles.
- Bind tooltip handlers once, at node creation. Re-binding inside the refresh
  loop added a pointerover listener every 10s, so a single hover after an hour
  fired thousands of handlers.

Two correctness bugs fixed while in there:
- init() race. The Topology tab button calls showView('topo') AND loadTopo(), and
  showView() itself calls loadTopo(), so two ran concurrently. The re-entry guard
  tested `topoGraph && topoGraph.app`, but topoGraph was assigned BEFORE the
  awaited init(), so app was still null and a second renderer was built. Both
  cleared host.innerHTML and appended their own canvas, so the last init to
  finish won the DOM while the bridge still pointed at the other — a live canvas
  that was no longer on the page. Now guarded by a single-flight promise, and
  the instance is published only after init() resolves.
- Dropped the dead topoLoaded flag left over from the SVG renderer.

Hierarchical drag: dragging a team now carries its 16 challenge nodes with it.
The parent->children index is built from the edge list, the same source the
lines are drawn from, so the drag hierarchy cannot disagree with the picture;
challenges missing from the edge list still attach by id prefix. Children
translate rigidly (verified: 0.000px deviation across all 16).

Verified from a cold panel restart: all 5 topology tests pass, 73 FPS, 2ms
lag, 37 nodes drawn, graph framed at 4 viewport widths, no page errors, and
platform SLA unregressed at 32/32.
2026-09-27 01:17:56 +08:00
root 30f8052e79 Add viewport regression test for the topology layout fix
The off-screen-layout bug (nodes centred on the scrollable width instead of the
viewport) only reproduced at the default window size, so a single-width test
would have shipped it. Verifies the graph stays framed at 1920/1400/1100/820,
and records which assumption in the pixel census is safe: the measurement uses
the canvas buffer size, not the screen size, because the CSS min-width makes
narrow viewports scroll.
2026-09-27 00:39:07 +08:00
root 1461c874c3 Render the topology graph with PixiJS v8
Replaces the hand-rolled inline-SVG topology with a PixiJS 8 scene graph.

Why: the old renderer rebuilt all 37 nodes / 36 edges as one innerHTML string
every 10s, which tore down and recreated every DOM node. That restarted CSS
animations mid-flight and made dragging fight the browser's own hit-testing.
The scene graph gives per-node transforms, so pan/zoom is a single container
transform instead of getScreenCTM() matrix math.

Changes:
- static/topo_pixi.js: new self-contained renderer. Owns its Application and
  tears it down on tab exit so a second WebGL context cannot leak.
- static/index.html: the <svg id=topoSvg> host becomes a <div id=topoHost>;
  the 188-line SVG renderer is replaced by a bridge to the module.
- static/vendor/pixi.mjs: PixiJS 8.21.0 self-hosted (MIT). The .mjs build is
  required; the .js build exports no global. See vendor/README.md.
- main.py: mount /static. Pages were served as inline HTMLResponse, so the
  directory was never mounted and the module had no URL to load from.

Two real bugs found by measuring pixels rather than trusting init():
- preserveDrawingBuffer: without it WebGL clears the back buffer after
  compositing, so any readback or screenshot of the canvas is a coin flip
  depending on which frame it lands on. The graph rendered intermittently
  blank. Now enabled: cheap for a 2D scene, and it makes the view capturable.
- Layout was centred on the SCROLLABLE width (nodes.length * 130), not the
  viewport, so with 37 nodes every team and challenge node landed at
  x=2230-2650 on a 1310px canvas: entirely off-screen. Layout now centres on
  the visible width and reset() frames the whole graph to fit.

test_topo_pixels.js documents three wrong test designs it replaces, all of
which reported false failures against a working graph: counting scene-graph
children (passes on a blank canvas), diffing against the background colour
(the theme is dark by design, so a perfect render measures ~0%), and diffing
two Playwright screenshots (both can be captured after the scene was mutated).
The check now reads the GL back buffer via readPixels in one evaluate.

Verified: 37 nodes / 36 edges drawn (7.94% of frame, max channel delta 225),
graph bbox [437,46,881,476] inside the 1310x520 canvas, glGetError=0, no page
errors, zoom and frame-to-fit reset working. Platform unregressed: SLA 32/32.
2026-09-27 00:31:47 +08:00
root 50cb782ded fix(portal): per-challenge SSH user in web terminal + credential API
The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".

Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
  defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
  6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
  the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table

Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
  holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
  receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
  challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them

Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.

Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
2026-09-26 16:37:40 +08:00
root 3ef905b3bb feat: team activity feed (attacks in/out) + activity tab + auto-refresh + leaderboard tab on team portal 2026-09-23 22:56:07 +08:00
root 877f14ecf3 fix: editTeam endpoint + leaderboard live name resolution + domains in team set + apt-insecure.conf in all service dirs 2026-09-23 21:50:16 +08:00
root 029b0f809a tools in all containers + apt GPG fix for 2026 clock + target dropdown fixed
- all 6 Dockerfiles: vim curl wget netcat git python3-pip now installed
- apt-insecure.conf (AllowInsecureRepositories) copied into images so
  participants can apt-get install despite expired Ubuntu/Debian GPG keys
- warmup base ubuntu:20.04 (EOL, GPG expired) -> ubuntu:24.04
- installed vim+git live into all 18 running team containers
- team portal target dropdown reloads after login (was empty pre-auth)
- attack log endpoint + A/D submit (attacker vs target) verified e2e
2026-09-23 18:54:03 +08:00
root 24dbf0a662 draggable topology + attack visualizer + tools in containers
- topology nodes draggable (pointer events, SVG transform), layout hint shown
- attack visualizer: /api/attacks logs attacker->target events; red pulsing
  dashed arcs on recent attacks (60s hot), ⚔ counts ok/fail
- submit_flag now takes attacker_idx vs target_idx (A/D semantics); UI has
  target dropdown (enemy teams), leaderboard records target
- containers get vim+curl+wget+netcat+git+pip3 (Dockerfiles blogpost/cdn/
  phew/sheesh/warmup); warmup base ubuntu:20.04 EOL -> 24.04
- team portal: target dropdown refreshed after login (was empty pre-auth)
2026-09-23 18:05:44 +08:00
root 44dc1ac852 feat: target matrix + reset scores/environment buttons
- /api/targets (admin): matrix of all teams' domain:port targets
- /api/reset/scores: wipe leaderboard (admin, confirm dialog)
- /api/reset/environment: stop all teams, remove containers+receivers+
  team dirs+systemd units, wipe scores, drop domains (admin, confirm)
- portal targets now only domain+port (no ssh/labels)
- UI: tab Target Matrix, header buttons Reset Skor / Reset Environment
  with confirm() alerts 'apakah anda yakin ingin mereset...'
2026-09-23 17:14:42 +08:00
root b66530e7fd fix: per-team receivers as independent systemd services
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.
2026-09-23 17:06:24 +08:00
root 72382c43d0 fix: guide link IDOR, full team-api IDOR hardening, loading overlay
- guide link now server-side replaced to /team/<idx>/guide (no /team/0 403)
- _check_team_host() applied to ALL team endpoints (login, info, targets,
  status, guide, portal, ssh-ws): host must match team domain; panel/gemastik
  host only with admin session. Cross-domain session reuse -> 403.
- host check BEFORE auth on info/targets (no team-existence oracle)
- loading overlay (spinner + text) on start/stop all-team/set; JS util
  showLoading/hideLoading
2026-09-23 17:02:52 +08:00
root 01ffc05a11 fix: simplify team domain link (root = portal) 2026-09-23 16:38:30 +08:00
root 43ed3df557 feat: team portal with own login, SSH web terminal, targets & guide
- Portal tim punya login sendiri (password = ssh_pass, session team_token)
- SSH Web GUI: /api/team/{idx}/ssh/ws (WebSocket+paramiko) → xterm.js terminal
  (fix: ssh_to_ws non-blocking poll, chall_passwords per-container auth)
- Root <slug>.gemastik.imrnes.team → portal tim (bukan login admin)
- Host validation: /team/{idx} & /team/{idx}/guide 403 kalau host != team domain
- /api/team/{idx}/targets: daftar service tim musuh (attack target)
- guide.html: panduan SSH/attack/defense untuk peserta
- fix esc() String(s) (bug: (s||'').replace is not a function saat port number)
- set_ssh_passwords retry loop (container boot race)
2026-09-23 16:37:58 +08:00
root 4a65f24af3 chore: gitignore runtime flags + leaderboard 2026-09-23 16:18:59 +08:00
root 5e44a70049 feat: team-specific subdomains + portal tim
- create_team now takes label -> slug -> <slug>.gemastik.imrnes.team
- ensure_team_domains() writes Traefik dynamic config (gemastik-teams.yaml)
- team portal at /team/<idx> (public, shows chall ports, SSH, submit form)
- /api/team/<idx>/info + /api/team/<idx>/status (server-side receiver auth)
- UI: name inputs per team (set count -> labels), domain link in card
- delete team endpoint drops its domain
2026-09-23 16:18:11 +08:00
root 265159e9d8 feat: set per-team SSH passwords at runtime (chpasswd in shared images) 2026-09-23 15:47:22 +08:00
root 8c061ba1b7 fix: port scheme 30000 (avoid syncthing 22000), reuse base images (no per-team rebuild), recover corrupted receiver/main.py, compose -p project isolation
- PORT_BASE 20000->30000: team1=31xxx team2=32xxx; syncthing owns 22000
- create_team replaces build: with image: services-<name> so teams reuse base images (was rebuilding 6 images per team, disk 100%)
- recovery: receiver/main.py was corrupted by bad patch (write_file with read_file format); restored from team1 copy + original GitHub
- docker compose -p teamN: project isolation so team compose doesn't overlap (was showing team1 containers for team2)
2026-09-23 15:44:53 +08:00
root 1ca963b2b7 feat: multi-team orchestrator - topology UI, team management, flag randomizer, public submit + leaderboard
- panel/teams.py: create/start/stop teams with isolated ports+creds, per-team receivers with CHALLENGE_PORT/CONTAINER env, flag randomization, submit validation + leaderboard
- panel/main.py: /api/teams, /api/teams/{idx}/randomize, /api/flag/submit, /api/leaderboard, /api/public/teams, /submit page
- panel/static/submit.html: public flag submission UI for teams
- receiver: _ch_port/_ch_container read .env per team, container name overrides
- .gitignore: exclude teams/, .venv, __pycache__
2026-09-23 15:14:52 +08:00