Replace dead tera2.sylyt93.workers.dev (NXDOMAIN) with a local
terabox-resolver service (systemd, :4092) that uses Playwright +
residential proxy + user cookies to resolve TeraBox share links
end-to-end to real signed dlink. Proven path: navigate terabox.app
directly, duplicate cookies across .terabox.app domains, call
/share/list from browser context. Supports directories (child files).
live-probed: /stalk/genshin was NOT real — it just pass-through the raw
SvelteKit __data.json (only nickname/level/nameCardId/profilePicture),
not a full Genshin profile like the Shirokami source (which parses
characters/element/costume/stats via enka-network-api -> get_profile_for_uid).
That proper API is now 404 without an enka token. Removed.
stalk now has 3 real endpoints: github, youtube, twitter.
Ported from Shirokami-API scraper/ai/pollinationsai-{text,gemini,image}.js:
- /ai/chat - OpenAI-compatible chat via text.pollinations.ai (default model 'openai'; source's gpt-5-nano is dead upstream, 404 now triggers fallback)
- /ai/gemini - gemini model via same endpoint
- /ai/image - flux image gen via image.pollinations.ai (returns raw bytes, sniffs JPEG/PNG mime)
Note: pollinations is a free shared API - intermittently 402/502/empty from VPS IP (upstream flakiness, not code)
Shirokami-API source uploads to Ryzumi S3 (s3.ryzumi.vip) which is DNS-dead
from this VPS. Implemented a local-disk uploader keeping the source's response
shape {success, url, fileName, size}:
- POST /uploader/ryzencdn (multipart 'file' field, magic-byte ext detection)
- GET /uploader/file/{name} (path-traversal safe, MIME by extension)
Files stored under /var/lib/scraper/uploads (owned by scraper user).
Ported from Shirokami-API scraper/image/brat.js (v2 path):
- /image/brat - static brat PNG via brat.siputzx.my.id API
- /image/brat/animated - animated brat GIF
Both return raw image bytes with proper Content-Type (image/png, image/gif)
Ported from Shirokami-API scraper/tool/*.js:
- /tool/whois - RDAP lookup (whois.com HTML .df-raw is dead/JS-rendered since 2026; RDAP returns structured register/expiry/registrar data)
- /tool/iplocation - ipapi.co JSON
- /tool/tinyurl - tinyurl.com API
- /tool/check-hosting - hosting-checker.net API
- /tool/cek-resi - cekresi.com AES-128-CBC encrypted form POST
- /tool/hargapangan - Bapanas government API (upstream unreachable from VPS)
Fixes: scraper crate has no :has-text() pseudo-class -> replaced history-table scan with 2-cell-row detection; Html held across .await makes future non-Send -> scoped doc parsing
Ported all missing downloader endpoints from Shirokami-API source
(verified against /home/code/Shirokami-API/scraper/downloader/):
- /download/videy: direct CDN link build (cdn.videy.co/{id}.mp4|.mov)
matches videy.js exactly (id.len()==9 && id[8]=='2' -> .mov)
- /download/tiktok/v2: douyin.wtf hybrid API w/ fallback to embed scrape
- /download/twitter/v2: twitsave.com HTML scrape w/ fallback to Syndication API
- detect_platform now recognizes videy.co
- download_all_in_one dispatches videy
fetch_videy already existed as dead code (never wired); connected it.
== all others (bilibili/danbooru/dood/gdrive/instagram/kfiles/mediafire/
== mega/pinterest/pixeldrain/savetube/soundcloud/spotify/terabox/threads/
== tiktok/twitter/youtubeV2) already mapped to their Rust equivalents
New platforms from Shirokami-style downloadter set:
- /download/dailymotion: real MP4/HLS streams (288p up to 2160p)
- /download/reddit: real v.redd.it HLS media for public posts
- detect + all-in-one dispatch updated
Both verified working from VPS with real URLs.
Spotify (was returning DRM metadata stub):
- Resolve track title via Spotify oEmbed API (no auth, works from VPS)
- Search YouTube Music via yt-dlp ytsearch: -> return real audio stream URL
- Provider: spotify-oembed+ytsearch
Pinterest (was returning 'No download URLs found'):
- Replace dead pinterestdownloader.io API with pinterest-dl CLI scraping
- pinterest-dl handles guest-token/cookie dance, returns real media for
public pins: HLS video stream (v1.pinimg.com/videos/*.m3u8) + poster image
- Provider: pinterest-dl, with Playwright fallback retained
Both verified live: Spotify returns real googlevideo audio URL (itag=251,
3.4MB); Pinterest returns real v1.pinimg.com HLS streams (200 verified).
- fetch_pinterest: pinterestdownloader.io API failures now non-fatal, fall
through to Playwright scraping (captures v.pinimg.com video URLs); returns
graceful 200 'No download URLs found' instead of 502 on dead API.
- fetch_spotify: catch yt-dlp DRM error and return 200 with track metadata +
clear message (direct audio needs Spotify Premium). No more raw 502.
Scrape https://www.tiktok.com/embed/v2/{id} HTML and extract the direct
v16m.tiktokcdn.com MP4 from the <video data-testid=play-video> tag.
Works server-side (no auth) while main site/API are Cloudflare-blocked.
Verimplemented before tikwm/yt-dlp/Playwright fallbacks.
- Nix flake now installs scrape_media.py to $out/bin so CI deploys it with the binary
- run_playwright_scraper locates the script alongside the running binary (Nix store),
the Cargo manifest dir (dev), /home/code/scraper, or $SCRAPER_SCRIPT_DIR
- Fix var_os() Option match
- Add scrape_media.py using Playwright+Chromium to bypass Cloudflare/anti-bot
- Add run_playwright_scraper() Rust helper (spawns Python, probes venv)
- Add playwright_to_download_result() JSON->DownloadResult converter
- Instagram/Facebook: try downr.org, fall back to Playwright (returns real cdninstagram/fbcdn URLs)
- Twitter: scope scraper::Html/Selector parsing in a block so the future stays Send, then Playwright fallback
- Revert --impersonate chrome (unsupported on Linux) to --user-agent
- Install Chromium browsers to /usr/local/share/ms-playwright for all users
The current otakudesu site changed paths and card layout, breaking two APIs:
1. genre pages: singular /genre/{slug}/page/{n}/ 301-redirects to a dead
otakudesu.io placeholder → 0 items. Fix: plural /genres/{slug}/page/{n}/.
Also re-target parse_genre_anime_document from the retired .venz ul li /
.thumbz h2.jdlflm layout to the new .col-anime card layout
(.col-anime-title a, .col-anime-eps, .col-anime-rating, .col-anime-cover img).
2. latest: /latest-anime/ was removed (301 → otakudesu.io). The homepage IS
the latest-episodes feed (.venz ul li with .thumbz h2.jdlflm/.epz), so
fetch_latest_anime_page now uses base_url() for page 1 and /page/N/ for
later pages instead of the dead path.
Blanket-switching the whole komik module to komiku.org broke the manga
list: komiku.org/manga/?tipe=manga 301-redirects to /pustaka/ (the new
library layout, 0 .bge items), while api.komiku.org/manga/?tipe=manga
serves the .bge grid the parser expects unredirected.
Correct split:
- api_url() = api.komiku.org -> manga/manhua/manhwa/genre/search lists (.bge layout)
- genre_list fetches base_url() = komiku.org root, which is POPULATED
(#Genre .ls3 .ls3p h4 + /genre/<slug>/ links) — api.komiku.org root is empty
Reverts the api_url() half of a46adf1; keeps the underlying insight that
the genre-list needs the populated komiku.org root.
api.komiku.org returns an empty body for its root path, which broke the
komik genre-list endpoint (it fetches the api root and parsed 0 genres).
komiku.org serves identical sub-path content (manga lists, genre, search)
plus a populated root. Consolidate the komik module onto the canonical
working domain. KOMIK2_API_URL env override still honoured.
The search parser targeted the index layout (.venz ul li, .thumbz h2.jdlflm,
img, .epz, genre-tag/status/rating classes) but otakudesu search results live
in <ul class=chivsrc><li> with <h2><a>Title</a></h2> plus .set label/value
rows. Search silently returned 0 items.
Rewrote to parse .chivsrc li, extracting title + url from h2 a and
Status/Rating/Genre from .set rows. No poster on the search page (empty).
The scraper cached a single RedisCache (one multiplexed connection) in a
OnceCell forever. When that one connection broke (Redis restart, idle
timeout, network blip), every cache op failed with 'cache io: broken pipe',
and because Cache::get_or_set propagated the post-compute write error,
EVERY API (anime, anime2, komik) returned 500 until a process restart.
Fix:
- Build a fresh RedisCache from a freshly checked-out deadpool connection
per call, so a broken connection self-heals without a process restart
(deadpool recycles/drops dead connections and reconnects on checkout).
- Make Cache::get_or_set treat a cache-write failure as non-fatal: return
the freshly computed value (cache is best-effort), so a transient Redis
outage degrades to cache-less instead of 500.
Pre-existing clippy warnings (repositories/parsers) untouched — out of scope.
mytheclipse-queue and mytheclipse-tracing are now published on crates.io
(v1.21.2). Switch both from local path deps to version deps so the whole
mytheclipse family resolves from crates.io — no more path deps in the
dependency graph.
The user published all mytheclipse crates to crates.io at 1.21.2
(including the mytheclipse-cache redis-0.32 fix). Switch the scraper to
version deps for the released crates (mytheclipse, -cache, -config,
-event) so it tracks the published releases instead of local path deps.
mytheclipse-queue + mytheclipse-tracing are not yet on crates.io, so they
remain path deps to the local workspace.