Ported all missing downloader endpoints from Shirokami-API source
(verified against /home/code/Shirokami-API/scraper/downloader/):
- /download/videy: direct CDN link build (cdn.videy.co/{id}.mp4|.mov)
matches videy.js exactly (id.len()==9 && id[8]=='2' -> .mov)
- /download/tiktok/v2: douyin.wtf hybrid API w/ fallback to embed scrape
- /download/twitter/v2: twitsave.com HTML scrape w/ fallback to Syndication API
- detect_platform now recognizes videy.co
- download_all_in_one dispatches videy
fetch_videy already existed as dead code (never wired); connected it.
== all others (bilibili/danbooru/dood/gdrive/instagram/kfiles/mediafire/
== mega/pinterest/pixeldrain/savetube/soundcloud/spotify/terabox/threads/
== tiktok/twitter/youtubeV2) already mapped to their Rust equivalents
New platforms from Shirokami-style downloadter set:
- /download/dailymotion: real MP4/HLS streams (288p up to 2160p)
- /download/reddit: real v.redd.it HLS media for public posts
- detect + all-in-one dispatch updated
Both verified working from VPS with real URLs.
Spotify (was returning DRM metadata stub):
- Resolve track title via Spotify oEmbed API (no auth, works from VPS)
- Search YouTube Music via yt-dlp ytsearch: -> return real audio stream URL
- Provider: spotify-oembed+ytsearch
Pinterest (was returning 'No download URLs found'):
- Replace dead pinterestdownloader.io API with pinterest-dl CLI scraping
- pinterest-dl handles guest-token/cookie dance, returns real media for
public pins: HLS video stream (v1.pinimg.com/videos/*.m3u8) + poster image
- Provider: pinterest-dl, with Playwright fallback retained
Both verified live: Spotify returns real googlevideo audio URL (itag=251,
3.4MB); Pinterest returns real v1.pinimg.com HLS streams (200 verified).
- fetch_pinterest: pinterestdownloader.io API failures now non-fatal, fall
through to Playwright scraping (captures v.pinimg.com video URLs); returns
graceful 200 'No download URLs found' instead of 502 on dead API.
- fetch_spotify: catch yt-dlp DRM error and return 200 with track metadata +
clear message (direct audio needs Spotify Premium). No more raw 502.
Scrape https://www.tiktok.com/embed/v2/{id} HTML and extract the direct
v16m.tiktokcdn.com MP4 from the <video data-testid=play-video> tag.
Works server-side (no auth) while main site/API are Cloudflare-blocked.
Verimplemented before tikwm/yt-dlp/Playwright fallbacks.
- Nix flake now installs scrape_media.py to $out/bin so CI deploys it with the binary
- run_playwright_scraper locates the script alongside the running binary (Nix store),
the Cargo manifest dir (dev), /home/code/scraper, or $SCRAPER_SCRIPT_DIR
- Fix var_os() Option match
- Add scrape_media.py using Playwright+Chromium to bypass Cloudflare/anti-bot
- Add run_playwright_scraper() Rust helper (spawns Python, probes venv)
- Add playwright_to_download_result() JSON->DownloadResult converter
- Instagram/Facebook: try downr.org, fall back to Playwright (returns real cdninstagram/fbcdn URLs)
- Twitter: scope scraper::Html/Selector parsing in a block so the future stays Send, then Playwright fallback
- Revert --impersonate chrome (unsupported on Linux) to --user-agent
- Install Chromium browsers to /usr/local/share/ms-playwright for all users
The current otakudesu site changed paths and card layout, breaking two APIs:
1. genre pages: singular /genre/{slug}/page/{n}/ 301-redirects to a dead
otakudesu.io placeholder → 0 items. Fix: plural /genres/{slug}/page/{n}/.
Also re-target parse_genre_anime_document from the retired .venz ul li /
.thumbz h2.jdlflm layout to the new .col-anime card layout
(.col-anime-title a, .col-anime-eps, .col-anime-rating, .col-anime-cover img).
2. latest: /latest-anime/ was removed (301 → otakudesu.io). The homepage IS
the latest-episodes feed (.venz ul li with .thumbz h2.jdlflm/.epz), so
fetch_latest_anime_page now uses base_url() for page 1 and /page/N/ for
later pages instead of the dead path.
Blanket-switching the whole komik module to komiku.org broke the manga
list: komiku.org/manga/?tipe=manga 301-redirects to /pustaka/ (the new
library layout, 0 .bge items), while api.komiku.org/manga/?tipe=manga
serves the .bge grid the parser expects unredirected.
Correct split:
- api_url() = api.komiku.org -> manga/manhua/manhwa/genre/search lists (.bge layout)
- genre_list fetches base_url() = komiku.org root, which is POPULATED
(#Genre .ls3 .ls3p h4 + /genre/<slug>/ links) — api.komiku.org root is empty
Reverts the api_url() half of a46adf1; keeps the underlying insight that
the genre-list needs the populated komiku.org root.
api.komiku.org returns an empty body for its root path, which broke the
komik genre-list endpoint (it fetches the api root and parsed 0 genres).
komiku.org serves identical sub-path content (manga lists, genre, search)
plus a populated root. Consolidate the komik module onto the canonical
working domain. KOMIK2_API_URL env override still honoured.
The search parser targeted the index layout (.venz ul li, .thumbz h2.jdlflm,
img, .epz, genre-tag/status/rating classes) but otakudesu search results live
in <ul class=chivsrc><li> with <h2><a>Title</a></h2> plus .set label/value
rows. Search silently returned 0 items.
Rewrote to parse .chivsrc li, extracting title + url from h2 a and
Status/Rating/Genre from .set rows. No poster on the search page (empty).
The scraper cached a single RedisCache (one multiplexed connection) in a
OnceCell forever. When that one connection broke (Redis restart, idle
timeout, network blip), every cache op failed with 'cache io: broken pipe',
and because Cache::get_or_set propagated the post-compute write error,
EVERY API (anime, anime2, komik) returned 500 until a process restart.
Fix:
- Build a fresh RedisCache from a freshly checked-out deadpool connection
per call, so a broken connection self-heals without a process restart
(deadpool recycles/drops dead connections and reconnects on checkout).
- Make Cache::get_or_set treat a cache-write failure as non-fatal: return
the freshly computed value (cache is best-effort), so a transient Redis
outage degrades to cache-less instead of 500.
Pre-existing clippy warnings (repositories/parsers) untouched — out of scope.
mytheclipse-queue and mytheclipse-tracing are now published on crates.io
(v1.21.2). Switch both from local path deps to version deps so the whole
mytheclipse family resolves from crates.io — no more path deps in the
dependency graph.
The user published all mytheclipse crates to crates.io at 1.21.2
(including the mytheclipse-cache redis-0.32 fix). Switch the scraper to
version deps for the released crates (mytheclipse, -cache, -config,
-event) so it tracks the published releases instead of local path deps.
mytheclipse-queue + mytheclipse-tracing are not yet on crates.io, so they
remain path deps to the local workspace.
- Add AlqEpisode type with episode, title, url, date, download_url
- Add parse_episode_list() to extract episodes from .eplister
- Add parse_episode_download() to extract download URL from episode page
- Add episodes field to AlqDetailData
- Fetch download URLs for latest 5 episodes in parallel
alqanime.net is behind Cloudflare challenge - direct fetch returns
challenge page and proxy relays are down. alqanime.si works with
direct fetch.
Also change repository to use fetch_with_proxy (direct+relay fallback)
instead of fetch_with_proxy_only (relay only).
- Change BASE_URL from https://alqanime.si to https://alqanime.net
- Remove BASE_DETAIL_URL (now same as BASE_URL)
- Remove detail_image_url method (no longer needed)
- Simplify detail() use case: remove dual-fetch logic that merged
results from both domains (now same URL)
- Fix CLAUDE.md stale comment
- Delete infrastructure/browser/ directory entirely (browser pool is gone)
- Remove browser pool initialization from bootstrap
- Remove fetch_via_browserless fallback from proxy_fetch
- perform_proxy_chain now only uses relay endpoints
- Clean up log messages, README, and compose env var
The proxy-bun service at https://git.imrnes.team/MythEclipse/proxy-bun
handles all proxy needs now.
- Restructure from Modular MVC to Clean Architecture with 4 strict layers
- Domain: entities (anime/komik/proxy), Repository traits, typed errors
- Application: use case classes with proper error propagation
- Infrastructure: repository impls, parsers, Redis cache, HTTP/scraping, browser
- Presentation: Axum handlers, DTOs, AppState, router, AppError+IntoResponse
- Migrate all parsers (otakudesu, alqanime, komik) to native infra implementations
- Replace once_cell::sync::Lazy/OnceCell with std::sync::LazyLock/OnceLock
- Remove 150+ old files in modules/ and shared/ directories
- Remove once_cell from Cargo.toml dependencies
- Fix test/debug binaries to use new import paths
- Zero new clippy warnings
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>