Commit Graph
44 Commits
Author SHA1 Message Date
asepharyana 4dee99cc43 feat: Spotify via oEmbed+ytsearch, Pinterest via pinterest-dl CLI scraping
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Spotify (was returning DRM metadata stub):
- Resolve track title via Spotify oEmbed API (no auth, works from VPS)
- Search YouTube Music via yt-dlp ytsearch: -> return real audio stream URL
- Provider: spotify-oembed+ytsearch

Pinterest (was returning 'No download URLs found'):
- Replace dead pinterestdownloader.io API with pinterest-dl CLI scraping
- pinterest-dl handles guest-token/cookie dance, returns real media for
  public pins: HLS video stream (v1.pinimg.com/videos/*.m3u8) + poster image
- Provider: pinterest-dl, with Playwright fallback retained

Both verified live: Spotify returns real googlevideo audio URL (itag=251,
3.4MB); Pinterest returns real v1.pinimg.com HLS streams (200 verified).
2026-09-01 17:14:38 +07:00
asepharyana 6b3fad1ad4 feat: Pinterest Playwright fallback + Spotify graceful DRM message
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- fetch_pinterest: pinterestdownloader.io API failures now non-fatal, fall
  through to Playwright scraping (captures v.pinimg.com video URLs); returns
  graceful 200 'No download URLs found' instead of 502 on dead API.
- fetch_spotify: catch yt-dlp DRM error and return 200 with track metadata +
  clear message (direct audio needs Spotify Premium). No more raw 502.
2026-09-01 16:36:36 +07:00
asepharyana 563bce5211 fix: Twitter/X via Syndication API returns real MP4s (twitter-syndication provider)
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Use cargo tokenless https://cdn.syndication.twimg.com/tweet-result?id={id}&token=0
which returns tweet JSON with video variants (MP4 at multiple bitrates).
Works server-side with no auth, wired as primary before savetwitter/yt-dlp.
2026-09-01 16:22:43 +07:00
asepharyana 134fde4ac2 fix: TikTok via embed-page scraping returns real MP4 (tiktok-embed provider)
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Scrape https://www.tiktok.com/embed/v2/{id} HTML and extract the direct
v16m.tiktokcdn.com MP4 from the <video data-testid=play-video> tag.
Works server-side (no auth) while main site/API are Cloudflare-blocked.
Verimplemented before tikwm/yt-dlp/Playwright fallbacks.
2026-09-01 16:17:26 +07:00
asepharyana 4684a6a75c fix: make scrape_media.py path resolution robust + include it in Nix build
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- Nix flake now installs scrape_media.py to $out/bin so CI deploys it with the binary
- run_playwright_scraper locates the script alongside the running binary (Nix store),
  the Cargo manifest dir (dev), /home/code/scraper, or $SCRAPER_SCRIPT_DIR
- Fix var_os() Option match
2026-09-01 15:40:57 +07:00
asepharyana 4c5dfdb156 feat: add Playwright browser scraping fallback for Instagram/Facebook/Twitter
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- Add scrape_media.py using Playwright+Chromium to bypass Cloudflare/anti-bot
- Add run_playwright_scraper() Rust helper (spawns Python, probes venv)
- Add playwright_to_download_result() JSON->DownloadResult converter
- Instagram/Facebook: try downr.org, fall back to Playwright (returns real cdninstagram/fbcdn URLs)
- Twitter: scope scraper::Html/Selector parsing in a block so the future stays Send, then Playwright fallback
- Revert --impersonate chrome (unsupported on Linux) to --user-agent
- Install Chromium browsers to /usr/local/share/ms-playwright for all users
2026-09-01 15:35:20 +07:00
asepharyana dd490535fc fix: integrate yt-dlp for Bilibili/SoundCloud/TikTok fallbacks; fix Bilibili BV ID regex; fix stdout pipe buffer truncation (Stdio::piped); fix duplicate comment lines
Deploy Scraper / build-and-deploy (push) Canceled after 0s
2026-09-01 13:47:36 +07:00
asepharyana fb33b67a29 fix: replace dead YouTube APIs with yt-dlp subprocess; fix all JSON panic crashes; add yt-dlp fallback for TikTok/Bilibili/SoundCloud; fix Bilibili BV ID regex; fix stdout pipe buffer truncation (Stdio::piped)
Deploy Scraper / build-and-deploy (push) Canceled after 0s
2026-09-01 13:34:47 +07:00
asepharyana a41425a887 fix: replace dead YouTube APIs with yt-dlp subprocess; fix all JSON panic crashes; add yt-dlp helper functions
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- YouTube MP3/MP4: replaced dead savetube.media + ytdlpyton.nvlgroup.my.id with yt-dlp subprocess
- Added find_ytdlp() + run_ytdlp_json() helpers with spawn_blocking for >64KB stdout
- Fixed all resp['key'] -> resp.get('key') panics in Twitter, Spotify, KrakenFiles, Danbooru
- Removed dead fetch_youtube_v2_mp3/mp4 + decrypt_savetube functions
- ScraperError::Http -> AppError::ScraperError -> HTTP 502 instead of 500
- All endpoints now return graceful 502 with error message instead of crashing 500

Tested with real URLs:
- YouTube MP3: 200, success, real audio URL extracted
- YouTube MP4: 200, success, 48 format URLs extracted
- TikTok: 200 (tikwm provider)
- MediaFire: 200 (success)
2026-09-01 12:51:52 +07:00
asepharyana c3e6537a63 fix: map ScrapingError to 502 Bad Gateway instead of 500 Internal Server Error; fix downr.org Cloudflare block; fix YouTube savetube DNS failure; fix twitter/pinterest error handling 2026-09-01 11:26:29 +07:00
asepharyana 763a641593 feat(downloader): add 26 downloader APIs + detect endpoint
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Port downloaders from Shirokami-API (Node.js):
- Instagram/Facebook (snapsave.app, downr.org fallback)
- TikTok (tikwm.com)
- YouTube MP4/MP3 (savetube.media, ydlp.yard.id)
- Spotify, Twitter/X (api.lrm.tube)
- Bilibili, Pinterest, MediaFire, Mega, TeraBox
- PixelDrain, Threads, DoodStream, KrakenFiles
- Danbooru, SoundCloud, Google Drive

Domain layer: DownloadResult + MediaItem entities
Application layer: 7 use-cases + detect_platform auto-detect
Infrastructure: DownloaderRepository (22 fetch functions)
Presentation: 21 handlers + DTO + OpenAPI schemas

Also: upgrade OpenTelemetry to 0.28 (axum 0.7 conflict resolution)
2026-09-01 00:54:06 +07:00
asepharyana e6d75edf33 fix(anime): repair genre & latest endpoints for the new otakudesu layout
Deploy Scraper / build-and-deploy (push) Canceled after 0s
The current otakudesu site changed paths and card layout, breaking two APIs:

1. genre pages: singular /genre/{slug}/page/{n}/ 301-redirects to a dead
   otakudesu.io placeholder → 0 items. Fix: plural /genres/{slug}/page/{n}/.
   Also re-target parse_genre_anime_document from the retired .venz ul li /
   .thumbz h2.jdlflm layout to the new .col-anime card layout
   (.col-anime-title a, .col-anime-eps, .col-anime-rating, .col-anime-cover img).

2. latest: /latest-anime/ was removed (301 → otakudesu.io). The homepage IS
   the latest-episodes feed (.venz ul li with .thumbz h2.jdlflm/.epz), so
   fetch_latest_anime_page now uses base_url() for page 1 and /page/N/ for
   later pages instead of the dead path.
2026-08-31 20:59:20 +07:00
asepharyana 80af7b9460 fix(komik): keep api.komiku.org for lists, source genre-list from komiku.org root
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Blanket-switching the whole komik module to komiku.org broke the manga
list: komiku.org/manga/?tipe=manga 301-redirects to /pustaka/ (the new
library layout, 0 .bge items), while api.komiku.org/manga/?tipe=manga
serves the .bge grid the parser expects unredirected.

Correct split:
- api_url() = api.komiku.org  -> manga/manhua/manhwa/genre/search lists (.bge layout)
- genre_list fetches base_url() = komiku.org  root, which is POPULATED
  (#Genre .ls3 .ls3p h4 + /genre/<slug>/ links) — api.komiku.org root is empty

Reverts the api_url() half of a46adf1; keeps the underlying insight that
the genre-list needs the populated komiku.org root.
2026-08-31 20:42:12 +07:00
asepharyana a46adf1ce9 fix(komik): use komiku.org instead of dead api.komiku.org root
Deploy Scraper / build-and-deploy (push) Canceled after 0s
api.komiku.org returns an empty body for its root path, which broke the
komik genre-list endpoint (it fetches the api root and parsed 0 genres).
komiku.org serves identical sub-path content (manga lists, genre, search)
plus a populated root. Consolidate the komik module onto the canonical
working domain. KOMIK2_API_URL env override still honoured.
2026-08-31 20:25:56 +07:00
asepharyana 6bfffa5364 fix(anime): parse otakudesu search results from .chivsrc li
The search parser targeted the index layout (.venz ul li, .thumbz h2.jdlflm,
img, .epz, genre-tag/status/rating classes) but otakudesu search results live
in <ul class=chivsrc><li> with <h2><a>Title</a></h2> plus .set label/value
rows. Search silently returned 0 items.

Rewrote to parse .chivsrc li, extracting title + url from h2 a and
Status/Rating/Genre from .set rows. No poster on the search page (empty).
2026-08-31 20:22:39 +07:00
asepharyana 7a12f0e9c8 fix(cache): self-heal Redis connection + don't 500 on cache write failure
Deploy Scraper / build-and-deploy (push) Canceled after 0s
The scraper cached a single RedisCache (one multiplexed connection) in a
OnceCell forever. When that one connection broke (Redis restart, idle
timeout, network blip), every cache op failed with 'cache io: broken pipe',
and because Cache::get_or_set propagated the post-compute write error,
EVERY API (anime, anime2, komik) returned 500 until a process restart.

Fix:
- Build a fresh RedisCache from a freshly checked-out deadpool connection
  per call, so a broken connection self-heals without a process restart
  (deadpool recycles/drops dead connections and reconnects on checkout).
- Make Cache::get_or_set treat a cache-write failure as non-fatal: return
  the freshly computed value (cache is best-effort), so a transient Redis
  outage degrades to cache-less instead of 500.

Pre-existing clippy warnings (repositories/parsers) untouched — out of scope.
2026-08-31 20:11:04 +07:00
asepharyana 3d38008af3 build: use published mytheclipse crates for queue + tracing (1.21.2)
Deploy Scraper / build-and-deploy (push) Canceled after 0s
mytheclipse-queue and mytheclipse-tracing are now published on crates.io
(v1.21.2). Switch both from local path deps to version deps so the whole
mytheclipse family resolves from crates.io — no more path deps in the
dependency graph.
2026-08-30 20:02:01 +07:00
asepharyana e681bbbe51 build: use published mytheclipse crates (1.21.2) via version deps
Deploy Scraper / build-and-deploy (push) Canceled after 0s
The user published all mytheclipse crates to crates.io at 1.21.2
(including the mytheclipse-cache redis-0.32 fix). Switch the scraper to
version deps for the released crates (mytheclipse, -cache, -config,
-event) so it tracks the published releases instead of local path deps.
mytheclipse-queue + mytheclipse-tracing are not yet on crates.io, so they
remain path deps to the local workspace.
2026-08-30 19:45:05 +07:00
asepharyana 966d86ff67 refactor: migrate thread/async/queue/scheduler infra to mytheclipse
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- bootstrap: TracingLayer from mytheclipse-tracing (composed with scraper
  env filter), RuntimeConfig::auto() thread logging, init job queue + cron.
- proxy_fetch: leader task via ::mytheclipse::spawn_io, gzip decompression
  offloaded to ::mytheclipse::compute (sized rayon pool, panic-isolated),
  bounded by new async fetch limiter (tokio Semaphore bridge).
- queue: new infrastructure/queue module over mytheclipse-queue
  (InMemoryQueue + WorkerPool + BackpressureEnforcer) for repair jobs.
- scheduler: new infrastructure/scheduler.rs using mytheclipse::cron for
  the daily 02:00 UTC image-cache cleanup.
- deps: add mytheclipse-queue + mytheclipse-tracing path deps; mytheclipse
  -> full feature; rayon 1.12; keep path deps for unpublished crates.
- tests: infra_round2 runtime smoke tests (compute panic isolation,
  spawn_io, backpressure admission, cron parse, queue roundtrip).
- Fix: local cache::mytheclipse bridge module shadowed the mytheclipse
  crate name; use leading :: at spawn_io/compute call sites.
2026-08-30 19:40:01 +07:00
asepharyana c7ffd29e79 refactor: migrate infra to mytheclipse crates (retry, cache, event, ratelimit, config)
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Replace hand-rolled infrastructure with the custom mytheclipse library:

- retry (src/infrastructure/scraping/retry.rs): backoff crate -> mytheclipse
  RetryConfig + retry() with retry_all predicate; same call-site helpers.
- cache (src/infrastructure/cache): deadpool raw AsyncCommands ->
  mytheclipse_cache::RedisCache bridge (src/infrastructure/cache/mytheclipse.rs),
  typed JSON wrapper keeps the Cache<'a> API used by use cases.
- events (src/events/bus.rs): custom broadcast pub/sub ->
  mytheclipse_event InMemoryEventBus + TypedEventBus alias.
- ratelimit (src/presentation/middleware/ratelimit.rs): hand-rolled window ->
  mytheclipse::RateLimiter token bucket.
- config (src/config/mod.rs): config crate -> mytheclipse_config ConfigLoader,
  preserving env-only fallback (config files optional).

No route/API/Redis-key changes. mytheclipse-cache path-dep points at local
source (redis 0.32 aligned with deadpool). Removed backoff/config deps.
2026-08-30 18:53:00 +07:00
asepharyana 6ebff90d5b ci: deploy via Attic binary cache (attic.asepharyana.my.id) instead of ssh copy
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Push build result to attic cache directly from runner, then VPS
substitutes it via nix-store --realise. ssh copy retained only as fallback.
2026-08-27 20:38:36 +07:00
asepharyana d1b24f8f85 ci: add standalone Nix build + deploy workflow; remove parent notify
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Repo is now self-contained: flake.nix (cargo build) + deploy.yml
(nix build -> nix copy -> nix-env profile -> systemctl restart scraper).
No longer dispatches submodule-updated to asepharyana-hub.
2026-08-27 19:46:35 +07:00
asepharyana ad0d4f4992 chore: update dependabot-auto-merge.yml
Notify Parent Repo / dispatch (push) Canceled after 0s
2026-08-26 08:53:39 +07:00
asepharyana 7ff47a79bc chore: update dependabot.yml 2026-08-26 08:53:38 +07:00
asepharyana c709c84a4d feat: add dependabot auto-merge workflow 2026-08-25 20:32:55 +07:00
asepharyana 55cfd513e3 feat: update dependabot config 2026-08-25 18:16:42 +07:00
asepharyana 380ecfcd9b feat: update dependabot config 2026-08-25 18:12:56 +07:00
asepharyana a23698a6c4 feat: add dependabot config (npm auto-deps) 2026-08-25 18:08:40 +07:00
mytheclipsebotreview 62aa5b0e52 fix(config): remove list_separator, breaks string fields 2026-07-30 22:18:03 +07:00
asepharyana bc782ae4f8 ci: add workflow_dispatch to notify-parent for manual triggering 2026-07-25 18:13:25 +07:00
asepharyana 7640e66c76 chore: deny unused_mut, unreachable_code, trivial_casts, trivial_numeric_casts, explicit_outlives_requirements, unused_labels, unused_braces, noop_method_call, clippy::unnecessary_to_owned 2026-07-22 15:03:13 +07:00
asepharyana 96ce29b0f3 chore: set dead_code lint to deny 2026-07-22 14:52:19 +07:00
asepharyana 5c51ec895f feat: remove croxy proxy and image cache features 2026-07-22 14:51:08 +07:00
asepharyana 891064354b fix: disable extra_float_digits startup param for PG16 compatibility 2026-07-22 14:23:36 +07:00
asepharyana 4cd04410cf chore: remove 17 unused Rust dependencies from scraper 2026-07-22 14:12:58 +07:00
asepharyana cd4e3aee20 feat: add episode list with download URLs to detail endpoint
- Add AlqEpisode type with episode, title, url, date, download_url
- Add parse_episode_list() to extract episodes from .eplister
- Add parse_episode_download() to extract download URL from episode page
- Add episodes field to AlqDetailData
- Fetch download URLs for latest 5 episodes in parallel
2026-07-22 14:04:47 +07:00
asepharyana 052c6cfeb8 fix: revert to alqanime.si, use fetch_with_proxy for direct+relay
alqanime.net is behind Cloudflare challenge - direct fetch returns
challenge page and proxy relays are down. alqanime.si works with
direct fetch.

Also change repository to use fetch_with_proxy (direct+relay fallback)
instead of fetch_with_proxy_only (relay only).
2026-07-22 13:56:34 +07:00
asepharyana 1a54e8a127 docs: update CLAUDE.md alqanime.si -> alqanime.net 2026-07-22 13:43:11 +07:00
asepharyana c035601a36 feat: migrate anime2 scraper from alqanime.si to alqanime.net
- Change BASE_URL from https://alqanime.si to https://alqanime.net
- Remove BASE_DETAIL_URL (now same as BASE_URL)
- Remove detail_image_url method (no longer needed)
- Simplify detail() use case: remove dual-fetch logic that merged
  results from both domains (now same URL)
- Fix CLAUDE.md stale comment
2026-07-22 13:41:03 +07:00
asepharyana 562ada69ec fix: only connect OTLP metrics when OTEL_EXPORTER_OTLP_ENDPOINT is set
When unset, use a no-op meter provider instead of defaulting to
http://localhost:4317 and spamming connection errors every 5s.
2026-07-22 13:03:03 +07:00
asepharyana 1b195ac3f6 refactor: remove EXTERNAL_BROWSERLESS_WS / browserless fallback
- Delete infrastructure/browser/ directory entirely (browser pool is gone)
- Remove browser pool initialization from bootstrap
- Remove fetch_via_browserless fallback from proxy_fetch
- perform_proxy_chain now only uses relay endpoints
- Clean up log messages, README, and compose env var
The proxy-bun service at https://git.imrnes.team/MythEclipse/proxy-bun
handles all proxy needs now.
2026-07-22 12:55:21 +07:00
asepharyana 1ddcd15540 fix: update notify-parent target to asepharyana/asepharyana-hub
The old target MythEclipse/ultimate-asepharyana.tech is dead.
Repository dispatch now goes to the correct parent repo.
2026-07-22 07:06:58 +07:00
asepharyanaandClaude Opus 4.8 776fa828d8 feat: full clean architecture refactor — domain → application → infrastructure → presentation
- Restructure from Modular MVC to Clean Architecture with 4 strict layers
- Domain: entities (anime/komik/proxy), Repository traits, typed errors
- Application: use case classes with proper error propagation
- Infrastructure: repository impls, parsers, Redis cache, HTTP/scraping, browser
- Presentation: Axum handlers, DTOs, AppState, router, AppError+IntoResponse
- Migrate all parsers (otakudesu, alqanime, komik) to native infra implementations
- Replace once_cell::sync::Lazy/OnceCell with std::sync::LazyLock/OnceLock
- Remove 150+ old files in modules/ and shared/ directories
- Remove once_cell from Cargo.toml dependencies
- Fix test/debug binaries to use new import paths
- Zero new clippy warnings

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 12:54:23 +07:00
asepharyana 80c96eaa42 chore: initial commit for asepharyana-hub-scraper 2026-07-09 22:07:03 +07:00