Commit Graph
8 Commits
Author SHA1 Message Date
asepharyana 0157029da8 fix: pin autofix agent to claude-sonnet-5 via custom:9router provider 2026-09-23 22:00:02 +07:00
asepharyana 66b567a6bc feat(worker): per-fork session header + 3600s sync timeout + merge salvage
- _hermes_api_post sends X-Hermes-Session-Id so each fork sync uses a
  fresh gateway transcript (was: derived session reused 200k-token history
  across upstream tips, making later runs pay growing context)
- SYNC_CLAUDE_TIMEOUT 2700→3600: a fresh session needs ~25min/120 calls
  to resolve 17 files + verify + commit; 2700s cut the HTTP call while the
  agent was still committing
- salvage-on-timeout: when the agent call ends early (timeout/transport)
  but the merge is already committed (MERGE_HEAD gone), finish+push it
  instead of aborting — verified live: 17 conflicts resolved by agent,
  push ok, fork 27 ahead / 0 behind
- _sync_finish_merge blocks committing files that still carry conflict
  markers (guard against git add of unresolved files)
- tests: 68/68 (session header, salvage completed merge, abort unfinished,
  marker guard; fake-leak hardening in _install_git_fake)
2026-09-21 21:02:33 +07:00
asepharyana 99d47c8ff8 feat(worker): drive Hermes gateway API server instead of claude -p
Claude Code could not complete AI fixes / conflict resolution against this
host's provider setup: it hung 900s spawning an MCP server, then failed with
'body is JSON but not a Message' (Anthropic-Messages transport mismatch), then
exited 1 with empty stderr. The Hermes gateway already runs continuously with
the working 9router provider config and a full toolset, so drive it directly:

- POST http://127.0.0.1:8642/v1/chat/completions (OpenAI-compatible API server)
- bearer auth from API_SERVER_KEY (env or ~/.hermes/.env), overridable via
  API_SERVER_URL
- model_options.max_turns caps a runaway run; 900s timeout for sync, 600s for
  PR fixes
- every failure maps to an [INFRA] string so the existing skip-once logic works
- the agent commits locally; the WORKER pushes (agents must never push)

run_ai_fix no longer shells out to claude; it calls the API server, then pushes
the agent's commit itself and reports push failures explicitly. Sync call sites
keep their contract via _run_claude_sync -> _run_hermes_sync alias, with labels
renamed hermes_sync_conflicts / hermes_sync_quality.

Tests: 59/59 (8 new assertions exercise a real local HTTP round-trip: path,
bearer auth, OpenAI message shape, max_turns cap, HTTP-error/missing-key/
unreachable -> [INFRA]).
2026-09-21 18:24:47 +07:00
asepharyana 0653f6c508 fix(worker): pure --dry runs + Claude Code MCP-server hang
- --dry now touches nothing: no state file writes (run_upstream_sync uses a
  throwaway state dict; sync_fork_repo refuses to mutate in dry), no Discord,
  no push, no PR (protected-branch dry prints would-open instead). The earlier
  dry run polluted /tmp/pr-queue-sync-state.json and posted a false 'skip'
  notification — both are gone.
- Claude Code conflict/quality runs now use --strict-mcp-config (with
  --mcp-config ''): --mcp-config '' alone still lets claude -p spawn MCP
  servers from settings.json/managed/plugins (observed ouroboros mcp serve
  hanging 15+ min with zero output until the 900s timeout). These calls only
  read/edit a throwaway clone and run git — no MCP server is ever needed.
- tests: 51/51 (added dry-purity cases: clean-merge, conflict-failure).
2026-09-21 17:31:16 +07:00
asepharyana c48adea6c3 feat(worker): upstream fork auto-sync — pull+merge fork repos hourly
Merge new upstream (parent) commits into every fork in the App installation,
gated by a per-repo interval (default 1h), inside the existing 5-minute tick
(STEP 0, max 2 forks/tick, oldest-first).

- Conflicted merges are resolved by Claude Code (merge-reconciler rules:
  never wholesale --ours/--theirs, verify with the repo's own
  typecheck+tests, commit --no-edit; Claude never pushes — harness does).
- Clean merges get a single Claude Code quality pass commit.
- Push path: owner PAT (gh CLI) first — the App lacks workflows:write and a
  workflows-touching merge is rejected for the App token; App token fallback.
- Protected default branch: detected from the push result (GH006 /
  required-status-check) → upstream-sync-<ts> branch + PR through the normal
  pipeline; duplicate open sync PRs are skipped.
- CI safety: after a direct push, ticks verify the fork CI at our merge sha;
  red CI at OUR merge (still the tip) → sha-guarded force-revert to
  pre-merge sha + Discord notify; never reverts foreign commits.
- Discord: synced / PR opened / reverted / skipped-once on pr-agent-ops.
- merge_pr gains the same PAT fallback (a PR merge touching workflows is a
  workflow-file push).
- CLI: --sync-status, --sync-only <repo> [--dry].
- Tests: scripts/test_pr_queue_sync.py (46 assertions, monkeypatched, no
  network); py_compile clean.
- Plan: .hermes/plans/2026-09-21-upstream-auto-sync.md
2026-09-21 17:04:39 +07:00
asepharyana 1dc4d098fa feat(worker): skip-on-error no loop + conflict auto-fix + deduped notif
1. Skip ONCE on infra errors (Claude Code CLI missing/timeout/unreachable):
   - run_ai_fix returns '[INFRA] ...' reasons; worker marks the PR permanently
     skipped at that head SHA in fix-state (no more retry every 5 min)
   - skip is recorded per {repo,pr,sha}; cleared when head SHA changes
2. Merge conflict auto-fix via Claude Code:
   - mergeable=False or merge HTTP 409 now trigger run_ai_fix (prompt already
     merges base + resolves conflicts) once per head SHA
   - success → next tick re-checks mergeable and merges; failure → skip once
3. Discord skip notification dedupe:
   - notify_skip_once(): posts '⏭️ Skipped: <reason>' exactly once per
     PR+head_sha (state.notified flag); no repeated spam every cron tick
   - skip reason + which PR is visible in the notification
4. State migration: legacy {repo:{pr:'sha'}} → dict form handled in _pr_entry
Verified: py_compile clean, state-helper unit tests pass (skip/fixed/notify
dedupe/legacy migration). Cron wrapper execs repo copy — no manual sync.
2026-09-21 15:58:52 +07:00
asepharyana bc69998d8e feat(server): production cut-over to Bun — notify_review, setup/callback, key resolution
- index.ts: add POST /api/v1/notify_review (queue-worker Discord bridge),
  /setup/callback (GitHub App manifest conversion), error logging on review
  failure, PR_AGENT_APP_DIR-based private key resolution
- config.ts: API key falls back to on-disk omniroute_key (same source as
  run_server.py) when no env key present — fixes 401 in systemd context
- markdown.ts: don't hyperlink non-URL ticket values (N/A)
- secrets.ts: fallback private key from ~/.hermes/keys + omni key ~/.hermes
- pr-queue-worker.py / auto_merge_bot.py: notify_review default port 4002→4023

Deploy: pr-agent-bun.service (Bun binary, port 4023) replaces
pr-agent-server.service (Python, port 4002, disabled). Caddy route updated in
asepharyana/infra (proxy 4023). Verified live: webhook → review → claude-opus-5
→ persistent GitHub comment published after Python shutdown.
2026-09-21 13:47:47 +07:00
asepharyana 835552c5ca feat(ops): add PR queue worker with toolchain pin guard
- pr-queue-worker.py: cron orchestrator (review trigger → AI fix → safety →
  CI gate → approve/merge) now versioned in-repo
- TOOLCHAIN_PINS: close dependabot PRs bumping pinned majors
  (typescript/eslint/@tsparticles/eslint-config-next/eslint-plugin-react)
- STALE_CI_CLOSE_DAYS=2: close dependabot PRs stuck failing CI
- Secrets externalized to env (PR_AGENT_*), hydrated from ~/.hermes/.env —
  file is safe for the public repo; no inline secrets
2026-09-21 11:36:14 +07:00