Files
plex-playlist/scripts/gitea-actions/diagnose_runner_startup.xsh
Xlorep DarkHelm f3698b095b
All checks were successful
CICD / Build and Publish CICD Base Image (push) Successful in 6m8s
CICD / Build and Push CICD Image (push) Successful in 23m14s
CICD / Build CICD Image Failure Postmortem (push) Has been skipped
CICD / Backend Tests (push) Successful in 7m10s
CICD / Frontend Tests (push) Successful in 45s
CICD / Backend Doctests (push) Successful in 18s
CICD / Pre-commit Checks (push) Successful in 14m57s
CICD / Source Lanes Failure Postmortem (push) Has been skipped
CICD / CICD Tests Complete (push) Successful in 3s
CICD / Build Backend Base Image (push) Successful in 18s
CICD / Build Integration Tester Image (push) Successful in 1m5s
CICD / Build Backend Main Image (push) Successful in 1m52s
CICD / Build Frontend Base Image (push) Successful in 10m42s
CICD / Build Frontend Main Image (push) Successful in 33s
CICD / Build E2E Tester Image (push) Successful in 32m17s
CICD / Production Images Complete (push) Successful in 5s
CICD / Production Image Failures Postmortem (push) Has been skipped
CICD / Runtime Black-Box Integration Tests (push) Successful in 1m13s
CICD / Integration Tests Failure Postmortem (push) Has been skipped
CICD / End-to-End Tests (push) Successful in 11m23s
CICD / E2E Tests Failure Postmortem (push) Has been skipped
Stabilize self-hosted CI workflows and resolve issue #62 (#73)
## Summary

Hardens CI workflows for self-hosted Gitea runners by stabilizing E2E execution and Renovate behavior across internal/external network paths.

Closes #62

## What Changed

### E2E workflow reliability
- Fixed E2E workspace handoff to ensure expected repository contents are present during test execution.
- Added stricter preflight checks for required frontend files before running E2E.
- Reduced mount/path fragility while preserving runtime image pull and compose flow.

### Renovate workflow hardening
- Added internal-first endpoint reachability selection with fallback handling.
- Added token preflight checks for repository access.
- Added explicit host-rule auth handling for API/git paths.
- Added container-level connectivity preflight diagnostics.
- Added git URL override aligned with selected endpoint context.
- Removed incorrect forced Dogar host-IP pinning that broke HTTPS clone routing.

## Why

CI behavior was sensitive to runner networking and Renovate clone/auth interactions. These changes make the workflow deterministic in our runner topology and address recurring CI failures.

## Scope

- Workflow logic only (`cicd.yaml`, `renovate.yml`)
- No app feature or API behavior changes

## Validation

- Workflow YAML validation passed during updates.
- Changes were applied and verified iteratively from real failing run diagnostics.

Co-authored-by: copilotcoder <copilotcoder@darkhelm.org>
Reviewed-on: #73
2026-07-13 11:16:16 -04:00

105 lines
3.7 KiB
Plaintext
Executable File

#!/usr/bin/env xonsh
import sys
hosts = {
'kankali.darkhelm.lan': '/home/darkhelm/Projects/DarkHelm.org/gitea',
'zhokq.darkhelm.lan': '/home/darkhelm/Projects/DarkHelm.org/gitea-runner',
'urtzul.darkhelm.lan': '/home/darkhelm/Projects/DarkHelm.org/gitea-runner',
'pi-desktop.darkhelm.lan': '/home/darkhelm/Projects/DarkHelm.org/gitea-runner',
}
mirror_image = 'kankali.darkhelm.lan:3001/darkhelm.org/act-ubuntu:act-latest'
ghcr_runner_image = 'ghcr.io/catthehacker/ubuntu:act-latest'
renovate_image = 'ghcr.io/renovatebot/renovate:41'
do_fix = '--fix' in sys.argv[1:]
remote_diag = f"""
print('pwd=' + $(pwd).strip())
print('[runner_containers]')
r = !(docker ps -a --filter name=gitea-act-runner --format '{{{{.Names}}}} {{{{.Status}}}}')
print(str(r.out).strip())
r = !(docker ps -a --filter name=gitea-act-runner --format '{{{{.Names}}}}')
runner_names = [line.strip() for line in str(r.out).splitlines() if line.strip()]
print('[registry_config]')
r = !(docker info 2> /dev/null | grep -qi 'kankali.darkhelm.lan:3001')
print('REGISTRY_CONFIG:ok' if r.returncode == 0 else 'REGISTRY_CONFIG:missing')
print('[image_state]')
r = !(docker image inspect {mirror_image} > /dev/null 2> /dev/null)
print('MIRROR_LOCAL:present' if r.returncode == 0 else 'MIRROR_LOCAL:missing')
r = !(docker pull {mirror_image} > /tmp/runner-mirror-pull.log 2>&1)
if r.returncode == 0:
print('MIRROR_REMOTE:pull-ok')
else:
mismatch = !(grep -qi 'http response to https client' /tmp/runner-mirror-pull.log)
print('MIRROR_REMOTE:https-mismatch' if mismatch.returncode == 0 else 'MIRROR_REMOTE:pull-failed')
tail_out = !(tail -n 20 /tmp/runner-mirror-pull.log)
print(str(tail_out.out).strip())
if not runner_names:
print('[recent_logs] no-runner-containers-found')
for name in runner_names:
print(f'[recent_logs_{name}]')
log = !(docker logs --tail 60 @(name) 2>&1)
log_lines = [
line for line in str(log.out).splitlines()
if any(tok in line.lower() for tok in ['error', 'failed', 'timeout', 'panic', 'unregister', 'forbidden', 'unauthorized', 'connection refused', 'context deadline', 'no route to host', 'network is unreachable'])
]
print('\\n'.join(log_lines[-60:]))
"""
remote_fix = """
print('[fix] restarting runners')
r = !(docker compose up -d --force-recreate act_runner_1 act_runner_2 > /tmp/runner-restart.log 2>&1)
if r.returncode != 0:
r = !(docker compose up -d --force-recreate runner1 runner2 > /tmp/runner-restart.log 2>&1)
if r.returncode != 0:
tail_out = !(tail -n 60 /tmp/runner-restart.log)
print(str(tail_out.out).strip())
raise SystemExit(1)
$(sleep 3)
print('[fix] warming images')
$(docker pull ubuntu:22.04)
$(docker pull python:3.14-slim)
$(docker pull node:20-bookworm-slim)
$(docker pull __GHCR_RUNNER_IMAGE__)
$(docker pull __MIRROR_IMAGE__)
$(docker pull __RENOVATE_IMAGE__)
r = !(docker ps --filter name=gitea-act-runner --format '{{{{.Names}}}} {{{{.Status}}}}')
print(str(r.out).strip())
"""
remote_fix = (
remote_fix
.replace("__MIRROR_IMAGE__", mirror_image)
.replace("__GHCR_RUNNER_IMAGE__", ghcr_runner_image)
.replace("__RENOVATE_IMAGE__", renovate_image)
)
for host, path in hosts.items():
print(f"\n===== {host} =====")
cmd = f"cd {path}\n{remote_diag}"
r = !(ssh @(host) @(cmd))
if r.returncode != 0:
print('DIAG_FAILED')
print(str(r.out).strip())
continue
print(str(r.out).strip())
if do_fix:
print('[fix] applying restart')
fix_cmd = f"cd {path}\n{remote_fix}"
f = !(ssh @(host) @(fix_cmd))
if f.returncode != 0:
print('FIX_FAILED')
print(str(f.out).strip())
else:
print(str(f.out).strip())