Files
plex-playlist/docs/GITEA_ACTIONS_TROUBLESHOOTING.md
Xlorep DarkHelm f3698b095b
All checks were successful
CICD / Build and Publish CICD Base Image (push) Successful in 6m8s
CICD / Build and Push CICD Image (push) Successful in 23m14s
CICD / Build CICD Image Failure Postmortem (push) Has been skipped
CICD / Backend Tests (push) Successful in 7m10s
CICD / Frontend Tests (push) Successful in 45s
CICD / Backend Doctests (push) Successful in 18s
CICD / Pre-commit Checks (push) Successful in 14m57s
CICD / Source Lanes Failure Postmortem (push) Has been skipped
CICD / CICD Tests Complete (push) Successful in 3s
CICD / Build Backend Base Image (push) Successful in 18s
CICD / Build Integration Tester Image (push) Successful in 1m5s
CICD / Build Backend Main Image (push) Successful in 1m52s
CICD / Build Frontend Base Image (push) Successful in 10m42s
CICD / Build Frontend Main Image (push) Successful in 33s
CICD / Build E2E Tester Image (push) Successful in 32m17s
CICD / Production Images Complete (push) Successful in 5s
CICD / Production Image Failures Postmortem (push) Has been skipped
CICD / Runtime Black-Box Integration Tests (push) Successful in 1m13s
CICD / Integration Tests Failure Postmortem (push) Has been skipped
CICD / End-to-End Tests (push) Successful in 11m23s
CICD / E2E Tests Failure Postmortem (push) Has been skipped
Stabilize self-hosted CI workflows and resolve issue #62 (#73)
## Summary

Hardens CI workflows for self-hosted Gitea runners by stabilizing E2E execution and Renovate behavior across internal/external network paths.

Closes #62

## What Changed

### E2E workflow reliability
- Fixed E2E workspace handoff to ensure expected repository contents are present during test execution.
- Added stricter preflight checks for required frontend files before running E2E.
- Reduced mount/path fragility while preserving runtime image pull and compose flow.

### Renovate workflow hardening
- Added internal-first endpoint reachability selection with fallback handling.
- Added token preflight checks for repository access.
- Added explicit host-rule auth handling for API/git paths.
- Added container-level connectivity preflight diagnostics.
- Added git URL override aligned with selected endpoint context.
- Removed incorrect forced Dogar host-IP pinning that broke HTTPS clone routing.

## Why

CI behavior was sensitive to runner networking and Renovate clone/auth interactions. These changes make the workflow deterministic in our runner topology and address recurring CI failures.

## Scope

- Workflow logic only (`cicd.yaml`, `renovate.yml`)
- No app feature or API behavior changes

## Validation

- Workflow YAML validation passed during updates.
- Changes were applied and verified iteratively from real failing run diagnostics.

Co-authored-by: copilotcoder <copilotcoder@darkhelm.org>
Reviewed-on: #73
2026-07-13 11:16:16 -04:00

220 lines
6.4 KiB
Markdown

# Gitea Actions Troubleshooting Guide
This document contains solutions to common issues with Gitea Actions CI/CD pipeline.
## Critical Issue: Jobs Stuck in "Waiting" State Forever
### Symptoms
- Workflows are created but jobs show "Waiting" indefinitely
- Runners are online and healthy
- No tasks appear in `action_task` database table
- Jobs get cancelled immediately (0-second duration)
- UI shows "Waiting" but database shows status 5 (cancelled)
### Root Cause
**Docker syntax in `runs-on` labels** causes Gitea Actions to immediately cancel jobs.
### Problem Syntax (BROKEN)
```yaml
jobs:
setup:
runs-on: ubuntu-latest:docker://ubuntu:22.04
backend:
runs-on: python-latest:docker://python:3.14-slim
frontend:
runs-on: node-latest:docker://node:20-bookworm-slim
```
### Solution Syntax (WORKING)
```yaml
jobs:
setup:
runs-on: ubuntu-latest
backend:
runs-on: python-latest
frontend:
runs-on: node-latest
```
### Why This Works
The runners are configured with Docker images in their labels:
```bash
GITEA_RUNNER_LABELS=ubuntu-latest:docker://ubuntu:22.04,node-latest:docker://node:20-bookworm-slim,python-latest:docker://python:3.14-slim
```
So jobs still run in the correct Docker containers, but Gitea can properly parse and dispatch them.
### Diagnosis Steps
1. **Check if new runs are created:**
```sql
SELECT id, status, title FROM action_run ORDER BY id DESC LIMIT 3;
```
2. **Check job status and duration:**
```sql
SELECT arj.id, arj.job_id, arj.status, ar.created, ar.updated, (ar.updated - ar.created) as duration_seconds
FROM action_run_job arj
JOIN action_run ar ON arj.run_id = ar.id
WHERE ar.id = (SELECT MAX(id) FROM action_run);
```
3. **Check if tasks are created:**
```sql
SELECT * FROM action_task ORDER BY id DESC LIMIT 5;
```
4. **Verify runners are online:**
```sql
SELECT id, name, last_online, agent_labels FROM action_runner WHERE last_online > (EXTRACT(epoch FROM NOW()) - 300)::bigint;
```
### Key Indicators
- **Duration = 0 seconds** → Immediate cancellation due to syntax issue
- **Empty action_task table** → Jobs never converted to executable tasks
- **Status 5 jobs with Status 7 dependents** → Setup job cancelled, others skipped
### Test Procedure
Create a minimal test workflow to isolate issues:
```yaml
# .gitea/workflows/test-simple.yml
name: Simple Test
on: push
jobs:
test:
name: Simple Test
runs-on: ubuntu-latest
steps:
- name: Echo
run: echo "Hello World"
```
If this works but your main workflow doesn't, the issue is likely syntax-related.
## Other Common Issues
### Cache/UI Synchronization Problems
If UI shows different status than database:
1. Restart Gitea: `docker compose restart server`
2. Clear browser cache
3. Check database vs UI status discrepancies
### Stuck Runs from Previous Sessions
Clean up stuck runs:
```sql
-- Clear stuck pending jobs
UPDATE action_run_job SET status = 5 WHERE status IN (1, 2);
UPDATE action_run SET status = 5 WHERE status IN (1, 2);
```
### Runner Registration Issues
If runners show "unregistered runner" errors:
1. Delete runner registrations: `DELETE FROM action_runner;`
2. Restart all runner containers
3. Let them auto-register with fresh state
## Infrastructure Overview
### Current Setup
- **Gitea Server**: Docker container with PostgreSQL backend
- **Runners**: 8 Raspberry Pi runners across 4 servers
- pi-desktop: Pi 400 4GB (2 runners)
- kankali: Pi with local Gitea (2 runners)
- urtzul: Pi 4B 8GB (2 runners)
- zhokq: Pi 4B 8GB (2 runners)
### Runner Configuration
Each runner supports multiple Docker environments:
- `ubuntu-latest``ubuntu:22.04`
- `python-latest``python:3.14-slim`
- `node-latest``node:20-bookworm-slim`
- `ubuntu-act``catthehacker/ubuntu:act-latest`
### Mirroring the `ubuntu-act` Runner Image
If GHCR pulls are flaky, mirror the runner image into your local registry and point the label at that mirror instead of the upstream tag.
Example mirror flow:
```bash
docker pull ghcr.io/catthehacker/ubuntu:act-latest
docker tag ghcr.io/catthehacker/ubuntu:act-latest kankali.darkhelm.lan:3001/darkhelm.org/act-ubuntu:act-latest
docker push kankali.darkhelm.lan:3001/darkhelm.org/act-ubuntu:act-latest
```
Recommended runner label once mirrored:
```bash
GITEA_RUNNER_LABELS=ubuntu-latest:docker://ubuntu:22.04,node-latest:docker://node:20-bookworm-slim,python-latest:docker://python:3.14-slim,ubuntu-act:docker://kankali.darkhelm.lan:3001/darkhelm.org/act-ubuntu:act-latest
```
If you want a failover strategy, keep the cached image tagged in the local registry and only refresh it when the upstream digest changes. That way the runner never depends on GHCR at job start.
Automation scripts for this workflow live in `scripts/gitea-actions/`:
- `scripts/gitea-actions/repair_runner_mirror.xsh`
Repairs mirrors by pulling upstream images, tagging/pushing to local registry, and verifying pullability on each runner host.
Current mirror targets:
- `kankali.darkhelm.lan:3001/darkhelm.org/act-ubuntu:act-latest`
- `kankali.darkhelm.lan:3001/darkhelm.org/renovate:41`
- `kankali.darkhelm.lan:3001/darkhelm.org/playwright-browsers:v1.56.1-jammy`
- `scripts/gitea-actions/check_runner_images.xsh`
Verifies host-by-host image presence and pullability for `ubuntu:22.04`, upstream act image (`GHCR`), mirrored act image (`MIRROR`), and Renovate image.
- `scripts/gitea-actions/runner-prechange-capture.xsh`
Captures restart counts, OOM flags, memory state, and runner logs across all hosts.
- `scripts/gitea-actions/diagnose_runner_startup.xsh`
Focused startup diagnostics for runner containers and mirror pull behavior.
Run from repository root (xonsh):
```sh
source scripts/gitea-actions/repair_runner_mirror.xsh
source scripts/gitea-actions/check_runner_images.xsh
```
### Workflow Design
Multi-stage pipeline with artifact passing:
1. **Setup**: Checkout code, create artifacts
2. **Parallel Setup**: Backend (Python/uv) + Frontend (Node.js/Yarn)
3. **Parallel Tests**: Backend tests + Frontend tests
## Lessons Learned
1. **Gitea Actions syntax is stricter than GitHub Actions**
2. **Runner labels must match exactly** - no Docker syntax in workflow files
3. **Database debugging is essential** - UI can show cached/incorrect status
4. **Job cancellation happens immediately** for syntax errors
5. **Empty action_task table** is the key indicator of dispatch failure
---
_Last updated: June 2, 2026_
_Issue resolved after extensive database-level debugging and syntax isolation_