The idea
The homelab started as a media server and turned into a full private cloud. The governing rule is infrastructure-as-code: every host config, container definition, firewall note, and alert rule lives in a Git repository and is deployed by Ansible. If it can't be reproduced with a playbook run, it doesn't belong on a host. The lab runs real workloads — family media, budgeting, recipes, password management, household inventory, a personal CRM, local AI inference — under the same discipline I bring to production engineering.
Architecture
The network is split into three zones: an Oracle Cloud Always-Free VM acting as an internet-facing DMZ, a self-hosted WireGuard mesh (Headscale) connecting everything, and the home LAN behind OPNsense where the real hardware lives.
The AI ops layer
The piece I'm proudest of: alerts don't just page me — an AI agent investigates first. Prometheus rules route through Alertmanager to a webhook adapter on the agent host, which wakes an LLM-driven session with SSH access to the fleet and the runbook repository. The agent reproduces the symptom, forms a hypothesis, applies low-risk remediation (service restarts, cache cleanup, config re-deploy), and posts a write-up to Discord. Incidents leave behind markdown post-mortems committed to the same repo, so the next occurrence starts with documentation already written.
The fleet
| Host | Platform | Role |
|---|---|---|
| hexz690 | Proxmox VE (Debian) | Hypervisor — 9 LXCs, Arc Pro B70 passthrough |
| opnsense | OPNsense (FreeBSD) | Firewall, gateway, DNS |
| hexztruenas | TrueNAS (FreeBSD) | ZFS storage: 11TB data + 500GB backup pool, NFSv3 |
| hexzproxy | Oracle Cloud ARM VM | DMZ: Caddy, Headscale + DERP, CrowdSec, Uptime Kuma |
| hexzdocker | LXC · Docker | Media stack (VPN'd qBit + *arr suite) and monitoring stack (Prometheus, Grafana, Alertmanager) |
| hexzjellyfin | LXC | Jellyfin with Intel iGPU transcode, 4K HDR direct play |
| hexzai | LXC · GPU | llama.cpp SYCL serving a 27B Q4 model on Arc Pro B70, OpenAI-compatible API |
| hexzhermes | LXC | Hermes agent — Discord gateway, Alertmanager remediation adapter |
| hexzpbs | LXC | Proxmox Backup Server, datastore on TrueNAS NFS |
| hexzactual · hexztandoor · hexzmonica · hexzgrocy | LXCs | Family apps: budgeting, recipes, personal CRM, household inventory |
| hexzhome | Home Assistant OS | Smart home + local voice pipeline (STT/TTS), Hermes as conversation agent |
Notable decisions
-
DMZ in the cloud, not at home. An Oracle
Always-Free ARM VM takes all inbound traffic. The house has zero
open ports; Caddy terminates TLS and routes
*.jankowski.aito LAN services over the mesh. - Self-hosted WireGuard mesh (Headscale) replacing ZeroTier. Static mesh IPs (100.64.0.0/10) act as the canonical address plan, survives ISP renumbering, and DERP relays ride the proxy's 443. Migration included a documented outage RCA and a recovery runbook.
- Ansible vault + full IaC. Secrets are encrypted in-repo; playbooks cover baseline config, updates, media/monitoring deploys, hardening, and a fleet-wide health check. Check-mode dry runs and idempotency are part of the workflow.
- 3-2-1-ish backups. Proxmox Backup Server on its own LXC, datastore living on TrueNAS ZFS over NFS, snapshots scrubbed on schedule.
- Local AI on a budget GPU. An Intel Arc Pro B70 (32GB) serves a 27B Q4 model through llama.cpp SYCL with speculative decoding, behind an OpenAI-compatible API — the inference backend the ops agent runs on.
- Hardware reliability as software problems. A daily NIC-hang crash loop was root-caused to driver/offload interaction and retired with a layered kernel mitigation playbook (sysctl, offload disabling, watchdog, logrotate).
What it taught me
Running this lab is the closest thing to being an on-call SRE for a company of one: capacity planning at 2 a.m., blast-radius thinking before a deploy, and the discipline to write the post-mortem even when nobody else would read it. It also became a proving ground for agentic AI — the alert-remediation pipeline maps directly onto the agent-orchestration work I do professionally at Juno, with the same failure modes (hallucinated fixes, scope creep) needing the same guardrails (approval gates, verification loops, read-only tools by default).