Homelab Infrastructure

A personal data center run like a production environment: everything defined as code, everything monitored, and an AI agent on call.

The idea

The homelab started as a media server and turned into a full private cloud. The governing rule is infrastructure-as-code: every host config, container definition, firewall note, and alert rule lives in a Git repository and is deployed by Ansible. If it can't be reproduced with a playbook run, it doesn't belong on a host. The lab runs real workloads — family media, budgeting, recipes, password management, household inventory, a personal CRM, local AI inference — under the same discipline I bring to production engineering.

Architecture

The network is split into three zones: an Oracle Cloud Always-Free VM acting as an internet-facing DMZ, a self-hosted WireGuard mesh (Headscale) connecting everything, and the home LAN behind OPNsense where the real hardware lives.

INTERNET ORACLE CLOUD — DMZ (ARM Always-Free) hexzproxy Caddy (TLS, *.jankowski.ai) CrowdSec + fail2ban + UFW Uptime Kuma (external probes) Headscale control plane DERP relay + STUN vpn.jankowski.ai 100.64.0.0/10 tailnet WireGuard mesh (no port forwarding) HOME LAN — 10.213.45.0/24 (behind OPNsense) opnsense firewall / gateway DNS · VLAN hexztruenas ZFS · 11TB data 500GB backup pool NFSv3 exports scrub + snapshot plans hexz690 — Proxmox VE hypervisor hexzdocker media + monitoring LXCs hexzjellyfin 4K HDR · iGPU transcode hexzai Arc Pro B70 · LLM hexzhermes AI ops agent hexzpbs backup server + 4 more LXCs budget·recipes·CRM·grocery 9 LXCs · kernels patched via Ansible · PBS snapshots → TrueNAS NFS hexzhome Home Assistant + voice pipeline NFS mounts · management VLAN · wireguard peers into mesh
Figure 1 — Network topology. Public traffic terminates at the Oracle DMZ; everything reaches the LAN through the Headscale mesh, so no services are port-forwarded at home.

The AI ops layer

The piece I'm proudest of: alerts don't just page me — an AI agent investigates first. Prometheus rules route through Alertmanager to a webhook adapter on the agent host, which wakes an LLM-driven session with SSH access to the fleet and the runbook repository. The agent reproduces the symptom, forms a hypothesis, applies low-risk remediation (service restarts, cache cleanup, config re-deploy), and posts a write-up to Discord. Incidents leave behind markdown post-mortems committed to the same repo, so the next occurrence starts with documentation already written.

Metric breach node/cadvisor/blackbox Alertmanager routing · dedupe · silence Hermes agent LLM + SSH + runbooks Remediate restart · redeploy Discord write-up post-mortem docs → back into the loop
Figure 2 — The alert-to-resolution loop: humans see the summary, not the pager noise.

The fleet

Host Platform Role
hexz690 Proxmox VE (Debian) Hypervisor — 9 LXCs, Arc Pro B70 passthrough
opnsense OPNsense (FreeBSD) Firewall, gateway, DNS
hexztruenas TrueNAS (FreeBSD) ZFS storage: 11TB data + 500GB backup pool, NFSv3
hexzproxy Oracle Cloud ARM VM DMZ: Caddy, Headscale + DERP, CrowdSec, Uptime Kuma
hexzdocker LXC · Docker Media stack (VPN'd qBit + *arr suite) and monitoring stack (Prometheus, Grafana, Alertmanager)
hexzjellyfin LXC Jellyfin with Intel iGPU transcode, 4K HDR direct play
hexzai LXC · GPU llama.cpp SYCL serving a 27B Q4 model on Arc Pro B70, OpenAI-compatible API
hexzhermes LXC Hermes agent — Discord gateway, Alertmanager remediation adapter
hexzpbs LXC Proxmox Backup Server, datastore on TrueNAS NFS
hexzactual · hexztandoor · hexzmonica · hexzgrocy LXCs Family apps: budgeting, recipes, personal CRM, household inventory
hexzhome Home Assistant OS Smart home + local voice pipeline (STT/TTS), Hermes as conversation agent

Notable decisions

  • DMZ in the cloud, not at home. An Oracle Always-Free ARM VM takes all inbound traffic. The house has zero open ports; Caddy terminates TLS and routes *.jankowski.ai to LAN services over the mesh.
  • Self-hosted WireGuard mesh (Headscale) replacing ZeroTier. Static mesh IPs (100.64.0.0/10) act as the canonical address plan, survives ISP renumbering, and DERP relays ride the proxy's 443. Migration included a documented outage RCA and a recovery runbook.
  • Ansible vault + full IaC. Secrets are encrypted in-repo; playbooks cover baseline config, updates, media/monitoring deploys, hardening, and a fleet-wide health check. Check-mode dry runs and idempotency are part of the workflow.
  • 3-2-1-ish backups. Proxmox Backup Server on its own LXC, datastore living on TrueNAS ZFS over NFS, snapshots scrubbed on schedule.
  • Local AI on a budget GPU. An Intel Arc Pro B70 (32GB) serves a 27B Q4 model through llama.cpp SYCL with speculative decoding, behind an OpenAI-compatible API — the inference backend the ops agent runs on.
  • Hardware reliability as software problems. A daily NIC-hang crash loop was root-caused to driver/offload interaction and retired with a layered kernel mitigation playbook (sysctl, offload disabling, watchdog, logrotate).

What it taught me

Running this lab is the closest thing to being an on-call SRE for a company of one: capacity planning at 2 a.m., blast-radius thinking before a deploy, and the discipline to write the post-mortem even when nobody else would read it. It also became a proving ground for agentic AI — the alert-remediation pipeline maps directly onto the agent-orchestration work I do professionally at Juno, with the same failure modes (hallucinated fixes, scope creep) needing the same guardrails (approval gates, verification loops, read-only tools by default).