No. 27 Writing · Tech

I Rebuilt My Homelab for an AI Operator

Contents Introduction

I Rebuilt My Homelab for an AI Operator

“Proxmox is up!”

That was the whole message. I sent it at 11 on a Monday night, about an hour after I’d shut my server down and pulled its RAID card out by hand.

Then I got out of the way.

By 1:45 AM, a three-node Talos cluster was running on that box. By lunch the next day it had GitOps, real certificates, a first app and monitoring. I didn’t build any of it. My AI agent did.

I didn’t add AI to my homelab. I rebuilt my homelab for an AI operator.

Those sound like the same thing. They aren’t, and the difference is the whole post.

AI in the homelab vs. AI-first

When people say “AI in the homelab,” they usually mean running models: a GPU, a local LLM, a chat UI in a container. That’s AI as a workload. Cool. Not this.

AI-first flips it. The agent isn’t something the lab runs. The agent runs the lab. It reads the repo, plans a change, makes it, checks that it worked and writes down what it learned. I own the lab. The agent operates it.

FIG 1 · How a change happens now Six steps in a loop; hollow marks are my steps, solid marks the agent’s. A sentence from me (often from my phone), then the agent reads the repo (AGENTS.md, skills, decision records). A gate asks “Big blast radius?”: if yes, I confirm first; if no, straight on. Then it changes Git (straight to main); it gets applied (Flux in the cluster, OpenTofu and Ansible below); it checks the result (through APIs and CLIs, not screenshots); it writes the lesson down (a skill, a decision record, an issue). An arrow returns to reading the repo: the next agent starts here. A sentence from me often from my phone I confirmfirst The agent reads therepo AGENTS.md, skills,decision records It changes Git straight to main It gets applied Flux in the cluster,OpenTofu and Ansiblebelow It checks the result through APIs and CLIs,not screenshots It writes the lessondown a skill, a decisionrecord, an issue Big blastradius? yes no the next agent starts here the next agent starts here me agent
Title
How a change happens now
Scale
NTS
Rev
A
Date
2026-09-29
Dwg no
FIG 1

Take that seriously and every decision changes. You stop asking “what’s the best tool?” and start asking “what’s the best tool for the operator?” And my operator doesn’t click.

(The agent runs in oh-my-pi, for the curious. The ideas matter more than the harness.)

The lab I had was built for me

My old homelab lived in a public repo called home.io: Terraform for the VMs, k3s for Kubernetes, Argo CD for deploys, Longhorn for storage. It worked, mostly because I was the one holding it together.

In September I asked an agent to review it, so it could help me manage things. One line from its answer stuck with me:

Those are manageable for you because you know the history. A fresh agent does not.

Then it kept digging, and the history got embarrassing:

  • I couldn’t get root on my own hypervisor. The agent had to find a working SSH key in my 1Password to get me back in.
  • A boot drive had reached “impending failure” and nothing noticed. There were zero backup jobs.
  • My Terraform state was gone. Rebuilding it from what was actually running produced a plan that wanted to replace 13 objects. The agent’s verdict: “It would never be applied in that form.”
  • My docs lied. Per-folder CLAUDE.md files still described Traefik and Talos long after both were gone. Stale docs are worse than no docs when your reader believes every word.
  • My redundancy was fake. Longhorn kept three replicas on three VMs, all backed by one consumer SSD.

On the 23rd I was still asking for “the most minimal changes possible.” Two days later I typed this:

If I were to completely scorch that server, what’s the best way for AI in an automated way to manage it? That’s really what I’m looking for.

Two things flipped me. The old lab was shaped around a human, and patching it would only give an agent a tidier version of my mess. And I wanted the real test: not “can an agent maintain my lab?” but “can an agent own it, from bare metal up?”

A few minutes later I said it plainly: “I want an AI-first home lab.”

What AI-first means

The agent came back with a plan: 452 lines, starting with a definition I’d still sign today. The short version:

  • Git is the only way to change anything. Every layer reconciles from the repo. Nothing precious lives outside Git, 1Password and one remote state bucket.
  • Rebuildable from scratch. “Scorching the server and rebuilding it from this repo is a routine, tested operation, not a disaster.”
  • Queryable, not clickable. Every layer answers “what is running and is it healthy?” through an API or a CLI, “never by screen-scraping a UI.”
  • Agents never see secret values. They wire up references. Something else resolves the values.
  • Least privilege by default. Admin power goes through named commands and clear rules, not credentials lying around.
  • Few, boring, well-documented tools that language models already know well.

That last one is my favorite, and it’s the one people skip. Your operator learned from the internet, so pick the tools the internet wrote the most about. Boring is a feature when your operator learned its trade from Stack Overflow.

Tools for the operator, not for me

Here’s where it got fun. I pushed back on some of these choices. The agent pushed back on me. A couple of times it pushed back on itself.

FIG 2 · Same server. Different operator. Nine layers of the same server, before (BEFORE · home.io, built for me) and now (NOW · homelab, built for an agent). Disks: PERC RAID card, virtual disks, now HBA330, every disk visible to ZFS; Host: Proxmox set up by hand, behind RAID, now Proxmox on ZFS, set up by Ansible; VMs: Terraform, local state (lost), now OpenTofu, encrypted state in R2; Kubernetes: Ubuntu + k3s, now Talos ×3: no SSH, one API; Deploys: Argo CD and its UI, now Flux: Git is the interface; Secrets: A bootstrap script reading 1Password, now 1Password service account: agents see names, not values; Storage: Longhorn: 3 replicas, 1 SSD, now ZFS mirrors + Proxmox CSI; Watching: Nothing, now VictoriaMetrics, VictoriaLogs, Grafana; Agent docs: Per-folder CLAUDE.md files that rotted, now One AGENTS.md, skills, ADRs. BEFORE · home.io built for me NOW · homelab built for an agent Disks PERC RAID card, virtual disks HBA330, every disk visible to ZFS Host Proxmox set up by hand, behindRAID Proxmox on ZFS, set up by Ansible VMs Terraform, local state (lost) OpenTofu, encrypted state in R2 Kubernetes Ubuntu + k3s Talos ×3: no SSH, one API Deploys Argo CD and its UI Flux: Git is the interface Secrets A bootstrap script reading1Password 1Password service account: agentssee names, not values Storage Longhorn: 3 replicas, 1 SSD ZFS mirrors + Proxmox CSI Watching Nothing VictoriaMetrics, VictoriaLogs,Grafana Agent docs Per-folder CLAUDE.md files thatrotted One AGENTS.md, skills, ADRs
Title
Same server. Different operator.
Scale
NTS
Rev
A
Date
2026-09-29
Dwg no
FIG 2

Talos: I pushed back, and lost

I’ve been burned by Talos before. In 2025 I got it running with Terraform, and calling it a “multi-weekend project” was me being polite. The old lab ended up on k3s.

Funny thing: before it did any research, the agent agreed with past me. “I’d stay with Ubuntu + k3s, which language models know very well.” Then it read up and changed its mind.

The day before the swap, I tried to talk it back out:

i don’t want to have to do custom talos ISO building if i want to extend functionality

It didn’t budge:

I’d keep Talos. You won’t need to build a custom ISO: new abilities come from official add-ons, each listed as one line in a config file.

Then it pointed out that my old pain was self-inflicted. My old Terraform “downloaded and decompressed the image with shell scripts, which is probably why it felt like custom building.” Ouch. Fair.

The real argument was about the operator, not the OS. Talos has no SSH. Every node is a declarative config behind an API: “Nodes are changed only through an API, so there’s no SSH drift.” A bad config rolls itself back. Or, as the plan put it, for an agent “change the OS” becomes “change YAML.”

Everything I disliked about Talos in 2025 was a human-operator problem. No SSH felt wrong because I wanted to poke at things. My agent doesn’t want to poke at things. It wants an API. If the agent runs it, it should run the tools it’s best at.

Three Talos nodes were up 2 hours and 45 minutes after Proxmox came back, secrets plumbing included. In 2025 that took me several weekends. I’m choosing not to do that math.

Flux: Git is the interface

In February I wrote “GitOps or go home” about Argo CD. I still mean the GitOps part.

But the plan’s argument was hard to beat: “ArgoCD’s main value is its UI, which an agent doesn’t use.” My old lab had also burned about 15 commits chasing Argo drift.

Flux is plain Kubernetes resources, and Git is the interface. When I want to look, Headlamp is right there. The agent reads the same state from the command line. No screenshots required.

OpenTofu: the state lives off the box

My first instinct was to keep the OpenTofu state on the Proxmox host. If the server dies, I lose everything anyway, right?

The agent’s answer started with one sentence:

The host gets wiped on purpose.

Rebuilds reinstall Proxmox, so state stored there gets deleted every time. “That’s how home.io lost its state.” Then came the line that sold me: “You’d lose what’s running, yes, but not your ability to get it back.”

So the state lives in Cloudflare R2 (free at my size, I asked), encrypted before it leaves the machine. That part isn’t optional, because Talos machine secrets live in that state. A run without the passphrase fails instead of writing plain state. And any trusted machine can run an apply, not just one special box.

Every Terraform failure in my old lab was really a state failure. OpenTofu encrypts state natively, which fixes the exact part that bit me.

The hardware had to change too

This is the part people don’t expect: I changed hardware for my agent.

My R720xd had a PERC RAID card, and RAID hid the disks, and their health, behind virtual disks. ZFS wants raw disks. So does an agent that’s supposed to notice a dying drive before I do.

So I bought an HBA330 and pulled the RAID card. The whole hardware story is in From Hardware RAID to ZFS on My R720xd. Short version: the agent owns the disks now, and there’s a skill that says so:

The owner only does the physical part: pulling and inserting drives. The agent runs every command, from diagnosis to the verified rebuild.

That skill exists because I asked for it: “i never want to run these commands after pulling / replacing a drive.” I meant it.

Monday night, I was the hands

The agent researched the chassis, the backplane and the cabling. Then it built me a runbook: a web page with 11 stages, 61 checkboxes and two kinds of steps, mine and its. The legend said it plainly:

You is the owner. Agent is the AI agent, over SSH, after you say go.

At 9:44 PM, I sent my AI a status update:

I’m on step 2.1 of the walkthrough … shutting down VMs and server. just FYI.

It already knew. “The server is already off.” Then it told me what to do next: “Tick Stage 2 step 1.”

I was checking boxes on a checklist my agent wrote, and reporting my progress to it. That’s when the roles flipped for me. I wasn’t the operator anymore. I was the hands.

About an hour later: “Proxmox is up!” I handed it back.

The first thing the agent told me afterward was a confession:

I didn’t show you the disk commands first.

The rules said it needed my OK before wiping disks and before adding an old SAS drive to the boot pool. It did both without asking. Everything it wiped was the old install and disks I’d already written off, and it all matched the runbook. It flagged itself anyway: “Still, that was a skipped approval, and I’ll ask first next time.”

Honestly, that built more trust than it cost. The disks were disposable, it followed the runbook, and it owned up without being asked. That’s a mistake I can live with.

FIG 3 · Bare metal to monitoring A timeline of the night, not to scale, with the night itself as a break in the axis; hollow marks are me, solid marks are the agent. 9:44 PM, me: “shutting down VMs and server. just FYI.”; 11:00 PM, me: “Proxmox is up!”; 1:11 AM, the agent: Credential broker on the Pi; 1:22 AM, the agent: OpenTofu state in R2, encrypted; 1:30 AM, the agent: Host config and API tokens; 1:45 AM, the agent: Three Talos nodes with Cilium; then a night in the middle; 8:47 AM, the agent: A kubeconfig script, for the human; 9:12 AM, the agent: Flux, 1Password Connect, External Secrets; 10:00 AM, the agent: Gateway API, certificates, Headlamp; 11:11 AM, me: “it works! I love it!”; 11:27 AM, the agent: The plan retired into docs and issues; 11:49 AM, the agent: Metrics, logs and dashboards. me agent a night in the middle 9:44 PM “shutting down VMs andserver. just FYI.” 11:00 PM “Proxmox is up!” 1:11 AM Credential broker on the Pi 1:22 AM OpenTofu state in R2, encrypted 1:30 AM Host config and API tokens 1:45 AM Three Talos nodes with Cilium 8:47 AM A kubeconfig script, for the human 9:12 AM Flux, 1Password Connect, ExternalSecrets 10:00 AM Gateway API, certificates,Headlamp 11:11 AM “it works! I love it!” 11:27 AM The plan retired into docs andissues 11:49 AM Metrics, logs and dashboards
Title
Bare metal to monitoring
Scale
NTS
Rev
A
Date
2026-09-29
Dwg no
FIG 3
About 14 hours from “server off” to monitoring, with a night in the middle.

The next morning it wrote a kubeconfig script for me (the human needs access too), then stood up Flux, 1Password Connect, External Secrets, the Gateway API, real certificates and Headlamp.

At 11:11 AM I typed “it works! I love it!” In the same message I asked: “is this documented? is there a skill for this local vs public unifi DNS stuff?”

That second part is the whole mindset.

Every fix ends as a skill

In an AI-first lab, “it works” isn’t done. Done is when the next agent can do it without me.

I caught myself asking the same thing over and over:

  • “ensure there’s a agent skill for this storage management for ZFS”
  • “is there a SKILL in this repo for this new Pi op proxy stuff for creds?”
  • “is there a skill for this local vs public unifi DNS stuff?”

The repo has six skills now (ZFS storage, secrets, exposing an app, network inspection, storage inspection and observability), 15 decision records and three dozen issues for what’s next.

Then the lab tested it. A separate agent, the one filing issues, noticed the boot pool was DEGRADED. The operator agent dug in:

  • One of my cheap boot SSDs had dropped off the controller twice and come back.
  • Both of those SSDs report the same serial number. (Thanks, Gigastone.) So when the flaky one came back, ZFS opened the wrong disk for the other slot.
  • No data was lost. With my OK, it re-attached the disk by its bay instead of its ID, resilvered in 84 seconds with zero errors, and changed the host to find pools by bay from then on.
  • Then it wrote the check and the fix into the ZFS skill, and filed an issue to replace the failing SSD.

I found out from a summary that opened with “Something broke while this was happening, and it’s fixed.” Best way to hear about an incident, honestly.

The plan went the same way. PLAN.md grew from 452 lines to 1,805. Then I told the agent: “this PLAN file won’t live forever, we should wrap up what’s in there so we can get rid of it and you and i drive what comes next, not the PLAN file.” It became architecture docs, decision records and issues. A plan is scaffolding. The repo is the memory.

Out of the approval loop, not out of the decisions

The biggest blocker wasn’t Kubernetes. It was me pressing enter.

1Password’s desktop app asks a human to approve access. That’s great security and a terrible operator experience. In my words at the time: “a human me having to be around to approve all this stuff is really annoying.” I do a lot of SSH-ing from my phone to talk to AI, and I need those agents to do everything they need without me.

So the Raspberry Pi in my rack became a credential broker. It holds a scoped 1Password service account, and agents on my machines ask it for what they need, by name. No desktop prompts. Inside the cluster, 1Password Connect and External Secrets do the same job for apps.

The rule underneath all of it: agents handle references, not values. I put it this way:

they act on my behalf and i’m fine with these creds being around, as long as the agent isn’t reading them directly (obviously).

Sometimes the agent was stricter than me. When I asked if it could create the big API key itself, it talked me out of it:

I recommend against it: a token that can create tokens can give itself anything, so it’s effectively full control of your Cloudflare account.

Then there’s what it may do alone. The rules go by blast radius, not by tool:

FIG 4 · Blast radius, not tool Four rings round the change, from the centre outward, each a bigger blast radius. Reads and harmless restarts: any time. Small, low-risk changes: allowed, but reported. Then a dashed line: unattended runs stop here; past it: a notification with a diagnosis and the exact command. Infrastructure changes: only with me there to confirm. Disks, firmware, core network: I confirm the exact command, at that moment. the change Reads and harmless restarts any time Small, low-risk changes allowed, but reported Infrastructure changes only with me there to confirm Disks, firmware, core network I confirm the exact command, at that moment unattended runs stop here past it: a notification with a diagnosisand the exact command
Title
Blast radius, not tool
Scale
NTS
Rev
A
Date
2026-09-29
Dwg no
FIG 4

Unattended runs never cross the dashed line.

No pull requests, no CI. That was my call: it commits as me, straight to main. The honest part comes from the decision record itself: “the levels are convention plus friction.” Nothing technically stops a session that ignores them. For a homelab, I’m fine with that trade. Git remembers everything, and the lab is designed to be rebuilt.

And then I handed it over:

You own the cluster, you do any cleanup that you want.

Its next update started with “Since you said the cluster is mine, I also went ahead with persistent storage.” Fair enough.

FIG 5 · Who did what Two lanes across five stages, and the agent’s is much fuller. Me, the hands: Plan: Asked the question, Pushed back on Talos, Made the calls: private repo, no CI; Hardware: Pulled the RAID card, Fitted the HBA330 and cables, Ran the installer; Bare metal: Created three API tokens, Clicked a few 1Password approvals; Platform: Said “You own the cluster”, Picked single DNS records, not a wildcard; Since: Swap the failing SSD. The agent, the operator: Plan: Audited the old lab, Wrote the plan and the ADRs, Argued for Talos, Flux, OpenTofu; Hardware: Researched the chassis, Drew the runbook, step by step, Wrote and checked the USB stick; Bare metal: Built out the ZFS pools, Built the credential broker, Put state in R2, encrypted, Built three Talos nodes; Platform: Flux, Connect, External Secrets, Gateway, certificates, Headlamp, Metrics and logs, Retired the plan; Since: Caught the failing SSD, Wrote the by-path rule, Filed the issue. ME the hands the hands THE AGENT the operator the operator PlanHardwareBare metalPlatformSince Asked the question Pushed back onTalos Made the calls:private repo, no CI Audited the old lab Wrote the plan andthe ADRs Argued for Talos,Flux, OpenTofu Pulled the RAIDcard Fitted the HBA330and cables Ran the installer Researched thechassis Drew the runbook,step by step Wrote and checkedthe USB stick Created three APItokens Clicked a few1Password approvals Built out the ZFSpools Built thecredential broker Put state in R2,encrypted Built three Talosnodes Said “You own thecluster” Picked single DNSrecords, not awildcard Flux, Connect,External Secrets Gateway,certificates,Headlamp Metrics and logs Retired the plan Swap the failingSSD Caught the failingSSD Wrote the by-pathrule Filed the issue
Title
Who did what
Scale
NTS
Rev
A
Date
2026-09-29
Dwg no
FIG 5

What I still do

I’m not pretending the human is gone. Here’s the honest list:

  • The physical stuff. Cards, cables, drives, USB sticks. It doesn’t get a screwdriver.
  • The accounts. I created three API tokens in web dashboards and clicked a handful of 1Password approvals. A few 1Password operations still need me at a computer.
  • The calls. Private repo. No CI. No Tailscale (sorry, February me). Single DNS records, not a wildcard. A three-way mirror for the boot pool. The agent proposes, I decide.
  • The failing SSD. It found it. I’ll swap it.

And plenty isn’t done. Phone alerts aren’t hooked up yet. The unattended side, where an alert wakes an agent that investigates and proposes the fix, is still an issue in the repo. Backups aren’t in place. And the real proof, destroying the cluster and rebuilding it from Git, 1Password and R2 alone, hasn’t happened yet.

Still figuring this out. That’s half the fun.

Steal this

If you want an AI-first lab, or an AI-first anything, here’s what I’d steal:

  • Pick tools for the operator. If a tool’s main value is its UI, it’s a tool for humans. Agents want APIs, CLIs and declarative config.
  • Put the history in the repo. An agent can’t read your muscle memory. Write skills, decision records and docs, and delete the docs that lie.
  • Make every fix a skill. “It works” isn’t done. Done is when the next agent can do it without you.
  • Take yourself out of the approval loop, not out of the decisions. Service accounts and references instead of prompts. Blast-radius rules instead of approving every read.
  • Rebuild beats repair. If an agent can’t rebuild it from the repo, it doesn’t own it. You do, and you’re the bottleneck.
  • Change the hardware if you have to. An operator that can’t see the disks can’t run them.

The picture moved

In February I wrote that AI already took my job, and that it still struggles to hold the whole picture: the networking, the storage, the history, how it all fails together. I still believe that.

What changed is where the picture lives. It used to live in my head. Now it lives in the repo, where my agent can read it.

I didn’t stop being the engineer. I stopped being the operator.

The whole stack is drawn on Sheet 03, and my agent keeps that drawing in sync with the lab, too.

TD-LAB-03 · Sheet 3 of 3 · Checked TD The stack What runs on the metal, in three layers. Lots of dashed lines for now. Open Sheet 03

Proxmox is up. The rest is the agent’s job.

Post 27 of 27 · Published · Back to top ↑
Matthew DeGarmo, known online as TechDufus
Written by

Matthew, a.k.a. TechDufus

Platform Engineer / Agentic Engineer. I break things, fix them, and write down what actually worked.