This morning I woke up at five, not because of a nightmare, but because of the fan noise from both my 16GB GPUs working so hard they were almost singing along with my homelab life. That hum told me the system was burning through its brainpower generating images and processing large models simultaneously. I walked out to check the dashboard showing the status of the machines, render nodes, and personal storage. The numbers made me feel both proud and wary at the same time.
Overview of the Night GPUs Worked Hard
When I looked at the usage graphs, the first GPU showed 100% utilization, while memory had been used up to 15,745 MiB out of a total 16,311 MiB, leaving very little breathing room. The second GPU was similar at 98% with 14,475 MiB of memory. These numbers aren't trivial—they mean the system was working almost to its physical limits. I saw jobs queued in ComfyUI running FLUX models alongside qwen_image_edit simultaneously, which we normally wouldn't do because we're afraid the system would crash. But tonight I decided to experiment and see what would happen if we pushed it to full capacity.
The fact that both GPUs were using nearly 16GB of memory was a silent warning sign. I've encountered OOM (Out of Memory) issues many times, almost enough to make me give up. The system still standing without crashing shows that the architecture we designed is becoming more stable. But anxiety remained every second those numbers blinked on the screen. It's like driving too fast on wet roads—you know it's dangerous but you want to test the limits of your tires and brakes.
Job Queue and Image Generation Pace
What made me wake up and check was the image generation rate of about 10 per hour, which is quite high for a homelab at this level. I looked at the job queue in n8n connected to ComfyUI and saw that job #9403 was running in the M3-M6 stage, including seed, prompt, and various gates. Previous jobs like #9402 and #9401 had all passed VERIFIED status, showing our pipeline was working continuously.
The fact that images kept being generated without interruption was commendable, but I felt worried about the word "simultaneously" in the job queue. Running FLUX and qwen_image_edit at the same time forced the GPU to switch contexts frequently, which could affect overall efficiency in the long run. I noticed that while the number of images produced was high, some pieces showed slight instability—a side effect of resources being shared so densely.
Lessons from GPU Law and Borrow/TTL Concepts
While watching the system work, I thought about the "GPU Law" principle I established for this homelab, which emphasizes dynamic resource management. Having a borrow and TTL (Time To Live) system for jobs is crucial because without this mechanism, some jobs might occupy the GPU too long, causing others to wait or fail.
Tonight I clearly saw the value of TTL. When one job finished, it immediately released resources back to the pool, allowing the next job to use them without waiting long. The borrow system helps us temporarily borrow resources from other parts when needed, which can increase throughput. But the risk is that if borrowing goes wrong, the system could enter a deadlock state. So I had to check logs closely to see if there was any abnormal waiting.
Failures and Lessons Learned from Trading on My Own
Beyond technical matters, tonight I also felt discouraged by my own trading results. The outcomes weren't as stable as hoped, leading me to decide to separate a small trading portfolio to reduce risk. But in terms of the homelab, I found that trying to control everything myself carries high risk. GPUs running at 100% for too long could cause heat buildup and affect long-term lifespan.
The important lesson is "don't push the system beyond what it can handle." Even if the numbers look impressive, machine health matters more than temporary throughput. I learned that good system design must be flexible enough to stop or slow down when necessary, not just always pushing for maximum speed. Small failures tonight revealed weaknesses in our monitoring process that still lacks clear enough alerting.
Next Steps and Feelings After a Heavy Night
After watching until morning, I decided to reduce ComfyUI concurrency by half, trading speed for stability and GPU lifespan. Accepting some speed loss in exchange for sustainability is something homelab engineers must learn. I'll add stricter thermal throttling and check the logs of qwen38-chat:latest on the Mac Mini M4 Pro, which was also working hard with 32768 token context.
This night taught me that running an AI lab at home isn't just about accumulating equipment—it's about understanding limitations and respecting them. I feel proud that the system still stood, but I'm aware there's much more to improve so this homelab can operate sustainably in the long term without having to wake up and watch every night like this.
!docker ps screen from personal storage machine, dated 2026-09-06, our actual system containers
!nvidia-smi screen from our render machine, dated 2026-09-06, all cards running nearly full

