GPU sharing is moving from soft allocation to governance—scheduling decisions and runtime isolation can finally be reconciled.
AI Engineering
LLM, AI Native Infra and Agentic AI.
Olares and HAMi: Desktop AI Workstation Inflection
HAMi moves from cluster to desktop with Olares.
Why GPU Is the Foundation of AI
A GPU explainer for Kubernetes veterans new to AI. Maps token, model, training, inference, Transformer, Tensor Core, HBM, and KV cache to concepts you already know.
From GPU utilization to productive GPU-hours.
Agentic AI Infrastructure Reliability
A practical AI Infra review of Agentic AI reliability, covering a five-dimension framework, fault tolerance, recovery, observability, and hybrid architecture design.
From GPU to Token: The 8-Layer Observability Stack for AI Infrastructure
From GPU hardware, Kubernetes scheduling, inference engines to token cost — understanding the 8-layer observability architecture for modern AI infrastructure.
How I built a personal AI infrastructure using ChatGPT, OpenClaw, Obsidian, GitHub, Lark, GLM-5.1, and a Mac mini M4.
AI Infra Industry Trends: From Compute Bottlenecks to Ecosystem Evolution
A practitioner’s perspective on AI infrastructure trends: evolving bottlenecks, roles of CPU/GPU/scheduling, ecosystem shifts, and compute demand across training, inference, and Agent workloads.
Kubernetes as the GPU Control Plane for AI
Observations on the evolution of AI infrastructure control planes, focusing on HAMi v2.9, GPU scheduling, and Kubernetes resource models.
When GPUs Move Toward Open Scheduling: Structural Shifts in AI Native Infrastructure
A CTO/VP view on open GPU scheduling: CDI, Kubernetes DRA, virtualization data planes, ecosystem governance, and lock-in risk.