GPU Scheduling Landscape and mutex Semantics
Why GPU scheduling matters: how the default scheduler misses node, card, and topology placement, how HAMi v2.10 fills the gaps, and what mutex really means on kind.
In-depth articles and insights on open source, AI, cloud-native, DevOps, and software engineering.
GPU Scheduling Landscape and mutex Semantics
Why GPU scheduling matters: how the default scheduler misses node, card, and topology placement, how HAMi v2.10 fills the gaps, and what mutex really means on kind.
GPU sharing is moving from soft allocation to governance—scheduling decisions and runtime isolation can finally be reconciled.
After HAMi's CNCF Incubating: From a Community of Code to a Network of Consensus
Code is cheap; consensus is the new scarce good.
Olares and HAMi: Desktop AI Workstation Inflection
HAMi moves from cluster to desktop with Olares.
Every Nation Begins with Textiles
From Anji bamboo weaving to industrialization
Why GPU Is the Foundation of AI
A GPU explainer for Kubernetes veterans new to AI. Maps token, model, training, inference, Transformer, Tensor Core, HBM, and KV cache to concepts you already know.
From GPU utilization to productive GPU-hours.
Agentic AI Infrastructure Reliability
A practical AI Infra review of Agentic AI reliability, covering a five-dimension framework, fault tolerance, recovery, observability, and hybrid architecture design.
From GPU to Token: The 8-Layer Observability Stack for AI Infrastructure
From GPU hardware, Kubernetes scheduling, inference engines to token cost — understanding the 8-layer observability architecture for modern AI infrastructure.
How I built a personal AI infrastructure using ChatGPT, OpenClaw, Obsidian, GitHub, Lark, GLM-5.1, and a Mac mini M4.