<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Jimmy Song – Jimmy Song's Blog</title><link>https://jimmysong.io/</link><description>Recent content in Jimmy Song's Blog on Jimmy Song</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>Jimmy Song</managingEditor><webMaster>Jimmy Song</webMaster><follow_challenge><feedId>51621818828612637</feedId><userId>59800919738273792</userId></follow_challenge><lastBuildDate>Wed, 15 Jul 2026 09:17:07 +0800</lastBuildDate><atom:link href="https://jimmysong.io/index.xml" rel="self" type="application/rss+xml"/><item><title>Core Model Overview</title><link>https://jimmysong.io/book/ai-infra-dao/model-overview/</link><pubDate>Tue, 10 Feb 2026 13:56:12 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/model-overview/</guid><description>A four-layer model (Yin-Yang, Five Elements, Yun, Qi) for understanding AI infrastructure as an evolving organic system</description><content:encoded>
&lt;p&gt;The &lt;strong&gt;Yin-Yang - Five Elements - Yun - Qi Model&lt;/strong&gt; views AI Infrastructure as an organic whole, revealing its operational mechanisms from four dimensions. Each layer focuses on different fundamental questions:&lt;/p&gt;
&lt;h2 id="four-layer-model"&gt;Four-Layer Model&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Focus Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin-Yang&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;State Layer&lt;/td&gt;
&lt;td&gt;The system&amp;rsquo;s internal unity of opposites tension structure, revealing how dual elements like performance vs. constraints, innovation vs. governance coexist&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Role Layer&lt;/td&gt;
&lt;td&gt;Five basic role elements in the system and their collaborative relationships, breaking down complex infrastructure into data, models, compute, platforms, and hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yun&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Time Layer&lt;/td&gt;
&lt;td&gt;The development stage the system is in and its cyclical patterns, describing the evolution cycle from exploration to platformization, then scaling and rebalancing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Flow Layer&lt;/td&gt;
&lt;td&gt;The effective &amp;ldquo;field&amp;rdquo; of flow within the system, characterizing the conduction and feedback of signals and resources, reflecting the overall smoothness of operation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Four-Layer Model
&lt;/figcaption&gt;
&lt;h2 id="model-interactions"&gt;Model Interactions&lt;/h2&gt;
&lt;p&gt;The four-layer model is not isolated but an interconnected organic whole:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The tension of &lt;strong&gt;Yin-Yang&lt;/strong&gt; permeates the dynamic balance of &lt;strong&gt;Five Elements&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;The development of &lt;strong&gt;Five Elements&lt;/strong&gt; roles is constrained by their &lt;strong&gt;Yun&lt;/strong&gt; stage&lt;/li&gt;
&lt;li&gt;The flow of &lt;strong&gt;Qi&lt;/strong&gt; connects the above elements into a self-adaptive cyclic system&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The overview diagram below illustrates each layer of the model and their interactions:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/model-overview/aa67f442a69fb3a6e2b135543dffc745.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/model-overview/aa67f442a69fb3a6e2b135543dffc745.svg" alt="Figure 3: AI Infrastructure ‘Yin-Yang - Five Elements - Yun - Qi’ Model Overview. The Yin-Yang layer embodies the system’s internal tension and unity of opposites, the Five Elements layer defines core role elements, the Yun layer describes system stage cycles, and Qi as a flow element permeates and drives the entire system." data-caption="Figure 3: AI Infrastructure ‘Yin-Yang - Five Elements - Yun - Qi’ Model Overview. The Yin-Yang layer embodies the system’s internal tension and unity of opposites, the Five Elements layer defines core role elements, the Yun layer describes system stage cycles, and Qi as a flow element permeates and drives the entire system."
width="652"
height="1031"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: AI Infrastructure ‘Yin-Yang - Five Elements - Yun - Qi’ Model Overview. The Yin-Yang layer embodies the system’s internal tension and unity of opposites, the Five Elements layer defines core role elements, the Yun layer describes system stage cycles, and Qi as a flow element permeates and drives the entire system.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="model-application-value"&gt;Model Application Value&lt;/h2&gt;
&lt;p&gt;This four-layer model provides a unique perspective for the design, operations, and governance of AI Infrastructure:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Holistic Cognitive Framework&lt;/strong&gt;: Transcend the limitations of single technical metrics to grasp system state as a whole&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamic Balance Thinking&lt;/strong&gt;: Understand unity of opposites relationships and avoid extremes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evolutionary Stage Awareness&lt;/strong&gt;: Grasp the system&amp;rsquo;s development stage and act in accordance with the situation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flow Insights&lt;/strong&gt;: Focus on energy flow within the system to anticipate problems&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Next, we will delve into the connotation, engineering mapping, and mechanism of each layer.&lt;/p&gt;</content:encoded></item><item><title>The Yin-Yang Layer: Dynamic Balance of System States</title><link>https://jimmysong.io/book/ai-infra-dao/yin-yang/</link><pubDate>Tue, 10 Feb 2026 13:56:33 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/yin-yang/</guid><description>Understanding system tensions: expansion vs. constraint, innovation vs. governance, speed vs. stability in AI infrastructure</description><content:encoded>
&lt;p&gt;&lt;strong&gt;Yin-Yang&lt;/strong&gt; is originally a fundamental concept in Chinese philosophy, representing two opposing yet interdependent forces present in all things in the universe. Everything in the world can be classified as either Yin or Yang, and their continuous movement and change generate the various transformations we observe. In the context of systems, Yin-Yang represents the unity of opposites through &lt;strong&gt;tension&lt;/strong&gt;—a pair of attributes or tendencies that pull against yet depend on each other.&lt;/p&gt;
&lt;h2 id="three-typical-pairs-of-yin-yang-tensions"&gt;Three Typical Pairs of Yin-Yang Tensions&lt;/h2&gt;
&lt;p&gt;In AI infrastructure, we identify three typical pairs of &lt;strong&gt;Yin-Yang tensions&lt;/strong&gt;:&lt;/p&gt;
&lt;h2 id="expansion--constraint"&gt;Expansion ↔ Constraint&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Expansion ↔ Constraint&lt;/strong&gt;: The tension between &lt;strong&gt;growth&lt;/strong&gt; trends and &lt;strong&gt;limiting&lt;/strong&gt; forces.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Yang (Expansion)&lt;/strong&gt;: System expansion speed, such as continuously adding tasks and scaling resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Yin (Constraint)&lt;/strong&gt;: Limiting forces, such as cost controls, regulatory constraints, and hardware limits&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;System expansion speed and constraint intensity always coexist. For example, continuously adding tasks and scaling resources in GPU clusters (the Yang of expansion) is constrained by costs, regulations, or hardware limits (the Yin of constraint).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imbalance manifestations&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Pursuing expansion without regard for constraints → Resource contention and crashes&lt;/li&gt;
&lt;li&gt;Excessive constraint → Stifling system vitality&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="innovation--governance"&gt;Innovation ↔ Governance&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Innovation ↔ Governance&lt;/strong&gt;: The tension between &lt;strong&gt;creative&lt;/strong&gt; capability and &lt;strong&gt;control&lt;/strong&gt; requirements.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Yang (Innovation)&lt;/strong&gt;: Technical innovation, introduction of new features&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Yin (Governance)&lt;/strong&gt;: Security reviews, rule-making&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The faster technical innovation progresses, the more easily governance gaps are exposed. For example, introducing new Agent features (innovation, Yang) may outpace security reviews and rule-making (governance, Yin), leading to potential risks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imbalance manifestations&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Innovation outpaces governance → Potential security risks&lt;/li&gt;
&lt;li&gt;Excessively strict governance → Slowing innovation momentum&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="speed--stability"&gt;Speed ↔ Stability&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Speed ↔ Stability&lt;/strong&gt;: The tension between &lt;strong&gt;performance&lt;/strong&gt; advancement and &lt;strong&gt;reliable&lt;/strong&gt; operation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Yang (Speed)&lt;/strong&gt;: Performance improvements, increased throughput&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Yin (Stability)&lt;/strong&gt;: Reliable operation, system stability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When we pursue speed improvements single-mindedly, the cost to stability will eventually manifest. For example, pushing GPU utilization to the limit during model training (speed, Yang) easily leads to more frequent failures or delays (decline in stability, Yin).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imbalance manifestations&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Extreme pursuit of speed → Decline in stability&lt;/li&gt;
&lt;li&gt;Excessive conservatism → Performance waste&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-art-of-yinyang-balance"&gt;The Art of Yin–Yang Balance&lt;/h2&gt;
&lt;p&gt;The Yin–Yang poles described above are not simple trade-offs where you choose one and sacrifice the other, but rather inherent relationships of unity of opposites in systems. Both Yin and Yang sides are opposed yet complementary, neither can be dispensed with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Expansion without constraints is difficult to sustain, constraints without expansion lose meaning&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;As the ancient saying goes, &amp;ldquo;One Yin and one Yang constitute the Way&amp;rdquo; (一阴一阳之谓道). Balancing Yin and Yang is the &amp;ldquo;Way&amp;rdquo; of healthy system operation. For architects, the key lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Insight into dominant tensions&lt;/strong&gt;: Determine which pair of tensions is currently dominant&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Introducing the opposite&lt;/strong&gt;: Introduce the complementary side at the right time to restore balance&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamic adjustment&lt;/strong&gt;: Dynamically transform based on changes in system environment and stage&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practical-cases"&gt;Practical Cases&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Case: GPU Cluster Expansion&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When the cluster is in a state of rapid expansion (Yang exuberant, Yin deficient):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Add scheduling policies and resource quotas (supplement Yin)&lt;/li&gt;
&lt;li&gt;✓ Establish cost control mechanisms (supplement Yin)&lt;/li&gt;
&lt;li&gt;✗ Do not pursue expansion speed single-mindedly&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Case: Agent Feature Innovation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When introducing new Agent features:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Simultaneously establish monitoring and sandboxing mechanisms (supplement Yin)&lt;/li&gt;
&lt;li&gt;✓ Improve security review processes (supplement Yin)&lt;/li&gt;
&lt;li&gt;✗ Do not let innovation outpace governance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Case: Model Training Performance Optimization&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When optimizing model training performance:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Simultaneously strengthen fault tolerance mechanisms and testing (supplement the Yin of stability)&lt;/li&gt;
&lt;li&gt;✓ Set performance baselines and rollback mechanisms (supplement Yin)&lt;/li&gt;
&lt;li&gt;✗ Do not infinitely compress fault tolerance time&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dynamic-transformation-of-yinyang-states"&gt;Dynamic Transformation of Yin–Yang States&lt;/h2&gt;
&lt;p&gt;It&amp;rsquo;s important to note that Yin–Yang states are not static and unchanging, but dynamically transform with system environment and stage.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The same capability may transform from an advantage to a risk at different stages&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For example, a &amp;ldquo;rapid development&amp;rdquo; strategy that drives rapid iteration during the startup stage, if applied without restraint during the scaling stage, can instead become a major threat to stability.&lt;/p&gt;
&lt;p&gt;The analysis of the Yin–Yang layer reminds us to constantly pay attention to the ebb and flow of these opposing forces, and to keep the system in a state of elastic tension through adjustments, rather than snapping or becoming slack and ineffective.&lt;/p&gt;</content:encoded></item><item><title>Five Elements Layer: Classification and Collaboration of System Roles</title><link>https://jimmysong.io/book/ai-infra-dao/five-elements/</link><pubDate>Tue, 10 Feb 2026 13:56:06 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/five-elements/</guid><description>Five system roles: data, models, compute, platforms, and hardware—how they interact and balance in AI infrastructure</description><content:encoded>
&lt;p&gt;&lt;strong&gt;Five Elements (Wǔxíng, Five Elements or Five Phases) theory&lt;/strong&gt; divides everything in the world into five basic elements: Wood, Fire, Earth, Metal, Water. Each element represents a fundamental attribute or functional role, with the five elements generating and overcoming each other in an endless cycle.&lt;/p&gt;
&lt;p&gt;In AI infrastructure, we use &amp;ldquo;Five Elements&amp;rdquo; to characterize the system&amp;rsquo;s five core elements and their responsibilities:&lt;/p&gt;
&lt;h2 id="engineering-mapping-of-five-elements"&gt;Engineering Mapping of Five Elements&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Five Elements&lt;/th&gt;
&lt;th&gt;Symbol&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Engineering Correspondence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Water&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🌊&lt;/td&gt;
&lt;td&gt;Flow and containment&lt;/td&gt;
&lt;td&gt;Data flow and quality: data pipelines, data assets, and quality control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Wood&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🌲&lt;/td&gt;
&lt;td&gt;Growth and creation&lt;/td&gt;
&lt;td&gt;Model growth and capability expansion: model architecture iteration, parameter scale expansion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fire&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🔥&lt;/td&gt;
&lt;td&gt;Energy and execution&lt;/td&gt;
&lt;td&gt;Compute conversion and work efficiency: GPU/TPU computing, job scheduling efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Earth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🏔️&lt;/td&gt;
&lt;td&gt;Support and stability&lt;/td&gt;
&lt;td&gt;Platform support and orchestration governance: distributed coordination, middleware, scheduling systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;⚙️&lt;/td&gt;
&lt;td&gt;Strength and standardization&lt;/td&gt;
&lt;td&gt;Hardware constraints and physical boundaries: GPU/CPU performance, storage capacity, network bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: Engineering Mapping of Five Elements
&lt;/figcaption&gt;
&lt;h2 id="water--data-flow-and-quality"&gt;Water – Data Flow and Quality&lt;/h2&gt;
&lt;p&gt;Corresponds to &lt;strong&gt;data pipelines, data assets, and quality control&lt;/strong&gt; in the system.&lt;/p&gt;
&lt;p&gt;Water symbolizes flow and containment, analogous to the circulation and nourishing role of data in the system, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training data acquisition&lt;/li&gt;
&lt;li&gt;Real-time data input&lt;/li&gt;
&lt;li&gt;Feedback signal transmission&lt;/li&gt;
&lt;li&gt;Data cleaning and quality assurance&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="wood--model-growth-and-capability-expansion"&gt;Wood – Model Growth and Capability Expansion&lt;/h2&gt;
&lt;p&gt;Corresponds to &lt;strong&gt;the evolution and growth of machine learning models and algorithms&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Wood represents growth and creation, mapped to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model architecture iteration&lt;/li&gt;
&lt;li&gt;Parameter scale expansion&lt;/li&gt;
&lt;li&gt;Cultivation of new capabilities&lt;/li&gt;
&lt;li&gt;Algorithm optimization and improvement&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="fire--compute-conversion-and-work-efficiency"&gt;Fire – Compute Conversion and Work Efficiency&lt;/h2&gt;
&lt;p&gt;Corresponds to &lt;strong&gt;computing processes and the utilization of compute resources&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Fire symbolizes energy and execution, reflected as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Using GPU/TPU and other compute resources for calculation&lt;/li&gt;
&lt;li&gt;Converting electrical energy into model training and inference work&lt;/li&gt;
&lt;li&gt;Parallel computing capability&lt;/li&gt;
&lt;li&gt;Job scheduling efficiency&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="earth--platform-support-and-orchestration-governance"&gt;Earth – Platform Support and Orchestration Governance&lt;/h2&gt;
&lt;p&gt;Corresponds to &lt;strong&gt;the support and governance capabilities of the platform layer&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Earth represents support and stability, analogous to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Infrastructure platform support for upper-layer applications&lt;/li&gt;
&lt;li&gt;Distributed system coordination and orchestration&lt;/li&gt;
&lt;li&gt;Middleware services&lt;/li&gt;
&lt;li&gt;Scheduling systems and policy management&lt;/li&gt;
&lt;li&gt;Permission systems, service quality assurance&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="metal--hardware-constraints-and-physical-boundaries"&gt;Metal – Hardware Constraints and Physical Boundaries&lt;/h2&gt;
&lt;p&gt;Corresponds to &lt;strong&gt;underlying hardware and system hard limits&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Metal represents strength and standardization, mapped to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPU/CPU hardware performance&lt;/li&gt;
&lt;li&gt;Storage capacity&lt;/li&gt;
&lt;li&gt;Network bandwidth&lt;/li&gt;
&lt;li&gt;Physical conditions and hard rules (power consumption, safety specifications, etc.)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="five-elements-generation-relationships"&gt;Five Elements Generation Relationships&lt;/h2&gt;
&lt;p&gt;The Five Elements form a &lt;strong&gt;positive cycle&lt;/strong&gt; through &amp;ldquo;generation&amp;rdquo; relationships:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Data (Water) spawns model growth (Wood), model requirements stimulate compute investment (Fire), compute development drives platform thickening (Earth), platform capabilities utilize hardware to push the boundaries of (Metal), and hardware progress in turn supports greater data acquisition (Water)&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/8c375bf76fabf035907532a52a97107e.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/8c375bf76fabf035907532a52a97107e.svg" alt="Figure 7: Five Elements generation relationship diagram. Water generates Wood, Wood generates Fire, Fire generates Earth, Earth generates Metal, Metal generates Water, representing the mutually reinforcing cycle between data, models, compute, platforms, and hardware." data-caption="Figure 7: Five Elements generation relationship diagram. Water generates Wood, Wood generates Fire, Fire generates Earth, Earth generates Metal, Metal generates Water, representing the mutually reinforcing cycle between data, models, compute, platforms, and hardware."
width="1040"
height="217"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: Five Elements generation relationship diagram. Water generates Wood, Wood generates Fire, Fire generates Earth, Earth generates Metal, Metal generates Water, representing the mutually reinforcing cycle between data, models, compute, platforms, and hardware.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="five-elements-overcoming-relationships"&gt;Five Elements Overcoming Relationships&lt;/h2&gt;
&lt;p&gt;At the same time, &lt;strong&gt;overcoming&lt;/strong&gt; relationships also exist among the Five Elements, meaning when one element is too strong or imbalanced, it will suppress or weaken another element:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Wood overcomes Earth&lt;/strong&gt;: Excessive model expansion increases the burden on the platform (Earth), potentially even crushing the existing architecture&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Earth overcomes Water&lt;/strong&gt;: Overly heavy platforms and rules will hinder the free flow of data (Water)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Water overcomes Fire&lt;/strong&gt;: Data bottlenecks will limit the performance of compute&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fire overcomes Metal&lt;/strong&gt;: Excessive compute demand may break through hardware (Metal) limits&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metal overcomes Wood&lt;/strong&gt;: Strict hardware and rule limitations will curb the expansion of models (Wood)&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/a97764daf5c160b53c8b0426159cc16f.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/a97764daf5c160b53c8b0426159cc16f.svg" alt="Figure 8: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element." data-caption="Figure 8: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element."
width="908"
height="333"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 8: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/b7ac6707c367b63143436e20483fc83c.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/five-elements/b7ac6707c367b63143436e20483fc83c.svg" alt="Figure 9: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element." data-caption="Figure 9: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element."
width="729"
height="273"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 9: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system’s internal checks and balances mechanism: any element becoming excessively strong will constrain another element.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;em&gt;Figure 3: Five Elements generation and overcoming relationship diagram. Dashed arrows indicate overcoming relationships, reflecting the system&amp;rsquo;s internal checks and balances mechanism: any element becoming excessively strong will constrain another element.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="five-elements-balance-diagnosis"&gt;Five Elements Balance Diagnosis&lt;/h2&gt;
&lt;p&gt;Through the Five Elements model, engineering teams can systematically check the &lt;strong&gt;role completeness and balance&lt;/strong&gt; of infrastructure.&lt;/p&gt;
&lt;h2 id="common-imbalance-patterns"&gt;Common Imbalance Patterns&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Imbalance Pattern&lt;/th&gt;
&lt;th&gt;Manifestation&lt;/th&gt;
&lt;th&gt;Consequence&lt;/th&gt;
&lt;th&gt;Solution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strong Wood, Weak Water&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Focus on model algorithm iteration, neglect data quality&lt;/td&gt;
&lt;td&gt;Model performance hits bottlenecks&lt;/td&gt;
&lt;td&gt;Strengthen data pipelines and quality control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strong Metal, Weak Earth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stack hardware, insufficient platform governance capability&lt;/td&gt;
&lt;td&gt;Poor resource utilization, lack of vitality&lt;/td&gt;
&lt;td&gt;Improve platform governance and scheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Vigorous Fire, Broken Wood&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Large compute investment, models can&amp;rsquo;t keep up&lt;/td&gt;
&lt;td&gt;Resource waste&lt;/td&gt;
&lt;td&gt;Optimize model architecture, improve compute utilization efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: Common Imbalance Patterns
&lt;/figcaption&gt;
&lt;h2 id="balance-principles"&gt;Balance Principles&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Successful large-scale systems require coordinated cooperation of all five elements&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Let each of the five elements fulfill its duties in their respective roles&lt;/li&gt;
&lt;li&gt;Maintain generation as primary, overcoming as secondary&lt;/li&gt;
&lt;li&gt;Prevent any side from excessive expansion or shrinkage&lt;/li&gt;
&lt;li&gt;Regularly check the balance state of Five Elements&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Only by letting the five elements fulfill their respective roles and mutually promote each other, while preventing any side from excessive expansion or shrinkage, can the entire system maintain &lt;strong&gt;robustness and evolutionary capability&lt;/strong&gt;.&lt;/p&gt;</content:encoded></item><item><title>The Yun Layer: Stages and Cycles of System Evolution</title><link>https://jimmysong.io/book/ai-infra-dao/yun/</link><pubDate>Tue, 10 Feb 2026 13:56:38 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/yun/</guid><description>System evolution stages: exploration, platform, scale, and rebalancing phases in AI infrastructure growth</description><content:encoded>
&lt;p&gt;&lt;strong&gt;Yun (运)&lt;/strong&gt; here refers to the developmental stages and temporal rhythms experienced by a system, which can be understood as the lifecycle cycles or &amp;ldquo;fortune&amp;rdquo; of infrastructure.&lt;/p&gt;
&lt;p&gt;Large-scale infrastructure is not static but evolves cyclically through the &lt;strong&gt;Exploration Period&lt;/strong&gt;, &lt;strong&gt;Platform Period&lt;/strong&gt;, &lt;strong&gt;Scale Period&lt;/strong&gt;, and &lt;strong&gt;Rebalancing Period&lt;/strong&gt;, with each stage having its primary contradictions and tasks.&lt;/p&gt;
&lt;p&gt;Below are the four evolutionary stages.&lt;/p&gt;
&lt;h2 id="exploration-period-initial-stage"&gt;Exploration Period (Initial Stage)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Characteristics&lt;/strong&gt;: High variance, low structure, rapid trial and error&lt;/p&gt;
&lt;p&gt;At this stage, new technologies and requirements emerge constantly, system architecture is loose, and diverse experiments coexist.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Primary Tasks&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Explore effective paths&lt;/li&gt;
&lt;li&gt;Rapidly validate model and functional directions&lt;/li&gt;
&lt;li&gt;Collect data and preliminary stability signals&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Characteristics&lt;/strong&gt;: &lt;strong&gt;Wood and Fire in Command&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model innovation (Wood) and computing experimentation (Fire) are core drivers&lt;/li&gt;
&lt;li&gt;Expansion (Yang) outweighs constraints (Yin)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Architecture Strategy&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Tolerate some chaos&lt;/li&gt;
&lt;li&gt;✓ Encourage innovation and iteration&lt;/li&gt;
&lt;li&gt;✓ Focus on collecting data and preliminary stability signals&lt;/li&gt;
&lt;li&gt;✗ Don&amp;rsquo;t prematurely introduce heavy processes and restrictions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="platform-period-growth-stage"&gt;Platform Period (Growth Stage)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Characteristics&lt;/strong&gt;: Standardization emerges, interfaces and processes converge&lt;/p&gt;
&lt;p&gt;After exploration, the system enters a stage of integration and regulation, beginning to establish unified platforms, standard interfaces, and governance processes, consolidating scattered results into platform capabilities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Primary Tasks&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Establish unified platforms&lt;/li&gt;
&lt;li&gt;Define standard interfaces&lt;/li&gt;
&lt;li&gt;Consolidate governance processes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Characteristics&lt;/strong&gt;: &lt;strong&gt;Fire Generates Earth&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Successful practices in computing and functionality (Fire) give rise to platform support requirements (Earth)&lt;/li&gt;
&lt;li&gt;Governance and standards gradually strengthen&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Architecture Strategy&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Extract common requirements&lt;/li&gt;
&lt;li&gt;✓ Build support platforms (Yin increases)&lt;/li&gt;
&lt;li&gt;✓ Lay the foundation for next-stage scaling&lt;/li&gt;
&lt;li&gt;✗ Don&amp;rsquo;t remain in disordered exploration&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="scale-period-mature-stage"&gt;Scale Period (Mature Stage)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Characteristics&lt;/strong&gt;: Efficiency, throughput, and cost become the main battlefield&lt;/p&gt;
&lt;p&gt;The system is deployed at scale, and focus shifts to optimizing efficiency and costs, improving throughput and reliability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Primary Tasks&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Optimize efficiency&lt;/li&gt;
&lt;li&gt;Improve throughput&lt;/li&gt;
&lt;li&gt;Reduce costs&lt;/li&gt;
&lt;li&gt;Ensure reliability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Characteristics&lt;/strong&gt;: &lt;strong&gt;Heavy Earth Breaks Wood&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Platforms (Earth) and hard constraints begin to dominate&lt;/li&gt;
&lt;li&gt;Overly idealistic model expansion (Wood) will encounter setbacks from realistic conditions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Architecture Strategy&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Strengthen monitoring and automated operations&lt;/li&gt;
&lt;li&gt;✓ Control overly strong &amp;ldquo;Yang&amp;rdquo; through governance means&lt;/li&gt;
&lt;li&gt;✓ Ensure robust system operation&lt;/li&gt;
&lt;li&gt;✗ Don&amp;rsquo;t continue with startup-era casual practices&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="rebalancingsubstitution-period-renewal-stage"&gt;Rebalancing/Substitution Period (Renewal Stage)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Characteristics&lt;/strong&gt;: Old structures are corrected or replaced by new structures&lt;/p&gt;
&lt;p&gt;When the previous stage&amp;rsquo;s patterns reach their limits, the system either enters self-correction by introducing new elements to rebalance, or gets disrupted and replaced by a new paradigm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Primary Tasks&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Introduce new elements to rebalance&lt;/li&gt;
&lt;li&gt;Or accept substitution by a new paradigm&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Characteristics&lt;/strong&gt;: &lt;strong&gt;Metal and Water Rise Again&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Suppressed hardware/rule innovations (Metal) and new data potentials (Water) rise again&lt;/li&gt;
&lt;li&gt;Driving system transformation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Architecture Strategy&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Be forward-looking, dare to break through&lt;/li&gt;
&lt;li&gt;✓ Transition smoothly, avoid severe volatility&lt;/li&gt;
&lt;li&gt;✗ Don&amp;rsquo;t cling to the status quo&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="evolutionary-cycle"&gt;Evolutionary Cycle&lt;/h2&gt;
&lt;p&gt;The above stages form a cyclical pattern, where the endpoint of each stage is also the starting point of the next &lt;strong&gt;↻&lt;/strong&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/yun/f0b211635d09b8bc69e8540c27060c22.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/yun/f0b211635d09b8bc69e8540c27060c22.svg" alt="Figure 3: The “Yun” cycle of AI infrastructure evolution. Systems start from the exploration period, undergo platform period standardization, enter the scale period for efficiency optimization, and ultimately move toward a new cycle of rebalancing or substitution.*" data-caption="Figure 3: The “Yun” cycle of AI infrastructure evolution. Systems start from the exploration period, undergo platform period standardization, enter the scale period for efficiency optimization, and ultimately move toward a new cycle of rebalancing or substitution.*"
width="1692"
height="282"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: The “Yun” cycle of AI infrastructure evolution. Systems start from the exploration period, undergo platform period standardization, enter the scale period for efficiency optimization, and ultimately move toward a new cycle of rebalancing or substitution.*&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-art-of-following-the-momentum"&gt;The Art of Following the Momentum&lt;/h2&gt;
&lt;p&gt;A mature infrastructure organization should be able to determine its current stage based on internal and external signals and adjust its strategy accordingly.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If stage transitions are ignored or excessively rushed, the system will experience disturbances or even crises&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="error-examples"&gt;Error Examples&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Erroneous Behavior&lt;/th&gt;
&lt;th&gt;Manifestation&lt;/th&gt;
&lt;th&gt;Consequence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pulling Up Seedlings to Help Them Grow&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Managing systems still in exploration period as scaled systems, prematurely suppressing change&lt;/td&gt;
&lt;td&gt;Stifling innovation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Going Against the Momentum&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Remaining in disordered exploration when it&amp;rsquo;s time to enter the platform period&lt;/td&gt;
&lt;td&gt;Missing the window for structured growth and creating hidden risks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Clinging to the Status Quo&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Unwilling to change when rebalancing period is needed&lt;/td&gt;
&lt;td&gt;System rigidity and aging&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Development Stage Characteristics
&lt;/figcaption&gt;
&lt;h2 id="stage-assessment-checklist"&gt;Stage Assessment Checklist&lt;/h2&gt;
&lt;p&gt;Through the &amp;ldquo;Yun&amp;rdquo; layer perspective, teams can examine the current macro stage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Are we validating new concepts or expanding our achievements?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; What is the system&amp;rsquo;s primary contradiction?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; When might the next stage arrive?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Does our strategy align with the current stage?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Example Questions&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Are we in the exploration period?
&lt;ul&gt;
&lt;li&gt;If yes → Focus on rapid trial and error and validation&lt;/li&gt;
&lt;li&gt;If no → Consider whether to enter the platform period&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Does our system need standardization?
&lt;ul&gt;
&lt;li&gt;If yes → Enter platform period, establish platforms and standards&lt;/li&gt;
&lt;li&gt;If no → Continue exploration&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Qi Layer: Effective System Flow and Pressure Fields</title><link>https://jimmysong.io/book/ai-infra-dao/qi/</link><pubDate>Tue, 10 Feb 2026 13:56:17 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/qi/</guid><description>Effective flow and pressure distribution in systems—data flow, signal propagation, and system health monitoring</description><content:encoded>
&lt;p&gt;&lt;strong&gt;Qi (气)&lt;/strong&gt; in Chinese culture refers to the energy and flow field that permeates all things. In AI infrastructure, we borrow the concept of &amp;ldquo;Qi&amp;rdquo; to describe the effective flow and pressure distribution within systems.&lt;/p&gt;
&lt;p&gt;This includes the circulation of data, tasks, and signals throughout the system, as well as how various explicit or implicit &lt;strong&gt;system pressures&lt;/strong&gt; accumulate, propagate, and release.&lt;/p&gt;
&lt;h2 id="the-essence-of-qi-overall-state-of-affairs"&gt;The Essence of Qi: Overall State of Affairs&lt;/h2&gt;
&lt;p&gt;Unlike traditional single-point metric monitoring, the concept of &amp;ldquo;Qi&amp;rdquo; reminds us to focus on the overall &lt;strong&gt;state of affairs&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Signals are not isolated events, but rather gather and flow like a field&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A sudden spike in GPU utilization may not be abnormal&lt;/li&gt;
&lt;li&gt;But if multiple metrics (job queue length, response latency, memory usage, etc.) show a simultaneous trend of increase and persistence → this indicates a change in the &amp;ldquo;Qi field&amp;rdquo;&lt;/li&gt;
&lt;li&gt;This signals the system entering a high-pressure state&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This &lt;strong&gt;signal field&lt;/strong&gt; manifests as the gathering and stretching of Qi, indicating the accumulation of some form of system tension.&lt;/p&gt;
&lt;h2 id="two-states-of-qi"&gt;Two States of Qi&lt;/h2&gt;
&lt;h2 id="qi-flow-system-active"&gt;Qi Flow: System Active&lt;/h2&gt;
&lt;p&gt;When all elements coordinate well, data and instructions flow smoothly, producing value efficiently:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Processing rates across all stages are basically matched&lt;/li&gt;
&lt;li&gt;No long-term backlogs or idle resources&lt;/li&gt;
&lt;li&gt;Timely system responses&lt;/li&gt;
&lt;li&gt;Balanced resource utilization&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="qi-stagnation-system-pathological"&gt;Qi Stagnation: System Pathological&lt;/h2&gt;
&lt;p&gt;If a bottleneck or imbalance occurs somewhere, Qi&amp;rsquo;s flow is obstructed, causing local pressure to surge:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Jobs queue for long periods&lt;/li&gt;
&lt;li&gt;CPU/GPU long-term idle or 100% utilization&lt;/li&gt;
&lt;li&gt;Serious message queue backlog&lt;/li&gt;
&lt;li&gt;Frequent anomaly alerts&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Ultimately, this may trigger failures or performance collapse at weak points.&lt;/p&gt;
&lt;h2 id="qis-flow-path"&gt;Qi&amp;rsquo;s Flow Path&lt;/h2&gt;
&lt;p&gt;To intuitively understand Qi&amp;rsquo;s flow path, we can view the system as a closely connected network:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/qi/045a20987f2f172d95c9136ea84ea722.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/qi/045a20987f2f172d95c9136ea84ea722.svg" alt="Figure 3: Diagram of system ‘Qi’ flow path. Data (Water) Qi enters Model (Wood), triggering Computing Power (Fire) operation, coordinated via Platform (Earth), executed on Hardware (Metal), producing results that feed back to the data layer, forming a closed loop." data-caption="Figure 3: Diagram of system ‘Qi’ flow path. Data (Water) Qi enters Model (Wood), triggering Computing Power (Fire) operation, coordinated via Platform (Earth), executed on Hardware (Metal), producing results that feed back to the data layer, forming a closed loop."
width="2558"
height="382"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Diagram of system ‘Qi’ flow path. Data (Water) Qi enters Model (Wood), triggering Computing Power (Fire) operation, coordinated via Platform (Earth), executed on Hardware (Metal), producing results that feed back to the data layer, forming a closed loop.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Qi&amp;rsquo;s Cycle&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Data (Water) Qi enters Model (Wood)&lt;/li&gt;
&lt;li&gt;Drives Computing Power (Fire) to operate&lt;/li&gt;
&lt;li&gt;Coordinated via Platform (Earth)&lt;/li&gt;
&lt;li&gt;Executes computation on Hardware (Metal)&lt;/li&gt;
&lt;li&gt;Outputs results, producing new data or signals&lt;/li&gt;
&lt;li&gt;Feeds back into the data pool (Water)&lt;/li&gt;
&lt;li&gt;Cycle repeats&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="two-forms-of-qi"&gt;Two Forms of Qi&lt;/h2&gt;
&lt;h2 id="healthy-flow"&gt;Healthy Flow&lt;/h2&gt;
&lt;p&gt;Qi circulates ceaselessly among the five elements, maintaining system functionality:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If every step flows smoothly → system operates smoothly&lt;/li&gt;
&lt;li&gt;If any step is obstructed → Qi flow slows or even reverses, damaging system performance and stability&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="pressure-propagation"&gt;Pressure Propagation&lt;/h2&gt;
&lt;p&gt;Qi refers not only to healthy flow, but also to &lt;strong&gt;pressure propagation&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Example: Data Inflow Surge&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Data inflow surges but model processing capacity cannot keep up&lt;/li&gt;
&lt;li&gt;Unprocessed data continuously accumulates&lt;/li&gt;
&lt;li&gt;Manifests as excessive pressure in the data layer (Water)&lt;/li&gt;
&lt;li&gt;Leading to suppression of computing power performance (Fire weakens)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Example: Hardware Resource Exhaustion&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Hardware (Metal) resources exhausted&lt;/li&gt;
&lt;li&gt;Computing requests cannot be satisfied&lt;/li&gt;
&lt;li&gt;Obstructed Qi transforms into queuing pressure&lt;/li&gt;
&lt;li&gt;Feeds back to platform (Earth) scheduling layer and user experience&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="application-of-qi-layer-in-operations"&gt;Application of Qi Layer in Operations&lt;/h2&gt;
&lt;p&gt;Through the lens of &amp;ldquo;Qi&amp;rdquo;, operations and architecture teams can more sensitively detect sub-optimal system states:&lt;/p&gt;
&lt;h2 id="not-just-whether-theres-a-problem-but-how-its-trending"&gt;Not Just Whether There&amp;rsquo;s a Problem, But How It&amp;rsquo;s Trending&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Qi State&lt;/th&gt;
&lt;th&gt;Manifestation&lt;/th&gt;
&lt;th&gt;Warning Significance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Stagnation Emerging&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Latency jitter gradually worsening&lt;/td&gt;
&lt;td&gt;System entering sub-stable state, needs 疏导&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Flow Obstruction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Request failure rate rising, retries increasing&lt;/td&gt;
&lt;td&gt;某环节阻塞，needs investigation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Scattering&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Metrics fluctuating severely, irregular&lt;/td&gt;
&lt;td&gt;System severely imbalanced, needs overall adjustment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Deficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Resource utilization long-term low&lt;/td&gt;
&lt;td&gt;Configuration unreasonable, needs optimization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: Qi State and Warning Significance
&lt;/figcaption&gt;
&lt;h2 id="qi-disorder-precedes-major-incidents"&gt;Qi Disorder Precedes Major Incidents&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Latency jitter gradually worsening → signals system entering sub-stable state&lt;/li&gt;
&lt;li&gt;If no measures are taken to resolve (scaling resources, optimizing algorithms, or rate limiting) → may evolve to complete failure&lt;/li&gt;
&lt;li&gt;Agent task interaction rhythm (Qi) slows or stops → may indicate poor communication between agents or deadlock&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="strategies-for-guiding-qi-flow"&gt;Strategies for Guiding Qi Flow&lt;/h2&gt;
&lt;p&gt;Maintaining &lt;strong&gt;smooth Qi flow&lt;/strong&gt; requires building resilience:&lt;/p&gt;
&lt;h2 id="architecture-level"&gt;Architecture Level&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Peak shaving and valley filling mechanisms&lt;/strong&gt;: Absorb 突发流量&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Message queue backpressure protection&lt;/strong&gt;: Prevent pressure backflow&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Elastic buffer design&lt;/strong&gt;: Reserve margin to handle impacts&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="strategy-level"&gt;Strategy Level&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Slack capacity&lt;/strong&gt;: Maintain certain redundancy&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Elastic scaling strategies&lt;/strong&gt;: Dynamically adjust resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate limiting and degradation mechanisms&lt;/strong&gt;: Protect core functionality&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="agent-system-special-attention"&gt;Agent System Special Attention&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Monitor task queues and communication latency&lt;/li&gt;
&lt;li&gt;Ensure information flow (Qi) between agents is unobstructed&lt;/li&gt;
&lt;li&gt;Introduce coordinator agents or reduce concurrency when necessary to smooth Qi flow&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="qi-layer-monitoring-practices"&gt;Qi Layer Monitoring Practices&lt;/h2&gt;
&lt;p&gt;Establish system-wide observability:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Monitoring Dimension&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;Tool Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Traffic Distribution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Request flow across stages&lt;/td&gt;
&lt;td&gt;Distributed Tracing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Queue Backlog&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Queue length trends&lt;/td&gt;
&lt;td&gt;Message Queue Monitoring&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Resource Utilization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;CPU/GPU/Memory/Storage&lt;/td&gt;
&lt;td&gt;Prometheus + Grafana&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Latency Distribution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;P50/P95/P99 latency&lt;/td&gt;
&lt;td&gt;APM Tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Anomaly Trends&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Error rate, retry rate changes&lt;/td&gt;
&lt;td&gt;Log Aggregation Analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: Qi Layer Monitoring Dimensions
&lt;/figcaption&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Qi layer provides an effective liquidity metric, helping us pulse-check whether the system&amp;rsquo;s &amp;ldquo;blood and Qi&amp;rdquo; are abundant and flowing smoothly&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Qi&amp;rsquo;s operation can be understood as whether the system&amp;rsquo;s &amp;ldquo;meridians&amp;rdquo; are unobstructed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Qi flow means system active&lt;/strong&gt;: Data and instructions flow smoothly, producing value efficiently&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qi stagnation means system pathological&lt;/strong&gt;: Flow obstructed, local pressure surges, ultimately triggering failures&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Just as in Traditional Chinese Medicine&amp;rsquo;s four examination methods, by observing &amp;ldquo;Qi&amp;rsquo;s&amp;rdquo; operation, we can predict the trajectory of system problems and apply targeted remedies.&lt;/p&gt;</content:encoded></item><item><title>Dynamic Relationship Modeling: Five Elements Flow Under Yin-Yang Balance</title><link>https://jimmysong.io/book/ai-infra-dao/dynamic-modeling/</link><pubDate>Tue, 10 Feb 2026 13:55:47 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/dynamic-modeling/</guid><description>Integrating Yin-Yang, Five Elements, Yun, and Qi layers to explain complex AI infrastructure system behavior</description><content:encoded>
&lt;h2 id="yin-yang--five-elements-intrinsic-tension-of-elements"&gt;Yin-Yang × Five Elements: Intrinsic Tension of Elements&lt;/h2&gt;
&lt;p&gt;Each &lt;strong&gt;Five Elements&lt;/strong&gt; component contains both &lt;strong&gt;Yin&lt;/strong&gt; and &lt;strong&gt;Yang&lt;/strong&gt; aspects, manifesting with different polarities in different contexts:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/dynamic-modeling/a179955c34047cc8d9fb31f92c6fb594.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/dynamic-modeling/a179955c34047cc8d9fb31f92c6fb594.svg" alt="Figure 3: Yin-Yang states of Five Elements. Each element includes Yin (potential, static, introverted) and Yang (explicit, dynamic, extroverted) aspects, with transformation possible between them depending on context." data-caption="Figure 3: Yin-Yang states of Five Elements. Each element includes Yin (potential, static, introverted) and Yang (explicit, dynamic, extroverted) aspects, with transformation possible between them depending on context."
width="4592"
height="408"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Yin-Yang states of Five Elements. Each element includes Yin (potential, static, introverted) and Yang (explicit, dynamic, extroverted) aspects, with transformation possible between them depending on context.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Yin-Yang Attributes of the Five Elements&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Five Elements&lt;/th&gt;
&lt;th&gt;Yin State&lt;/th&gt;
&lt;th&gt;Yang State&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Water (Data)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Potential data reserves, implicit patterns (static storage of historical data)&lt;/td&gt;
&lt;td&gt;Instant data flow, real-time feedback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Wood (Model)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dormant capabilities (unactivated parameters, backup algorithms)&lt;/td&gt;
&lt;td&gt;Explicit expansion (model architecture updates, parameter surge)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fire (Compute)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stored energy (idle compute, waiting for scheduling)&lt;/td&gt;
&lt;td&gt;High-load operation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Earth (Platform)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Static support (stable operation, non-intervention)&lt;/td&gt;
&lt;td&gt;Proactive scheduling and expanded governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metal (Hardware)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Implicit constraints (unused capacity)&lt;/td&gt;
&lt;td&gt;Explicit limits (resource hard caps maxed out)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 7: Dynamic Model Overview
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Signs of Yin-Yang Imbalance&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Fire Excessively Yin&lt;/strong&gt;: GPU compute idle for long periods while tasks backlog → Poor scheduling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fire Excessively Yang&lt;/strong&gt;: GPUs at 24-hour full load with no elasticity → Hidden crash risk&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Earth Excessively Yang&lt;/strong&gt;: Too many platform rules → Stifling innovation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Earth Excessively Yin&lt;/strong&gt;: Lack of platform control → Leading to chaos&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="five-elements--qi-dynamic-network-of-flow"&gt;Five Elements × Qi: Dynamic Network of Flow&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;Five Elements&lt;/strong&gt; framework provides tools to decompose systems, but system components are not static puzzles—rather, they connect into a dynamic network through &lt;strong&gt;the flow of Qi&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Generating Relationships&lt;/strong&gt;: Qi flows smoothly, forming positive feedback loops&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Controlling Relationships&lt;/strong&gt;: Qi stagnates at certain links or reverse effects strengthen&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Relationship Principles&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Generating primarily, Controlling secondarily&lt;/strong&gt;—main energy flows transmit successfully through each link, while balancing forces intervene moderately only to prevent extreme situations.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="yun--yin-yang-five-elements-boundary-conditions-for-stage-evolution"&gt;Yun × Yin-Yang Five Elements: Boundary Conditions for Stage Evolution&lt;/h2&gt;
&lt;p&gt;The stage-based nature of &lt;strong&gt;Yun&lt;/strong&gt; provides a perspective of &lt;strong&gt;boundary conditions&lt;/strong&gt; evolving over time for the aforementioned Yin-Yang Five Elements dynamics.&lt;/p&gt;
&lt;p&gt;Each stage strengthens or weakens certain elements and tensions:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Main Characteristics&lt;/th&gt;
&lt;th&gt;Five Elements Characteristics&lt;/th&gt;
&lt;th&gt;Yin-Yang Characteristics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exploration Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;High variance, low structure, rapid trial and error&lt;/td&gt;
&lt;td&gt;Wood and Fire dominant&lt;/td&gt;
&lt;td&gt;Expansion (Yang) outweighs Constraints (Yin)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Platform Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standardization emerges, interfaces and processes converge&lt;/td&gt;
&lt;td&gt;Fire generates Earth&lt;/td&gt;
&lt;td&gt;Governance (Yin increasing) gradually strengthens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scale Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Efficiency, throughput, cost become main battlegrounds&lt;/td&gt;
&lt;td&gt;Earth dominates Wood&lt;/td&gt;
&lt;td&gt;Stability (Yin) takes precedence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rebalancing Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Old structures corrected or replaced by new structures&lt;/td&gt;
&lt;td&gt;Metal and Water resurge&lt;/td&gt;
&lt;td&gt;Transformation (Yang) rises again&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 8: Typical Interaction Scenarios
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Stage Transitions&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;The Yun layer tells us &lt;strong&gt;when to shift focus&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;As stages change, the system needs to &amp;ldquo;allocate interests&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Previously dominant elements may become excessive and need convergence&lt;/li&gt;
&lt;li&gt;Previously minor elements need strengthening to address shortcomings&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Examples&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In Platform Stage/Scale Stage → Must strengthen governance (Earth&amp;rsquo;s Yang) and hardware optimization (Metal&amp;rsquo;s Yang)&lt;/li&gt;
&lt;li&gt;To curb the 野蛮 growth tendencies left over from early stages (excessive Wood-Fire Qi)&lt;/li&gt;
&lt;li&gt;In Rebalancing Stage → May need to reactivate suppressed innovation potential (Water-Wood Qi)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="comprehensive-analysis-case-gpu-scheduling-scenario"&gt;Comprehensive Analysis Case: GPU Scheduling Scenario&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s see how to apply the four-layer model to analyze a real GPU scheduling problem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Problem Scenario&lt;/strong&gt;: Cluster experiences task queues under high load&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Diagnosis&lt;/th&gt;
&lt;th&gt;Findings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Observe Qi flow state&lt;/td&gt;
&lt;td&gt;Compute Fire Qi is obstructed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Locate elements&lt;/td&gt;
&lt;td&gt;Data input too intense (Water Yang excessive) but scheduling (Platform Earth) strategy cannot keep up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin-Yang Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Analyze tensions&lt;/td&gt;
&lt;td&gt;Scheduling strategy blindly pursues maximizing utilization (excessively Yang) while lacking elastic buffers (Yin)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yun Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Assess stage&lt;/td&gt;
&lt;td&gt;This is an emerging business that just passed exploration stage and hasn&amp;rsquo;t perfected scheduling—Platform Stage&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 9: Four-Layer Diagnostic Analysis
&lt;/figcaption&gt;
&lt;h2 id="solutions"&gt;Solutions&lt;/h2&gt;
&lt;p&gt;Based on four-layer collaborative diagnosis, develop comprehensive solutions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Qi Layer&lt;/strong&gt;: Unblock Qi flow&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Expand resources or optimize algorithms&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Layer&lt;/strong&gt;: Balance elements&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Strengthen platform scheduling capabilities (Earth)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Yin-Yang Layer&lt;/strong&gt;: Restore balance&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Introduce elastic buffer mechanisms (supplement Yin)&lt;/li&gt;
&lt;li&gt;Avoid blindly pursuing high utilization&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Yun Layer&lt;/strong&gt;: Follow the trend&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Accelerate introduction of standardized scheduling and resource governance (Earth&amp;rsquo;s Yun is approaching)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="value-of-dynamic-modeling"&gt;Value of Dynamic Modeling&lt;/h2&gt;
&lt;p&gt;Through the multi-level dynamic modeling above, we can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Explain complex scenarios more comprehensively&lt;/strong&gt;: No longer limited to single perspectives&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Locate root causes of problems&lt;/strong&gt;: Find fundamental causes rather than surface phenomena&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Point improvement directions&lt;/strong&gt;: Obtain systematic solutions&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Predict system evolution&lt;/strong&gt;: Prepare in advance for stage transitions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practical-recommendations"&gt;Practical Recommendations&lt;/h2&gt;
&lt;p&gt;In daily architecture design and operations, you can establish these thinking habits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;When encountering problems&lt;/strong&gt;: Analyze layer by layer from a four-layer perspective&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When making decisions&lt;/strong&gt;: Consider impacts on all four layers&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When conducting post-mortems&lt;/strong&gt;: Check whether warning signals from the four-layer model were ignored&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The value of a system lies not in pursuing the extreme of a single performance indicator without limit, but in balancing all elements to achieve long-term coordinated development&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;</content:encoded></item><item><title>Engineering Practice Guide: Architecture Decisions Guided by Theory</title><link>https://jimmysong.io/book/ai-infra-dao/engineering-practice/</link><pubDate>Tue, 10 Feb 2026 13:55:59 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/engineering-practice/</guid><description>Practical principles for applying the Yin-Yang Five Elements Qi model in GPU scheduling, Agent Runtime, and platform governance</description><content:encoded>
&lt;p&gt;The theoretical models mentioned above are not 停留在停留在 the conceptual level, but directly provide guidance for the engineering practice of AI infrastructure. In specific scenarios such as GPU scheduling, Agent runtime, and platform governance, we can follow the principles below to apply the Yin-Yang Five Elements Qi Movement model.&lt;/p&gt;
&lt;h2 id="balance-yin-and-yang-avoid-extremes"&gt;Balance Yin and Yang, Avoid Extremes&lt;/h2&gt;
&lt;p&gt;Consider both propelling forces and restraining forces when making architecture decisions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GPU Cluster Scaling&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Satisfy business growth (expanding Yang)&lt;/li&gt;
&lt;li&gt;✓ Set quota and priority policies (constraining Yin)&lt;/li&gt;
&lt;li&gt;✓ Prevent resource abuse&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agent Runtime Design&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;✓ Give agents more autonomy (innovation, Yang)&lt;/li&gt;
&lt;li&gt;✓ Introduce monitoring and sandboxing mechanisms (governance, Yin)&lt;/li&gt;
&lt;li&gt;✓ Prevent loss of control&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Practice Checklist&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After every major adjustment, ask yourself: &lt;strong&gt;Have I introduced corresponding counter-forces to stabilize the system?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="complete-the-five-elements-identify-and-fill-weaknesses"&gt;Complete the Five Elements, Identify and Fill Weaknesses&lt;/h2&gt;
&lt;p&gt;Regularly review whether the five types of elements in the system are balanced.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GPU Infrastructure Check&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Do data pipelines keep up with computing power improvements? (Water and Fire matching)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Does model optimization fully utilize hardware? (Wood and Metal matching)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Can the scheduling platform handle peak loads? (Earth supporting Fire)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Has hardware resources become a bottleneck? (Metal not holding back)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agent Platform Check&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is there high-quality knowledge base or real-time data support? (Water)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is there strong model capability? (Wood)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is there sufficient computing resources? (Fire)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is there a good orchestration framework? (Earth)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is there a reliable environment and interfaces? (Metal)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Practice Strategy&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Once a bottleneck or overload is discovered in a certain link, decisively invest resources to &lt;strong&gt;fill the weakness or reduce the burden on the overloaded part&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem Discovered&lt;/th&gt;
&lt;th&gt;Solution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Insufficient data quality (&amp;ldquo;Water&amp;rdquo; weak)&lt;/td&gt;
&lt;td&gt;Prioritize data governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-term low hardware utilization (Metal strong, Fire weak)&lt;/td&gt;
&lt;td&gt;Optimize algorithms or scheduling to better utilize hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Problem Discovery and Solutions
&lt;/figcaption&gt;
&lt;h2 id="follow-the-trend-align-with-the-movement"&gt;Follow the Trend, Align with the Movement&lt;/h2&gt;
&lt;p&gt;Develop reasonable strategies based on the stage of the system.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategies for Different Stages&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Should Do&lt;/th&gt;
&lt;th&gt;Should Not Do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exploration Phase&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rapid trial and error, validate value&lt;/td&gt;
&lt;td&gt;Prematurely introduce heavy processes and constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Platform Phase&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standardized management, MLOps tools&lt;/td&gt;
&lt;td&gt;Remain in disordered exploration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scale Phase&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Strengthen governance and efficiency optimization&lt;/td&gt;
&lt;td&gt;Still use the casual practices of the startup period&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rebalancing Phase&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Architecture innovation, introduce new technologies&lt;/td&gt;
&lt;td&gt;Refuse to move forward&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: Strategies for Different Stages
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Regular Assessment&lt;/strong&gt;:
At each quarter or important milestone, assess:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; &lt;strong&gt;Which stage are we currently in&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; What is the main contradiction in this stage?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; When might the next stage arrive?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Prepare in advance for the transition&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Practice Cases&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An AI training cluster after validating the concept → Should consider entering standardized management (transitioning from exploration phase to platform phase)&lt;/li&gt;
&lt;li&gt;When system scale expansion encounters bottlenecks → Consider whether to enter the rebalancing phase and break through through architecture innovation&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="observe-qi-field-optimize-flow"&gt;Observe Qi Field, Optimize Flow&lt;/h2&gt;
&lt;p&gt;Establish global observability of the system, focusing on trends and correlations rather than single-point metrics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monitoring Methods&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Distributed tracing&lt;/li&gt;
&lt;li&gt;Metric correlation analysis&lt;/li&gt;
&lt;li&gt;Full-link monitoring&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Signals of Qi Disorder&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Possible Cause&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Frequent occurrence of various abnormal logs&lt;/td&gt;
&lt;td&gt;Global investigation needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A metric&amp;rsquo;s periodic fluctuations becoming increasingly intense&lt;/td&gt;
&lt;td&gt;The system may be approaching a limit internally&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Signals of Qi Disorder
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Strategies to Keep Qi Flowing Smoothly&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Architecture Level&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Peak clipping and valley filling mechanisms&lt;/li&gt;
&lt;li&gt;Message queue backpressure protection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Strategy Level&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Slack capacity&lt;/li&gt;
&lt;li&gt;Elastic scaling strategies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agent System Special Attention&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Monitor task queues and communication latency&lt;/li&gt;
&lt;li&gt;Ensure smooth information flow (Qi) between agents&lt;/li&gt;
&lt;li&gt;Introduce coordinator agents or reduce concurrency when necessary&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dynamic-adjustment-continuous-rebalancing"&gt;Dynamic Adjustment, Continuous Rebalancing&lt;/h2&gt;
&lt;p&gt;Integrate the Yin-Yang Five Elements Qi Movement model into the team&amp;rsquo;s continuous improvement process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core Questions in Architecture Reviews or Incident Retrospectives&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is the current main contradiction more inclined toward expansion or constraint, speed or stability?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is any Five Elements element overloaded (Yang excess) or missing (Yin deficiency)?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is &lt;strong&gt;System Qi&lt;/strong&gt; congested somewhere?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Do our strategies align with the current stage?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Continuous Improvement Process&lt;/strong&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Problem Discovery → Four-Layer Model Diagnosis → Strategy Formulation → Implementation Adjustment → Effect Evaluation → Continuous Optimization&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="practice-case-large-scale-gpu-training-cluster-optimization"&gt;Practice Case: Large-Scale GPU Training Cluster Optimization&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Background&lt;/strong&gt;: A team encountered stability issues while operating a large-scale GPU training cluster.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Four-Layer Model Diagnosis&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Diagnosis&lt;/th&gt;
&lt;th&gt;Findings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin-Yang Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Speed vs Stability&lt;/td&gt;
&lt;td&gt;Continuously compressing fault tolerance and testing time in pursuit of efficiency (speed Yang), leading to frequent online failures (stability Yin damaged)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Five Elements Check&lt;/td&gt;
&lt;td&gt;Data pipeline latency gradually increasing (Water weaker than Fire)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Movement Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stage Judgment&lt;/td&gt;
&lt;td&gt;System has moved from barbaric growth period to maturity period&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Qi Flow State&lt;/td&gt;
&lt;td&gt;Qi stagnation phenomenon obvious&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 4: Monitoring Methods
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Comprehensive Solution&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Yin-Yang Balance&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Suspend performance optimization&lt;/li&gt;
&lt;li&gt;Invest time to strengthen fault tolerance mechanisms and testing (supplement stability Yin)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Five Elements Completion&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Add data preprocessing nodes and caching (strengthen Water)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Movement Adjustment&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Change mindset, shift focus from feature expansion to optimization and governance&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Qi Flow Regulation&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Build full-link tracing system&lt;/li&gt;
&lt;li&gt;Monitor the time of each link from training job submission to completion&lt;/li&gt;
&lt;li&gt;Identify Qi stagnation points and clear them&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;: While maintaining high utilization, the cluster&amp;rsquo;s stability was greatly improved, and no serious downtime occurred again.&lt;/p&gt;
&lt;h2 id="scenario-application-quick-reference-table"&gt;Scenario Application Quick Reference Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Yin-Yang Focus&lt;/th&gt;
&lt;th&gt;Five Elements Check&lt;/th&gt;
&lt;th&gt;Movement Judgment&lt;/th&gt;
&lt;th&gt;Qi Flow Monitoring&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPU Scheduling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Utilization vs Elasticity&lt;/td&gt;
&lt;td&gt;Fire - Earth - Metal Balance&lt;/td&gt;
&lt;td&gt;Scale Phase Efficiency Optimization&lt;/td&gt;
&lt;td&gt;Task queues, resource utilization curves&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent Runtime&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Autonomy vs Governance&lt;/td&gt;
&lt;td&gt;Water - Wood - Fire Coordination&lt;/td&gt;
&lt;td&gt;Exploration Phase Rapid Iteration&lt;/td&gt;
&lt;td&gt;Communication latency, task interaction rhythm&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Platform Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Innovation Risk Control vs Process Efficiency&lt;/td&gt;
&lt;td&gt;Earth - Metal Constraints&lt;/td&gt;
&lt;td&gt;Platform Phase Standardization&lt;/td&gt;
&lt;td&gt;Rule execution rate, change frequency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost Optimization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Performance vs Cost&lt;/td&gt;
&lt;td&gt;Fire - Metal Matching&lt;/td&gt;
&lt;td&gt;Scale Phase Refinement&lt;/td&gt;
&lt;td&gt;Resource waste, idle time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: Signals of Qi Disorder
&lt;/figcaption&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Through the Yin-Yang Five Elements Qi Movement model, we can in practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Avoid Extremes&lt;/strong&gt;: Not blindly pursuing single metrics&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systematic Thinking&lt;/strong&gt;: Analyzing problems from multiple dimensions&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Follow the Trend&lt;/strong&gt;: Adjust strategies based on stages&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Predict Problems&lt;/strong&gt;: Early warning of risks through Qi field changes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Continuous Improvement&lt;/strong&gt;: Establish systematic optimization processes&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;The value of this system lies in: combining Eastern wisdom with engineering practice to provide a unique and effective thinking framework for complex AI infrastructure&lt;/p&gt;
&lt;/blockquote&gt;</content:encoded></item><item><title>System Diagnosis Principles: Criteria for Health Status</title><link>https://jimmysong.io/book/ai-infra-dao/system-diagnosis/</link><pubDate>Tue, 10 Feb 2026 13:56:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/system-diagnosis/</guid><description>Five-dimensional diagnosis framework for AI infrastructure health: element balance, flow smoothness, tension dynamics, stage alignment, and runaway warnings</description><content:encoded>
&lt;p&gt;To maintain the long-term healthy evolution of AI infrastructure, post-mortem summaries are far from sufficient. We need a set of &lt;strong&gt;system diagnosis principles&lt;/strong&gt; to detect hidden risks early and correct deviations.&lt;/p&gt;
&lt;p&gt;Based on the Yin-Yang Five Elements Yun model, diagnosis can be conducted from the following five dimensions:&lt;/p&gt;
&lt;h2 id="five-dimensional-diagnosis-framework"&gt;Five-Dimensional Diagnosis Framework&lt;/h2&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/system-diagnosis/14c23d6ef51af9a7b4c30cb6da07e94a.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/system-diagnosis/14c23d6ef51af9a7b4c30cb6da07e94a.svg" alt="Figure 5: Five-Dimensional Diagnosis Framework Diagram" data-caption="Figure 5: Five-Dimensional Diagnosis Framework Diagram"
width="2899"
height="615"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Five-Dimensional Diagnosis Framework Diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="five-elements-balance-check"&gt;Five Elements Balance Check&lt;/h2&gt;
&lt;p&gt;Assess the current status of five aspects: Data (Water), Models (Wood), Compute (Fire), Platform (Earth), and Hardware (Metal).&lt;/p&gt;
&lt;h2 id="diagnosis-method"&gt;Diagnosis Method&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Checklist&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Can data pipelines keep up with demands? (Water)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Are model capabilities fully utilized? (Wood)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Are compute resources effectively used? (Fire)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Can the platform support current load? (Earth)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Is hardware becoming a bottleneck? (Metal)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="identify-problems"&gt;Identify Problems&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem Type&lt;/th&gt;
&lt;th&gt;Manifestation&lt;/th&gt;
&lt;th&gt;Solution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Short Board&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One element significantly weaker than others&lt;/td&gt;
&lt;td&gt;Prioritize strengthening that element&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Overload&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One element consumes excessive resources or frequently becomes a bottleneck&lt;/td&gt;
&lt;td&gt;Introduce limits or expand other elements to share pressure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 19: Problem Types and Solutions
&lt;/figcaption&gt;
&lt;h2 id="typical-symptoms"&gt;Typical Symptoms&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Water Level Too Low&lt;/strong&gt;: Data pipelines always lag behind training needs → Replenish data processing capacity&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metal Overload&lt;/strong&gt;: Hardware often runs at full capacity or even triggers limit alarms → Expand capacity or impose constraints on upper layers&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Most failures do not stem from missing components, but from long-term role imbalance&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="qi-flow-smoothness-check"&gt;Qi Flow Smoothness Check&lt;/h2&gt;
&lt;p&gt;Analyze whether Qi flows smoothly through the system via full-link monitoring.&lt;/p&gt;
&lt;h2 id="diagnosis-method-1"&gt;Diagnosis Method&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Key Metrics&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Latency distribution of key processes&lt;/li&gt;
&lt;li&gt;Queue backlogs&lt;/li&gt;
&lt;li&gt;Resource utilization curves&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="qi-smooth-vs-qi-not-smooth"&gt;Qi Smooth vs. Qi Not Smooth&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Characteristics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Smooth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Processing rates across stages basically match, without long-term backlogs or idle resources&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Not Smooth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One stage remains a bottleneck for long periods, or large amounts of resources sit idle&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 20: Qi Flow: Smooth vs Obstructed
&lt;/figcaption&gt;
&lt;h2 id="diagnosis-points"&gt;Diagnosis Points&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Distinguish temporary fluctuations from persistent trends: brief peaks don&amp;rsquo;t necessarily indicate Qi blockage, but persistent deviations must be addressed&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Tool Support&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Dashboards and automated alerts&lt;/li&gt;
&lt;li&gt;Timely capture of &amp;ldquo;stagnant Qi&amp;rdquo; locations&lt;/li&gt;
&lt;li&gt;Further investigation of causes (which Five Elements imbalance corresponds)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="yin-yang-dynamics-check"&gt;Yin-Yang Dynamics Check&lt;/h2&gt;
&lt;p&gt;Assess whether current strategy and state are &lt;strong&gt;Yang Excess Yin Deficiency&lt;/strong&gt; or &lt;strong&gt;Yin Excess Yang Deficiency&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="diagnosis-method-2"&gt;Diagnosis Method&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Qualitative Analysis&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Look at whether recent architecture decisions overly favor one extreme&lt;/li&gt;
&lt;li&gt;Have you been continuously expanding and adding new features while ignoring stability?&lt;/li&gt;
&lt;li&gt;Or conversely, multiple layers of approval and strict constraints but lack innovation momentum?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Metrics&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Yang Excess&lt;/th&gt;
&lt;th&gt;Yin Excess&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Change Frequency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Extremely high&lt;/td&gt;
&lt;td&gt;Extremely low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Incident Rate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frequent&lt;/td&gt;
&lt;td&gt;Extremely low but no change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Release Rhythm&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Continuous&lt;/td&gt;
&lt;td&gt;Long-term stagnation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 21: Yin-Yang Status
&lt;/figcaption&gt;
&lt;h2 id="balance-strategy"&gt;Balance Strategy&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Symptoms&lt;/th&gt;
&lt;th&gt;Solution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yang Excess Yin Deficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frequent changes with frequent incidents&lt;/td&gt;
&lt;td&gt;Pause releases, focus on addressing hazards (replenish Yin)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin Excess Yang Deficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Long-term no change and stagnation&lt;/td&gt;
&lt;td&gt;Introduce challenges and innovation (add Yang)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 22: Balance Strategies
&lt;/figcaption&gt;
&lt;h2 id="yun-alignment-check"&gt;Yun Alignment Check&lt;/h2&gt;
&lt;p&gt;Determine whether the organization&amp;rsquo;s actions match the system&amp;rsquo;s current stage, preventing &lt;strong&gt;counter-Yun operation&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="diagnosis-method-3"&gt;Diagnosis Method&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Combine Business Development and Technical Maturity&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Error Pattern&lt;/th&gt;
&lt;th&gt;Manifestation&lt;/th&gt;
&lt;th&gt;Consequences&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Premature Standardization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Spending 大量精力 on process management and cost optimization for emerging projects&lt;/td&gt;
&lt;td&gt;These are typically scale stage concerns, but the project is still in exploration stage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Counter-Yun Exploration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frequently changing underlying architecture for widely used platforms without rigorous testing&lt;/td&gt;
&lt;td&gt;Inconsistent with scaling stage&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 23: Error Patterns
&lt;/figcaption&gt;
&lt;h2 id="stage-strategy-reference-table"&gt;Stage-Strategy Reference Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Should Focus On&lt;/th&gt;
&lt;th&gt;Should Not Do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exploration Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Diversity, flexibility, rapid trial and error&lt;/td&gt;
&lt;td&gt;Premature pursuit of efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Platform Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standardization, process norms&lt;/td&gt;
&lt;td&gt;Frequent arbitrary changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scale Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Optimization, stability, efficiency&lt;/td&gt;
&lt;td&gt;Still growing wildly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rebalancing Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Transformation, breakthrough, innovation&lt;/td&gt;
&lt;td&gt;Clinging to the past&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 24: Stage-Strategy Mapping
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Checklist&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Which stage are we currently in?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Do our actions match the stage?&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Do we need to adjust strategy?&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When discovering actions don&amp;rsquo;t match the stage, immediately adjust strategy to avoid working at cross-purposes&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="yang-runaway-warning"&gt;Yang Runaway Warning&lt;/h2&gt;
&lt;p&gt;Pay special attention to whether there are signs of &lt;strong&gt;Yang state runaway&lt;/strong&gt; in the system.&lt;/p&gt;
&lt;h2 id="what-is-yang-runaway"&gt;What is Yang Runaway?&lt;/h2&gt;
&lt;p&gt;Exponential explosion or collapse risk caused by unconstrained positive feedback.&lt;/p&gt;
&lt;h2 id="typical-scenarios"&gt;Typical Scenarios&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service Call Volume Surge&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Bug or abuse → Resource strain → Queuing and retry storms → Further increase in calls&lt;/td&gt;
&lt;td&gt;Resource exhaustion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Training Task Self-Replication&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tasks unlimitedly self-replicate to accelerate → Cluster resource exhaustion&lt;/td&gt;
&lt;td&gt;System collapse&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 25: Typical Scenarios
&lt;/figcaption&gt;
&lt;h2 id="diagnosis-signals"&gt;Diagnosis Signals&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A metric shows &lt;strong&gt;exponential explosive growth&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Lack of slowing mechanisms&lt;/li&gt;
&lt;li&gt;Formation of vicious cycles&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="response-strategy"&gt;Response Strategy&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Means&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Establish Hard Limits&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Metal&amp;rsquo;s constraints&lt;/td&gt;
&lt;td&gt;Immediate shutdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Introduce Negative Feedback&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Earth&amp;rsquo;s governance (rate limiting, quotas)&lt;/td&gt;
&lt;td&gt;Braking and deceleration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Break Positive Feedback Chain&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Activate emergency plan&lt;/td&gt;
&lt;td&gt;Pull back to steady state&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 26: Response Strategies
&lt;/figcaption&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When discovering a metric showing exponential explosive growth without slowing mechanisms, intervene immediately&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="diagnosis-implementation-process"&gt;Diagnosis Implementation Process&lt;/h2&gt;
&lt;h2 id="regular-diagnosis-mechanism"&gt;Regular Diagnosis Mechanism&lt;/h2&gt;
&lt;p&gt;Recommend establishing a periodic diagnosis process:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-infra-dao/system-diagnosis/884491832caca1efa808558d60e611b9.svg" data-img="https://assets.jimmysong.io/images/book/ai-infra-dao/system-diagnosis/884491832caca1efa808558d60e611b9.svg" alt="Figure 6: Regular Diagnosis Mechanism Flowchart" data-caption="Figure 6: Regular Diagnosis Mechanism Flowchart"
width="3418"
height="361"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Regular Diagnosis Mechanism Flowchart&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="diagnosis-meeting-agenda"&gt;Diagnosis Meeting Agenda&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Fixed Session of Weekly Operations Review Meeting&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Check Five Elements scores for each module&lt;/li&gt;
&lt;li&gt;Browse global Qi flow diagram&lt;/li&gt;
&lt;li&gt;Analyze Yin-Yang dynamics&lt;/li&gt;
&lt;li&gt;Discuss current Yun&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;This systematic examination makes hidden risks 无处遁形，thus achieving prevention before problems occur&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="diagnosis-action-matrix"&gt;Diagnosis Action Matrix&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Diagnosis Result&lt;/th&gt;
&lt;th&gt;Action Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements: One Element Too Weak&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Concentrate resources to strengthen the weakness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements: One Element Overloaded&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Expand capacity or introduce constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi Stagnation at One Stage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Clear bottlenecks, optimize processes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yang Excess Yin Deficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Strengthen governance and stability mechanisms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin Excess Yang Deficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Activate innovation and boost vitality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Counter-Yun Operation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Adjust strategy and go with the flow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yang Runaway Warning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Immediate intervention, break positive feedback&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 27: Diagnosis Action Matrix
&lt;/figcaption&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Through the above diagnosis principles, architects and operations teams can periodically take the pulse of infrastructure like TCM pulse diagnosis.&lt;/p&gt;
&lt;p&gt;When diagnosis indicates imbalance in some aspect, immediately prescribe remedy based on the theory: &lt;strong&gt;replenish what needs replenishing, purge what needs purging&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Long-term adherence will keep the system on a healthy evolutionary trajectory.&lt;/p&gt;</content:encoded></item><item><title>Conclusion and Outlook</title><link>https://jimmysong.io/book/ai-infra-dao/summary/</link><pubDate>Tue, 10 Feb 2026 13:56:22 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-infra-dao/summary/</guid><description>Core value and applications of the Yin-Yang Five Elements Qi-Yun model for AI infrastructure architects</description><content:encoded>
&lt;p&gt;This paper systematically presents the four-layer model of &amp;ldquo;Yin-Yang - Five Elements - Yun - Qi&amp;rdquo; for AI infrastructure, providing a comprehensive cognitive map from theory to practice.&lt;/p&gt;
&lt;h2 id="review-of-theoretical-model"&gt;Review of Theoretical Model&lt;/h2&gt;
&lt;p&gt;Through four dimensions, we have constructed a global framework for understanding AI infrastructure:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Core Value&lt;/th&gt;
&lt;th&gt;Key Insights&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yin-Yang&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Understanding the tension and balance within systems&lt;/td&gt;
&lt;td&gt;Expansion and constraint, innovation and governance, speed and stability—these three are opposites yet unified, all indispensable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Five Elements&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Organizing the fundamental role elements of systems&lt;/td&gt;
&lt;td&gt;Data, models, computing power, platforms, hardware—these five generate and restrain each other in endless cycles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Yun&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Grasping the periodic patterns of system evolution&lt;/td&gt;
&lt;td&gt;Exploration phase, platform phase, scale phase, rebalancing phase—act in accordance with the trends&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qi&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Insight into the flow state of system operation&lt;/td&gt;
&lt;td&gt;When Qi flows, the system is active; when Qi stagnates, the system becomes pathological&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Layer, Core Value, and Key Insights
&lt;/figcaption&gt;
&lt;p&gt;More importantly, we have demonstrated how this theory combining Eastern wisdom with engineering practice can provide insights and guidance for real-world problems such as GPU scheduling, Agent Runtime, and platform governance.&lt;/p&gt;
&lt;h2 id="core-value-of-the-model"&gt;Core Value of the Model&lt;/h2&gt;
&lt;h2 id="holistic-view"&gt;Holistic View&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Traditional fragmented perspectives often see trees but not the forest, making it difficult to provide timely warnings of systemic risks&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The Yin-Yang Five Elements Qi-Yun model, with its &lt;strong&gt;holistic view&lt;/strong&gt;, helps architects:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Break free from the constraints of pure technical metrics&lt;/li&gt;
&lt;li&gt;Grasp the principal contradictions and driving forces of system evolution&lt;/li&gt;
&lt;li&gt;Extract meaningful patterns from complex signals&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dynamic-view"&gt;Dynamic View&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The value of a system lies not in pursuing the extreme of a single performance indicator without limit, but in balancing all elements to achieve long-term coordinated development&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The model&amp;rsquo;s &lt;strong&gt;dynamic view&lt;/strong&gt; reminds us:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Yin-Yang dynamics transform dynamically with environment and stage&lt;/li&gt;
&lt;li&gt;The same capability may shift from advantage to risk at different stages&lt;/li&gt;
&lt;li&gt;Strategies need timely adjustment as Yun changes&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="balance-view"&gt;Balance View&lt;/h2&gt;
&lt;p&gt;The core philosophy of the model is &lt;strong&gt;balance&lt;/strong&gt; rather than extreme:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not pursuing the limit of a single metric&lt;/li&gt;
&lt;li&gt;But pursuing system coordination and sustainability&lt;/li&gt;
&lt;li&gt;Finding dynamic balance points within unity of opposites&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practical-application-value"&gt;Practical Application Value&lt;/h2&gt;
&lt;h2 id="during-architecture-design"&gt;During Architecture Design&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Consider the completeness and balance of the Five Elements&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Reserve Yin-Yang constraint mechanisms&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Design evolution paths that align with Yun trends&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Plan channels for Qi flow&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="during-operations-and-governance"&gt;During Operations and Governance&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Regularly check Five Elements balance&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Monitor Qi circulation status&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Assess Yin-Yang dynamic changes&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Determine Yun phase transitions&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Provide early warning of Yang loss-of-control risks&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="during-decision-review"&gt;During Decision Review&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Analyze root causes from the four-layer model perspective&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Check whether basic principles of any layer were violated&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Develop systematic solutions&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Establish long-term improvement mechanisms&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="insights-for-architects"&gt;Insights for Architects&lt;/h2&gt;
&lt;p&gt;In an era of flourishing large models and autonomous agents, infrastructure has become unprecedentedly complex and active.&lt;/p&gt;
&lt;h2 id="cognitive-upgrade"&gt;Cognitive Upgrade&lt;/h2&gt;
&lt;p&gt;From &amp;ldquo;managing machines and applications&amp;rdquo; to &amp;ldquo;managing intelligence and knowledge&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not only focus on application logic itself&lt;/li&gt;
&lt;li&gt;But more on how knowledge and intelligence integrate into systems&lt;/li&gt;
&lt;li&gt;View models as dynamically evolving components&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="mindset-shift"&gt;Mindset Shift&lt;/h2&gt;
&lt;p&gt;From single-metric optimization to system balance:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not pursuing the extreme of a single element&lt;/li&gt;
&lt;li&gt;But pursuing overall coordination and sustainability&lt;/li&gt;
&lt;li&gt;Finding dynamic balance within unity of opposites&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="capability-development"&gt;Capability Development&lt;/h2&gt;
&lt;p&gt;From technical expert to systems philosopher:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;While mastering technical tools&lt;/li&gt;
&lt;li&gt;Cultivate systems thinking and philosophical reflection&lt;/li&gt;
&lt;li&gt;Apply holistic frameworks like Yin-Yang and Five Elements&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="limitations-of-the-model"&gt;Limitations of the Model&lt;/h2&gt;
&lt;p&gt;It must be noted that this theory is not a panacea:&lt;/p&gt;
&lt;h2 id="not-a-rigid-formula"&gt;Not a Rigid Formula&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Its value lies not in providing a rigid formula, but in guiding us to return to reality and think about problems from a more comprehensive perspective&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;The model provides a thinking framework, not standard answers&lt;/li&gt;
&lt;li&gt;Specific applications need to consider actual scenarios&lt;/li&gt;
&lt;li&gt;Architects ultimately must make judgments based on specific context&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="requires-continuous-validation"&gt;Requires Continuous Validation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Theory needs continuous validation and refinement in practice&lt;/li&gt;
&lt;li&gt;Different scenarios may require adjustment and extension&lt;/li&gt;
&lt;li&gt;Feedback and improvement in practice are encouraged&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="supplement-not-replace"&gt;Supplement, Not Replace&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The model is a tool to assist decision-making&lt;/li&gt;
&lt;li&gt;Cannot replace professional judgment and experience&lt;/li&gt;
&lt;li&gt;Should be used in combination with other methodologies&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="future-outlook"&gt;Future Outlook&lt;/h2&gt;
&lt;h2 id="theory-development"&gt;Theory Development&lt;/h2&gt;
&lt;p&gt;This model has significant room for development:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Quantitative Metrics&lt;/strong&gt;: Develop more precise quantitative indicators to make the theory more actionable&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool Support&lt;/strong&gt;: Develop analysis tools and automated diagnostic systems based on the model&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Case Accumulation&lt;/strong&gt;: Collect more practical cases to validate and enrich the theory&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-Domain Application&lt;/strong&gt;: Explore applications of the model in other complex system domains&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice-promotion"&gt;Practice Promotion&lt;/h2&gt;
&lt;p&gt;We hope this framework can help:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CTOs, infrastructure architects, and platform R&amp;amp;D teams&lt;/li&gt;
&lt;li&gt;When facing increasingly complex AI infrastructure&lt;/li&gt;
&lt;li&gt;Make wiser decisions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="ultimate-vision"&gt;Ultimate Vision&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Standing with sword in the midst of waves of change, embracing both the Yang of innovation and the Yin of governance, riding the system&amp;rsquo;s Qi above the currents&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;AI infrastructure stands at the starting point of a new era. We need not only technological innovation but also conceptual innovation.&lt;/p&gt;
&lt;p&gt;The Yin-Yang Five Elements Qi-Yun model offers a unique perspective—combining Eastern philosophical wisdom with modern engineering practice—helping us find simplicity in complexity, stability in change, and unity in opposition.&lt;/p&gt;
&lt;p&gt;We hope this model becomes a powerful tool for your thinking about AI infrastructure, helping you find your own &amp;ldquo;Way&amp;rdquo; in the balance and evolution of systems.&lt;/p&gt;</content:encoded></item><item><title>What Is AI-Native Infrastructure?</title><link>https://jimmysong.io/book/ai-native-infra/definition/</link><pubDate>Sun, 18 Jan 2026 05:43:57 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/definition/</guid><description>Core definition, boundaries, and evaluation criteria for AI-native infrastructure, focusing on model behavior, compute scarcity, and uncertainty governance.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The essence of AI-native infrastructure is to make model behavior, compute scarcity, and uncertainty governable system boundaries.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;AI-native infrastructure is not a simple checklist of technologies, but rather a new operating order designed for a world where &amp;ldquo;models become actors, compute becomes scarce, and uncertainty is the system default.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The core of AI-native infrastructure is not faster inference or cheaper GPUs, but providing governable, measurable, and evolvable system boundaries for model behavior, compute scarcity, and uncertainty—making AI systems deliverable, governable, and evolvable in production environments.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="why-we-need-a-more-rigorous-definition"&gt;Why We Need a More Rigorous Definition&lt;/h2&gt;
&lt;p&gt;The term &amp;ldquo;AI-native infrastructure/architecture&amp;rdquo; is being adopted by an increasing number of vendors, but its meaning is often oversimplified as &amp;ldquo;data centers better suited for AI&amp;rdquo; or &amp;ldquo;more complete AI platform delivery.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;In practice, different vendors emphasize different aspects of AI-native infrastructure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cisco&lt;/strong&gt; emphasizes delivering AI-native infrastructure across &lt;strong&gt;edge/cloud/data center&lt;/strong&gt; domains, highlighting delivery paths where &amp;ldquo;open &amp;amp; disaggregated&amp;rdquo; and &amp;ldquo;fully integrated systems&amp;rdquo; coexist (e.g., Cisco Validated Designs).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HPE&lt;/strong&gt; emphasizes an &lt;strong&gt;open, full-stack AI-native architecture&lt;/strong&gt; for the entire AI lifecycle, model development, and deployment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NVIDIA&lt;/strong&gt; explicitly proposes an &lt;strong&gt;AI-native infrastructure tier&lt;/strong&gt; to support inference context reuse for long-context and agentic workloads.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For CTOs/CEOs, a definition that can guide strategy and organizational design must meet two criteria:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Clarify &lt;strong&gt;how the first-principles constraints of infrastructure have changed in the AI era&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Converge &amp;ldquo;AI-native&amp;rdquo; from a marketing adjective into &lt;strong&gt;verifiable architectural properties and operating mechanisms&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="authoritative-one-sentence-definition"&gt;Authoritative One-Sentence Definition&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;AI-native infrastructure&lt;/strong&gt; is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An infrastructure system and operating mechanism premised on &amp;ldquo;models/agents as execution subjects, compute as scarce assets, and uncertainty as the norm,&amp;rdquo; which closes the loop on &amp;ldquo;intent (API/Agent) → execution (Runtime) → resource consumption (Accelerator/Network/Storage) → economic and risk outcomes&amp;rdquo; through compute governance.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This definition contains two layers of meaning:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;: Not just a software/hardware stack, but also includes scaled delivery and systemic capabilities (consistent with vendors&amp;rsquo; emphasis on &amp;ldquo;full-stack integration/reference architectures/lifecycle delivery&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Operating Model&lt;/strong&gt;: It inevitably rewrites organizational and operational methods, not just a technical upgrade—budget, risk, and release rhythm are strongly bound to the same governance loop.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="a-more-precise-layered-model-paradigm-above-refactoring-below"&gt;A More Precise Layered Model: Paradigm Above, Refactoring Below&lt;/h2&gt;
&lt;p&gt;AI-native infrastructure should not be framed as a replacement of cloud-native infrastructure, but as a combined model of &lt;strong&gt;higher-level paradigm + downward refactoring&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Layer 0 — Cloud-Native Substrate (retained)&lt;/strong&gt;: Kubernetes, container runtime, CNI/CSI, and observability primitives remain the execution substrate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Layer 1 — Resource Model Rewrite (core breakpoint)&lt;/strong&gt;: Scheduling semantics shift from CPU/memory-centric to GPU/token/context-centric; context and token become first-class resources. GPU virtualization, slicing, sharing, and oversubscription (such as HAMi) belong here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Layer 2 — Runtime and Control Plane Shift&lt;/strong&gt;: Workload form shifts from request/response services to agentic loops, async multi-step workflows, and fused inference/training/evaluation execution paths.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Layer 3 — Governance and Scheduling First&lt;/strong&gt;: Scheduling becomes the primary control plane; deployment becomes a secondary concern under budget, risk, and policy constraints.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words: &lt;strong&gt;Cloud Native is the substrate, AI Native is the semantic override and resource-model shift on top of that substrate.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="three-premises"&gt;Three Premises&lt;/h2&gt;
&lt;p&gt;The core premises of AI-native infrastructure are as follows. The diagram below illustrates the correspondence between these three premises and governance boundaries.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/definition/ai-native-infra-constitution-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/definition/ai-native-infra-constitution-en.svg" alt="Figure 1: Three constitutional premises of AI-native infrastructure" data-caption="Figure 1: Three constitutional premises of AI-native infrastructure"
width="1576"
height="696"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Three constitutional premises of AI-native infrastructure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model-as-Actor&lt;/strong&gt;: Models/agents become &amp;ldquo;execution subjects&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compute-as-Scarcity&lt;/strong&gt;: Compute (accelerators, interconnects, power consumption, bandwidth) becomes the core scarce asset&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Uncertainty-by-Default&lt;/strong&gt;: Behavior and resource consumption are highly uncertain (especially in agentic and long-context scenarios)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These three points collectively determine: the core task of AI-native infrastructure is not to &amp;ldquo;make systems more elegant,&amp;rdquo; but to &lt;strong&gt;make systems controllable, sustainable, and capable of scaled delivery under uncertain behavior.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="boundaries-what-ai-native-infrastructure-manages-and-what-it-doesnt"&gt;Boundaries: What AI-Native Infrastructure Manages and What It Doesn&amp;rsquo;t&lt;/h2&gt;
&lt;p&gt;In practical engineering, defining boundaries helps focus resources and capability development. The table below summarizes what AI-native infrastructure focuses on versus what it doesn&amp;rsquo;t:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Not focused on:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Prompt design and business-level agent logic&lt;/li&gt;
&lt;li&gt;Individual model capabilities and training secrets&lt;/li&gt;
&lt;li&gt;Application-layer product features themselves&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Focused on:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Compute Governance&lt;/strong&gt;: Quotas, budgets, isolation/sharing, topology and interconnects, preemption and priorities, throughput/latency versus cost tradeoffs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Execution Form Engineering&lt;/strong&gt;: Unified operation, scheduling, and observability for training/fine-tuning/inference/batch processing/agentic workflows&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Closed-Loop Mechanisms&lt;/strong&gt;: How intent is constrained, measured, and mapped to controllable resource consumption and economic/risk outcomes&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="verifiable-architectural-properties-three-planes--one-loop"&gt;Verifiable Architectural Properties: Three Planes + One Loop&lt;/h2&gt;
&lt;p&gt;To facilitate understanding, the following sections introduce the core architectural properties of AI-native infrastructure.&lt;/p&gt;
&lt;p&gt;The diagram below shows the visualization of the three planes and the closed loop, facilitating rapid boundary alignment during reviews.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/definition/three-planes-one-loop-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/definition/three-planes-one-loop-en.svg" alt="Figure 2: Three Planes and One Loop reference architecture" data-caption="Figure 2: Three Planes and One Loop reference architecture"
width="1576"
height="476"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Three Planes and One Loop reference architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Three Planes:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Intent Plane&lt;/strong&gt;: APIs, MCP, Agent workflows, policy expressions&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Execution Plane&lt;/strong&gt;: Training/inference/serving/runtime (including tool calls and state management)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance Plane&lt;/strong&gt;: Accelerator orchestration, isolation/sharing, quotas/budgets, SLO and cost control, risk policies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The Loop:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Only with an &amp;ldquo;intent → consumption → cost/risk outcome&amp;rdquo; closed loop can it be called AI-native.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is also why NVIDIA elevates the sharing and reuse of &amp;ldquo;new state assets&amp;rdquo; like inference context to an independent AI-native infrastructure layer: essentially bringing the resource consequences of agentic/long-context into governable system boundaries.&lt;/p&gt;
&lt;h2 id="ai-native-vs-cloud-native-where-the-differences-lie"&gt;AI-Native vs Cloud Native: Where the Differences Lie&lt;/h2&gt;
&lt;p&gt;Cloud Native focuses on delivering services in distributed environments with portability, elasticity, observability, and automation. Its governance objects are primarily &lt;strong&gt;service/instance/request&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;AI-native infrastructure addresses a different set of structural problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Execution unit shift&lt;/strong&gt;: From service request/response to agent action/decision/side effect&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resource constraint shift&lt;/strong&gt;: From elastic CPU/memory to hard GPU/throughput/token constraints and cost ceilings&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reliability pattern shift&lt;/strong&gt;: From &amp;ldquo;reliable delivery of deterministic systems&amp;rdquo; to &amp;ldquo;controllable operation of non-deterministic systems&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, AI-native is not a parallel track that replaces cloud-native, but a reinterpretation of cloud-native under AI constraints: the substrate remains, while resource, runtime, and governance semantics are rewritten.&lt;/p&gt;
&lt;h2 id="bringing-it-to-engineering-what-capabilities-ai-native-infrastructure-must-have"&gt;Bringing It to Engineering: What Capabilities AI-Native Infrastructure Must Have&lt;/h2&gt;
&lt;p&gt;To avoid &amp;ldquo;right concept, misaligned execution,&amp;rdquo; the following minimum closed-loop capabilities are listed.&lt;/p&gt;
&lt;h3 id="resource-model-making-gpu-context-and-token-first-class-resources"&gt;Resource Model: Making GPU, Context, and Token First-Class Resources&lt;/h3&gt;
&lt;p&gt;Cloud native abstracts CPU/memory into schedulable resources; AI-native must further bring the following resources under governance:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPU/Accelerator Resources&lt;/strong&gt;: Scheduled and governed by partitioning, sharing, isolation, and preemption&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context Resources&lt;/strong&gt;: Context windows, retrieval paths, cache hits, KV/inference state asset reuse, etc., which directly affect tokens and costs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Token/Throughput&lt;/strong&gt;: Become measurable capacity and cost carriers (can enter budgets, SLOs, and product strategies)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When tokens become &amp;ldquo;capacity units,&amp;rdquo; the platform is no longer just running services, but operating an &amp;ldquo;AI factory.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="budgets-and-policies-binding-costrisk-to-organizational-decisions"&gt;Budgets and Policies: Binding &amp;ldquo;Cost/Risk&amp;rdquo; to Organizational Decisions&lt;/h3&gt;
&lt;p&gt;AI systems cannot operate with a &amp;ldquo;ship and done&amp;rdquo; approach. Budgets and policies must become the control plane:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Trigger rate limiting/degradation when budgets are exceeded&lt;/li&gt;
&lt;li&gt;Trigger stricter verification or disable high-risk tools when risk increases&lt;/li&gt;
&lt;li&gt;Version releases and experiments are constrained by &amp;ldquo;budget/risk headroom&amp;rdquo; (institutionalizing release rhythm)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key is &lt;strong&gt;infrastructure solidifying organizational rules into executable policies&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="observability-and-audit-making-model-behavior-accountable-and-observable"&gt;Observability and Audit: Making Model Behavior Accountable and Observable&lt;/h3&gt;
&lt;p&gt;Traditional observability focuses on latency/error/traffic; AI-native must add at least three types of signals:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Behavior Signals&lt;/strong&gt;: Which tools the model called, which systems it read/wrote, what actions it took, what side effects it caused&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost Signals&lt;/strong&gt;: Tokens, GPU time, cache hits, queue wait, interconnect bottlenecks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality and Safety Signals&lt;/strong&gt;: Output quality, violation/over-privilege risks, rollback frequency and reasons&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Without &amp;ldquo;behavior observability,&amp;rdquo; governance cannot be implemented.&lt;/p&gt;
&lt;h3 id="risk-governance-bringing-high-risk-capabilities-under-continuous-assessment-and-control"&gt;Risk Governance: Bringing High-Risk Capabilities Under Continuous Assessment and Control&lt;/h3&gt;
&lt;p&gt;When model capabilities approach thresholds that can &amp;ldquo;cause serious harm,&amp;rdquo; organizations need a systematic risk governance framework, not relying on single-point prompts or manual reviews.&lt;/p&gt;
&lt;p&gt;Can be split into two layers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;System-Level Trustworthiness Goals&lt;/strong&gt;: Organizational-level requirements for security, transparency, explainability, and accountability&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Frontier Capability Readiness Assessment&lt;/strong&gt;: Tiered assessment of high-risk capabilities, launch thresholds, and mitigation measures&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The value lies in: transforming &amp;ldquo;safety/risk&amp;rdquo; from concepts into executable launch thresholds and operational policies.&lt;/p&gt;
&lt;h2 id="takeaways--checklist"&gt;Takeaways / Checklist&lt;/h2&gt;
&lt;p&gt;The following checklist can be used to determine whether an organization has entered the AI-native stage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Do we treat models as &amp;ldquo;agents that act,&amp;rdquo; not as replaceable APIs?&lt;/li&gt;
&lt;li&gt;Do we bring compute and budgets into business SLAs and decision processes?&lt;/li&gt;
&lt;li&gt;Do we treat uncertainty as the default premise, not as an exception?&lt;/li&gt;
&lt;li&gt;Do we have audit, rollback, and accountability for model behavior?&lt;/li&gt;
&lt;li&gt;Do we have cross-team AI governance mechanisms, not single-point engineering optimizations?&lt;/li&gt;
&lt;li&gt;Can we explain the system&amp;rsquo;s operating boundaries, cost boundaries, and risk boundaries?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The essence of AI-native infrastructure lies in: taking models as behavior subjects, compute as scarce assets, and uncertainty as the norm, achieving deliverable, governable, and evolvable AI systems through governance and closed-loop mechanisms. Only by engineering these capabilities can organizations truly step into the AI-native stage.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.cisco.com/site/us/en/solutions/artificial-intelligence/infrastructure/index.html" target="_blank" rel="noopener"&gt;Cisco AI-Native Infrastructure - cisco.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.hpe.com/us/en/newsroom/blog-post/2023/12/introducing-an-ai-native-architecture-for-ai-driven-transformation.html" target="_blank" rel="noopener"&gt;HPE AI-native architecture - hpe.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/" target="_blank" rel="noopener"&gt;NVIDIA Rubin: AI-native infrastructure tier - developer.nvidia.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://lfnetworking.org/40362-2/" target="_blank" rel="noopener"&gt;LF Networking: becoming AI-native is a redefinition of the operating model - lfnetworking.org&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.nist.gov/itl/ai-risk-management-framework" target="_blank" rel="noopener"&gt;NIST AI Risk Management Framework - nist.gov&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sre.google/workbook/error-budget-policy/" target="_blank" rel="noopener"&gt;Google SRE Workbook - Error Budgets - sre.google&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/safety/preparedness/" target="_blank" rel="noopener"&gt;OpenAI Preparedness Framework - openai.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>AI-Native Infrastructure One-Page Reference Architecture: Three Planes + One Loop</title><link>https://jimmysong.io/book/ai-native-infra/reference-architecture/</link><pubDate>Sun, 18 Jan 2026 00:00:00 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/reference-architecture/</guid><description>Three planes (Intent, Execution, Governance) + closed-loop feedback for AI-native infrastructure architecture alignment.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The true value of architecture is enabling organizational consensus on complex systems within five minutes, not creating another new technology stack.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Industry-leading vendors emphasize different aspects: Cisco focuses more on AI-native infrastructure and reference designs such as Cisco Validated Designs; HPE emphasizes an open, full-stack AI-native architecture across the AI full lifecycle; NVIDIA explicitly proposes adding a new AI-native infrastructure tier for inference context reuse in long-context and agentic workloads. This chapter converges these perspectives into a &lt;strong&gt;verifiable architecture framework&lt;/strong&gt;: Three Planes + One Loop.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Note
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
The &amp;ldquo;architecture&amp;rdquo; in this chapter serves as a review framework, not a component checklist. The goal is to unify organizational language and review boundaries, not to reinvent the technology stack.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="one-page-architecture-overview"&gt;One-Page Architecture Overview&lt;/h2&gt;
&lt;p&gt;Below is the detailed reference architecture diagram for the three planes and closed loop of AI-native infrastructure, helping readers quickly establish an overall understanding:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/reference-architecture/three-planes-detailed-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/reference-architecture/three-planes-detailed-en.svg" alt="Figure 1: Three planes detailed architecture diagram" data-caption="Figure 1: Three planes detailed architecture diagram"
width="1783"
height="1485"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Three planes detailed architecture diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This diagram can be understood as: &lt;strong&gt;The new control plane (Intent) of AI-native infrastructure must be constrained by the Governance Plane and produce measurable resource consequences in the Execution Plane.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is also the common ground behind different vendor narratives: Cisco uses reference designs and delivery frameworks to make infrastructure capabilities scalable and replicable; HPE uses open/full-stack to cover lifecycle delivery; NVIDIA elevates the reuse of &amp;ldquo;context state assets&amp;rdquo; to an independent infrastructure layer. All three point to the same issue: &lt;strong&gt;incorporating AI resource consequences into governable system boundaries.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="core-capabilities-of-the-three-planes"&gt;Core Capabilities of the Three Planes&lt;/h2&gt;
&lt;p&gt;This section details the core capabilities of each of the three planes to help clarify focus areas during architecture reviews.&lt;/p&gt;
&lt;h3 id="intent-plane"&gt;Intent Plane&lt;/h3&gt;
&lt;p&gt;The Intent Plane is responsible for expressing &amp;ldquo;what I want,&amp;rdquo; including the following capabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Inference/Training APIs (entry points and contracts)&lt;/li&gt;
&lt;li&gt;MCP/Tool calling protocols and tool catalog (standardizing tool access as &amp;ldquo;declaratable capability boundaries&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;Agent/Workflow (breaking down tasks into executable steps)&lt;/li&gt;
&lt;li&gt;Policy as Intent: priorities, budgets, quotas, compliance/security constraints (front-loaded in the form of &amp;ldquo;intent&amp;rdquo;)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Point&lt;/strong&gt;: The Intent Plane is not the starting point itself; the real starting point is—&lt;strong&gt;whether intent can be translated into executable and governable plans.&lt;/strong&gt; Otherwise, Agents/MCP will only amplify uncertainty: more tools, longer chains, larger state spaces, and more uncontrollable resource consumption.&lt;/p&gt;
&lt;p&gt;During architecture reviews, focus on these questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Is intent declarable (contract) and rejectable (admission)?&lt;/li&gt;
&lt;li&gt;Does intent carry budget/priority/compliance constraints (policy as intent)?&lt;/li&gt;
&lt;li&gt;Is the translation from intent to execution traceable?&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="execution-plane"&gt;Execution Plane&lt;/h3&gt;
&lt;p&gt;The Execution Plane is responsible for landing intent into &amp;ldquo;actual execution,&amp;rdquo; mainly including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training, fine-tuning, inference serving, batch processing, agentic runtime&lt;/li&gt;
&lt;li&gt;&amp;ldquo;State and Context&amp;rdquo; services: cache/KV/vector/context memory, etc., for carrying inference context, retrieval results, and session state&lt;/li&gt;
&lt;li&gt;Full-chain observability hooks: token metering, GPU time, video memory, network traffic, storage I/O, queue wait times, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;An industry trend worth emphasizing: as long-context and agentic workloads become widespread, &amp;ldquo;context&amp;rdquo; itself becomes a critical state asset and may even rise to become an independent infrastructure layer. NVIDIA explicitly proposes inference context memory storage in the Rubin platform, establishing an AI-native infrastructure tier to provide shared, low-latency inference context at the pod level to support reuse (for long-context and agentic workloads).&lt;/p&gt;
&lt;p&gt;Review points focus on three things:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Is execution measurable&lt;/strong&gt;: Can attribution be done across token/GPU/network/storage dimensions?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Is state governable&lt;/strong&gt;: What are the lifecycle, reuse boundaries, and isolation strategies for context and cache?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Is observability closed-loop oriented&lt;/strong&gt;: Observability is not for &amp;ldquo;seeing,&amp;rdquo; but for &amp;ldquo;enabling governance to correct deviations.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="governance-plane"&gt;Governance Plane&lt;/h3&gt;
&lt;p&gt;The Governance Plane is the &amp;ldquo;core differentiator&amp;rdquo; of AI-native infrastructure, responsible for transforming resource scarcity and uncertainty into a controllable system:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Budget/quotas/billing: governing consumption across teams, tenants, projects, models, and agent tasks&lt;/li&gt;
&lt;li&gt;Isolation and sharing strategies: same-card sharing, video memory isolation, preemption, priorities, fairness&lt;/li&gt;
&lt;li&gt;Topology-aware scheduling: incorporating GPU, interconnect, network, and storage topology into placement (especially in training and high-throughput inference)&lt;/li&gt;
&lt;li&gt;Risk and compliance control: audits, policy enforcement points, sensitive data and access control&lt;/li&gt;
&lt;li&gt;Integration with FinOps/SRE/SecOps: incorporating cost, reliability, and risk into a single operational mechanism&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a vendor narrative perspective, this layer typically corresponds to &amp;ldquo;reference architecture + full-stack delivery&amp;rdquo;:
Cisco emphasizes accelerating and scaling delivery in AI infrastructure through &amp;ldquo;fully integrated systems + Cisco Validated Designs&amp;rdquo;; HPE emphasizes end-to-end delivery with &amp;ldquo;open, full-stack AI-native architecture&amp;rdquo; to support model development and deployment.&lt;/p&gt;
&lt;p&gt;The baseline question for Governance Plane reviews is: &lt;strong&gt;Can you make explainable resource allocation and degradation decisions under budget/risk constraints?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="closed-loop-mechanism-explained"&gt;Closed-Loop Mechanism Explained&lt;/h2&gt;
&lt;p&gt;This section introduces the core workflow of the closed-loop mechanism to help understand the essential difference between AI-native and AI-ready.&lt;/p&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
Significance of the Closed-Loop Mechanism
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
The closed loop is the most confusing yet critical dividing line between AI-native and &amp;ldquo;AI-ready.&amp;rdquo;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The minimal implementation of the closed loop includes four steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Admission&lt;/strong&gt;: Bind intent with policy at the entry point (budget, priority, compliance)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Translation&lt;/strong&gt;: Translate intent into executable plans (select runtime, resource specifications, topology preferences)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metering&lt;/strong&gt;: End-to-end metering and attribution across tokens/GPU/network/storage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enforcement&lt;/strong&gt;: Budget triggers degradation/rate limiting/preemption; risk triggers isolation/audits; SLO triggers scaling/routing&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In other words: &lt;strong&gt;The closed loop is not a &amp;ldquo;monitoring dashboard,&amp;rdquo; but a &amp;ldquo;governance-driven real-time correction mechanism.&amp;rdquo;&lt;/strong&gt;
If there is no closed loop for &amp;ldquo;intent → consumption → cost/risk outcomes,&amp;rdquo; systems can easily spin out of control across cost, risk, quality, and other dimensions.&lt;/p&gt;
&lt;p&gt;This is also why &amp;ldquo;AI-native&amp;rdquo; is often accompanied by changes in operating model: when system execution speed and resource consumption are amplified by models/agents, organizations must front-load governance mechanisms and institutionalize them. LF Networking also explicitly points out: becoming AI-native is not just a technical migration, but a redefinition of the operating model.&lt;/p&gt;
&lt;h2 id="practical-usage-of-the-one-page-architecture"&gt;Practical Usage of the One-Page Architecture&lt;/h2&gt;
&lt;p&gt;In subsequent chapters, this &amp;ldquo;one-page architecture&amp;rdquo; can be repeatedly reused as a review template:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Discussing MCP/Agent: Position them in the &lt;strong&gt;Intent Plane&lt;/strong&gt; and constrain with the closed loop (admission/translation) to avoid &amp;ldquo;intent proliferation&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Discussing runtime and platforms: Place in the &lt;strong&gt;Execution Plane&lt;/strong&gt;, focusing on observable, attributable, governable state assets (context/cache/KV/vector)&lt;/li&gt;
&lt;li&gt;Discussing GPUs, scheduling, costs: Ground in the &lt;strong&gt;Governance Plane&lt;/strong&gt;, using budget/isolation/topology/metering as leverage points&lt;/li&gt;
&lt;li&gt;Discussing enterprise implementation: Use the &lt;strong&gt;closed loop&lt;/strong&gt; to examine if it&amp;rsquo;s &amp;ldquo;truly AI-native&amp;rdquo; (whether cost/risk outcomes can be written back as executable policies)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you can only remember one sentence:
&lt;strong&gt;The determination of AI-native is not in &amp;ldquo;how many AI components are used,&amp;rdquo; but in &amp;ldquo;whether there exists an executable governance closed loop that constrains intent to controllable resource consequences and economic/risk outcomes.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The one-page reference architecture provides a unified systems language and review framework for AI-native infrastructure. Through the three planes of Intent, Execution, and Governance, combined with the closed-loop mechanism, organizations can achieve efficient collaboration in architecture design, resource governance, and risk control. Looking ahead, as AI-native capabilities continue to mature, the governance closed loop will become a core competitive advantage for enterprises implementing AI.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/products/ai-enterprise/" target="_blank" rel="noopener"&gt;NVIDIA AI Enterprise Reference Architecture - nvidia.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/ai-infrastructure" target="_blank" rel="noopener"&gt;Google Cloud AI Infrastructure - cloud.google.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aws.amazon.com/architecture/well-architected/" target="_blank" rel="noopener"&gt;AWS Well-Architected Framework - aws.amazon.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Why Start with Compute Governance, Not API Design</title><link>https://jimmysong.io/book/ai-native-infra/compute-governance/</link><pubDate>Sun, 18 Jan 2026 05:45:13 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/compute-governance/</guid><description>Discussing Intent vs Consequence, why compute and cost are the first-order constraints of AI-native infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Compute and governance boundaries are the true foundation of AI-native infrastructure architecture.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The previous chapter presented a &amp;ldquo;Three Planes + One Closed Loop&amp;rdquo; reference architecture. This chapter focuses on a core CTO/CEO-level question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How should AI-native infrastructure be layered? What belongs in the &amp;ldquo;control plane&amp;rdquo; of APIs/Agents, what belongs in the &amp;ldquo;execution plane&amp;rdquo; of runtime, and what must be pushed down to the &amp;ldquo;governance plane (compute and economic constraints)&amp;rdquo;?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This question is critical because over the past year, many platform companies &amp;ldquo;pivoting to AI&amp;rdquo; have fallen into a common trap: &lt;strong&gt;treating AI as an API morphology change rather than a system constraint change&lt;/strong&gt;. When your system shifts from &amp;ldquo;serving requests&amp;rdquo; to &amp;ldquo;model behavior&amp;rdquo; (multi-step Agent actions with side effects), what truly determines system boundaries is often not the elegance of API design, but rather: &lt;strong&gt;whether compute, context, and economic constraints are institutionalized as enforceable governance boundaries&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The core argument of this chapter can be summarized as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI-native infrastructure must be designed starting from &amp;ldquo;Consequence&amp;rdquo; rather than stacking capabilities from &amp;ldquo;Intent&amp;rdquo;; the control plane is responsible for expressing intent, but the governance plane is responsible for bounding consequences.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="the-purpose-of-layering-engineering-the-binding-between-intent-and-resource-consequences"&gt;The Purpose of Layering: Engineering the Binding Between &amp;ldquo;Intent&amp;rdquo; and &amp;ldquo;Resource Consequences&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;In AI-native infrastructure, mechanisms like MCP, Agents, and Tool Calling enhance system capabilities while also introducing higher risks. These risks are not abstract &amp;ldquo;uncontrollability,&amp;rdquo; but rather engineering &amp;ldquo;unbudgetable consequences&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Path explosion in behavior, long contexts, and multi-round reasoning bring long-tail resource consumption;&lt;/li&gt;
&lt;li&gt;The same &amp;ldquo;intent&amp;rdquo; can lead to orders-of-magnitude differences in tokens, GPU time, and network/storage pressure;&lt;/li&gt;
&lt;li&gt;Without governance closed loops, systems will move toward &amp;ldquo;cost and risk runaway&amp;rdquo; while becoming &amp;ldquo;more capable.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, the fundamental purpose of layering is not abstract aesthetics, but achieving a hard constraint goal:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Ensure each layer can translate upper-layer &amp;ldquo;intent&amp;rdquo; into executable plans and produce measurable, attributable, and constrainable resource consequences.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In other words, layering is not about making architecture diagrams clearer, but about encoding &amp;ldquo;who expresses intent, who executes, and who bears consequences&amp;rdquo; into system structure.&lt;/p&gt;
&lt;h2 id="ai-native-infrastructure-five-layer-structure-and-three-planes-mapping"&gt;AI-Native Infrastructure Five-Layer Structure and &amp;ldquo;Three Planes&amp;rdquo; Mapping&lt;/h2&gt;
&lt;p&gt;To help understand the layering logic, the diagram below refines the &amp;ldquo;Three Planes&amp;rdquo; architecture from the previous chapter, proposing a more actionable &amp;ldquo;five-layer structure&amp;rdquo;:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/compute-governance/intent-consequence-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/compute-governance/intent-consequence-en.svg" alt="Figure 1: Layered governance relationship from intent to consequence" data-caption="Figure 1: Layered governance relationship from intent to consequence"
width="1576"
height="156"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Layered governance relationship from intent to consequence&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul&gt;
&lt;li&gt;Top two layers = &lt;strong&gt;Intent Plane&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Middle two layers = &lt;strong&gt;Execution Plane&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Bottom layer = &lt;strong&gt;Governance Plane&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Below is a detailed expansion of the five-layer architecture, showing the primary responsibilities and typical capabilities of each layer:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/compute-governance/five-layers-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/compute-governance/five-layers-en.svg" alt="Figure 2: Five-layer architecture diagram" data-caption="Figure 2: Five-layer architecture diagram"
width="2131"
height="1476"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Five-layer architecture diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It is important to note that &lt;strong&gt;MCP belongs to Layer 4 (Intent and Orchestration Layer), not Layer 1&lt;/strong&gt;. The reason is that MCP primarily defines &amp;ldquo;how capabilities are exposed to models/Agents and how they are invoked,&amp;rdquo; addressing control plane consistency and composability, but does not directly take responsibility for &amp;ldquo;how the resource consequences of capability invocations are metered, constrained, and attributed.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="mcpagent-is-the-new-control-plane-but-must-be-constrained-by-the-governance-layer"&gt;MCP/Agent is the &amp;ldquo;New Control Plane,&amp;rdquo; But Must Be Constrained by the Governance Layer&lt;/h2&gt;
&lt;p&gt;MCP/Agent is called the &amp;ldquo;new control plane&amp;rdquo; because it moves system &amp;ldquo;decisions&amp;rdquo; from static code to dynamic processes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;Tool catalogs + schemas + invocations&amp;rdquo; form a composable capability surface;&lt;/li&gt;
&lt;li&gt;Agents complete tasks by selecting tools, invoking tools, and iterating reasoning;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Policy&amp;rdquo; is no longer just in code branches but expressed as routing, priorities, budgets, and compliance intent.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;However, it is crucial to emphasize an infrastructure stance, which is also the foundation of this chapter:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;MCP/Agent can express intent, but the key to AI-native is: intent must be translated into governable execution plans and metered and constrained within economically viable boundaries.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This statement aims to correct two common misconceptions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Control plane is not the starting point&lt;/strong&gt;: Treating MCP/Agent as &amp;ldquo;the entry point for AI platform upgrades&amp;rdquo; easily leads systems down a &amp;ldquo;capability-first&amp;rdquo; path;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance plane is the baseline&lt;/strong&gt;: When compute and tokens become capacity units, any unconstrained &amp;ldquo;intent expression&amp;rdquo; will leak as cost, latency, or risk.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, system layering should be clear: Layer 4 is responsible for &amp;ldquo;expression,&amp;rdquo; Layers 1/2/3 are responsible for &amp;ldquo;fulfillment and bearing consequences,&amp;rdquo; and the governance loop is responsible for &amp;ldquo;correction.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="context-is-rising-to-a-new-infrastructure-layer"&gt;&amp;ldquo;Context&amp;rdquo; Is Rising to a New Infrastructure Layer&lt;/h2&gt;
&lt;p&gt;In traditional cloud-native systems, request states are mostly short-lived, relying more on application-layer state management. Infrastructure typically only handles &amp;ldquo;computation and networking&amp;rdquo; without needing to understand the economic value of &amp;ldquo;request context.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;AI-native infrastructure is different. Long-context, multi-turn dialogue, and multi-agent reasoning mean &lt;strong&gt;inference state often survives across requests&lt;/strong&gt; and directly determines throughput and cost. In particular, KV cache and context reuse are evolving from &amp;ldquo;performance optimization techniques&amp;rdquo; to &amp;ldquo;platform capacity structures.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This can be summarized as an infrastructure law:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When a state asset (context/state) becomes a determinant variable of system cost and throughput, it rises from application detail to infrastructure layer.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This trend is gradually appearing in the industry: inference context and KV reuse are explicitly elevated to &amp;ldquo;infrastructure layer&amp;rdquo; capability development directions. Future expansion will include distributed KV, parameter caching, inference routing state, Agent memory, and a series of &amp;ldquo;state assets.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-foundation-of-ai-native-infrastructure-reference-designs-and-delivery-systems"&gt;The Foundation of AI-Native Infrastructure: Reference Designs and Delivery Systems&lt;/h2&gt;
&lt;p&gt;AI-native infrastructure is far more than &amp;ldquo;buying a few GPUs.&amp;rdquo; Compared to traditional internet services, AI workloads have three characteristics that make the &amp;ldquo;foundation&amp;rdquo; more engineered and productized:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Stronger topology dependencies&lt;/strong&gt;: Network fabric, interconnects, storage tiers, and GPU affinity determine available throughput;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Harder scarcity constraints&lt;/strong&gt;: GPU and token throughput boundaries are less &amp;ldquo;elastic&amp;rdquo; than CPU/memory;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Higher delivery complexity&lt;/strong&gt;: Multi-cluster, multi-tenant, multi-model/multi-framework coexistence means only &amp;ldquo;replicable delivery&amp;rdquo; can scale.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, AI Infra is not just a component list, but must include &amp;ldquo;scalable delivery and repeatable operation&amp;rdquo; system capabilities:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reference Designs (validated designs)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Codify &amp;ldquo;correct topology and ratios&amp;rdquo; into reusable solutions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Automated Delivery&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Institutionalize deployment, upgrade, scaling, rollback, and capacity planning.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Governance Implementation&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Make budgeting, isolation, metering, and auditing default capabilities rather than after-the-fact patches.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a CTO/CEO perspective, this means: what you purchase is not &amp;ldquo;hardware&amp;rdquo; but a &amp;ldquo;delivery system for predictable capacity.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="layered-responsibility-boundaries-from-a-ctoceo-perspective"&gt;&amp;ldquo;Layered Responsibility Boundaries&amp;rdquo; from a CTO/CEO Perspective&lt;/h2&gt;
&lt;p&gt;To facilitate internal alignment on &amp;ldquo;who is responsible for what and what is the cost of failure,&amp;rdquo; the table below maps &amp;ldquo;technical layers&amp;rdquo; to &amp;ldquo;organizational responsibilities,&amp;rdquo; avoiding the scenario where platform teams only build control planes while no one bears consequence boundaries.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Typical Capabilities&lt;/th&gt;
&lt;th&gt;Primary Owner (Recommended)&lt;/th&gt;
&lt;th&gt;Cost of Failure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Layer 5 Business Interface&lt;/td&gt;
&lt;td&gt;SLA, product experience, business goals&lt;/td&gt;
&lt;td&gt;Product / Business&lt;/td&gt;
&lt;td&gt;Customer experience and revenue impact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layer 4 Intent/Orchestration (MCP/Agent)&lt;/td&gt;
&lt;td&gt;Capability catalogs, workflow, policy expression&lt;/td&gt;
&lt;td&gt;App / Platform / AI Eng&lt;/td&gt;
&lt;td&gt;Behavior runaway, tool abuse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layer 3 Execution (Runtime)&lt;/td&gt;
&lt;td&gt;Serving, batching, routing, caching policies&lt;/td&gt;
&lt;td&gt;AI Platform / Infra&lt;/td&gt;
&lt;td&gt;Insufficient throughput, latency jitter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layer 2 Context/State&lt;/td&gt;
&lt;td&gt;KV/cache/context tier&lt;/td&gt;
&lt;td&gt;Infra + AI Platform&lt;/td&gt;
&lt;td&gt;Token cost spike, throughput collapse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layer 1 Compute/Governance&lt;/td&gt;
&lt;td&gt;Quotas, isolation, topology scheduling, metering&lt;/td&gt;
&lt;td&gt;Infra / FinOps / SRE&lt;/td&gt;
&lt;td&gt;Budget explosion, resource contention, incident spillover&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: AI-Native Infrastructure Layer and Organizational Responsibility Mapping
&lt;/figcaption&gt;
&lt;p&gt;As you can see, &lt;strong&gt;the organizational challenge of AI-native is not in &amp;ldquo;whether we have agents,&amp;rdquo; but in &amp;ldquo;whether inter-layer closed loops are established&amp;rdquo;&lt;/strong&gt;. When model-driven amplification of consequences occurs, organizations must institutionalize governance mechanisms as platform capabilities: executable budgets, explainable consequences, attributable anomalies, and rewritable policies. This is the true meaning of &amp;ldquo;starting from compute governance&amp;rdquo; rather than &amp;ldquo;starting from API design.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The layered design of AI-native infrastructure centers on engineering the binding between &amp;ldquo;intent&amp;rdquo; and &amp;ldquo;resource consequences.&amp;rdquo; The control plane is responsible for expressing intent, while the governance plane is responsible for bounding consequences. Only by institutionalizing governance mechanisms as platform capabilities can we ensure cost, risk, and capacity remain controllable while enhancing capabilities. As context, state assets, and other new variables become infrastructure, AI Infra delivery systems will continue to evolve, becoming the foundation for sustainable enterprise innovation.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/capacity-planning/" target="_blank" rel="noopener"&gt;Google SRE - Capacity Planning - sre.google&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aws.amazon.com/architecture/well-architected/" target="_blank" rel="noopener"&gt;AWS Well-Architected - Cost Optimization - aws.amazon.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/architecture/guide/cost-management/finops/" target="_blank" rel="noopener"&gt;Microsoft FinOps for AI - learn.microsoft.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Operating and Governing AI-Native Infrastructure: Metrics, Budget, Isolation, Sharing, SLO to Cost</title><link>https://jimmysong.io/book/ai-native-infra/metrics-budget-isolation/</link><pubDate>Sun, 18 Jan 2026 04:17:45 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/metrics-budget-isolation/</guid><description>Analyzing the closed-loop governance of metrics, budgets, isolation, and sharing in AI-native infrastructure, and explaining how SLO maps to cost and risk.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The key to governing AI-native infrastructure lies in how to institutionalize the closed-loop management of costs and risks arising from uncertainty.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In the cloud-native era, system operations were typically considered &amp;ldquo;basically deterministic&amp;rdquo;: request paths were predictable, resource curves were relatively stable, and scaling could respond promptly to load changes. However, entering the AI era, this assumption no longer holds—&lt;strong&gt;uncertainty has become the norm&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This chapter aims to provide CTOs/CEOs with key conclusions for architecture reviews:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The starting point of AI-native infrastructure is to treat uncertainty as the default input;&lt;/strong&gt;
&lt;strong&gt;The goal is to achieve closed-loop governance of the resource consequences (cost, risk, experience) arising from uncertainty.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is also why &amp;ldquo;becoming AI-native&amp;rdquo; in organizational contexts increasingly points to the reshaping of operational methods and governance models: when system consequences are amplified, governance must be institutionalized.&lt;/p&gt;
&lt;h2 id="what-is-an-uncertain-system"&gt;What is an &amp;ldquo;Uncertain System&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;In this handbook, &amp;ldquo;uncertainty&amp;rdquo; does not refer to randomness in the probabilistic sense, but to three types of phenomena in engineering practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Unpredictable behavior&lt;/strong&gt;: execution paths change dynamically with model inference, especially evident in agentic processes (Agent intelligent workflows).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unpredictable resource consumption&lt;/strong&gt;: tokens, KV cache, tool calls, I/O, and network overhead exhibit long-tail and burst characteristics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Non-linear consequences&lt;/strong&gt;: the same &amp;ldquo;intent&amp;rdquo; can produce cost and risk outcomes differing by orders of magnitude.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, the infrastructure problem of AI-native infrastructure has shifted from &amp;ldquo;how to make the system more elegant&amp;rdquo; to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How to ensure the system maintains economic viability, controllability, and recoverability when worst-case scenarios occur.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;During architecture reviews, if you cannot answer &amp;ldquo;what is the worst case, where are the upper bounds, and how to degrade/rollback when triggered,&amp;rdquo; you are still reviewing the inertia extension of deterministic systems, not true AI-native systems.&lt;/p&gt;
&lt;h2 id="major-sources-of-uncertainty"&gt;Major Sources of Uncertainty&lt;/h2&gt;
&lt;p&gt;The following table summarizes common sources of uncertainty in AI-native infrastructure and their specific manifestations, facilitating quick reference for CTOs/CEOs.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Manifestations&lt;/th&gt;
&lt;th&gt;Impact Areas&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Behavior Uncertainty&lt;/td&gt;
&lt;td&gt;Agent task decomposition path changes, tool selection and call sequence changes, failure retry and reflection&lt;/td&gt;
&lt;td&gt;Cost, Risk, Resilience&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Demand Uncertainty&lt;/td&gt;
&lt;td&gt;Concurrency and burst, long-tail requests, multi-tenant interference (noisy neighbor)&lt;/td&gt;
&lt;td&gt;Resource pools, Experience, Isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State Uncertainty&lt;/td&gt;
&lt;td&gt;Context reuse across requests, KV cache migration and sharing&lt;/td&gt;
&lt;td&gt;Performance, Cost, Governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infrastructure Uncertainty&lt;/td&gt;
&lt;td&gt;High sensitivity to network/storage/interconnect, congestion and jitter amplified into tail latency&lt;/td&gt;
&lt;td&gt;Experience, Cost, Stability&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Sources and Manifestations of Uncertainty in AI-Native Infrastructure
&lt;/figcaption&gt;
&lt;h3 id="behavior-uncertainty"&gt;Behavior Uncertainty&lt;/h3&gt;
&lt;p&gt;Behavior uncertainty is mainly reflected in changes to agent task decomposition paths, dynamic adjustment of tool selection and call sequences, and path explosion caused by failure retry, reflection, and multi-round planning. Tools and contexts are combined through standard interfaces (such as MCP protocol integration), significantly expanding system capability surfaces while making branch space a governance challenge.&lt;/p&gt;
&lt;p&gt;More critically, tool calls are not &amp;ldquo;free external functions&amp;rdquo; - they occupy context windows and consume token budgets, amplifying cost and tail latency pressures. Therefore, behavior uncertainty is not merely &amp;ldquo;feature flexibility&amp;rdquo; at the product layer, but &amp;ldquo;cost and risk elasticity&amp;rdquo; at the platform layer, which must be budgeted, capped, and made auditable.&lt;/p&gt;
&lt;h3 id="demand-uncertainty"&gt;Demand Uncertainty&lt;/h3&gt;
&lt;p&gt;Demand uncertainty includes concurrency and burst (peaks), long-tail requests (ultra-long contexts, complex reasoning), and mutual interference under multi-tenancy (noisy neighbor). This drives capacity planning from &amp;ldquo;average capacity&amp;rdquo; to &amp;ldquo;tail capacity + governance strategies.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;In AI-native infrastructure, experience and cost are often determined not by average requests, but by the combination of tail requests: a small number of long-chain, long-context, tool-intensive requests can overwhelm shared resource pools. Therefore, demand uncertainty requires answering: &lt;strong&gt;which requests deserve guarantees, which must be throttled, and which should be isolated.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="statecontext-uncertainty"&gt;State/Context Uncertainty&lt;/h3&gt;
&lt;p&gt;State uncertainty is the most underestimated category in the AI era: &lt;strong&gt;context is a state asset&lt;/strong&gt;, and it often exists across requests. When inference state / KV cache is elevated to a reusable, shareable, migratable system capability, it is no longer an application detail but a decisive variable for throughput and unit cost. NVIDIA in public materials identifies &lt;em&gt;Inference Context Memory Storage&lt;/em&gt; as a new infrastructure layer, pointing to state reuse and sharing requirements for long-context and agentic workloads.&lt;/p&gt;
&lt;p&gt;The conclusion is: &lt;strong&gt;&amp;ldquo;context/state&amp;rdquo; has changed from optional optimization to a critical infrastructure asset that must be meterable, allocable, and governable.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="infrastructure-uncertainty"&gt;Infrastructure Uncertainty&lt;/h3&gt;
&lt;p&gt;AI workloads are far more sensitive to network, interconnect, and storage than traditional microservice workloads. Congestion, packet loss, and I/O jitter are amplified into tail latency and job completion time instability, creating &amp;ldquo;non-linear consequences&amp;rdquo; for experience and cost.&lt;/p&gt;
&lt;p&gt;This type of uncertainty usually cannot be solved through &amp;ldquo;component selection&amp;rdquo; but requires &lt;strong&gt;end-to-end path engineering constraints&lt;/strong&gt;: from topology, bandwidth, and queuing, to transport protocols, isolation strategies, and congestion control—all must be incorporated into the governance plane, not just the operations plane.&lt;/p&gt;
&lt;h2 id="how-uncertainty-amplifies-across-layers"&gt;How Uncertainty Amplifies Across Layers&lt;/h2&gt;
&lt;p&gt;The diagram below illustrates the closed-loop relationship between metrics, budgets, and isolation strategies, emphasizing that governance must be rewritable.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/metrics-budget-isolation/slo-to-cost-loop-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/metrics-budget-isolation/slo-to-cost-loop-en.svg" alt="Figure 1: SLO to cost feedback loop" data-caption="Figure 1: SLO to cost feedback loop"
width="1656"
height="456"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: SLO to cost feedback loop&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The flowchart below demonstrates the cross-layer amplification path of uncertainty in AI-native infrastructure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/metrics-budget-isolation/uncertainty-amplification-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/metrics-budget-isolation/uncertainty-amplification-en.svg" alt="Figure 2: Uncertainty amplification across layers" data-caption="Figure 2: Uncertainty amplification across layers"
width="1579"
height="1305"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Uncertainty amplification across layers&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Typical phenomena include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agent branch explosion&lt;/strong&gt;: more tools and composable paths make tail costs increasingly uncontrollable.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context inflation&lt;/strong&gt;: long contexts and multi-round reasoning make KV cache a performance bottleneck and cost black hole.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resource contention distortion&lt;/strong&gt;: GPU/network contention under multi-tenancy makes &amp;ldquo;average performance&amp;rdquo; meaningless—tails must be governed.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, the core of AI-native is not &amp;ldquo;making execution stronger,&amp;rdquo; but enabling you to stably answer three questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Where are the upper bounds&lt;/strong&gt; (budgets, steps, call counts, state occupancy)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What to do when crossing boundaries&lt;/strong&gt; (degradation, rollback, isolation, blocking)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How results are rewritten&lt;/strong&gt; (policy iteration and cost correction)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="engineering-response-of-ai-native-infrastructure"&gt;Engineering Response of AI-Native Infrastructure&lt;/h2&gt;
&lt;p&gt;Enterprises can refer to the following five &amp;ldquo;hard standards&amp;rdquo; during reviews—missing any one means inability to achieve closed-loop governance of uncertainty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Admission: Ingress Admission Control&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Implement tiered admission for requests with ultra-long contexts, oversized tool graphs, or ultra-high budgets&lt;/li&gt;
&lt;li&gt;Bind &amp;ldquo;budget, priority, compliance&amp;rdquo; as part of intent (policy as intent)&lt;/li&gt;
&lt;li&gt;Clearly communicate rejection reasons and explain why requests are denied&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Key Point
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
The responsibility of admission is not &amp;ldquo;to allow features,&amp;rdquo; but to write consequence constraints into the contract.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Translation: Intent Translation to Governable Execution Plans&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Select runtime, routing/batching strategies, and caching strategies for requests&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Cap&amp;rdquo; agent workflows: maximum steps, maximum tool calls, maximum tokens&lt;/li&gt;
&lt;li&gt;Include fallback paths: deterministic alternatives, cached answers, manual/rule-based fallbacks&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Key Point
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Upgrade from &amp;ldquo;prompt-driven execution&amp;rdquo; to &amp;ldquo;plan-driven execution&amp;rdquo;—plans must be understandable and constrainable by the governance plane.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Metering: End-to-End Metering and Attribution&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Meter tokens, GPU time, KV cache footprint, I/O, and network for each request/agent task&lt;/li&gt;
&lt;li&gt;Attribute by tenant, project, model, and tool to form cost and quality metrics&lt;/li&gt;
&lt;li&gt;Separately label &amp;ldquo;tail overhead&amp;rdquo; so long-tail costs no longer hide in averages&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Key Point
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
No ledger means no budget; no attribution means no governance, let alone ROI.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Enforcement: Budget and Degradation Mechanisms&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Budget triggers: rate limiting, degradation, preemption, queuing (by priority and tenant isolation)&lt;/li&gt;
&lt;li&gt;Risk triggers: isolation&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The core of AI-native infrastructure governance lies in front-loading uncertainty, layered metering, policy feedback, and institutionalized constraints to form a closed loop of cost and risk. Only with engineering mechanisms such as Admission, Translation, Metering, and Enforcement can systems achieve economically viable, controllable, and recoverable operations under normalized uncertainty.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/service-level-objectives/" target="_blank" rel="noopener"&gt;Google SRE Book - Service Level Objectives - google&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.finops.org/" target="_blank" rel="noopener"&gt;FinOps Foundation - AI Cost Management - finops.org&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/policies/usage-policies/" target="_blank" rel="noopener"&gt;OpenAI Usage Policies - openai.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Organization and Culture: How the Operating Model Changes</title><link>https://jimmysong.io/book/ai-native-infra/operating-model/</link><pubDate>Sun, 18 Jan 2026 04:18:02 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/operating-model/</guid><description>Redrawing boundaries across platform, infra, ML, and security, and transforming accountability and collaboration in the AI era.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The compute governance closed loop is the foundational safeguard for sustainable innovation in AI-native organizations.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Perspective
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
API/Agent/MCP solve &amp;ldquo;how intent is expressed,&amp;rdquo; while compute governance addresses &amp;ldquo;whether the resource consequences of intent are economically viable and risk-controllable.&amp;rdquo; In the AI era, the latter becomes a prerequisite for the former. API-first without governance only amplifies costs and uncertainty, pushing organizations into the trap of &amp;ldquo;functional but unsustainable.&amp;rdquo;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The FinOps Foundation states directly in &amp;ldquo;Scaling Kubernetes for AI/ML Workloads with FinOps&amp;rdquo; that Kubernetes elasticity can easily evolve into a &lt;strong&gt;runaway cost problem&lt;/strong&gt;. Therefore, FinOps should not be just cost reporting, but must become a shared operating model where every scaling decision simultaneously answers two questions: &lt;strong&gt;Are performance SLOs met?&lt;/strong&gt;, and &lt;strong&gt;Is it affordable?&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="the-challenge-of-api-first-implicit-assumptions-in-the-ai-era"&gt;The Challenge of API-first &amp;ldquo;Implicit Assumptions&amp;rdquo; in the AI Era&lt;/h2&gt;
&lt;p&gt;The diagram below shows the boundary relationships and accountability chains between platform, ML, and security teams.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/operating-model/operating-model-boundaries-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/operating-model/operating-model-boundaries-en.svg" alt="Figure 1: Organizational boundaries and accountability chain" data-caption="Figure 1: Organizational boundaries and accountability chain"
width="1616"
height="456"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Organizational boundaries and accountability chain&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The intuitive path of API-first is: first make the interfaces and workflows work, then gradually optimize performance and cost through engineering. In AI-native infrastructure, this path often fails because it relies on three implicit assumptions that no longer hold in the AI era.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 1: Resources are not the core scarcity&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Traditional software bets scarcity on engineering efficiency, throughput, and stability; whereas AI-native infrastructure scarcity comes primarily from asset boundaries like &lt;strong&gt;GPU/interconnect/power consumption&lt;/strong&gt;. Scarcity is no longer &amp;ldquo;slow to scale,&amp;rdquo; but &amp;ldquo;hard to scale and expensive,&amp;rdquo; constrained by both supply chain and datacenter conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 2: Request costs are predictable&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Traditional request cost distribution is relatively stable; AI requests are inherently long-tailed: branching in agentic tasks, inflation of long contexts, and chain amplification of tool calls all make tokens and GPU time into random variables that cannot be linearly extrapolated. You think you&amp;rsquo;re scaling &amp;ldquo;QPS,&amp;rdquo; but actually you&amp;rsquo;re scaling &amp;ldquo;total cost of tail probability events.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 3: State is ephemeral and discardable&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The cloud-native era emphasized stateless scaling with externalized state; but on the inference side, &lt;strong&gt;inference state/context reuse&lt;/strong&gt; often determines whether unit costs are controllable. NVIDIA describes this in Rubin&amp;rsquo;s ICMS (Inference Context Memory Storage) as the &amp;ldquo;context storage challenge brought by new inference paradigms&amp;rdquo;: KV cache needs reuse across sessions/services, sequence length growth causes linear KV cache inflation, forcing persistence and shared access, forming a &amp;ldquo;new context tier,&amp;rdquo; and proving with TPS and energy efficiency gains that this is not a nice-to-have, but a threshold for scalability.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Conclusion
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
In AI-native infrastructure, state and compute governance have become prerequisites for &amp;ldquo;whether it can run,&amp;rdquo; not post-optimization items.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="the-nature-of-compute-governance-what-is-being-governed"&gt;The Nature of Compute Governance: What is Being Governed&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;Compute governance&amp;rdquo; is often misunderstood as &amp;ldquo;managing GPUs,&amp;rdquo; but what truly needs governance is &lt;strong&gt;the resource consequences of intent&lt;/strong&gt;. More precisely, it&amp;rsquo;s governing the combined effects of four types of objects:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Token Economics&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Each request/task&amp;rsquo;s token consumption, context inflation, implicit token tax from tool definitions and intermediate results, ultimately directly mapping to cost and latency.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Accelerator Time&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPU time, memory footprint, batching strategies, and the impact of routing and cache hits on effective throughput. The key is not &amp;ldquo;whether there are GPUs,&amp;rdquo; but &amp;ldquo;whether output per unit GPU time is controllable.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Interconnect and Storage (Fabric &amp;amp; Storage)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Network and storage pressures from training all-reduce, inference KV/cache sharing, and cross-service data movement. AI performance and cost are often amplified by fabric, not by APIs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Organizational Budget and Risk (Budget &amp;amp; Risk)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-tenant isolation, fairness, audit, compliance, and accountability. These determine whether the system can scale to multiple teams/business lines, not just scaling demos to more instances.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The FinOps Foundation also emphasizes: AI/ML cost drivers are not just GPUs; storage (checkpoints/embeddings/artifacts), network (distributed training/cross-AZ), and additional licensing and marketplace fees often &amp;ldquo;quietly exceed compute.&amp;rdquo; Therefore, governance objects must cover end-to-end, not just stare at inference bills.&lt;/p&gt;
&lt;h2 id="mcpagent-amplification-effects-under-governance-gaps"&gt;MCP/Agent: Amplification Effects Under Governance Gaps&lt;/h2&gt;
&lt;p&gt;MCP/Agent expand the &amp;ldquo;capability surface,&amp;rdquo; but simultaneously make cost curves steeper, especially showing exponential amplification when governance is missing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;More tools, more branches&lt;/strong&gt;: Planning space expands, tail probability rises, cost volatility becomes uncontrollable.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool definitions and intermediate results consume context&lt;/strong&gt;: Directly consuming context window and tokens, translating to cost and latency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stronger tool usage triggers more external I/O&lt;/strong&gt;: External system calls, network round trips, and data movement all enter the overall cost function.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Anthropic explicitly states in &amp;ldquo;Code execution with MCP&amp;rdquo; that direct tool calls increase cost and latency due to tool definitions and intermediate results consuming context window; when tool numbers rise to hundreds or thousands, this becomes a scalability bottleneck, thus proposing code execution forms to improve efficiency and reduce token consumption.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Conclusion
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
In the MCP/Agent era, governance is not about suppressing innovation, but making innovation sustainable within budget boundaries. Without governance, agents are not productivity tools, but cost amplifiers.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="minimal-implementation-path-for-compute-governance-first"&gt;Minimal Implementation Path for &amp;ldquo;Compute Governance First&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;You don&amp;rsquo;t have to bind to any vendor, but you must implement a &amp;ldquo;minimum viable governance stack.&amp;rdquo; The goal is not perfection, but giving the system controllable boundary conditions from day one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Admission and Budget (Admission + Budget)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Set budgets and priorities for workload types (training/inference/agent tasks).&lt;/li&gt;
&lt;li&gt;Include budget, max steps, max tokens, max tool calls in policy-as-intent, and enforce at the entry point.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
Practice Recommendation
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
FinOps&amp;rsquo; core view is: embed FinOps early into architecture, making every scaling decision simultaneously answer &amp;ldquo;performance&amp;rdquo; and &amp;ldquo;affordability,&amp;rdquo; otherwise bills only get attention when incidents occur.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;End-to-End Metering and Attribution (Metering + Attribution)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;At minimum achieve one traceable chain: request/agent → tokens → GPU time/memory → network/storage → cost attribution (tenant/project/model/tool).&lt;/li&gt;
&lt;li&gt;Without attribution, there is no governance; without governance, enterprise scaling is impossible, because costs and responsibilities cannot align, and organizations will internally waste energy on &amp;ldquo;who consumed the budget.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Isolation and Sharing (Isolation + Sharing)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sharing&lt;/strong&gt; for improving utilization; &lt;strong&gt;isolation&lt;/strong&gt; for reducing risk. Both must exist simultaneously, not either/or.&lt;/li&gt;
&lt;li&gt;CNCF&amp;rsquo;s Cloud Native AI report notes: GPU virtualization and sharing (like MIG, MPS, DRA, etc.) can improve utilization and reduce costs, but requires careful orchestration and management, and demands collaboration between AI and cloud-native engineering teams.&lt;/li&gt;
&lt;li&gt;The key to governance is not choosing sharing or isolation, but making it an executable policy: who shares under what conditions, who isolates under what conditions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Topology and Network as First-Class Citizens (Topology + Fabric First)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI training and high-throughput inference are highly sensitive to network characteristics.&lt;/li&gt;
&lt;li&gt;Cisco&amp;rsquo;s AI-ready infrastructure design guides and related CVD/Design Zone emphasize: building high-performance, lossless Ethernet fabric for AI/ML workloads, and delivering reference architectures and deployment guides through validated designs.&lt;/li&gt;
&lt;li&gt;This means topology is not &amp;ldquo;the datacenter team&amp;rsquo;s business,&amp;rdquo; but a core variable determining whether JCT, tail latency, and capacity models hold.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Context/State Becomes a Governance Object (Context as a Governed Asset)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When long-context and agentic become mainstream, KV cache and inference context reuse will directly determine unit costs.&lt;/li&gt;
&lt;li&gt;NVIDIA&amp;rsquo;s ICMS defines this as a &amp;ldquo;new context tier&amp;rdquo; for solving KV cache reuse and shared access, emphasizing TPS/energy efficiency gains.&lt;/li&gt;
&lt;li&gt;In this era, treating context as a temporary variable is actively relinquishing cost control.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="anti-pattern-checklist"&gt;Anti-Pattern Checklist&lt;/h2&gt;
&lt;p&gt;The following anti-patterns are not &amp;ldquo;engineering inelegance,&amp;rdquo; but will trigger organizational loss of control, worth vigilance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;API-first, treating governance as post-optimization&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Result: System launches first, only to discover unit costs and tail latency are uncontrollable, can only &amp;ldquo;hard brake&amp;rdquo; through feature limiting/rate limiting, ultimately locking the product roadmap.&lt;/li&gt;
&lt;li&gt;Contrast: FinOps points out elasticity easily becomes runaway costs, must advance cost governance into architecture decisions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Treating MCP/Agent as capability accelerators, not cost amplifiers&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Result: More tools make it &amp;ldquo;smarter,&amp;rdquo; but token and external call costs rise exponentially, engineering teams forced to fight systemic amplification with &amp;ldquo;more complex prompts and rules.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Contrast: Anthropic notes tool definitions and intermediate results consume context, increase cost and latency, proposing more efficient execution forms as the scalability path.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Only buying GPUs, without sharing/isolation and orchestration&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Result: Low utilization, severe contention, budget explosion, organizations internally blame each other &amp;ldquo;who&amp;rsquo;s grabbing resources, who&amp;rsquo;s burning money.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Contrast: CNCF Cloud Native AI report emphasizes sharing/virtualization improves utilization, but must match orchestration and collaboration mechanisms.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Ignoring network and topology, treating AI as ordinary microservices&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Result: Training JCT and inference tail latency amplified by network, capacity planning and cost models fail, more scaling makes it more unstable.&lt;/li&gt;
&lt;li&gt;Contrast: Cisco in AI-ready network design and validated designs makes requirements like lossless Ethernet fabric critical foundations for AI/ML.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The first-principle entry point for AI-native is the compute governance closed loop: budget and admission, metering and attribution, sharing and isolation, topology and network, context assetization. API/Agent/MCP remain important, but must be constrained by this closed loop, otherwise the system can only oscillate between &amp;ldquo;smarter&amp;rdquo; and &amp;ldquo;more bankrupt.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://sloanreview.mit.edu/" target="_blank" rel="noopener"&gt;MIT Sloan - sloanreview.mit.edu&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/ai/responsible-ai" target="_blank" rel="noopener"&gt;Google Cloud - cloud.google.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://oecd.ai/en/ai-principles" target="_blank" rel="noopener"&gt;OECD AI Principles - oecd.ai&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Migration Roadmap: From Cloud Native to AI Native</title><link>https://jimmysong.io/book/ai-native-infra/migration-roadmap/</link><pubDate>Sun, 18 Jan 2026 04:20:11 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/migration-roadmap/</guid><description>An actionable roadmap for AI-native migration, covering bypass pilot, domain isolation, AI-first refactoring, and anti-patterns, with focus on governance loops and organizational contracts.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Migration is not &amp;ldquo;rebuilding the platform,&amp;rdquo; but using governance loops and organizational contracts to transform uncertainty into controllable engineering capabilities.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The previous five chapters have established: &lt;strong&gt;AI-native infrastructure is Uncertainty-by-Default&lt;/strong&gt;. Therefore, the architectural starting point must be &lt;strong&gt;compute governance loops&lt;/strong&gt;, not &amp;ldquo;connect a model and call migration complete.&amp;rdquo; Otherwise, systems easily spiral out of control in three dimensions: &lt;strong&gt;cost&lt;/strong&gt; (runaway cost), &lt;strong&gt;risk&lt;/strong&gt; (unauthorized actions/side effects), and &lt;strong&gt;tail performance&lt;/strong&gt; (P95/P99 and queue tail behavior).&lt;/p&gt;
&lt;p&gt;This explains why the FinOps Foundation emphasizes: running AI/ML on Kubernetes, &amp;ldquo;elasticity&amp;rdquo; easily evolves into uncontrollable cost overflow. &lt;strong&gt;FinOps must be incorporated into architecture and organization upfront as a shared operating model, not as an after-the-fact reconciliation exercise.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This article presents an &lt;strong&gt;actionable migration roadmap&lt;/strong&gt;, covering both technical evolution paths and organizational implementation approaches. You don&amp;rsquo;t need to &amp;ldquo;rebuild an AI platform&amp;rdquo; all at once, but you must establish working &lt;strong&gt;governance loops&lt;/strong&gt; at each stage: budget/admission, metering/attribution, sharing/isolation, topology/networking, and context assetization.&lt;/p&gt;
&lt;h2 id="the-north-star-from-platform-delivery-to-governance-loops"&gt;The North Star: From Platform Delivery to Governance Loops&lt;/h2&gt;
&lt;p&gt;The diagram below shows the migration path from bypass pilot to AI-first refactoring.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/migration-roadmap/migration-roadmap-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/migration-roadmap/migration-roadmap-en.svg" alt="Figure 1: AI-native migration roadmap" data-caption="Figure 1: AI-native migration roadmap"
width="1663"
height="203"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: AI-native migration roadmap&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Cloud-native migration typically centers on &amp;ldquo;capability delivery&amp;rdquo;: CI/CD, self-service platforms, service governance, and auto-scaling. Its default assumptions: systems are deterministic, costs grow linearly with requests, and scaling doesn&amp;rsquo;t significantly alter system boundaries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI-native migration must center on &amp;ldquo;governance loops&amp;rdquo;&lt;/strong&gt;, focusing on cost, risk, tail performance, and state assets. Its default assumptions are precisely the opposite: systems are inherently uncertain, and the &amp;ldquo;actions and consequences&amp;rdquo; of inference/agents drive costs and risks into nonlinear territory.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
North Star Definition
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
&lt;strong&gt;AI-Native Migration = Establish an AI Landing Zone + Compute Governance Loop + Context Tier, and ensure all agents/APIs/runtimes operate within this loop.&lt;/strong&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Elevating &amp;ldquo;Landing Zone&amp;rdquo; to North Star level here isn&amp;rsquo;t chasing trends—it&amp;rsquo;s because it naturally serves an organizational-level task: &lt;strong&gt;delineating responsibility boundaries between platform teams and workload teams&lt;/strong&gt;. Major cloud providers universally use Landing Zones to host &amp;ldquo;shared governance baselines&amp;rdquo; (networking, identity, policies, auditing, quota/budget), while business teams iteratively build applications within controlled boundaries. For AI, this boundary is the carrier of the governance loop.&lt;/p&gt;
&lt;h2 id="migration-prerequisites-build-three-foundations-first-then-scale-applications"&gt;Migration Prerequisites: Build Three Foundations First, Then Scale Applications&lt;/h2&gt;
&lt;p&gt;You can run PoCs and build applications in parallel, but if these three foundations are missing, any &amp;ldquo;application explosion&amp;rdquo; can easily transform into platform firefighting and financial disputes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Foundation A: FinOps / Quotas as Control Plane (Finance and Quotas as Control Plane)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The first migration step is not &amp;ldquo;launch the first agent,&amp;rdquo; but incorporating &lt;strong&gt;budgets, alerts, showback/chargeback, and quotas&lt;/strong&gt; into the infrastructure control plane:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Budgets and alerts are not just financial reports, but triggers for runtime policies (rate limiting, degradation, queuing, preemption).&lt;/li&gt;
&lt;li&gt;showback/chargeback is not just accounting, but binding &amp;ldquo;cost consequences&amp;rdquo; to organizational decisions and product boundaries.&lt;/li&gt;
&lt;li&gt;Quotas are not static limits, but evolvable governance instruments (dynamic budgets and priorities by tenant/team/use-case).&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Migration Threshold
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
If you cannot attribute the primary consumption of each agent/job to team/project/model/use-case (at minimum covering tokens, GPU time, KV footprint, key network/storage), you haven&amp;rsquo;t reached the &amp;ldquo;scale&amp;rdquo; starting line. Piloting is acceptable, but expansion is not advisable.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Foundation B: Resource Governance (GPU Sharing/Isolation and Orchestration Capabilities)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;elasticity&amp;rdquo; of AI-native infrastructure is constrained by how scarce compute is governed. Treating GPUs as ordinary resources typically results in &lt;strong&gt;low utilization&lt;/strong&gt; and &lt;strong&gt;uncontrolled contention&lt;/strong&gt;. Therefore, you need viable combinations of sharing/isolation and orchestration capabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sharing/partitioning&lt;/strong&gt;: MIG/MPS/vGPU paths transform &amp;ldquo;exclusive&amp;rdquo; into &amp;ldquo;pooled.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling upgrades&lt;/strong&gt;: Introduce explicit modeling of topology, queues, fairness, preemption, and cost tiers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Orchestration loop&lt;/strong&gt;: Solidify isolation, preemption, and priority policies into executable rules.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key is not which partitioning technology you choose, but whether you can elevate GPUs from &amp;ldquo;machine assets&amp;rdquo; to &lt;strong&gt;first-class governance resources&lt;/strong&gt; and incorporate them into budget and admission systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Foundation C: Fabric as a First-Class Constraint (Network/Interconnect as First-Class Constraint)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Training and high-throughput inference are extremely sensitive to congestion, packet loss, and tail latency. Ignoring networking and topology leads to &amp;ldquo;seemingly sporadic but actually structural&amp;rdquo; problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training JCT is amplified by tail behavior, invalidating capacity planning;&lt;/li&gt;
&lt;li&gt;Inference P99 and queue tails are amplified, making SLOs difficult to honor.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, you need to build reusable AI-ready network baselines: capacity assumptions, lossless strategies, isolation domain partitioning, measurement and acceptance criteria. Networking is not &amp;ldquo;optimize later,&amp;rdquo; but baseline engineering that must land in Days 31–60.&lt;/p&gt;
&lt;h2 id="migration-path-selection-layered-by-organizational-risk-and-technical-debt"&gt;Migration Path Selection: Layered by Organizational Risk and Technical Debt&lt;/h2&gt;
&lt;p&gt;Migration isn&amp;rsquo;t &amp;ldquo;pick one path and see it through,&amp;rdquo; but mapping organizations with different risk appetites and debt structures to different starting approaches and exit criteria. Paths can advance in parallel, but each needs defined &lt;strong&gt;applicable conditions&lt;/strong&gt; and &lt;strong&gt;exit criteria&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Path 1: Bypass Pilot / Skunkworks&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Applicable when cloud-native platforms are running stably, but AI demand is just emerging, organizational uncertainty is high, and governance mechanisms are not yet mature.&lt;/p&gt;
&lt;p&gt;The approach is establishing an &amp;ldquo;AI minimum closed-loop sandbox&amp;rdquo; alongside the existing platform. The goal is not &amp;ldquo;feature completeness,&amp;rdquo; but &amp;ldquo;making the loop work&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Independent GPU pool (or at least independent queue) + basic admission and budget&lt;/li&gt;
&lt;li&gt;Minimal token/GPU metering and attribution&lt;/li&gt;
&lt;li&gt;Controlled inference/agent entry points (max context / max steps / max tool calls)&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Failure-acceptable&amp;rdquo; SLOs and cost caps (define boundaries first, then discuss experience)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Exit criteria:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cost curve is explainable (at minimum attributable to team/use-case)&lt;/li&gt;
&lt;li&gt;GPU utilization and isolation strategies form reusable templates&lt;/li&gt;
&lt;li&gt;Pilot capabilities can be absorbed into platform capabilities (enter Path 2)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Path 2: Domain-Isolated Platform&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Applicable when AI has entered multi-team, multi-tenant stages, requiring &amp;ldquo;pilot assets&amp;rdquo; to be solidified into platform capabilities to prevent cost and risk from spreading across domains.&lt;/p&gt;
&lt;p&gt;The approach is building an AI Landing Zone, where the platform team centrally manages shared governance capabilities, and workload teams iteratively build applications within controlled boundaries.&lt;/p&gt;
&lt;p&gt;Platform-side essential modules (recommend organizing by &amp;ldquo;governance loop&amp;rdquo;):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Identity/Policy&lt;/strong&gt;: Unified identity, policy distribution, and auditing (policy-as-intent)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Network/Fabric baseline&lt;/strong&gt;: AI-ready network baseline and automated acceptance&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compute governance&lt;/strong&gt;: Quotas, budgets, preemption, fairness, isolation/sharing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Observability &amp;amp; Chargeback&lt;/strong&gt;: End-to-end metering, alerts, showback/chargeback&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runtime catalog&lt;/strong&gt;: &amp;ldquo;Golden paths&amp;rdquo; and templated delivery for inference/training runtimes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Exit criteria: Platform provides &amp;ldquo;replicable AI workload landing approaches&amp;rdquo; and can scale use case count under budget constraints, rather than relying on manual firefighting to maintain stability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Path 3: AI-First Refactor (AI Factory / Replatform)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Applicable when AI is core business, requiring infrastructure to be treated as a &amp;ldquo;production line&amp;rdquo; rather than a &amp;ldquo;cluster,&amp;rdquo; and optimization objectives to switch from &amp;ldquo;shipping features&amp;rdquo; to &amp;ldquo;throughput/unit cost/energy efficiency.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The approach centers on &amp;ldquo;state assets + unit cost&amp;rdquo; refactoring:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Context/state&lt;/strong&gt; of inference/agents is explicitly governed and reused (no longer application-level tricks)&lt;/li&gt;
&lt;li&gt;Introduce &lt;strong&gt;Context Tier&lt;/strong&gt; architectural assumptions: long context and agentic inference require inference state / KV cache to be reusable across nodes and sessions&lt;/li&gt;
&lt;li&gt;Drive platform evolution with &amp;ldquo;unit token cost, tail latency, throughput/energy efficiency,&amp;rdquo; not &amp;ldquo;number of new components&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Exit criteria: Can consistently make engineering decisions using &amp;ldquo;unit cost and tail performance,&amp;rdquo; and treat context reuse as a platform capability rather than application team trick caching.&lt;/p&gt;
&lt;h2 id="90-day-actionable-plan-ai-landing-zone--minimum-governance-loop"&gt;90-Day Actionable Plan: AI Landing Zone + Minimum Governance Loop&lt;/h2&gt;
&lt;p&gt;The goal is to establish &amp;ldquo;AI Landing Zone + minimum governance loop&amp;rdquo; within 90 days, forming a replicable template. The key is not covering all scenarios, but connecting the &lt;strong&gt;admission—metering—enforcement—feedback&lt;/strong&gt; loop.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Day 0–30: Establish the Ledger (Cost &amp;amp; Usage Ledger)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;First, define attribution dimensions, establish budgets/alerts and baseline reports, and implement quotas/usage controls.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Attribution dimensions: tenant/team/project/model/use-case/tool&lt;/li&gt;
&lt;li&gt;Establish budgets and alerts, baseline reports (cost + business value metrics)&lt;/li&gt;
&lt;li&gt;Implement quotas and usage controls (at minimum covering GPU quotas and key service quotas)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Deliverables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cost and usage dashboard (weekly-level, traceable)&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Admission Policy v0&amp;rdquo; (max context / max steps / max budget)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Day 31–60: Establish Resource Governance (GPU Governance + Scheduling)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This phase requires evaluating GPU sharing/isolation strategies, introducing topology/networking constraints, and forming two golden paths for inference and training.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPU sharing/isolation strategy: MIG/MPS/vGPU/DRA path evaluation and PoC (executable strategy as acceptance criteria)&lt;/li&gt;
&lt;li&gt;Introduce topology/networking constraints, form AI-ready network baseline and capacity assumptions (including acceptance criteria)&lt;/li&gt;
&lt;li&gt;Form two templated delivery paths for inference/training&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Deliverables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Workload templates (1 each for inference and training)&lt;/li&gt;
&lt;li&gt;Scheduling and isolation strategies (whitelisted, auditable)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Day 61–90: Establish the Loop (Enforcement + Feedback)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The final phase requires executing budget policies, migrating pilot use cases to the landing zone, and solidifying organizational interfaces.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Execute budgets: rate limiting/queuing/preemption/degradation strategies, linked to SLOs&lt;/li&gt;
&lt;li&gt;Migrate pilot use cases to landing zone (or service landing zone capabilities)&lt;/li&gt;
&lt;li&gt;Solidify &amp;ldquo;organizational interface&amp;rdquo;: platform team vs workload team responsibility boundaries (forming executable contracts)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Deliverables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;AI Platform Runbook v1&amp;rdquo; (including oncall, changes, cost auditing)&lt;/li&gt;
&lt;li&gt;Two replicable use case landing paths (new use cases &amp;lt;= 30 minutes to golden path)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="operating-model-the-contract-between-platform-teams-and-workload-teams"&gt;Operating Model: The &amp;ldquo;Contract&amp;rdquo; Between Platform Teams and Workload Teams&lt;/h2&gt;
&lt;p&gt;Migration success depends on establishing clear, executable &amp;ldquo;organizational contracts.&amp;rdquo; The contract essence: who is responsible for &amp;ldquo;capability provision,&amp;rdquo; who is responsible for &amp;ldquo;behavioral consequences.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Platform teams provide (must be stable)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Landing zone, network baseline, identity and policies, budget/quota systems, metering/attribution, GPU governance capabilities, runtime golden paths&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Workload teams own (must be self-service)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Model selection, prompt/agent logic, tool integration, SLO definition, business value measurement, use case risk classification and rollback paths&lt;/p&gt;
&lt;p&gt;This is also why the FinOps Framework emphasizes operating model (personas, capabilities, maturity) rather than just tools: without &amp;ldquo;contracts,&amp;rdquo; budgets are difficult to execute; if budgets cannot execute, loops cannot form.&lt;/p&gt;
&lt;h2 id="migration-anti-patterns"&gt;Migration Anti-Patterns&lt;/h2&gt;
&lt;p&gt;Below are common migration anti-patterns and their consequences:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Anti-Pattern&lt;/th&gt;
&lt;th&gt;Typical Consequences&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Build only API/Agent platform, without ledger and budget&lt;/td&gt;
&lt;td&gt;runaway cost (most common, and difficult to remediate afterwards)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Treat GPUs as ordinary resources, without sharing/isolation and scheduling upgrades&lt;/td&gt;
&lt;td&gt;Low utilization + uncontrolled contention, platform forced to allocate compute via &amp;ldquo;administrative means&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ignore networking and topology&lt;/td&gt;
&lt;td&gt;Tail latency and training JCT amplified, capacity planning fails, SLOs difficult to honor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context not assetized (only &amp;ldquo;tricky caching&amp;rdquo; within applications)&lt;/td&gt;
&lt;td&gt;Unit cost out of control in long context/agentic era, reuse capabilities difficult to solidify as platform capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Common Migration Anti-Patterns and Consequences
&lt;/figcaption&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The core of AI-native migration is not a &amp;ldquo;migration checklist,&amp;rdquo; but &lt;strong&gt;under uncertainty premises, incorporating cost, risk, and tail performance into a unified governance loop, using Landing Zone to carry organizational contracts, and using Context Tier to implement state reuse infrastructure capabilities&lt;/strong&gt;. Only in this way can platform and business maintain controllability and efficiency during scaled evolution.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights" target="_blank" rel="noopener"&gt;McKinsey on AI Strategy - mckinsey.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.thoughtworks.com/radar" target="_blank" rel="noopener"&gt;Thoughtworks Technology Radar - thoughtworks.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/adoption-framework" target="_blank" rel="noopener"&gt;Google Cloud Adoption Framework - cloud.google.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Glossary</title><link>https://jimmysong.io/book/ai-native-infra/glossary/</link><pubDate>Sun, 18 Jan 2026 05:19:24 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/glossary/</guid><description>Bilingual glossary of core AI-native infrastructure terminology for aligning organizational language.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Unified terminology is the first step toward organizational consensus.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Conclusion first: in the context of AI-native infrastructure, key terms must remain consistent; otherwise, both governance and communication lose focus.&lt;/p&gt;
&lt;p&gt;The following glossary serves to align cross-team terminology.&lt;/p&gt;
&lt;h2 id="core-terms"&gt;Core Terms&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;AI Native Infrastructure&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;An infrastructure system premised on &amp;ldquo;models/agents as execution entities, compute as scarce assets, and uncertainty as the norm,&amp;rdquo; closed-looping &amp;ldquo;intent → execution → resource consumption → economic and risk outcomes&amp;rdquo; through compute governance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model-as-Actor&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Models/agents become &amp;ldquo;execution entities&amp;rdquo; with action capabilities, capable of invoking tools, modifying system state, and producing side effects, thus requiring governance and audit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compute-as-Scarcity&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Compute (GPU, interconnects, power consumption, bandwidth) becomes the core scarce asset, with expansion constrained by supply chain and data center conditions, and costs that cannot be made elastic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Uncertainty-by-Default&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Behavior and resource consumption are highly uncertain (especially in agentic and long-context scenarios), requiring verification and fallback mechanisms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intent Plane&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The API, Agent, and policy expression layer responsible for expressing &amp;ldquo;what I want,&amp;rdquo; including priorities, budgets, compliance, and other policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Execution Plane&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The training/inference/serving/runtime layer responsible for translating intent into actual execution, including state management, tool invocation, model routing, and so on.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Governance Plane&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The quota/budget, isolation/sharing, and cost control layer responsible for bounding resource consequences, including topology-aware scheduling, SLO and risk policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Loop&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Possessing a closed loop of &amp;ldquo;intent → consumption → cost/risk outcomes,&amp;rdquo; comprising four steps: Admission, Translation, Metering, and Enforcement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compute Governance&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Governing the resource consequences of intent, including four categories of objects: token economics, accelerator time, interconnect and storage, and organizational budgets and risks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;FinOps / Financial Operations&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Embedding cost governance early into architecture so that every scaling decision simultaneously answers &amp;ldquo;whether performance meets requirements&amp;rdquo; and &amp;ldquo;whether it is affordable.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;An execution entity that completes tasks by selecting tools, invoking tools, and iterating reasoning, with uncertain behavioral paths and resource consumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MCP / Model Context Protocol&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A protocol that standardizes tool access as &amp;ldquo;declaratable capability boundaries,&amp;rdquo; defining how capabilities are exposed to models/agents and how they are invoked.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Operating Model&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Institutional design for organization and operational methods, including responsibility boundaries, collaboration mechanisms, and decision-making processes, answering &amp;ldquo;who is responsible for what and what are the costs of failure.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.iso.org/standard/74296.html" target="_blank" rel="noopener"&gt;ISO/IEC 22989 AI Concepts and Terminology&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.nist.gov/ai" target="_blank" rel="noopener"&gt;NIST AI Glossary&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://oecd.ai/en/ai-glossary" target="_blank" rel="noopener"&gt;OECD AI Glossary&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Executive Checklist (10 Questions)</title><link>https://jimmysong.io/book/ai-native-infra/executive-checklist/</link><pubDate>Sun, 18 Jan 2026 05:21:59 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/book/ai-native-infra/executive-checklist/</guid><description>Ten critical questions for CEO/CTO to evaluate AI-native infrastructure readiness.</description><content:encoded>
&lt;p&gt;The following 10 questions assess whether an organization possesses the strategic and execution readiness for AI-native infrastructure. The diagram below categorizes these questions into three domains: strategy, governance, and execution, facilitating executive discussions.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/book/ai-native-infra/executive-checklist/executive-checklist-en.svg" data-img="https://assets.jimmysong.io/images/book/ai-native-infra/executive-checklist/executive-checklist-en.svg" alt="Figure 1: Executive checklist structure diagram" data-caption="Figure 1: Executive checklist structure diagram"
width="1616"
height="416"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Executive checklist structure diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol&gt;
&lt;li&gt;Can you clearly define the &lt;strong&gt;unit cost&lt;/strong&gt; for each major AI workload (e.g., per 1M tokens, per agent task, per batch job)?&lt;/li&gt;
&lt;li&gt;Do you have &lt;strong&gt;budget/quota mechanisms&lt;/strong&gt; that can constrain team/project/tenant compute consumption within controllable bounds?&lt;/li&gt;
&lt;li&gt;Can you make &lt;strong&gt;explicit policy trade-offs&lt;/strong&gt; between &amp;ldquo;performance (throughput/latency)—cost—risk&amp;rdquo; (rather than relying on verbal constraints)?&lt;/li&gt;
&lt;li&gt;Can your platform handle &lt;strong&gt;uncertainty&lt;/strong&gt;: spikes, long-tail effects, and resource fluctuations caused by agent path explosions?&lt;/li&gt;
&lt;li&gt;Are agent/MCP &amp;ldquo;intents&amp;rdquo; mapped to &lt;strong&gt;actionable and billable/auditable resource consequences&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;Do you have clear &lt;strong&gt;resource isolation and sharing strategies&lt;/strong&gt; (same-card sharing, memory isolation, preemption, prioritization) to improve utilization?&lt;/li&gt;
&lt;li&gt;Can you achieve cross-layer observability: end-to-end tracing from &lt;strong&gt;request/agent → runtime → GPU/network/storage → cost&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;Does your infrastructure support rapid adoption of new hardware/interconnect/topology changes (heterogeneity and evolution are the norm)?&lt;/li&gt;
&lt;li&gt;Has the organization established &lt;strong&gt;&amp;ldquo;AI SRE/ModelOps + FinOps&amp;rdquo;&lt;/strong&gt; collaboration mechanisms and accountability boundaries (who owns cost and reliability)?&lt;/li&gt;
&lt;li&gt;When you say &amp;ldquo;we are AI-native,&amp;rdquo; can you provide &lt;strong&gt;three planes + one closed loop&lt;/strong&gt; architecture diagram and governance strategy on a single page?&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://hbr.org/topic/artificial-intelligence" target="_blank" rel="noopener"&gt;Harvard Business Review - AI Strategy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sloanreview.mit.edu/tag/artificial-intelligence/" target="_blank" rel="noopener"&gt;MIT Sloan - Executive Guide to AI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.weforum.org/agenda/archive/artificial-intelligence/" target="_blank" rel="noopener"&gt;World Economic Forum - AI Governance&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>HAMi Community Evolution: When AI Writes Code, What Makes an Open Source Community Valuable?</title><link>https://jimmysong.io/slide/hami-community-evolution-slide/</link><pubDate>Wed, 15 Jul 2026 09:17:07 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/slide/hami-community-evolution-slide/</guid><description>A long-term participant&amp;#39;s observation on HAMi&amp;#39;s growth (2021 open source, 2024 CNCF Sandbox, 2026 CNCF Incubating) and what makes an open source community valuable in the AI era — AI lowers the cost of producing code, but not the cost of building consensus. Tomorrow&amp;#39;s maintainers are consensus builders.</description><content:encoded>
&lt;p&gt;This deck uses HAMi&amp;rsquo;s growth story to explore what makes an open source community valuable in the AI-coding era. HAMi is a GPU virtualization and sharing middleware for Kubernetes, evolving from its 2021 open-source debut to the 2024 CNCF Sandbox and 2026 CNCF Incubating milestones. The core thesis: AI reduces the cost of producing code, but it does not reduce the cost of building consensus — the moat of an open source community is shifting from code production toward consensus and trust.&lt;/p&gt;
&lt;p&gt;Use the controls or keyboard shortcuts below to navigate the embedded interactive slides.&lt;/p&gt;
&lt;div class="slide-embed-container" data-height="auto" data-style=""&gt;
&lt;iframe
loading="lazy"
src="https://jimmysong.io/slides/hami-community-evolution/index-en.html"
allowfullscreen="allowfullscreen"
allow="fullscreen"
title="Slide Presentation"
class="slide-embed-iframe"&gt;
&lt;/iframe&gt;
&lt;/div&gt;
&lt;figcaption&gt;Slide: HAMi Community Evolution&lt;/figcaption&gt;
&lt;h2 id="key-topics"&gt;Key Topics&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A journey from code to community&lt;/strong&gt; — 2021 open source → 2024 CNCF Sandbox → 2026 CNCF Incubating&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community growth is the real infrastructure&lt;/strong&gt; — 3,709 stars · 129 contributors · 16 releases · 609 forks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Organizations and geography&lt;/strong&gt; — 863 organizations across 44 countries&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI changed software development&lt;/strong&gt; — &amp;ldquo;Before AI coding&amp;rdquo; vs. the &amp;ldquo;AI coding era&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The maintainer&amp;rsquo;s new role&lt;/strong&gt; — technical direction, community trust, coordination; future maintainers are consensus builders&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance must evolve with AI&lt;/strong&gt; — HAMi PR #2019 on disclosing AI-generated contributions&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Cloud Native Community China Is Officially Established</title><link>https://jimmysong.io/notice/cloud-native-community-china-established/</link><pubDate>Sat, 11 Jul 2026 10:00:00 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/notice/cloud-native-community-china-established/</guid><description>Cloud Native Community China, part of the CNCF global community network, is now live.</description><content:encoded>
&lt;p&gt;On July 11, 2026, &lt;a href="https://ocgroups.dev/cncf/group/cncf-china" target="_blank" rel="noopener"&gt;Cloud Native Community China&lt;/a&gt; was officially established. It is the CNCF-recognized cloud native community for China, hosted on the &lt;a href="https://ocgroups.dev" target="_blank" rel="noopener"&gt;Open Community Groups&lt;/a&gt; platform.&lt;/p&gt;
&lt;p&gt;China has no shortage of cloud native developers, contributors, and end-user companies, yet we&amp;rsquo;ve never had an official, unified, continuously-synced community entry point. This launch fills exactly that gap.&lt;/p&gt;
&lt;h2 id="about-cloud-native-community-china"&gt;About Cloud Native Community China&lt;/h2&gt;
&lt;p&gt;CNCF has communities across many countries and regions worldwide. Each one connects its local cloud native developers, open source projects, end-user companies, and in-person events. Cloud Native Community China sits alongside these as a peer, together forming CNCF&amp;rsquo;s global community map.&lt;/p&gt;
&lt;p&gt;Cloud native activity in China used to be scattered across city-level meetups, project SIGs, and individual company blogs. It was hard to know where the next worthwhile talk was happening, or what peers were paying attention to. The point of an official community site is to wire all those scattered nodes into one network.&lt;/p&gt;
&lt;h2 id="what-the-site-is-for"&gt;What the Site Is For&lt;/h2&gt;
&lt;p&gt;The newly launched &lt;a href="https://ocgroups.dev/cncf/group/cncf-china" target="_blank" rel="noopener"&gt;cncf-china community page&lt;/a&gt; does three core things.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Aggregating China&amp;rsquo;s cloud native events&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Meetups, online talks, KubeCon meetups, project SIG meetings, and training sessions all get published here in one place. You no longer have to dig through separate channels to find what&amp;rsquo;s happening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Syncing community updates&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Project progress, releases, governance updates, and volunteer calls-to-action get pushed out through the community site, giving anyone following cloud native a stable information source.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Receiving notifications and staying connected&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Once you join, you can subscribe to notifications and get new events and updates delivered first. For anyone who wants to contribute, meet peers, or keep up with the latest in cloud native, this is the most direct entry point today.&lt;/p&gt;
&lt;h2 id="founding-members"&gt;Founding Members&lt;/h2&gt;
&lt;p&gt;A community&amp;rsquo;s launch always rests on a group of people willing to carry it at the start. The founding members of Cloud Native Community China come from maintainers of active Chinese cloud native open source projects, end-user company representatives, and community organizers who have run local meetups for years. Most have been deeply involved in the CNCF ecosystem for a long time, across areas like Kubernetes, Istio, HAMi, observability, and AI infrastructure.&lt;/p&gt;
&lt;p&gt;For the full, up-to-date founding member roster, check the &lt;a href="https://ocgroups.dev/cncf/group/cncf-china" target="_blank" rel="noopener"&gt;cncf-china community page&lt;/a&gt;. To become a contributor or co-organizer, you can reach them through the community page as well.&lt;/p&gt;
&lt;h2 id="join-the-community"&gt;Join the Community&lt;/h2&gt;
&lt;p&gt;If you&amp;rsquo;re a developer, SRE, or architect in the cloud native space, or your company runs on Kubernetes, Istio, or serverless, joining this community gives you at least three things:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Never miss an event&lt;/strong&gt;: All official events are published in one place; just subscribe to notifications.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meet peers and project maintainers&lt;/strong&gt;: A community is fundamentally a human network, and a lot of collaboration starts at a single meetup.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find a path into open source&lt;/strong&gt;: If you want to contribute to CNCF projects but don&amp;rsquo;t know where to start, the community has ready-made guidance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Visit &lt;a href="https://ocgroups.dev/cncf/group/cncf-china" target="_blank" rel="noopener"&gt;Cloud Native Community China&lt;/a&gt;, join the community on the page, and subscribe to notifications. Once you&amp;rsquo;re in, you&amp;rsquo;ll receive pushes about cloud native events, updates, and community news in China.&lt;/p&gt;
&lt;p&gt;Cloud native adoption in China is already large in scale, but the community&amp;rsquo;s organizational density has room to grow. Establishing Cloud Native Community China is the first step in wiring that momentum into CNCF&amp;rsquo;s global network. You&amp;rsquo;re welcome to join.&lt;/p&gt;</content:encoded></item><item><title>After HAMi Became a CNCF Incubating Project: Open Source Is Moving from Code to Consensus</title><link>https://jimmysong.io/blog/hami-cncf-incubating-consensus/</link><pubDate>Wed, 08 Jul 2026 13:51:44 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/hami-cncf-incubating-consensus/</guid><description>Code is cheap; consensus is the new scarce good.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;After HAMi became a CNCF Incubating project, I want to talk about something overlooked: AI is shifting the scarce resource of open source communities from code to consensus.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;On July 2, 2026, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; officially became a CNCF Incubating project (see &lt;a href="https://project-hami.io/blog/hami-cncf-incubating" target="_blank" rel="noopener"&gt;the announcement&lt;/a&gt;). For an open source project, this means more than recognition of its technical capability; it means that community governance, ecosystem building, and real-world adoption have all entered a new phase.&lt;/p&gt;
&lt;p&gt;But if you only read HAMi&amp;rsquo;s growth as &amp;ldquo;a GPU virtualization project succeeded,&amp;rdquo; you might miss the more important shift.&lt;/p&gt;
&lt;p&gt;My time building the HAMi community has left me with one increasingly strong feeling: AI is changing how open source communities produce. As AI coding drives the cost of producing code lower and lower, the core of competition in open source will no longer be who wrote the most code, but who can build stronger technical consensus, attract more contributors, and form a sustainable ecosystem network.&lt;/p&gt;
&lt;p&gt;This article is my attempt to lay out that argument clearly.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-incubating.webp" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-incubating.webp" alt="Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project" data-caption="Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project"
width="1024"
height="1536"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="hamis-growth-shows-that-a-projects-real-asset-is-its-community"&gt;HAMi&amp;rsquo;s Growth Shows That a Project&amp;rsquo;s Real Asset Is Its Community&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s start with the data. Here is where HAMi stands today:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Stars&lt;/td&gt;
&lt;td&gt;3,700+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contributors&lt;/td&gt;
&lt;td&gt;Nearly 500, from 27 countries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Participating organizations&lt;/td&gt;
&lt;td&gt;Multiple, and growing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release cadence&lt;/td&gt;
&lt;td&gt;Once every three months&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: HAMi community key metrics (as of July 2026)
&lt;/figcaption&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-contributors-map.webp" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-contributors-map.webp" alt="Figure 10: HAMi community contributors map, contributors from 27 countries" data-caption="Figure 10: HAMi community contributors map, contributors from 27 countries"
width="2050"
height="870"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 10: HAMi community contributors map, contributors from 27 countries&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I&amp;rsquo;m not listing these numbers to show off &amp;ldquo;growth metrics.&amp;rdquo; I&amp;rsquo;m making a different point: an open source project is shifting from &amp;ldquo;software maintained by a team&amp;rdquo; into &amp;ldquo;a technical community of people gathered around a shared goal.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;That distinction matters. Software can be forked, rewritten, or generated overnight by AI. But a community with a shared goal, trust, and rhythm cannot be forked. That is the irreplaceable asset of an open source project.&lt;/p&gt;
&lt;p&gt;From a governance-maturity perspective, HAMi has passed through three milestones:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/a221c9ebe476891231ef44e3de11b783.svg" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/a221c9ebe476891231ef44e3de11b783.svg" alt="Figure 11: Evolution of HAMi’s governance maturity" data-caption="Figure 11: Evolution of HAMi’s governance maturity"
width="1372"
height="168"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 11: Evolution of HAMi’s governance maturity&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Note the middle stretch: from entering Sandbox in August 2024 to reaching Incubating in July 2026, roughly two years. During those two years the code certainly grew, but what actually convinced the CNCF Technical Oversight Committee (TOC) was the diversification of the community, the formalization of governance, and real production adoption. None of that is something you produce by writing code.&lt;/p&gt;
&lt;h2 id="when-hami-open-sourced-in-2021-there-was-no-ai-coding"&gt;When HAMi Open-Sourced in 2021, There Was No AI Coding&lt;/h2&gt;
&lt;p&gt;HAMi was first open-sourced in 2021. Back then, most developers did not see AI coding the way we do today. Whether an open source project survived depended on developers genuinely investing their time, discussing problems in issues, submitting code through pull requests, and building trust through code review.&lt;/p&gt;
&lt;p&gt;Today, the environment has changed.&lt;/p&gt;
&lt;p&gt;In a recent HAMi community livestream (&lt;a href="https://www.bilibili.com/video/BV1r6EC6zEWS/" target="_blank" rel="noopener"&gt;Mastering HAMi DRA, Yang Shouren, HAMi Community Livestream Episode 2&lt;/a&gt;), someone asked the maintainers a question: &amp;ldquo;How much of HAMi&amp;rsquo;s code now comes from AI assistance?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The answer from &lt;a href="https://github.com/shouren" target="_blank" rel="noopener"&gt;Yang Shouren&lt;/a&gt; stuck with me: &lt;strong&gt;about half of the code in the HAMi community today is already AI-assisted.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Half. And that share is still rising.&lt;/p&gt;
&lt;p&gt;This immediately raises a sharp question: if AI can write more and more code, what is the value of an open source community? Could one person plus a few AI agents just fork a &amp;ldquo;new HAMi&amp;rdquo;?&lt;/p&gt;
&lt;p&gt;My answer is no, because what AI lowers is the cost of producing code, not the cost of building technical consensus.&lt;/p&gt;
&lt;h2 id="in-the-ai-era-code-is-no-longer-scarce-consensus-is"&gt;In the AI Era, Code Is No Longer Scarce; Consensus Is&lt;/h2&gt;
&lt;p&gt;Let me sharpen that point.&lt;/p&gt;
&lt;p&gt;An AI agent can already do a lot today: write code, fix bugs, add tests, generate docs, produce migration scripts. These capabilities are getting stronger fast. But there are a few things in a community that AI cannot replace today, and in my judgment will not replace soon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, deciding which problems are worth solving.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Take HAMi. Why does GPU sharing matter? Not because &amp;ldquo;slicing cards finely&amp;rdquo; is a cool technique. It matters because once AI infrastructure scales, reality looks like this: GPU costs are enormous, heterogeneous hardware keeps multiplying, and Kubernetes&amp;rsquo; native resource model is no longer enough.&lt;/p&gt;
&lt;p&gt;The community has to first agree that &amp;ldquo;this problem is worth investing in&amp;rdquo; before anyone writes any code. That agreement is a human-to-human matter, supported by real scenarios, real costs, and real pain. AI can solve a problem you have already defined, but &amp;ldquo;which problem is worth defining&amp;rdquo; is decided by community consensus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, choosing a technical path.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In GPU virtualization there are many paths: MIG, MPS, time-slicing, vGPU, DRA. Each has trade-offs, and choosing wrong can cost you two years of detours.&lt;/p&gt;
&lt;p&gt;Code can be generated, but architectural choice is fundamentally a value judgment. HAMi&amp;rsquo;s decision on Ascend 910C to move from hardware SR-IOV to userspace HAMi-core was not about someone writing a better piece of code; it was about the maintainers holding to a judgment that &amp;ldquo;hardware partitioning is too coarse, software partitioning is more flexible.&amp;rdquo; That kind of judgment is ground out through repeated discussion, failure, and validation in the community, not prompted out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Third, trust.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Users don&amp;rsquo;t choose HAMi because of &amp;ldquo;how much AI-generated code is in this repo.&amp;rdquo; They care about: who maintains it? Who reviews it? Are there real production cases? Does the community respond when something breaks?&lt;/p&gt;
&lt;p&gt;Each of these is a relationship between people, a product of community organization, not a product of code quality.&lt;/p&gt;
&lt;p&gt;Put these three together, and the production model of open source communities in the AI era is shifting:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/3cc15ad018c9b11909feea79769e166c.svg" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/3cc15ad018c9b11909feea79769e166c.svg" alt="Figure 12: How the open source production model is changing in the AI era" data-caption="Figure 12: How the open source production model is changing in the AI era"
width="2333"
height="357"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 12: How the open source production model is changing in the AI era&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the past, code came first and the community sedimented out of the code; in the future, consensus comes first, AI rapidly turns consensus into code, and code flows back to test the consensus. The center of gravity of the scarce resource moves from &amp;ldquo;code&amp;rdquo; on the left to &amp;ldquo;consensus&amp;rdquo; on the right.&lt;/p&gt;
&lt;h2 id="in-the-ai-era-open-source-governance-itself-has-to-level-up"&gt;In the AI Era, Open Source Governance Itself Has to Level Up&lt;/h2&gt;
&lt;p&gt;Since AI has become a new category of contributor, a community&amp;rsquo;s governance rules have to keep up.&lt;/p&gt;
&lt;p&gt;My advice is: don&amp;rsquo;t treat AI merely as a tool, treat it as a new type of contributor. It used to be &amp;ldquo;developers write code, humans review&amp;rdquo;; in the future it will be &amp;ldquo;humans set intent, AI generates code, the community reviews, and shared knowledge is distilled.&amp;rdquo; There is an extra layer in the middle, and an extra layer of governance complexity.&lt;/p&gt;
&lt;p&gt;HAMi is already responding to this. Its &lt;a href="https://github.com/Project-HAMi/HAMi/blob/master/CONTRIBUTING.md" target="_blank" rel="noopener"&gt;CONTRIBUTING.md&lt;/a&gt; is explicit:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you are using any kind of AI assistance to contribute to HAMi, it must be disclosed in the pull request.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In other words, if you use AI to help with a contribution, you must declare it in the PR. But the community also knows that a norm without a gate is not enough (see the discussion in &lt;a href="https://github.com/Project-HAMi/HAMi/issues/1998" target="_blank" rel="noopener"&gt;Issue #1998&lt;/a&gt;), and there is already ongoing discussion about how to give that norm real enforcement.&lt;/p&gt;
&lt;p&gt;This is actually a problem every AI-era open source project will run into. I&amp;rsquo;d break it into a few questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Must AI-generated code always be declared?&lt;/li&gt;
&lt;li&gt;Which model, and what context, did the contributor use?&lt;/li&gt;
&lt;li&gt;How do you ensure the security of AI code, avoiding injection and licensing risks?&lt;/li&gt;
&lt;li&gt;What process should maintainers use to review an AI diff they may not be able to fully trace themselves?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Whoever figures out and operationalizes these rules first will keep contribution quality stable in the AI era. HAMi&amp;rsquo;s exploration here is worth a look for every open source project.&lt;/p&gt;
&lt;h2 id="what-cncf-incubating-really-means"&gt;What CNCF Incubating Really Means&lt;/h2&gt;
&lt;p&gt;Back to the promotion itself.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d lean against treating it as an &amp;ldquo;honor.&amp;rdquo; What CNCF Incubating really validates is not code quality, but whether a project has the capacity to become infrastructure. It examines a whole package: technical maturity, community governance, production adoption, and ecosystem building.&lt;/p&gt;
&lt;p&gt;HAMi&amp;rsquo;s case, in one sentence, is not &amp;ldquo;a Chinese team built a GPU project.&amp;rdquo; It is this:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI Infrastructure community, jointly shaped by developers from around the world, is taking shape.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The first is a product story; the second is an ecosystem story. The Incubating recognition from the CNCF is recognizing the latter, because competition over infrastructure is never competition between individual products; it is competition between ecosystem networks.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The open source competition of the next decade will not be just a competition of code, but a competition of communities.&lt;/p&gt;
&lt;p&gt;Once AI gives everyone near-infinite capacity to produce code, the truly scarce capabilities will be three: finding the right problem, building technical consensus, and organizing developers worldwide to solve a problem together.&lt;/p&gt;
&lt;p&gt;HAMi&amp;rsquo;s path to CNCF Incubating is just one snapshot of how open source communities are evolving in the AI era. Code will keep getting cheaper, and consensus will keep getting more expensive. Whoever understands this inversion will be the one who can build open source communities with real depth in the AI era.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Join the HAMi Community
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Add me on WeChat (&lt;code&gt;jimmysong&lt;/code&gt;) or follow &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi on GitHub&lt;/a&gt; to join the community focused on GPU virtualization and heterogeneous compute scheduling. Let&amp;rsquo;s talk about open source governance in the AI era.
&lt;/div&gt;
&lt;/div&gt;</content:encoded></item><item><title>Olares and HAMi: A New Inflection Point for Desktop AI Workstations</title><link>https://jimmysong.io/blog/olares-hami-edge-ai-cloud/</link><pubDate>Wed, 24 Jun 2026 22:30:00 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/olares-hami-edge-ai-cloud/</guid><description>HAMi moves from cluster to desktop with Olares.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;HAMi used to save cards in the cluster. Now it decides how good a desktop AI workstation feels.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/banner.webp" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/banner.webp" alt="Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations" data-caption="Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations"
width="1983"
height="793"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="an-old-friend-mentions-a-name"&gt;An Old Friend Mentions a Name&lt;/h2&gt;
&lt;p&gt;A few days ago an old friend came by to chat. We used to run the cloud-native community scene together back home, so we go way back. She recently joined a company called Olares, and somewhere in the conversation she dropped this: their project had integrated HAMi.&lt;/p&gt;
&lt;p&gt;I run the HAMi community day to day, so whenever I hear someone using it for something, I want to take a closer look. I went and dug through the &lt;a href="https://github.com/beclab/Olares" target="_blank" rel="noopener"&gt;Olares&lt;/a&gt; repo and website, and my first reaction was, huh, this is actually interesting.&lt;/p&gt;
&lt;p&gt;Isn&amp;rsquo;t this exactly the kind of local AI workstation I&amp;rsquo;d been eyeing forever but never pulled the trigger on? I wrote in &lt;a href="https://jimmysong.io/blog/personal-ai-stack"&gt;My Personal AI Stack&lt;/a&gt; that for someone like me who mainly works with my head, subscribing to models beats maintaining a high-end GPU. But how far can &amp;ldquo;one machine running the full AI stack&amp;rdquo; really go, and who has actually built it, I&amp;rsquo;ve wanted to see with my own eyes.&lt;/p&gt;
&lt;h2 id="what-is-this-thing-really"&gt;What Is This Thing, Really&lt;/h2&gt;
&lt;p&gt;Olares isn&amp;rsquo;t a &amp;ldquo;NAS with a GPU bolted on,&amp;rdquo; and it isn&amp;rsquo;t a &amp;ldquo;mini PC with Ollama installed.&amp;rdquo; It&amp;rsquo;s more like a personal cloud built on Kubernetes: local models, apps, identity, remote access, dev environment, storage, even GPU governance, all packed into one machine as a desktop cloud OS.&lt;/p&gt;
&lt;p&gt;HAMi isn&amp;rsquo;t filler here. It&amp;rsquo;s the layer that turns &amp;ldquo;one card&amp;rdquo; into &amp;ldquo;a resource pool you can share, isolate, and schedule.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="how-it-differs-from-the-ai-mini-pc-crowd"&gt;How It Differs from the AI Mini-PC Crowd&lt;/h2&gt;
&lt;p&gt;There&amp;rsquo;s no shortage of things calling themselves AI workstations, but most are just beefier mini PCs. Olares is different. It&amp;rsquo;s genuinely designed as a cloud.&lt;/p&gt;
&lt;p&gt;The company positions this machine outright as a 24/7 personal AI cloud, not a PC you sit in front of. That single framing explains almost everything that follows: why it has to use Kubernetes, why it cares so much about network traversal and remote access, why GPU governance suddenly matters here.&lt;/p&gt;
&lt;p&gt;Two other details tell you a lot: it supports Thunderbolt 5 external eGPUs, and two machines can cluster up. In other words, it was never a sealed box. It starts as a single node and grows upward.&lt;/p&gt;
&lt;h2 id="stuffing-a-whole-cloud-runtime-into-one-box"&gt;Stuffing a Whole Cloud Runtime into One Box&lt;/h2&gt;
&lt;p&gt;Architecturally, what Olares does is compress an entire cloud runtime into a single machine: auth and authorization, app lifecycle, tunnels and traversal, secrets, observability middleware, not a layer skipped.&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t &amp;ldquo;a few AI apps preinstalled.&amp;rdquo; It&amp;rsquo;s a complete cloud runtime stuffed into a desktop device.&lt;/p&gt;
&lt;p&gt;The diagram below is my own re-drawn layering based on its public docs, with the HAMi layer pulled out explicitly because it&amp;rsquo;s the point of this whole piece.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/olares-architecture-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/olares-architecture-en.svg" alt="Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane" data-caption="Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane"
width="1382"
height="1392"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="one-detail-that-caught-my-eye"&gt;One Detail That Caught My Eye&lt;/h2&gt;
&lt;p&gt;One thing jumped out while I was reading: &lt;strong&gt;Olares&amp;rsquo;s docs haven&amp;rsquo;t kept up with its own product.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Its 2025 architecture page still says &lt;code&gt;nvshare&lt;/code&gt;, noting that GPUs are limited to one card per node. But flip through the release notes and 1.12 already integrates the HAMi scheduler, with exclusive, time-slicing, and memory-slicing modes all there; 1.12.2 adds multi-GPU; 1.12.5 supports DGX Spark outright and rolls automatic scheduling across all three modes.&lt;/p&gt;
&lt;p&gt;The product is outrunning its docs. That alone tells you something: it&amp;rsquo;s shifting from a &amp;ldquo;personal cloud OS&amp;rdquo; toward a &amp;ldquo;local AI cloud OS.&amp;rdquo; That&amp;rsquo;s also why I wanted to write a whole piece on it.&lt;/p&gt;
&lt;h2 id="what-hami-actually-does-in-there"&gt;What HAMi Actually Does in There&lt;/h2&gt;
&lt;p&gt;HAMi is a CNCF Sandbox project, positioned as a heterogeneous AI compute virtualization middleware. On the path there&amp;rsquo;s a Webhook, a scheduler, a Device Plugin, and HAMi-core handling in-container resource control. I summed it up in one line in &lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes Is Becoming the GPU Control Plane of the AI Era&lt;/a&gt;: it turns GPU slicing from &amp;ldquo;a hardware capability&amp;rdquo; into &amp;ldquo;a control-plane capability.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;On Olares, that&amp;rsquo;s the real thing behind the GPU mode toggle in settings. On the surface it&amp;rsquo;s a UI option; underneath, it&amp;rsquo;s a policy switch on the resource plane.&lt;/p&gt;
&lt;p&gt;The most critical piece is almost certainly HAMi-core, which in one sentence: intercepts CUDA calls inside the container and does memory virtualization, compute throttling, and utilization monitoring. It doesn&amp;rsquo;t slice with hardware. It manages with software.&lt;/p&gt;
&lt;p&gt;The diagram below puts that injection path and the three GPU modes side by side.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-core-gpu-modes-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-core-gpu-modes-en.svg" alt="Figure 3: HAMi-core injection path and Olares’s three GPU modes" data-caption="Figure 3: HAMi-core injection path and Olares’s three GPU modes"
width="1421"
height="1022"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: HAMi-core injection path and Olares’s three GPU modes&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The evidence lines up too. In Olares&amp;rsquo;s GitHub issues you can see HAMi&amp;rsquo;s &lt;code&gt;libvgpu.so&lt;/code&gt; crashing under WSL2, the discussion mentions it getting injected into every process via &lt;code&gt;/etc/ld.so.preload&lt;/code&gt;, and the Olares team itself says this implementation &amp;ldquo;drew heavy inspiration&amp;rdquo; from the HAMi project.&lt;/p&gt;
&lt;p&gt;One thing I should be clear about, though: whether Olares uses upstream HAMi-core directly or maintains a forked version, there&amp;rsquo;s no single official answer in public. I won&amp;rsquo;t fill that in for them.&lt;/p&gt;
&lt;p&gt;The three modes make sense in order. Full-card exclusive is for heavy loads; time-slicing lets lightweight services take turns, and Olares even swaps inactive models into memory first, with roughly 5% switching overhead; memory-slicing cuts the memory into fixed quotas so multiple apps run together. On something like DGX Spark, where CPU and GPU share memory, it defaults to memory-slicing because there&amp;rsquo;s no traditional memory paging in and out to begin with.&lt;/p&gt;
&lt;h2 id="why-this-matters-for-hami"&gt;Why This Matters for HAMi&lt;/h2&gt;
&lt;p&gt;HAMi used to live mostly in big clusters: multi-tenant, inference serving, mixed training-and-inference, heterogeneous cards. NIO running it for sharing across 80 nodes and 600 cards is the textbook example. That line is well-trodden.&lt;/p&gt;
&lt;p&gt;What feels new about Olares is that it moves the same problem onto a single machine.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-scene-shift-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-scene-shift-en.svg" alt="Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience" data-caption="Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience"
width="1382"
height="1042"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What you actually run at once on Olares is never just one Ollama: a local model, a chat UI, a research agent, an image or video pipeline, plus a pile of OCR, speech-to-text, and automation tools. In a scenario like that, without a GPU scheduling layer, the GPU always degenerates into &amp;ldquo;whoever starts first grabs it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;So for HAMi this matters. It goes from &amp;ldquo;a tool that saves cards for platform engineers&amp;rdquo; to &amp;ldquo;something that directly decides whether an ordinary user has a good time.&amp;rdquo; In the cluster it manages utilization and cost; on Olares it manages concurrent experience, model switching, and whether this machine is actually pleasant to use. HAMi doesn&amp;rsquo;t only talk to platform engineers anymore.&lt;/p&gt;
&lt;h2 id="where-edge-ai-goes-from-here"&gt;Where Edge AI Goes from Here&lt;/h2&gt;
&lt;p&gt;My sense is, edge AI won&amp;rsquo;t settle at &amp;ldquo;a stronger local model box.&amp;rdquo; It&amp;rsquo;ll grow into a full stack of &amp;ldquo;control plane plus model plane plus application plane.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;NVIDIA pushing DGX Spark and RTX Spark onto the desktop has already moved the story from &amp;ldquo;run models locally&amp;rdquo; to &amp;ldquo;run agents locally.&amp;rdquo; Olares&amp;rsquo;s CUDA-plus-x86 is one path, DGX Spark&amp;rsquo;s unified memory is another, and below that there&amp;rsquo;s the DIY mini-PC-plus-eGPU route you assemble yourself. Interestingly, Olares lists eGPU and dual-node as supported paths itself, so the line between appliance and DIY isn&amp;rsquo;t that sharp.&lt;/p&gt;
&lt;p&gt;But my guess is, the one that breaks out won&amp;rsquo;t be the one with the most ferocious specs. It&amp;rsquo;ll be the one that fuses resource governance, dev experience, app ecosystem, security, and remote access into one closed loop. If desktop AI workstations really become a category, what people compete on isn&amp;rsquo;t GPU and memory, it&amp;rsquo;s the control plane.&lt;/p&gt;
&lt;p&gt;I went ahead and added Olares to the &lt;a href="https://landscape.jimmysong.io/projects/olares/" target="_blank" rel="noopener"&gt;AI Native Landscape&lt;/a&gt; I maintain, so I can keep watching how it grows.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In one line: Olares is a desktop AI OS with Kubernetes at its core, and HAMi is the layer that turns its single-machine GPU into a shareable, isolatable, schedulable resource pool.&lt;/p&gt;
&lt;p&gt;From cluster to desktop, HAMi&amp;rsquo;s story is expanding from &amp;ldquo;saving cards&amp;rdquo; to &amp;ldquo;making edge AI usable.&amp;rdquo; If desktop AI workstations become a real category, what people compete on is the control plane, not single-card performance.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/beclab/Olares" target="_blank" rel="noopener"&gt;Olares - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://landscape.jimmysong.io/projects/olares/" target="_blank" rel="noopener"&gt;Olares - AI Native Landscape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;HAMi official site - project-hami.io&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Every Nation Begins with Textiles</title><link>https://jimmysong.io/blog/every-nation-starts-with-textiles/</link><pubDate>Sat, 20 Jun 2026 05:43:26 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/every-nation-starts-with-textiles/</guid><description>From Anji bamboo weaving to industrialization</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;In a bamboo-weaving workshop in Anji, holding a strip of bamboo split hair-thin, I realized for the first time: a bolt of cloth, a machine, a country, all begin with this single thread in the hand.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/banner.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/banner.webp" alt="Figure 1: Bamboo strip" data-caption="Figure 1: Bamboo strip"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Bamboo strip&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;This is something I&amp;rsquo;d been thinking about for a long time, without ever quite sorting it out.&lt;/p&gt;
&lt;p&gt;In February this year, our company offsite took us to Anji, in Zhejiang province. It&amp;rsquo;s famous bamboo country in China, the mountains covered in moso bamboo, and the locals have been weaving with it for generations. Bamboo weaving is a national intangible cultural heritage. The offsite included a hands-on session for us to try it ourselves.&lt;/p&gt;
&lt;p&gt;I hadn&amp;rsquo;t expected much from this kind of &amp;ldquo;experiential activity.&amp;rdquo; But once I actually sat down, picked up a bamboo strip, and listened to the old master explain how to lift, press, and raise the threads, my mind started to wander. Up and down, lift and press, the warp and weft interlacing in his hands until a pattern slowly emerged. The motion was deeply repetitive, a fixed rhythm, a clear logic.&lt;/p&gt;
&lt;p&gt;And in that moment an almost absurd thought popped into my head: this is just 0 and 1.&lt;/p&gt;
&lt;p&gt;A warp thread raised is 1, pressed down is 0. Unfold a piece of cloth and it&amp;rsquo;s a two-dimensional structure made of 0s and 1s. And the thing the old master kept referring to as the &amp;ldquo;flower program&amp;rdquo; (花本), the pre-designed sequence of warp-lifting, is essentially a program: follow it, and you weave exactly the pattern you intended.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/red-horse.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/red-horse.webp" alt="Figure 2: A woven bamboo horse I finished in Anji, mounted for framing" data-caption="Figure 2: A woven bamboo horse I finished in Anji, mounted for framing"
width="1567"
height="1498"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: A woven bamboo horse I finished in Anji, mounted for framing&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;That thought flashed by and I didn&amp;rsquo;t make much of it. But then in March I went to the Netherlands and spent a week in Amsterdam. The two experiences sat together, and some feelings that had been vague began to sharpen: how could a single strip of bamboo, a single motion, connect to the looms of two centuries ago, to today&amp;rsquo;s computers, even to a country&amp;rsquo;s industrialization?&lt;/p&gt;
&lt;p&gt;I looked it up later and found I wasn&amp;rsquo;t the first to make the connection. Many historians of technology argue that weaving is one of the earliest forms of large-scale information encoding. Now, in June, I&amp;rsquo;m sitting down to try to lay out this thread. This isn&amp;rsquo;t a technical piece. It&amp;rsquo;s a personal, cultural reflection.&lt;/p&gt;
&lt;h2 id="a-cultural-contrast"&gt;A Cultural Contrast&lt;/h2&gt;
&lt;p&gt;Let me start with the Netherlands.&lt;/p&gt;
&lt;p&gt;What struck me most about that country wasn&amp;rsquo;t the windmills or the tulips, but a quality that the whole nation seems to exude: restraint, order, and an almost engineering-like way of being. The canals are dug, the land reclaimed from the sea is called a &amp;ldquo;polder,&amp;rdquo; and the windmills were never there to look pretty. They were built to pump water and drain the land. A country that sits below sea level literally engineered itself out of the sea. Though the windmills aren&amp;rsquo;t only about drainage. The wind in the Netherlands is genuinely strong. In March I got battered by it. It&amp;rsquo;s a wind that pushes at you constantly. So a below-sea-level country hauled itself out of the sea with engineering, and along the way put that fierce wind to work too.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/canal.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/canal.webp" alt="Figure 3: A canal and church in central Amsterdam" data-caption="Figure 3: A canal and church in central Amsterdam"
width="1800"
height="2699"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: A canal and church in central Amsterdam&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/windmill.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/windmill.webp" alt="Figure 4: Windmills at Zaanse Schans" data-caption="Figure 4: Windmills at Zaanse Schans"
width="1800"
height="2699"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Windmills at Zaanse Schans&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This quality seeps into every corner of daily life. Shops close around six in the evening, the streets slowly quiet down, and you almost never hear a car horn. Cars stop well ahead of time to let you cross. One evening I wandered a little too close to the bike lane, and a passing car actually rolled down its window to remind me to use the sidewalk. It surprised me at first, and then I understood: the rules there are clear, and everyone assumes you&amp;rsquo;ll follow them too.&lt;/p&gt;
&lt;p&gt;This temperament even shows up in how people get around. In the city center, not many people drive. Bicycles are the default. For one thing, the Netherlands is flat, so cycling takes no effort. For another, dedicated bike lanes are everywhere, and the streets are narrow, which makes them ill-suited to cars. Add that commutes are generally short, and cycling is just right. What&amp;rsquo;s interesting is that even in the southern suburbs, where the roads are wide and every household has a car, you still see a lot of people on bikes. For them, a bicycle isn&amp;rsquo;t exercise. It&amp;rsquo;s the default way to commute.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/street.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/street.webp" alt="Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists" data-caption="Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists"
width="1800"
height="2400"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I set that against the noise and speed of China, the loud, garish signboards back home. These two temperaments are really two paths of modernization. The Netherlands is an early, slow-and-steady kind of country, from the Age of Discovery to the world&amp;rsquo;s first stock exchange, from water engineering to today&amp;rsquo;s ASML and Philips. It has always moved forward in a &amp;ldquo;long-term, restrained, systematic&amp;rdquo; way. China is a late developer chasing from behind, relying on speed, scale, and a generation&amp;rsquo;s sheer exertion to cover in a few decades what took others centuries.&lt;/p&gt;
&lt;p&gt;Neither path is higher or lower than the other, but they settle into completely different urban textures and national characters. The ease of &amp;ldquo;closing at six&amp;rdquo; that the Dutch have is something China won&amp;rsquo;t reach for many years; and the Chinese speed of &amp;ldquo;just do it,&amp;rdquo; of laying a high-speed rail line overnight, is something the Netherlands probably couldn&amp;rsquo;t match either.&lt;/p&gt;
&lt;p&gt;But here&amp;rsquo;s the interesting part. If you trace both civilizations back through history, these two seemingly different cultures share a common starting point. Textiles.&lt;/p&gt;
&lt;h2 id="why-textiles-were-the-starting-point-of-industrialization"&gt;Why Textiles Were the Starting Point of Industrialization&lt;/h2&gt;
&lt;p&gt;Our generation is so familiar with the phrase &amp;ldquo;Industrial Revolution&amp;rdquo; that it&amp;rsquo;s almost gone numb. But have you ever asked: where did the first Industrial Revolution actually break out? Not in steel, not in coal, not in the railways. In textiles.&lt;/p&gt;
&lt;p&gt;The British Industrial Revolution is practically a history of textile machinery. The flying shuttle in 1733 made weaving faster. The spinning jenny in 1764 let one person spin many threads at once. The water frame in 1769 began replacing human power with water power. The power loom in 1785 turned weaving into an automated process. Every one of these inventions happened inside the textile industry.&lt;/p&gt;
&lt;p&gt;Why textiles of all things? At first it feels counterintuitive, since textiles seem too &amp;ldquo;light,&amp;rdquo; lacking the heft of making steel or guns. But think about it for a moment and it becomes obvious.&lt;/p&gt;
&lt;p&gt;Textiles are a basic need. Everyone has to wear clothes, clothes wear out, and they have to be replaced constantly. That&amp;rsquo;s an enormous market that already existed back in the agrarian age. Once the demand is there, any gain in efficiency turns immediately into profit, into capital that can be reinvested.&lt;/p&gt;
&lt;p&gt;More importantly, the processes of textile production are especially easy to mechanize. Spinning and weaving are deeply repetitive motions with a fixed rhythm and a clear logic, so machines can directly replace the human hand. Steelmaking, by contrast, has a far higher technical barrier, the science of chemistry wasn&amp;rsquo;t yet mature, and machine manufacturing itself needed an existing industrial base before it could even start.&lt;/p&gt;
&lt;p&gt;And textile production has a moderate investment threshold. Building a textile mill was far cheaper than building a steel plant, so merchant capital and money from colonial trade flowed in easily. A large share of early British industrial capital came from the cotton trade and colonial trade, and that money went first into textile mills.&lt;/p&gt;
&lt;p&gt;The most crucial point is that textiles don&amp;rsquo;t exist in isolation. They pull an entire supply chain along with them. To spin and weave, you first have to grow and transport cotton. To run machines, you have to build textile machinery. To power the machines, you need steam engines. To fuel steam engines, you have to mine coal. To move cotton and cloth, you have to build railways. So the steam engine was first deployed at scale in textile mills, the earliest railways carried cotton and cloth, and the earliest machine manufacturing served textile equipment. A supposedly &amp;ldquo;light&amp;rdquo; industry ended up dragging the entire mechanical and energy industries into being.&lt;/p&gt;
&lt;p&gt;So textiles became the progenitor of industry not because they were the most complex, but because they were the first to satisfy four conditions at once: large demand, mechanizable, low threshold, and able to pull others along. Textiles were humanity&amp;rsquo;s first trial plot on the way from handicraft into an industrial system.&lt;/p&gt;
&lt;h2 id="late-developers-all-come-up-this-same-road"&gt;Late Developers All Come Up This Same Road&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s a pattern I&amp;rsquo;d never noticed before: the industrialization of late developers, almost without exception, begins with textiles.&lt;/p&gt;
&lt;p&gt;Take Japan. Today when you think of Japanese industry, you think of Toyota cars. But how did Toyota begin? Its founder, Sakichi Toyoda, didn&amp;rsquo;t start by making cars. He invented automatic looms. In 1924 he invented the Type G automatic loom, and in 1926 he founded Toyoda Automatic Loom Works, the origin of the entire Toyota Group. It was his son Kiichiro who later moved the business into automobiles.&lt;/p&gt;
&lt;p&gt;So Toyota cars quite literally &amp;ldquo;grew out of a loom.&amp;rdquo; Its whole system of lean production, kanban management, total quality control, carries in its bones the obsession with precision, process, and zero defects that came from making looms.&lt;/p&gt;
&lt;p&gt;Then look at South Korea. Samsung today is a global electronics and semiconductor giant, but when it was founded it exported dried fish, vegetables, and fruit, only later moving into textiles and sugar. As South Korea industrialized after the war, textiles and garments were among its earliest foreign-exchange earners, and Samsung built its first fortune on trade and textiles in that period before pivoting to electronics in the 1970s. The Hyundai Group took a similar path, starting in engineering and heavy industry before expanding into automobiles, shipbuilding, and heavy industry.&lt;/p&gt;
&lt;p&gt;Britain was first; Japan and Korea are the late-developer cases. You&amp;rsquo;ll notice the sequence is strikingly consistent across all three: textiles first, then light industry, then machinery, then heavy industry, and finally high-tech. Nobody designed this path. It was forced by cost and by the market. Because textiles are the lowest-threshold, largest-market form of manufacturing, every nation that wants to turn from an agrarian country into an industrial one has to start from the easiest, most necessary step.&lt;/p&gt;
&lt;p&gt;Behind this lies a very plain truth: modernization doesn&amp;rsquo;t happen in one leap. It needs a starting point that can earn money, accumulate experience, and train the first generation of industrial workers. And textiles happen to be exactly that kind of starting point.&lt;/p&gt;
&lt;h2 id="where-china-sits-on-this-road"&gt;Where China Sits on This Road&lt;/h2&gt;
&lt;p&gt;Writing this far, I can&amp;rsquo;t help but turn and look back at China itself.&lt;/p&gt;
&lt;p&gt;China&amp;rsquo;s modern industrialization also began with textiles. The most famous Chinese-owned enterprises of the late Qing and early Republic were almost all cotton mills. Zhang Jian founded the Dasheng Cotton Mill in Nantong. The Rong brothers, Zongjing and Desheng, founded the Shenxin Cotton Mills in Wuxi. They were China&amp;rsquo;s first generation of national industrial capitalists, rising out of textiles and holding up half of modern Chinese industry. The thinking of Zhang Jian&amp;rsquo;s generation was plain and direct: foreigners were using machines to weave cloth and taking our money, so we had to set up our own mills, weave our own cloth, and keep that money at home.&lt;/p&gt;
&lt;p&gt;That was China&amp;rsquo;s first stretch. Back then, China was walking the very road that Britain, Japan, and Korea had all walked.&lt;/p&gt;
&lt;p&gt;Then the road broke. War, turmoil, the planned economy. Chinese industrialization took many detours. It wasn&amp;rsquo;t until reform and opening up that it reconnected. In the 1980s the coast was blanketed with processing factories for garments, shoes, and toys. In essence it was still that same old road of &amp;ldquo;starting with textiles and light industry,&amp;rdquo; only this time China relaunched it through exports, cheap labor, and the role of factory to the world.&lt;/p&gt;
&lt;p&gt;Further on, through the 1990s and 2000s, China pushed into heavy industry: steel, cement, shipbuilding, chemicals, catching up to the world&amp;rsquo;s front ranks one after another. Then electronics, home appliances, mobile phones. Then the internet, high-speed rail, new-energy vehicles, solar, semiconductors, and the AI that everyone talks about today.&lt;/p&gt;
&lt;p&gt;String this line together and you realize China is actually completing the same road that every late developer has walked, only faster, fiercer, and at a far greater scale. From Zhang Jian&amp;rsquo;s cotton mills to today&amp;rsquo;s new-energy vehicles and large AI models is barely over a hundred years. In just over a century we&amp;rsquo;ve run the entire course of industrialization that took Britain more than two hundred and Japan more than a hundred, and we&amp;rsquo;re still going.&lt;/p&gt;
&lt;p&gt;This often leaves me with a complicated feeling. On one hand, admiration: generation after generation, from Zhang Jian to today&amp;rsquo;s engineers, really did turn an agrarian country into the factory of the world and then into a country that now leads in quite a few high-tech fields. On the other hand, a faint unease: this road was run so fast that a lot of things got left behind.&lt;/p&gt;
&lt;h2 id="a-thread-that-never-broke"&gt;A Thread That Never Broke&lt;/h2&gt;
&lt;p&gt;Inside this long thread of industrialization, there&amp;rsquo;s one detail I never forgot: that moment in the bamboo-weaving workshop in Anji back in February.&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;flower program&amp;rdquo; the old master described actually has a very old tradition in China. The drawloom of the Han dynasty used a pre-designed system of cords to control which warp threads rose and fell, so a weaver following it could reproduce complex patterns. Then in the early nineteenth century, the Frenchman Joseph Marie Jacquard pushed the idea a big step forward, using punched cards to control the weaving pattern: where the card had a hole, a hook passed through and lifted the warp; where there was no hole, the hook was blocked and the warp stayed down. Hole or no hole, that&amp;rsquo;s 1 and 0.&lt;/p&gt;
&lt;p&gt;Jacquard&amp;rsquo;s punched cards were later borrowed by the computing pioneer Charles Babbage for his Analytical Engine, and Ada Lovelace used them to write the first computer program in history. Further on, IBM rose on punched-card machines, and if you trace today&amp;rsquo;s entire computing industry back to its source, it leads, improbably, to a loom.&lt;/p&gt;
&lt;p&gt;And that&amp;rsquo;s not the end of it. The Transformer, matrix operations, the weights of neural networks in today&amp;rsquo;s AI, at bottom are all about computing relationships across vast fields of &amp;ldquo;warp and weft.&amp;rdquo; A loom decides which warp threads to lift and when; a neural network decides which parameters relate to which others. The logic is the same.&lt;/p&gt;
&lt;p&gt;So the word &amp;ldquo;weaving&amp;rdquo; is more than the starting point of an industry. It&amp;rsquo;s a metaphor: for thousands of years, human beings have been learning how to encode complex things into structure, from a bolt of cloth, to a program, to a neural network. Every nation does this. They&amp;rsquo;re just at different stages.&lt;/p&gt;
&lt;p&gt;I plan to write the technical bloodline of this story separately, in another piece. Here I&amp;rsquo;ll only point to it, to make one thing clear: the thread that began with that strip of bamboo has never broken.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;From a bamboo-weaving session in Anji in February, to the canals and windmills of Amsterdam in March, and on to the industrialization of Britain, Japan, Korea, and China, what I want to say really comes down to one thing.&lt;/p&gt;
&lt;p&gt;Every nation&amp;rsquo;s modernization is a road from a single thread woven into a net. Textiles are the common starting point of that road, not because textiles are so lofty, but because they are the most necessary, the easiest, the quickest to earn that first money and train that first generation of workers. From that starting point, some walk with restraint, like the Netherlands; some move with ferocity, like China. Some take two hundred years, some a hundred, some only a few decades. But the starting point is the same, and so is the direction: from that thread in the hand, step by step, woven into the whole of industrial civilization as we know it today.&lt;/p&gt;
&lt;p&gt;The Dutch ease of closing at six, of no honking, of swans gliding through the canals, is a place China hasn&amp;rsquo;t reached yet. The Chinese speed of just doing it, of rolling something out overnight, is something the Netherlands couldn&amp;rsquo;t pull off either. Between these two temperaments of modernization there is no higher or lower, only trade-offs. And behind the trade-offs lies the question of where a country sits in its industrialization, and what rhythm it&amp;rsquo;s willing to pay for that net.&lt;/p&gt;
&lt;p&gt;That day in Anji, holding a strip of bamboo, I clumsily followed the old master, lifting and pressing, weaving a small, lopsided patch of pattern. I wasn&amp;rsquo;t thinking about any of this then. But looking back, from that single lift and press, you can trace upward to the Han dynasty drawloom, outward to a country&amp;rsquo;s industrialization, and forward, faintly, to today&amp;rsquo;s computers and AI. A single strip of bamboo, connected to so much.&lt;/p&gt;
&lt;p&gt;Perhaps the evolution of civilization, in the end, is just humanity learning, over and over, how to weave one thing into another.&lt;/p&gt;
&lt;p&gt;Starting from a single thread.&lt;/p&gt;</content:encoded></item><item><title>Why GPUs Became the Foundation of AI: A GPU Primer for K8s Veterans</title><link>https://jimmysong.io/blog/why-gpu-foundation-of-ai/</link><pubDate>Wed, 17 Jun 2026 14:04:42 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/why-gpu-foundation-of-ai/</guid><description>A GPU explainer for Kubernetes veterans new to AI. Maps token, model, training, inference, Transformer, Tensor Core, HBM, and KV cache to concepts you already know.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Note
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This post builds on Ameen Alam&amp;rsquo;s three-part GPU Architecture series and draws on TrendForce, NVIDIA (Slinky/Slurm, the GPUDirect data path) and Mirantis material on GPU infrastructure and agentic AI. It&amp;rsquo;s written for Kubernetes veterans who have never touched a GPU and never trained or served a model.
&lt;/div&gt;
&lt;/div&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/banner.webp" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/banner.webp" alt="Figure 1: GPU: the foundation of AI" data-caption="Figure 1: GPU: the foundation of AI"
width="1717"
height="916"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: GPU: the foundation of AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-a-cloud-native-veteran-is-writing-about-gpus"&gt;Why a cloud-native veteran is writing about GPUs&lt;/h2&gt;
&lt;p&gt;For the past decade my home turf has been containers and Kubernetes. From service mesh to the whole cloud-native ecosystem, I&amp;rsquo;ve spent almost every day on scheduling, networking, storage and observability, but always on the CPU side of cloud native. Last year I formally moved into AI infrastructure (AI Infra).&lt;/p&gt;
&lt;p&gt;Once I dove in, I found the concept density absurd. Token, Transformer, Tensor Core, HBM, KV Cache come at you one after another, and almost every doc and article assumes you already know them, which is deeply unfriendly to engineers who have never trained a model or run inference.&lt;/p&gt;
&lt;p&gt;I quickly hit on a trick: &lt;strong&gt;don&amp;rsquo;t learn from scratch, use what you already know by heart as a translator.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The mental model a Kubernetes veteran already carries, scheduling, Jobs, Deployments, controllers, caches, utilization, multi-tenancy, and then microservices, service mesh, distributed systems, event-driven, high-availability, maps onto AI Infra almost one to one. Once you build that mapping, the intimidating concepts become &amp;ldquo;oh, it&amp;rsquo;s just the GPU version of X&amp;rdquo;. This &amp;ldquo;translate AI through cloud-native eyes&amp;rdquo; approach got me up to speed fast, and I&amp;rsquo;m writing it down to help friends with the same background skip the detour.&lt;/p&gt;
&lt;p&gt;This is the first post in that translation series, tackling the most fundamental question: &lt;strong&gt;why does AI absolutely need GPUs?&lt;/strong&gt; Later posts will cover GPU resource management, scheduling and observability, topics a K8s veteran knows well. If you&amp;rsquo;re also crossing over from cloud native, I hope this saves you a few days of digging.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You type a sentence and the AI replies word by word. For someone who has used K8s, the most natural mental model is: AI is a &amp;ldquo;workload&amp;rdquo; that runs on GPUs, and a GPU is a special kind of &amp;ldquo;node&amp;rdquo; built for it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="translating-the-jargon-into-k8s"&gt;Translating the jargon into K8s&lt;/h2&gt;
&lt;p&gt;The AI world throws around terms that read like scripture to anyone who hasn&amp;rsquo;t done training or inference. Here&amp;rsquo;s a cheat sheet; I&amp;rsquo;ll re-explain each term in plain words below.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;AI concept&lt;/th&gt;
&lt;th&gt;In plain words&lt;/th&gt;
&lt;th&gt;K8s veteran&amp;rsquo;s analogy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Token&lt;/td&gt;
&lt;td&gt;A small unit text is chopped into; AI generates them one at a time&lt;/td&gt;
&lt;td&gt;A log line, a text chunk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;A &amp;ldquo;scoring program&amp;rdquo; holding billions of numbers (weights)&lt;/td&gt;
&lt;td&gt;A giant image whose &amp;ldquo;weights&amp;rdquo; aren&amp;rsquo;t code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training&lt;/td&gt;
&lt;td&gt;Feed it huge data, tune params repeatedly, build the model&lt;/td&gt;
&lt;td&gt;Run a Job to build an image; offline, batch, care about throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;Model is done, serve user requests and emit answers&lt;/td&gt;
&lt;td&gt;Run a Deployment taking traffic; online, care about latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformer&lt;/td&gt;
&lt;td&gt;The architecture shared by nearly all LLMs (GPT, Claude, LLaMA)&lt;/td&gt;
&lt;td&gt;A controller design pattern&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tensor Core&lt;/td&gt;
&lt;td&gt;A GPU unit that only does &amp;ldquo;matrix multiply&amp;rdquo; but insanely fast&lt;/td&gt;
&lt;td&gt;A sidecar worker that only does batched multiply&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HBM&lt;/td&gt;
&lt;td&gt;The GPU&amp;rsquo;s on-board high-bandwidth memory; model and cache live here&lt;/td&gt;
&lt;td&gt;Node-local RAM, but with bandwidth that crushes normal RAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV Cache&lt;/td&gt;
&lt;td&gt;Attention state stored at inference time, grows with the conversation&lt;/td&gt;
&lt;td&gt;A pod-local session notebook that gets thicker the longer you talk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: AI jargon → K8s veteran cheat sheet
&lt;/figcaption&gt;
&lt;p&gt;Memorize this table; the rest follows.&lt;/p&gt;
&lt;h2 id="the-question-most-people-never-ask"&gt;The question most people never ask&lt;/h2&gt;
&lt;p&gt;GPUs are everywhere in AI now, so accepted that most people skip past it. We care about which card to rent, which framework to use, which model to deploy, but rarely stop to ask the deeper question:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why GPUs?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Why not the more general CPU? Why did a chip originally built to render game graphics become the foundation of the most important technology shift in a generation?&lt;/p&gt;
&lt;p&gt;The answer isn&amp;rsquo;t just &amp;ldquo;it&amp;rsquo;s fast&amp;rdquo;, it&amp;rsquo;s about &lt;strong&gt;how computation itself is organized&lt;/strong&gt;, and why AI workloads demand an architecture fundamentally different from fifty years of software.&lt;/p&gt;
&lt;h2 id="two-philosophies-of-computing-the-craftsman-and-the-factory"&gt;Two philosophies of computing: the craftsman and the factory&lt;/h2&gt;
&lt;p&gt;You know CPUs well. K8s&amp;rsquo;s control plane, etcd, the scheduler all run on CPU; it excels at executing complex instructions one after another, with few cores (8 to 128) but each highly capable. &lt;strong&gt;A CPU optimizes for latency: how fast can I finish one complex task?&lt;/strong&gt; Like a master craftsman doing one intricate job at a time.&lt;/p&gt;
&lt;p&gt;A GPU takes the opposite path: &lt;strong&gt;execute simple instructions across massive amounts of data at once.&lt;/strong&gt; A modern data-center GPU has thousands of small cores, and &lt;strong&gt;it optimizes for throughput: how many simple tasks can I finish at the same moment?&lt;/strong&gt; Like a factory floor of thousands of workers, each doing one identical step simultaneously.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/cpu-vs-gpu-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/cpu-vs-gpu-en.svg" alt="Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing" data-caption="Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing"
width="1157"
height="340"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For any single complex job, the craftsman (CPU) is faster; but as long as the work is uniform and parallelizable, the factory (GPU) crushes the craftsman in total output per second. For decades CPU dominated because most software (web servers, databases, operating systems) is inherently serial. Then deep learning arrived.&lt;/p&gt;
&lt;h2 id="why-ai-broke-the-cpu"&gt;Why AI broke the CPU&lt;/h2&gt;
&lt;p&gt;The core of a neural network is, essentially, a giant pile of &lt;strong&gt;matrix multiplications&lt;/strong&gt; (an operation that batch-multiplies-and-adds two sets of numbers). When a model processes a token, it runs thousands or tens of thousands of these multiplies, and they &lt;strong&gt;don&amp;rsquo;t depend on each other&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The key insight in one line: &lt;strong&gt;these operations are naturally parallel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A 128-core CPU looks at this workload and sweeps through it sequentially; a many-thousand-core GPU chops the matrix into blocks and hands them to its small cores to run at once. Same work, orders of magnitude faster on GPU, not because a single GPU core is faster (it&amp;rsquo;s slower), but because &lt;strong&gt;the problem is parallel and the GPU is built for exactly that shape of computation.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the real reason AI runs on GPUs: not marketing, not legacy, architectural fit.&lt;/p&gt;
&lt;h2 id="the-three-things-inside-a-gpu-in-k8s-terms"&gt;The three things inside a GPU, in K8s terms&lt;/h2&gt;
&lt;h3 id="1-tensor-core-the-sidecar-that-only-does-batched-multiply"&gt;1. Tensor Core: the sidecar that only does batched multiply&lt;/h3&gt;
&lt;p&gt;A GPU has two kinds of cores. Regular CUDA cores are general-purpose workers that can compute anything; &lt;strong&gt;a Tensor Core is a dedicated unit that does exactly one thing: multiply two small matrices in a single shot.&lt;/strong&gt; But that one thing it does insanely fast, finishing in one cycle what would take a regular core thousands.&lt;/p&gt;
&lt;p&gt;AI&amp;rsquo;s core operation is matrix multiplication, so Tensor Cores are tailor-made for it. In K8s terms: regular CUDA cores are like general pods in a Deployment that do everything; a Tensor Core is like a highly specialized sidecar that only does &amp;ldquo;batched multiply&amp;rdquo;, single-function but with crushing throughput.&lt;/p&gt;
&lt;h3 id="2-hbm-the-nodes-high-speed-local-memory"&gt;2. HBM: the node&amp;rsquo;s high-speed local memory&lt;/h3&gt;
&lt;p&gt;Computing fast isn&amp;rsquo;t enough; you have to get the data. A GPU&amp;rsquo;s memory is layered just like a CPU node&amp;rsquo;s, and if you understand a K8s node&amp;rsquo;s L1/L2 cache, RAM and local disk, you understand the GPU&amp;rsquo;s:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/gpu-memory-hierarchy-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/gpu-memory-hierarchy-en.svg" alt="Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI" data-caption="Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI"
width="1037"
height="393"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The bottom layer, &lt;strong&gt;HBM (High Bandwidth Memory)&lt;/strong&gt;, is the GPU&amp;rsquo;s &amp;ldquo;main memory&amp;rdquo;; model weights, intermediate results and KV cache all live here. It&amp;rsquo;s characterized by large capacity (hundreds of GB per card) and extremely high bandwidth.&lt;/p&gt;
&lt;p&gt;Why build HBM at all? Because traditional VRAM (GDDR) lies flat on the circuit board, you run out of routing space and hit a bandwidth wall. HBM&amp;rsquo;s answer is to &lt;strong&gt;stack memory chips vertically&lt;/strong&gt; (connected by Through-Silicon Vias, TSVs), packed right against the GPU die with an interface thousands of lines wide running in parallel. By analogy: normal memory is like a warehouse spread across a parking lot where the movers can&amp;rsquo;t keep up; HBM is like building that warehouse into a dozens-story tower right next to the workshop, with TSV elevators running up and down at full speed.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a counter-intuitive fact: &lt;strong&gt;modern AI is memory-bound more than compute-bound.&lt;/strong&gt; Tensor Cores do matrix multiplication faster than HBM can feed them data, so the GPU is often waiting. That&amp;rsquo;s why each new GPU generation (H200 → Blackwell → Vera Rubin) sees its most important upgrade in HBM bandwidth, not compute (4.8 → 8 → 22 TB/s, &lt;a href="https://www.linkedin.com/pulse/hidden-technology-behind-modern-ai-gpus-ameen-alam-tzkre" target="_blank" rel="noopener"&gt;source&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Since data movement is the bottleneck, NVIDIA&amp;rsquo;s other move is to &lt;strong&gt;let data bypass the CPU and reach the GPU directly&lt;/strong&gt;, the three &amp;ldquo;expressways&amp;rdquo; collectively called &lt;a href="https://www.linkedin.com/pulse/why-nvidia-gpu-architecture-perfect-ai-insights-data-path-kawonise-a821e/" target="_blank" rel="noopener"&gt;GPUDirect&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPUDirect Storage&lt;/strong&gt;: data goes straight from NVMe to GPU memory, without detouring through host memory and the CPU.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPUDirect RDMA&lt;/strong&gt;: GPU memory talks directly to the NIC (InfiniBand), so cross-node gradient exchange no longer hops through the CPU.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NVLink&lt;/strong&gt;: GPUs inside one machine connect directly and share memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In one line: &lt;strong&gt;the CPU is no longer the traffic cop for every data movement.&lt;/strong&gt; In K8s terms, it&amp;rsquo;s like sending data over a direct data plane instead of routing every packet through the apiserver. NVIDIA&amp;rsquo;s edge isn&amp;rsquo;t just fast compute; it&amp;rsquo;s that the entire data highway from storage to GPU, GPU to GPU, and GPU to network has been straightened out.&lt;/p&gt;
&lt;h3 id="3-transformer--token-the-program-that-emits-answers-word-by-word"&gt;3. Transformer + Token: the program that emits answers word by word&lt;/h3&gt;
&lt;p&gt;A Transformer isn&amp;rsquo;t hardware, it&amp;rsquo;s a &lt;strong&gt;model architecture&lt;/strong&gt; (a program structure). GPT, Claude, LLaMA and other large models are all built on it. What it does is actually plain:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Read a piece of text, predict the next token.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A token is a small unit text is chopped into (a character, half a word). A Transformer feeds the input token sequence in, emits &amp;ldquo;the most likely next token&amp;rdquo;, appends it, feeds the new sequence back in, predicts the next-next, and so on, generating the whole answer one piece at a time.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/token-transformer-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/token-transformer-en.svg" alt="Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next" data-caption="Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next"
width="836"
height="285"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In K8s terms: a Transformer is a controller whose reconcile logic is &amp;ldquo;look at the current state (tokens so far) → emit the next action (the next token)&amp;rdquo;, looping continuously. That&amp;rsquo;s where the &amp;ldquo;word by word&amp;rdquo; effect of AI assistants comes from.&lt;/p&gt;
&lt;h2 id="training-vs-inference-writing-the-recipe-vs-serving-the-dish"&gt;Training vs Inference: writing the recipe vs serving the dish&lt;/h2&gt;
&lt;p&gt;This is the easiest pair to confuse when starting out, yet the most critical. &lt;strong&gt;Training and inference are two different things with completely different hardware needs.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Training&lt;/strong&gt;: feed huge data, repeatedly tune those billions of parameters, &amp;ldquo;teach&amp;rdquo; the model into existence. It runs for weeks to months as an offline batch job, &lt;strong&gt;cares about throughput, not per-request latency.&lt;/strong&gt; In K8s terms, it&amp;rsquo;s a long-running &lt;strong&gt;Job&lt;/strong&gt; whose goal is to build a model image.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference&lt;/strong&gt;: once the model is trained and online, it takes each user&amp;rsquo;s input, computes an answer and emits it back. The user is waiting, &lt;strong&gt;so it cares about latency&lt;/strong&gt;; it must serve thousands of users at once, &lt;strong&gt;so it cares about concurrency.&lt;/strong&gt; It&amp;rsquo;s like a &lt;strong&gt;Deployment&lt;/strong&gt; taking traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/training-vs-inference-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/training-vs-inference-en.svg" alt="Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)" data-caption="Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)"
width="1038"
height="347"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This directly affects which GPU metrics you should care about. The headline numbers (Tensor Core FLOPS, NVLink bandwidth, cluster size) are mostly &lt;strong&gt;training&lt;/strong&gt; metrics; what actually determines whether your AI assistant feels snappy (HBM capacity, memory bandwidth, KV cache management) are &lt;strong&gt;inference&lt;/strong&gt; metrics, and they get little airtime. Yet inference is where most production AI actually runs.&lt;/p&gt;
&lt;h2 id="multi-gpu-collaboration-an-ai-cluster-is-just-a-distributed-system"&gt;Multi-GPU collaboration: an AI cluster is just a distributed system&lt;/h2&gt;
&lt;p&gt;With single-card covered, back to reality: the biggest models don&amp;rsquo;t fit on one card, and training routinely needs hundreds or thousands of cards working together. At that point the GPU cluster is essentially a &lt;strong&gt;distributed system&lt;/strong&gt;, and almost everything you know from cloud native applies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Many GPUs = a microservice cluster.&lt;/strong&gt; A card is like a service instance; the model is sharded across cards, each computes a slice and the results are stitched back. That&amp;rsquo;s exactly the microservices playbook of splitting by responsibility and sharing load.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inter-card communication = the service-mesh data plane.&lt;/strong&gt; Cards constantly exchange data (gradients especially during training). Within a rack it&amp;rsquo;s NVLink (several TB/s), across racks it&amp;rsquo;s InfiniBand. It&amp;rsquo;s just the east-west traffic between microservices: in-node pod-to-pod direct is fastest (NVLink), cross-node cross-cluster goes over the network (IB). NVIDIA packaging GPUs, switches, DPUs and RDMA into a whole rack (&lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;Vera Rubin NVL72&lt;/a&gt;) is exactly a service mesh unifying the data plane, control plane and observability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distributed training&amp;rsquo;s all-reduce = distributed consensus.&lt;/strong&gt; Each training step synchronizes and averages the gradients across all cards, a step called all-reduce. Anyone who has done etcd or Raft gets it instantly: it&amp;rsquo;s distributed-system consensus and consistency, except here you&amp;rsquo;re syncing gradients, not a state-machine log.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Continuous batching = event-driven.&lt;/strong&gt; During inference, new requests are dropped into the running batch as they arrive, no waiting for the whole batch to finish. Requests are events, the batcher is a consumer, batch-then-process: that&amp;rsquo;s the event-driven / message-queue mindset, all to keep the GPU busy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disaggregated inference = microservices.&lt;/strong&gt; Split prefill (process input, compute-heavy) and decode (emit tokens, memory-bandwidth-heavy) into separate GPU pools that scale independently, with KV cache passed between them over NVLink/RDMA. That&amp;rsquo;s entirely the microservices pattern of splitting by responsibility and scaling each part; KV cache is the session context passed between services.&lt;/p&gt;
&lt;p&gt;See the pattern? An AI cluster isn&amp;rsquo;t a new species, it&amp;rsquo;s &lt;strong&gt;your familiar distributed-systems, microservices and service-mesh toolkit replayed on GPU hardware.&lt;/strong&gt; The instincts you built tuning traffic on Istio and Envoy, or consensus on etcd, are worth exactly as much here.&lt;/p&gt;
&lt;h2 id="running-gpus-on-k8s-slurm-for-training-k8s-for-inference"&gt;Running GPUs on K8s: Slurm for training, K8s for inference&lt;/h2&gt;
&lt;p&gt;By now, as a K8s veteran, you&amp;rsquo;re bound to ask: so what actually schedules a GPU cluster? The interesting answer: &lt;strong&gt;training and inference often run on two different schedulers.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On the training side, the HPC world&amp;rsquo;s heavyweight is Slurm.&lt;/strong&gt; It manages 65% of the world&amp;rsquo;s TOP500 supercomputers (&lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;NVIDIA&lt;/a&gt;), and large AI training teams have years invested in Slurm scripts, fair-share policies and accounting. Training is a long-running batch job that wants topology awareness (place chatty GPUs close together), long exclusive holds and high throughput, areas Slurm has refined for over a decade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On the inference side, the home field is Kubernetes.&lt;/strong&gt; Inference is an online service that wants fast autoscaling, per-request scheduling and integration with service mesh and observability, exactly K8s&amp;rsquo;s strengths.&lt;/p&gt;
&lt;p&gt;So what if a team needs both, maintain two environments? NVIDIA open-sourced the &lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;Slinky&lt;/a&gt; project to solve exactly this, in a very K8s-native way:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/SlinkyProject/slurm-operator" target="_blank" rel="noopener"&gt;slurm-operator&lt;/a&gt;&lt;/strong&gt;: turns each Slurm daemon (slurmctld for scheduling, slurmd for compute, slurmdbd for accounting) into a K8s CRD and Pod, with the control plane made highly available through Pod regeneration instead of Slurm&amp;rsquo;s native HA. Config changes sync automatically via ConfigMap/Secret, workers autoscale with HPA, scale-in drains running jobs first, and upgrades use PodDisruptionBudget to protect in-flight work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU Operator + DCGM Exporter&lt;/strong&gt;: auto-installs drivers and the device plugin, and can label metrics by Slurm job ID, giving you &lt;strong&gt;per-job&lt;/strong&gt; GPU metrics (scraped by Prometheus).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ComputeDomains + DRA&lt;/strong&gt;: for cross-node-NVLink machines like GB200 NVL72, K8s uses DRA (Dynamic Resource Allocation) to dynamically manage the cross-node GPU interconnect domain, so distributed training hits full NVLink bandwidth across nodes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;NVIDIA itself runs Slinky in production up to 8,000+ GPUs, with NCCL all-reduce/all-gather performance matching bare Slurm, the K8s layer adding almost no overhead (&lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;NVIDIA production data&lt;/a&gt;). &lt;strong&gt;K8s is becoming the substrate for GPU computing, with Slurm as the scheduling layer on top.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Training and inference have fundamentally different GPU-scheduling needs, so they can&amp;rsquo;t be managed the same way. That&amp;rsquo;s also why GPU resource management (covered later) has to be scenario-specific, and why solutions like &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; aim to be a unified GPU resource-management layer across Slurm + K8s hybrid environments.&lt;/p&gt;
&lt;h2 id="the-real-trap-of-inference-kv-cache"&gt;The real trap of inference: KV Cache&lt;/h2&gt;
&lt;p&gt;If there&amp;rsquo;s one concept that separates &amp;ldquo;GPU theory&amp;rdquo; from &amp;ldquo;AI inference reality&amp;rdquo;, it&amp;rsquo;s &lt;strong&gt;KV Cache&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When a Transformer generates each token, it has to reference the relationships among all previous tokens (this is &amp;ldquo;attention&amp;rdquo;). To avoid recomputing from scratch every time, it stores each token&amp;rsquo;s attention state, and that store is the KV Cache.&lt;/p&gt;
&lt;p&gt;The problem is it &lt;strong&gt;grows without bound&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user sends a 10,000-token conversation; the model stores a state entry for each.&lt;/li&gt;
&lt;li&gt;To generate the next token, &lt;strong&gt;it reads the entire KV Cache.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;A 70B model at 32K context can produce a KV Cache of 8 to 16 GB per request, larger than the model itself.&lt;/li&gt;
&lt;li&gt;50 concurrent users land, and KV Cache eats hundreds of GB of memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/kv-cache-growth-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/kv-cache-growth-en.svg" alt="Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook" data-caption="Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook"
width="980"
height="353"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In K8s terms: KV Cache is like a &lt;strong&gt;session notebook&lt;/strong&gt; local to a Pod. The longer the conversation, the thicker the notebook, and every new word forces a full read from cover to cover. So inference memory is often eaten not by the model but by this &amp;ldquo;notebook&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s also why we got paged attention (managing the notebook like virtual-memory pages to cut fragmentation, popularized by &lt;a href="https://github.com/vllm-project/vllm" target="_blank" rel="noopener"&gt;vLLM&lt;/a&gt;) and NVIDIA &lt;a href="https://github.com/ai-dynamo/dynamo" target="_blank" rel="noopener"&gt;Dynamo&lt;/a&gt;&amp;rsquo;s multi-tier cache (hot pages in HBM, warm offloaded to CPU memory, cold spilled to NVMe). &lt;strong&gt;Half an inference engineer&amp;rsquo;s job is managing this notebook.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="gpu-utilization-lies"&gt;GPU utilization lies&lt;/h2&gt;
&lt;p&gt;K8s folks know &amp;ldquo;utilization lies&amp;rdquo; best: a pod reporting Running isn&amp;rsquo;t necessarily doing work, it might be in CPU steal or waiting on IO. GPUs are worse.&lt;/p&gt;
&lt;p&gt;That &amp;ldquo;GPU utilization %&amp;rdquo; in monitoring usually only measures &amp;ldquo;is the GPU executing any kernel&amp;rdquo;, not how efficiently. A card can report 90% utilization while its Tensor Cores are actually busy only 30% of the time, with the rest spent on memory ops, kernel-launch overhead, or just waiting for data.&lt;/p&gt;
&lt;p&gt;The more honest metric is &lt;strong&gt;SM Efficiency (SM activity rate)&lt;/strong&gt;: it looks at how many SMs are doing useful work each clock cycle. A card showing 100% utilization in nvidia-smi may have an SM Efficiency of only 20-30%. Many companies think their &amp;ldquo;GPUs are maxed out&amp;rdquo; when in fact huge amounts of compute are spinning idle. So to judge whether a GPU is truly working, don&amp;rsquo;t look at utilization, look at SM Efficiency.&lt;/p&gt;
&lt;p&gt;The metrics that actually matter (mapping to the QPS, P99 latency and resource levels you watch in K8s):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Token throughput&lt;/strong&gt;: tokens generated per second per card, how many users you can serve (like QPS).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TTFT (Time To First Token)&lt;/strong&gt;: time from receiving a request to emitting the first token, sets the &amp;ldquo;responsiveness feel&amp;rdquo; (like cold-start time).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TPOT (Time Per Output Token, also called ITL / inter-token latency)&lt;/strong&gt;: average time to produce each token, sets streaming smoothness (like P99). Serving frameworks like vLLM and TGI generally use TPOT; NVIDIA more often calls it ITL, same thing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory breakdown&lt;/strong&gt;: how much is weights vs KV cache vs transient activations, tells you whether you&amp;rsquo;re memory-bound (like breaking down a pod&amp;rsquo;s memory usage).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="why-nvidia-is-so-hard-to-dislodge"&gt;Why NVIDIA is so hard to dislodge&lt;/h2&gt;
&lt;p&gt;You can&amp;rsquo;t talk GPUs without NVIDIA&amp;rsquo;s dominance. The hardware is excellent, but hardware alone can&amp;rsquo;t explain why AMD, Intel and a crowd of startups have failed to gain ground. The answer is &lt;strong&gt;ecosystem depth&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;CUDA isn&amp;rsquo;t just a parallel-computing platform, it&amp;rsquo;s a programming model, compiler, runtime and a stack of libraries, refined for 18 years. Nearly every AI framework (PyTorch, TensorFlow, JAX) grew up on CUDA first and was ported elsewhere as an afterthought. In K8s terms: CUDA is to GPUs roughly what the Linux kernel + containerd + the whole CNCF toolchain are to the container ecosystem, &lt;strong&gt;not something you replace by swapping a kernel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Switching to another card means re-validating every layer of your inference stack (kernels, libraries, frameworks, serving, monitoring, ops). The switching cost isn&amp;rsquo;t buying different hardware, it&amp;rsquo;s &lt;strong&gt;rebuilding an entire software ecosystem.&lt;/strong&gt; That&amp;rsquo;s the &amp;ldquo;CUDA moat&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="the-future-ai-factories"&gt;The future: AI factories&lt;/h2&gt;
&lt;p&gt;NVIDIA no longer describes its business in terms of &amp;ldquo;GPUs&amp;rdquo; or even &amp;ldquo;data centers&amp;rdquo;, but as &amp;ldquo;AI factories&amp;rdquo;: facilities that continuously convert electricity, silicon and data into intelligence. Beneath the language is a real architectural shift:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Power is now the primary constraint&lt;/strong&gt;: a single Vera Rubin rack draws over 100 kW (&lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;source&lt;/a&gt;), and a mid-size training cluster needs 10 to 50 MW, comparable to a small town. The GPU is no longer the hard part; securing reliable power is.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cooling moves off air&lt;/strong&gt;: liquid cooling is now standard for high-density GPUs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Networking decides GPU placement&lt;/strong&gt;: AI cluster design now starts with &lt;strong&gt;network topology and works backward to where GPUs go&lt;/strong&gt;, the reverse of traditional data centers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory bandwidth remains the scaling frontier&lt;/strong&gt;: each generation adds compute, but what actually moves the experience is HBM bandwidth.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s fundamentally HA + datacenter ops&lt;/strong&gt;: power redundancy, cooling failover, latency determined by network topology, single-point failure and DR are all old problems for anyone who has done distributed-systems HA. The AI factory isn&amp;rsquo;t new magic; it&amp;rsquo;s your HA-architecture skillset moved into a data center an order of magnitude denser in power and compute.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dont-retire-the-cpu-just-yet-agentic-ai-is-rebalancing-the-cpugpu-ratio"&gt;Don&amp;rsquo;t retire the CPU just yet: agentic AI is rebalancing the CPU:GPU ratio&lt;/h2&gt;
&lt;p&gt;After all this GPU praise, you might think the CPU is sidelined in the AI era. The opposite is true: &lt;strong&gt;agentic AI is putting the CPU back at center stage.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In traditional LLM inference the CPU mostly just compresses and routes data for the GPU, so AI data centers ran CPU:GPU ratios as low as 1:4 or even 1:8. Agents are different: they plan tasks autonomously, call tools, route data between sub-agents and decide whether a task is complete, and all of that &lt;strong&gt;orchestration logic lands squarely on the CPU.&lt;/strong&gt; Add that agents are often trained with reinforcement learning, where every action has to be evaluated by the CPU, and the CPU load gets heavier still.&lt;/p&gt;
&lt;p&gt;Research notes that in agentic scenarios the CPU-side tool processing (Python interpretation, web crawling, retrieval, database queries) can account for up to 90% of total latency. According to &lt;a href="https://insights.trendforce.com/p/agentic-ai-cpu-gpu" target="_blank" rel="noopener"&gt;TrendForce&lt;/a&gt;, Arm estimates that traditional AI data centers need about 30 million CPU cores per GW, a figure that will surge to 120 million in the agent era, &lt;strong&gt;a 4x increase&lt;/strong&gt;; the CPU:GPU ratio will shift from 1:4&lt;del&gt;1:8 toward **1:1&lt;/del&gt;1:2**. That&amp;rsquo;s also why in 2026 NVIDIA started selling the Vera CPU standalone and Arm got into the CPU business directly.&lt;/p&gt;
&lt;p&gt;The signal for K8s veterans is clear: &lt;strong&gt;the future AI node is a mixed-workload node where CPU and GPU are billed together&lt;/strong&gt;, and scheduling and resource management must handle both, not just stare at the GPU.&lt;/p&gt;
&lt;h2 id="what-this-means-for-me-a-k8s-veteran"&gt;What this means for me (a K8s veteran)&lt;/h2&gt;
&lt;p&gt;After all this hardware, it lands on what I work on: &lt;strong&gt;only by first understanding why the GPU is the foundation of AI can you understand why &amp;ldquo;GPU resource management&amp;rdquo; is a real problem.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When a card costs hundreds of thousands of yuan, a whole rack draws over a hundred kilowatts, and production GPUs still run below capacity because memory bandwidth, KV cache or batching aren&amp;rsquo;t tuned, just &amp;ldquo;handing out cards&amp;rdquo; is nowhere near enough. How to run multiple tenants safely on one card (like K8s scheduling many pods onto one node), how to partition resources between the very different workloads of training and inference, how to push utilization from &amp;ldquo;looks full&amp;rdquo; to &amp;ldquo;actually full&amp;rdquo;, that&amp;rsquo;s exactly what GPU virtualization and sharing (the work I do) solves.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a sobering number: per the &lt;a href="https://www.mirantis.com/blog/gpu-infrastructure-automation-and-strategy/" target="_blank" rel="noopener"&gt;ClearML AI infrastructure survey&lt;/a&gt;, only about &lt;strong&gt;7%&lt;/strong&gt; of enterprises hit over 85% GPU utilization at peak, more than half sit at 51-70%, and 15% are below 50%. In other words, a big chunk of the GPUs companies pay dearly for are spinning idle. That&amp;rsquo;s usually not a hardware shortage, it&amp;rsquo;s &lt;strong&gt;scheduling and management not keeping up.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;My own take is that GPU resource management is moving through three stages: &lt;strong&gt;Allocation → Utilization → Efficiency.&lt;/strong&gt; Stage one only answers &amp;ldquo;who gets this card&amp;rdquo; (device plugin, exclusive or time-slicing); stage two asks &amp;ldquo;is this card full&amp;rdquo; (dynamic batching, elastic autoscaling, where most enterprises are stuck); stage three asks &amp;ldquo;is this card&amp;rsquo;s compute producing maximum value&amp;rdquo; (SM Efficiency, tokens per watt, bandwidth utilization). The real challenge is stage three, and the future GPU scheduler shouldn&amp;rsquo;t be a resource allocator but a &lt;strong&gt;fine-grained orchestrator of the GPU&amp;rsquo;s internals&lt;/strong&gt;, reading how many SMs are active, the Tensor Core utilization, how much memory KV cache holds, and only then deciding whether one more request fits.&lt;/p&gt;
&lt;p&gt;Plainly put, &lt;strong&gt;the AI era is replaying the cloud-native story: from &amp;ldquo;single-machine exclusive&amp;rdquo; toward &amp;ldquo;multi-tenant sharing + scheduling + observability.&amp;rdquo;&lt;/strong&gt; And the GPU is the main stage of that play.&lt;/p&gt;
&lt;h2 id="this-is-only-the-first-half-who-else-is-at-the-table-besides-nvidia"&gt;This is only the first half: who else is at the table besides NVIDIA&lt;/h2&gt;
&lt;p&gt;By now you&amp;rsquo;ve probably noticed &lt;strong&gt;this post barely mentions anyone but NVIDIA.&lt;/strong&gt; Unavoidable, it&amp;rsquo;s the absolute protagonist today. But if you think the AI accelerator world begins and ends with NVIDIA, you&amp;rsquo;re very wrong.&lt;/p&gt;
&lt;p&gt;In fact, an &amp;ldquo;anti-NVIDIA alliance&amp;rdquo; is gathering from all sides:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cloud-vendor silicon&lt;/strong&gt;: Google&amp;rsquo;s TPU is on its sixth generation and backs almost all of its own AI; AWS&amp;rsquo;s Trainium (training) and Inferentia (inference) keep spreading; Meta and Microsoft aren&amp;rsquo;t sitting still.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Traditional chip giants&lt;/strong&gt;: AMD presses hard with the Instinct line and ROCm, Intel holds ground with Gaudi and oneAPI.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;China&amp;rsquo;s heterogeneous accelerator ecosystem&lt;/strong&gt;: Huawei Ascend, Hygon DCU, Cambricon, Moore Threads, Enflame, Kunlunxin, Metax, Biren&amp;hellip;, sprinting through the domestic-substitution window, with their software stacks climbing from &amp;ldquo;works&amp;rdquo; toward &amp;ldquo;works well&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The open ecosystem fights back&lt;/strong&gt;: the UXL Alliance, OpenAI Triton, and PyTorch&amp;rsquo;s native AMD/TPU backends are all gnawing at the walls of the CUDA moat.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What does that mean for a K8s veteran? It means &lt;strong&gt;the future AI cluster is almost certainly heterogeneous&lt;/strong&gt;: a rack might hold NVIDIA, AMD, TPU and domestic cards side by side, and one schedule has to manage several completely different kinds of hardware.&lt;/p&gt;
&lt;p&gt;So the really interesting questions follow:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How do you mix multiple accelerator types in one cluster and still schedule and observe them uniformly?&lt;/li&gt;
&lt;li&gt;Where do K8s device plugins and DRA (Dynamic Resource Allocation) evolve to, so they can elegantly describe this menagerie of hardware?&lt;/li&gt;
&lt;li&gt;Will the CUDA moat be breached by the open ecosystem, or will NVIDIA rule long-term like x86 did?&lt;/li&gt;
&lt;li&gt;What pieces are still missing before domestic heterogeneous accelerators are truly production-ready?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&amp;rsquo;ll dig into those in later posts. &lt;strong&gt;This one only nails &amp;ldquo;why GPUs&amp;rdquo;; the next will tackle the big chess game of how AI infrastructure should schedule things once &amp;ldquo;GPUs aren&amp;rsquo;t just one kind&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/why-gpus-became-foundation-modern-ai-ameen-alam-1oiee" target="_blank" rel="noopener"&gt;Why GPUs Became the Foundation of Modern AI&lt;/a&gt; (Part 1 of the trilogy)&lt;/li&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/hidden-technology-behind-modern-ai-gpus-ameen-alam-tzkre" target="_blank" rel="noopener"&gt;The Hidden Technology Behind Modern AI GPUs&lt;/a&gt; (Part 2)&lt;/li&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;Real AI Infrastructure and the Future of GPUs&lt;/a&gt; (Part 3)&lt;/li&gt;
&lt;li&gt;TrendForce, &lt;a href="https://insights.trendforce.com/p/agentic-ai-cpu-gpu" target="_blank" rel="noopener"&gt;The Great Rebalance: How Agentic AI Is Reshaping the CPU/GPU Ratio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Anton Polyakov (NVIDIA), &lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;Running Large-Scale GPU Workloads on Kubernetes with Slurm&lt;/a&gt; (Slinky / Slurm on K8s)&lt;/li&gt;
&lt;li&gt;Kawonise, &lt;a href="https://www.linkedin.com/pulse/why-nvidia-gpu-architecture-perfect-ai-insights-data-path-kawonise-a821e/" target="_blank" rel="noopener"&gt;Why NVIDIA GPU Architecture Is Perfect for AI: GPU Data Path for a Single Node&lt;/a&gt; (GPUDirect data path)&lt;/li&gt;
&lt;li&gt;Mirantis, &lt;a href="https://www.mirantis.com/blog/gpu-infrastructure-automation-and-strategy/" target="_blank" rel="noopener"&gt;GPU Infrastructure Automation and Strategy&lt;/a&gt; (GPU infra automation and utilization)&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>GPU Utilization Is Breaking: AI Infrastructure Needs a New Definition of Efficiency</title><link>https://jimmysong.io/blog/beyond-gpu-utilization-productive-gpu-hours/</link><pubDate>Wed, 17 Jun 2026 06:22:24 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/beyond-gpu-utilization-productive-gpu-hours/</guid><description>From GPU utilization to productive GPU-hours.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Every GPU should not just be used. It should create value.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="we-have-all-been-chasing-gpu-utilization"&gt;We have all been chasing GPU utilization&lt;/h2&gt;
&lt;p&gt;For the past few years, whether it is Kubernetes GPU scheduling, vGPU, MIG, or &lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;, everyone has really been doing the same thing: pushing one number up.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Utilization&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;It makes sense. GPUs are expensive. An H100 can run anywhere from a few dollars to over ten dollars per GPU-hour, and nobody can afford to let a GPU sit idle. So the entire AI Infra community&amp;rsquo;s narrative for the past few years has boiled down to one line:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Maximize GPU utilization.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I have been involved in the HAMi community for a long time, and a few posts I have written, like &lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes as the GPU Control Plane for AI&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;From GPU to Token: An Eight-Layer Observability Stack for AI Infrastructure&lt;/a&gt;, are really about the same thing: how to slice GPUs finer, share them more thoroughly, and schedule them more sensibly.&lt;/p&gt;
&lt;p&gt;But recently I read Arjun Kaarat&amp;rsquo;s piece in Towards Data Science, &lt;a href="https://towardsdatascience.com/when-gpu-utilization-lies-the-hidden-systems-problem-slowing-modern-ai/" target="_blank" rel="noopener"&gt;&lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI&lt;/em&gt;&lt;/a&gt;, and it made me rethink a question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Is GPU utilization really the metric we should be optimizing for?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There is one line in the article that made me pause for a few seconds the first time I read it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;GPUs can be busy without being productive.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Arjun Kaarat&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It punctures a layer of illusion. The number we have spent so much effort pushing up may have been answering the wrong question from the start.&lt;/p&gt;
&lt;h2 id="gpu-busy-is-not-gpu-productive"&gt;GPU Busy is not GPU Productive&lt;/h2&gt;
&lt;p&gt;Kaarat tells a representative story in the article.&lt;/p&gt;
&lt;p&gt;At 2 AM, an infrastructure team gets paged: inference latency just spiked 60%. They open the monitoring dashboard, and GPU utilization looks perfectly normal:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 79%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 82%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 84%&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Looks healthy. So the usual playbook kicks in: trigger autoscaling, add nodes, add GPUs. The cloud bill climbs, but latency barely improves.&lt;/p&gt;
&lt;p&gt;An hour later, they find the root cause: three nodes had quietly entered a RAID rebuild state, storage throughput was severely dragged down, and the inference tasks around them were starving. The scheduler kept treating these nodes as &amp;ldquo;still healthy enough&amp;rdquo; because the GPU and memory metrics looked fine, but the underlying disk performance had collapsed.&lt;/p&gt;
&lt;p&gt;What strikes me most about this story is that it is not a rare edge case. It is a failure mode that is becoming common.&lt;/p&gt;
&lt;p&gt;Many teams look at their monitoring and see:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 82%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 84%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 79%&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;and think &amp;ldquo;our cluster is busy and healthy.&amp;rdquo; But at the same time:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Latency ↑
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Queue ↑
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Throughput ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Cost ↑&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The real problem is that &lt;strong&gt;the GPU is waiting&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Waiting for what?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The retrieval pipeline to feed embeddings over;&lt;/li&gt;
&lt;li&gt;The SSD to read context out of storage;&lt;/li&gt;
&lt;li&gt;The CPU to prepare the data pipeline;&lt;/li&gt;
&lt;li&gt;The KV cache to have room for new requests;&lt;/li&gt;
&lt;li&gt;Storage I/O to not be squeezed out by background tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The GPU end is full, but the data path feeding it is empty, blocked, or collapsing. From the dashboard the GPU reads 84%, but the actual output may be less than half. Kaarat describes this state precisely in the original piece:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A GPU that appears active may still spend meaningful time waiting for the system around it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is the overlooked gap between &amp;ldquo;busy&amp;rdquo; and &amp;ldquo;productive&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="the-illusion-of-gpu-utilization"&gt;The &amp;ldquo;illusion&amp;rdquo; of GPU utilization&lt;/h2&gt;
&lt;p&gt;The RAID story above is still just a &amp;ldquo;point failure&amp;rdquo;. The more compelling part of Kaarat&amp;rsquo;s article describes a systemic phenomenon: &lt;strong&gt;Fragmentation&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Consider a cluster with three nodes, after running a mixed wave of GenAI workloads:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Node&lt;/th&gt;
&lt;th&gt;GPU compute&lt;/th&gt;
&lt;th&gt;HBM&lt;/th&gt;
&lt;th&gt;Storage bandwidth&lt;/th&gt;
&lt;th&gt;I/O CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;nearly full&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;saturated&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;limited&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;saturated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Residual resources on three nodes after a GenAI wave
&lt;/figcaption&gt;
&lt;p&gt;Now a new inference job arrives, with an ordinary footprint: a little GPU, a little VRAM, decent storage bandwidth, decent I/O capacity.&lt;/p&gt;
&lt;p&gt;In total, the cluster still has plenty of resources. A has GPU and bandwidth, B has VRAM, C has bandwidth and CPU. But no single node can take this job on its own.&lt;/p&gt;
&lt;p&gt;That is fragmentation. I drew it out, roughly like this:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/fragmentation-three-nodes-en.svg" data-img="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/fragmentation-three-nodes-en.svg" alt="Figure 1: The cluster is not short on resources. It is short on resources of the right shape." data-caption="Figure 1: The cluster is not short on resources. It is short on resources of the right shape."
width="702"
height="481"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: The cluster is not short on resources. It is short on resources of the right shape.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The cluster is not empty. It has just been carved into &amp;ldquo;leftovers&amp;rdquo; that can no longer be used productively. Kaarat sums up the phenomenon in one line, which I think is the most memorable sentence in the whole piece:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The cluster is not empty. It is fragmented into leftovers that are difficult to use productively.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Put another way:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The cluster does not lack resources. It lacks resources of the right shape.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This judgment matters a lot for a project like HAMi, which builds a GPU resource control plane. We used to think that &amp;ldquo;slicing the card finely and letting more people share it&amp;rdquo; was the answer to fragmentation. But Kaarat points to a deeper problem: &lt;strong&gt;fragmentation is not just a GPU-layer concern&lt;/strong&gt;. It spans GPU, HBM, storage bandwidth, and I/O CPU. You can slice the GPU as finely as you like, but if the storage dimension is choked, that node is still unavailable for the next genuinely useful task.&lt;/p&gt;
&lt;h2 id="what-hami-solves"&gt;What HAMi solves&lt;/h2&gt;
&lt;p&gt;Let me directly answer a question: what role does HAMi play in this chain?&lt;/p&gt;
&lt;p&gt;HAMi solves a very specific, and very foundational, problem:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Can the GPU be used by more people.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;What it does can be summarized in one line: &lt;strong&gt;reduce fragmentation at the GPU layer&lt;/strong&gt;. The concrete forms include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPU Sharing&lt;/strong&gt;: letting multiple Pods share one card instead of one card per Pod;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;vGPU / HAMi-core&lt;/strong&gt;: doing memory isolation and compute throttling in userspace, slicing one card into MB-level virtual devices;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MIG integration&lt;/strong&gt;: managing NVIDIA MIG hardware partitions in software;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Heterogeneous GPU abstraction&lt;/strong&gt;: abstracting more than a dozen device families, including NVIDIA, Ascend, Cambricon, Hygon, and Vastai, into semantics the scheduler can consume uniformly;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DRA compatibility&lt;/strong&gt;: keeping pace with the evolution of the Kubernetes resource model.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the first layer of efficiency. It answers the question &amp;ldquo;can the GPU be put to use&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;At this layer, HAMi&amp;rsquo;s value is already clear: in real environments with mixed domestic and foreign GPUs, mixed training and inference, and multi-tenant sharing, HAMi lets one card serve more workloads, so the cluster is no longer wasted by a coarse-grained &amp;ldquo;one card per Pod&amp;rdquo; model.&lt;/p&gt;
&lt;p&gt;But note: this is only the first chapter of the efficiency story.&lt;/p&gt;
&lt;h2 id="what-comes-after-hami"&gt;What comes after HAMi&lt;/h2&gt;
&lt;p&gt;If I step back and look at it from a higher vantage point, &amp;ldquo;improving GPU utilization&amp;rdquo; is actually solved across three distinct layers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer one: can the GPU be sliced, shared, and allocated?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is what a project like HAMi solves. It corresponds to the GPU resource control plane.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer two: who gets to use the GPU? Who runs first, who queues, who has priority?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is what schedulers like &lt;a href="https://volcano.sh/en/" target="_blank" rel="noopener"&gt;Volcano&lt;/a&gt;, &lt;a href="https://kueue.sigs.k8s.io/" target="_blank" rel="noopener"&gt;Kueue&lt;/a&gt;, and &lt;a href="https://github.com/NVIDIA/KAI-Scheduler" target="_blank" rel="noopener"&gt;KAI Scheduler&lt;/a&gt; solve. It corresponds to job queuing, fair share, priority, and gang scheduling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer three: can the GPU actually run? Are the data path, storage I/O, and KV cache keeping up?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is exactly where Kaarat&amp;rsquo;s article sounds the alarm. When a node&amp;rsquo;s RAID is rebuilding, its SSD queue is exploding, and its I/O CPU is eaten by background tasks, no matter how much GPU you allocate to it and no matter how elegantly the scheduler queues its tasks, it still cannot produce effective compute.&lt;/p&gt;
&lt;p&gt;Draw these three layers together, and you get what next-generation AI infrastructure should actually look like:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/productive-gpu-hours-stack-en.svg" data-img="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/productive-gpu-hours-stack-en.svg" alt="Figure 2: From GPU utilization to productive GPU-hours" data-caption="Figure 2: From GPU utilization to productive GPU-hours"
width="695"
height="552"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: From GPU utilization to productive GPU-hours&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I think this diagram is the single most important one for understanding the whole picture.&lt;/p&gt;
&lt;p&gt;For the past few years, almost all of the industry&amp;rsquo;s attention has been on the bottom two layers: Kubernetes and HAMi. Those two layers have essentially solved &amp;ldquo;can the GPU be put to use&amp;rdquo;. Volcano, Kueue, and KAI are also mature at layer two, solving queuing and priority.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;But layer three, storage / I/O-aware scheduling, is currently almost a blank.&lt;/strong&gt; And that is precisely the layer Kaarat&amp;rsquo;s article keeps emphasizing, the one that is becoming increasingly valuable in modern GenAI systems. Because for workloads like RAG, long context, and multimodal, the bottleneck has long since shifted from &amp;ldquo;is the GPU enough&amp;rdquo; to &amp;ldquo;is the data path feeding the GPU clear&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;To put it bluntly: HAMi slices the GPU finely, and Volcano queues the jobs well, but if a node assigned a task cannot feed its GPU, then all that upstream effort is just pumping blood into an idle endpoint.&lt;/p&gt;
&lt;h2 id="what-should-next-gen-ai-infrastructure-optimize-for"&gt;What should next-gen AI infrastructure optimize for&lt;/h2&gt;
&lt;p&gt;Based on the layering above, I want to make one clear point: &lt;strong&gt;we should upgrade the optimization target from &amp;ldquo;GPU utilization&amp;rdquo; to &amp;ldquo;Productive GPU-Hours&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is not wordplay. It is a shift that translates directly into money.&lt;/p&gt;
&lt;p&gt;Kaarat does the math in the article. A 1000-H100 cluster, at a blended cost of about $3 per GPU-hour, runs around $26 million a year. If fragmentation and I/O stall quietly waste 10% of the effective GPU time, that is roughly $2.6 million a year of wasted spend. Not because the GPUs are missing, but because the system failed to use them efficiently.&lt;/p&gt;
&lt;p&gt;That math can be translated into a simple contrast.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The past target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize GPU Utilization&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Today&amp;rsquo;s target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize Productive GPU-Hours&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;The future target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize Productive Compute
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Across Heterogeneous AI Clusters&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This evolution maps exactly onto the three-layer structure in the diagram above. That final line, &amp;ldquo;Across Heterogeneous AI Clusters&amp;rdquo;, is the heterogeneous narrative HAMi has been pushing all along: in the future you will not optimize just one kind of GPU. You will maximize effective compute uniformly across completely different cards from NVIDIA, Ascend, Cambricon, and Hygon.&lt;/p&gt;
&lt;p&gt;In other words, HAMi&amp;rsquo;s long-term value should not be boxed into the old &amp;ldquo;improve GPU utilization&amp;rdquo; narrative. Its real direction is: &lt;strong&gt;a resource control plane that lets Productive GPU-Hours be maximized across heterogeneous AI clusters.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="from-utilization-to-productive-gpu-hours"&gt;From utilization to productive GPU-hours&lt;/h2&gt;
&lt;p&gt;If I had to summarize this whole line of thinking in one sentence, I would put it like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;HAMi solves &amp;ldquo;how to let more GPUs be used&amp;rdquo;, and next-generation AI infrastructure has to solve &amp;ldquo;how to let every GPU actually create value&amp;rdquo;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is also the deepest line Kaarat&amp;rsquo;s article left me with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The real question is no longer &amp;ldquo;Are the GPUs busy?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;It is: &amp;ldquo;Are they productively busy?&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It also makes me rethink what the HAMi community&amp;rsquo;s tagline for the next phase should be. We used to say &amp;ldquo;let GPUs be shared by more people&amp;rdquo;. Next, we should probably move toward:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Turn GPU utilization into productive GPU-hours.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Let every GPU not just be used, but genuinely create value.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;GPU utilization was never the endpoint. It was only the first chapter of the efficiency story. Only when we start talking about Productive GPU-Hours, and start caring about storage, I/O, the data path, and the aggregate output of heterogeneous clusters, does AI infrastructure truly enter its second chapter.&lt;/p&gt;
&lt;h2 id="acknowledgments-and-references"&gt;Acknowledgments and references&lt;/h2&gt;
&lt;p&gt;Parts of this post were inspired by Arjun Kaarat&amp;rsquo;s piece in Towards Data Science, &lt;a href="https://towardsdatascience.com/when-gpu-utilization-lies-the-hidden-systems-problem-slowing-modern-ai/" target="_blank" rel="noopener"&gt;&lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI&lt;/em&gt;&lt;/a&gt;, and quote its published paper and article. My thanks to the author. Both diagrams in this post (resource fragmentation, the efficiency layering) were redrawn by the author based on the ideas in the original, not copied from it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Arjun Kaarat. &lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI.&lt;/em&gt; Towards Data Science, 2026.&lt;/li&gt;
&lt;li&gt;Kaarat, A., Batthula, V. J. R., &amp;amp; Segall, R. &lt;em&gt;Fitting the Void: Residual-Aware Geometric Packing for GenAI Workloads.&lt;/em&gt; IEEE, 2025.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Further reading:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes as the GPU Control Plane: HAMi v2.9 and Next-Gen AI Infra&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;From GPU to Token: An Eight-Layer Observability Stack for AI Infrastructure&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ai-inference-on-kubernetes"&gt;Why AI Inference Belongs on Kubernetes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>When an Agent Becomes a Distributed State Machine: Agentic AI Infrastructure Reliability</title><link>https://jimmysong.io/blog/agentic-ai-infrastructure-reliability/</link><pubDate>Tue, 16 Jun 2026 12:45:16 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/agentic-ai-infrastructure-reliability/</guid><description>A practical AI Infra review of Agentic AI reliability, covering a five-dimension framework, fault tolerance, recovery, observability, and hybrid architecture design.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Statement
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This is my reading and critique of the paper &amp;ldquo;AI Infrastructure Reliability Features and Architecture for Agentic AI&amp;rdquo; (June 2026), shared by Hesham ElBakoury in the &lt;a href="https://www.opencompute.org" target="_blank" rel="noopener"&gt;Open Compute Project (OCP)&lt;/a&gt; community. It blends my personal engineering perspective from the Kubernetes / AI Infra space, and is not a translation of the original paper.
&lt;/div&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;An Agent is not a single inference request but a long-running distributed state machine; therefore, Agent reliability is fundamentally a distributed-systems problem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="why-i-wanted-to-write-about-this-paper"&gt;Why I wanted to write about this paper&lt;/h2&gt;
&lt;p&gt;After years in cloud native, I have a stubborn instinct: &lt;strong&gt;whether a new technology is mature is not judged by how powerful its model is or how flashy the demo looks, but by whether it has a systematic language for reliability.&lt;/strong&gt; Kubernetes won not on the container runtime, but on liveness/readiness probes, the reconciler loop, PDs/PDBs, and the whole SRE vocabulary that made &amp;ldquo;long-running distributed state&amp;rdquo; legible.&lt;/p&gt;
&lt;p&gt;So when I came across the paper Hesham ElBakoury shared in the &lt;a href="https://www.opencompute.org" target="_blank" rel="noopener"&gt;Open Compute Project (OCP)&lt;/a&gt; community, it caught my eye. It proposes no new algorithm or framework; the entire paper does one thing: &lt;strong&gt;systematically translate the traditional SRE reliability vocabulary onto Agentic AI.&lt;/strong&gt; This is exactly what I have been circling around for the past six months in &lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt;, without a canonical reference to align against. So I decided to write a dedicated post: unpack its framework, then give it a critical evaluation from the AI Infra practitioner&amp;rsquo;s point of view.&lt;/p&gt;
&lt;p&gt;One-line positioning: &lt;strong&gt;it reads more like an SRE white paper for Agentic AI than a systems paper.&lt;/strong&gt; Its value is not in novelty but in &amp;ldquo;building consensus&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="one-line-summary"&gt;One-line summary&lt;/h2&gt;
&lt;p&gt;Traditional AI cares about model accuracy, while Agentic AI must care about &amp;ldquo;reliability during long-running operation&amp;rdquo;, so fault tolerance, recovery, monitoring, security, and state management must be elevated to first-class citizens of architectural design.&lt;/p&gt;
&lt;h2 id="the-fundamental-split-between-traditional-ai-and-agentic-ai"&gt;The fundamental split between Traditional AI and Agentic AI&lt;/h2&gt;
&lt;p&gt;The paper argues the fundamental difference between Traditional AI and Agentic AI lies in the &lt;strong&gt;execution model&lt;/strong&gt;. Traditional AI is one-shot request-response, focused on accuracy, latency, and throughput; Agentic AI is a continuous loop, where &lt;strong&gt;a single error is no longer just a wrong answer but a wrong decision that may change every subsequent action of the Agent.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/execution-model-comparison-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/execution-model-comparison-en.svg" alt="Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop" data-caption="Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop"
width="859"
height="290"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The comparison table below states this most clearly. I suggest you focus on the last row, &amp;ldquo;failure impact&amp;rdquo;, because that is the thesis of the whole paper:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Traditional AI&lt;/th&gt;
&lt;th&gt;Agentic AI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Execution model&lt;/td&gt;
&lt;td&gt;Request-response&lt;/td&gt;
&lt;td&gt;Continuous loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State management&lt;/td&gt;
&lt;td&gt;Stateless&lt;/td&gt;
&lt;td&gt;Stateful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger&lt;/td&gt;
&lt;td&gt;Human-triggered&lt;/td&gt;
&lt;td&gt;Self-initiated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time span&lt;/td&gt;
&lt;td&gt;Single interaction&lt;/td&gt;
&lt;td&gt;Long-running session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure impact&lt;/td&gt;
&lt;td&gt;Single wrong output&lt;/td&gt;
&lt;td&gt;Cascading behavior changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource usage&lt;/td&gt;
&lt;td&gt;Bursty&lt;/td&gt;
&lt;td&gt;Continuous&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Execution model: Traditional AI vs Agentic AI
&lt;/figcaption&gt;
&lt;p&gt;The point of this table is not the technical details but the shift in mental model: when the impact of failure escalates from &amp;ldquo;one wrong answer&amp;rdquo; to &amp;ldquo;behavior-level cascading errors&amp;rdquo;, the weight of reliability must be reassigned from the very bottom of the architecture. &lt;strong&gt;Engineering fault tolerance for a stateless API and for a self-directing, continuously running state machine are simply not the same problem.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-core-contribution-a-five-dimension-agent-reliability-framework"&gt;The core contribution: a five-dimension Agent reliability framework&lt;/h2&gt;
&lt;p&gt;This is the most memorable part of the paper. The author decomposes Agent reliability into five dimensions, forming a complete evaluation coordinate system. I drew it out: &amp;ldquo;Agent Reliability&amp;rdquo; sits in the center, with the five dimensions fanning out like petals:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/five-dim-framework-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/five-dim-framework-en.svg" alt="Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it" data-caption="Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it"
width="838"
height="525"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One by one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Functional Reliability&lt;/strong&gt;: can the Agent complete the task? Concerns correctness, accuracy, consistency, completeness. Plainly: &lt;em&gt;can this Agent get the job done?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Temporal Reliability&lt;/strong&gt;: can the Agent stay stable over time? Concerns timeliness, responsiveness, stability, durability. Plainly: &lt;em&gt;can it keep getting the job done?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environmental Reliability&lt;/strong&gt;: is it still effective when the environment changes? Concerns adaptability, robustness, portability. Plainly: &lt;em&gt;can it still work in a different environment?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Social Reliability&lt;/strong&gt;: this is an Agent-unique dimension. Concerns safety, trustworthiness, collaborativity, explainability. Plainly: &lt;em&gt;can it collaborate safely with humans and other Agents?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systemic Reliability&lt;/strong&gt;: the infrastructure layer. Concerns availability, fault tolerance, scalability, security. Plainly: &lt;em&gt;is the platform underneath the Agent solid?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
Why this framework matters
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
The real value of these five dimensions is that they turn &amp;ldquo;Agent reliability&amp;rdquo; from a vague concept into &lt;strong&gt;a decomposable, measurable, separately-owned engineering problem.&lt;/strong&gt; The first four describe behavioral attributes of the Agent itself; the fifth describes the platform that carries it, &lt;strong&gt;and that fifth dimension is exactly the landing point those of us doing AI Infra / Kubernetes should pick up.&lt;/strong&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="treat-the-agent-as-a-stateful-service"&gt;Treat the Agent as a &amp;ldquo;stateful service&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;The paper&amp;rsquo;s second core judgment is one I strongly agree with: &lt;strong&gt;the biggest infrastructure challenge for an Agent is not inference, it is state.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;An Agent simultaneously holds Memory, Goal, Plan, Context, and tool state. Treating it as &lt;code&gt;HTTP API + Model&lt;/code&gt; to operate is the root cause of why most demos today cannot survive production. Its true shape is closer to a composite of database + workflow engine + LLM + distributed system:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/stateful-runtime-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/stateful-runtime-en.svg" alt="Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system" data-caption="Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system"
width="892"
height="271"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is entirely consistent with where the industry is heading: LangGraph, Temporal, and the OpenAI Responses API are all promoting &amp;ldquo;stateful, long-running logic&amp;rdquo; to a first-class citizen. &lt;strong&gt;When an Agent runs for hours or even days, its state is more critical than the model parameters themselves, and far easier to lose for good on a restart.&lt;/strong&gt; A model can be reloaded, but a three-hour planning context, once lost, is usually lost.&lt;/p&gt;
&lt;h2 id="fault-tolerance-the-golden-trio-of-agent-systems"&gt;Fault tolerance: the &amp;ldquo;golden trio&amp;rdquo; of Agent systems&lt;/h2&gt;
&lt;p&gt;The paper provides a catalog of Agent fault-tolerance patterns:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Redundancy&lt;/td&gt;
&lt;td&gt;Standby Agent takes over at any time&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpointing&lt;/td&gt;
&lt;td&gt;Save state for recovery&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heartbeat&lt;/td&gt;
&lt;td&gt;Liveness detection&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Circuit Breaker&lt;/td&gt;
&lt;td&gt;Isolate abnormal Agents, prevent cascading&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rollback&lt;/td&gt;
&lt;td&gt;Revert a wrong decision&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quorum&lt;/td&gt;
&lt;td&gt;Multiple Agents vote to reach consensus&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-Healing&lt;/td&gt;
&lt;td&gt;Automatically detect and correct errors&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: Agent fault-tolerance pattern catalog: mechanism, effect, and implementation complexity
&lt;/figcaption&gt;
&lt;p&gt;The author values the &lt;strong&gt;Checkpointing + Redundancy + Heartbeat&lt;/strong&gt; trio most, calling it the &amp;ldquo;golden trio&amp;rdquo; of Agent systems. I drew a closed loop to show how they relate; none of the three can be missing, and drop any one link and the recovery chain is broken:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/golden-trio-loop-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/golden-trio-loop-en.svg" alt="Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance" data-caption="Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance"
width="838"
height="429"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
The closed-loop logic of the golden trio
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
Heartbeat &lt;strong&gt;detects problems fast&lt;/strong&gt;, checkpointing &lt;strong&gt;saves recoverable state&lt;/strong&gt;, and redundancy &lt;strong&gt;takes over seamlessly on failure.&lt;/strong&gt; Together they form the closed loop of &amp;ldquo;detect the problem → save the scene → switch the instance&amp;rdquo;. This is also why these mechanisms all sound like old friends from distributed systems: because Agent reliability is, by nature, a distributed-systems problem.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="recovery-porting-traditional-dr-thinking-onto-the-agent"&gt;Recovery: porting traditional DR thinking onto the Agent&lt;/h2&gt;
&lt;p&gt;This section essentially ports the RTO/RPO thinking of traditional DR (disaster recovery) onto the Agent scenario. The paper compares several recovery approaches:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Recovery approach&lt;/th&gt;
&lt;th&gt;RTO&lt;/th&gt;
&lt;th&gt;RPO&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cold Start&lt;/td&gt;
&lt;td&gt;minutes to hours&lt;/td&gt;
&lt;td&gt;High (total loss)&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm Start&lt;/td&gt;
&lt;td&gt;seconds to minutes&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hot Start&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;td&gt;Low (minimal loss)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpoint Recovery&lt;/td&gt;
&lt;td&gt;seconds to minutes&lt;/td&gt;
&lt;td&gt;Medium (since last checkpoint)&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State Reconstruction&lt;/td&gt;
&lt;td&gt;minutes to hours&lt;/td&gt;
&lt;td&gt;Low (fully recoverable)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Agent recovery approaches compared: RTO, RPO, and implementation complexity
&lt;/figcaption&gt;
&lt;p&gt;The author&amp;rsquo;s conclusion: &lt;strong&gt;the Agent fits Checkpoint Recovery best.&lt;/strong&gt; The reason was foreshadowed above: an Agent&amp;rsquo;s state is far more important than its model parameters. Checkpoint Recovery strikes the most balanced trade-off among RTO, RPO, and implementation complexity. That is also why checkpointing sits at the center of the &amp;ldquo;golden trio&amp;rdquo; earlier.&lt;/p&gt;
&lt;h2 id="agent-observability--infrastructure-metrics--behavior-analysis"&gt;Agent observability = infrastructure metrics + behavior analysis&lt;/h2&gt;
&lt;p&gt;This section resonates with me the most. The paper argues that traditional monitoring is far from enough: traditional monitoring watches CPU, memory, latency, and QPS, while an Agent must additionally monitor the behavior itself.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agent Health&lt;/strong&gt;: heartbeat, resource utilization, error rate, decision latency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Behavioral&lt;/strong&gt;: action frequency, decision patterns, state transitions, goal progress.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;System&lt;/strong&gt;: Agent count, communication volume, infrastructure health.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The &amp;ldquo;behavior&amp;rdquo; layer is unique to Agents.&lt;/p&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
The easiest trap to fall into
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
An Agent&amp;rsquo;s CPU may be low and its latency normal, but if its &amp;ldquo;action frequency&amp;rdquo; suddenly spikes or its &amp;ldquo;decision pattern&amp;rdquo; deviates from baseline, &lt;strong&gt;that is often the truly dangerous signal.&lt;/strong&gt; In other words: &lt;strong&gt;Agent Observability = Infra Metrics + Behavior Analytics.&lt;/strong&gt; Watching only infrastructure metrics means you will perfectly miss the Agent&amp;rsquo;s &amp;ldquo;behavioral runaway&amp;rdquo;.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;This lines up completely with the thinking I discussed in &lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;GPU to Token Observability&lt;/a&gt;: in the Agent era, behavioral signals must become first-class citizens of observability, rather than staying stuck at the hardware / inference-metric layer.&lt;/p&gt;
&lt;h2 id="the-recommended-hybrid-architecture"&gt;The recommended hybrid architecture&lt;/h2&gt;
&lt;p&gt;The paper compares four architectures:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;th&gt;Scalability&lt;/th&gt;
&lt;th&gt;Fault isolation&lt;/th&gt;
&lt;th&gt;Suitability for Agents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monolithic&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layered&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microservices&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Event-driven&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 4: Four architectures compared for Agent systems
&lt;/figcaption&gt;
&lt;p&gt;The final recommendation: &lt;strong&gt;production-grade Agent systems adopt a hybrid architecture&lt;/strong&gt;, the combination of &amp;ldquo;layered + microservices + event-driven&amp;rdquo;:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/hybrid-architecture-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/hybrid-architecture-en.svg" alt="Figure 5: The hybrid architecture for production-grade Agent systems: layered &amp;#43; microservices &amp;#43; event-driven" data-caption="Figure 5: The hybrid architecture for production-grade Agent systems: layered &amp;#43; microservices &amp;#43; event-driven"
width="818"
height="443"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: The hybrid architecture for production-grade Agent systems: layered + microservices + event-driven&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Layering provides clear separation of concerns, microservices provide independent scaling and fault isolation, and event-driven provides loosely-coupled async collaboration. &lt;strong&gt;No single architecture can satisfy an Agent system&amp;rsquo;s demands for clarity, elasticity, and collaboration at once, so hybrid is the inevitable conclusion.&lt;/strong&gt; For K8s veterans this layered stack should look very familiar: it is the cloud-native layering of governance, ported verbatim onto the Agent.&lt;/p&gt;
&lt;h2 id="the-biggest-value-of-this-paper-from-the-ai-infra-perspective"&gt;The biggest value of this paper, from the AI Infra perspective&lt;/h2&gt;
&lt;p&gt;From the Kubernetes / AI Infra standpoint, what this paper really conveys is a paradigm shift: &lt;strong&gt;Agent infrastructure will move from &amp;ldquo;Serving&amp;rdquo; to &amp;ldquo;Runtime&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/serving-to-runtime-shift-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/serving-to-runtime-shift-en.svg" alt="Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward" data-caption="Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward"
width="1095"
height="121"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The past cared about GPU utilization, throughput, and latency; the future cares about state recovery, long-task continuity, Agent isolation, Agent scheduling, Agent observability, and Agent security governance.&lt;/p&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
My read on the trend
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
&lt;strong&gt;The next phase of AI Infra is not better model Serving, but a more reliable Agent Runtime.&lt;/strong&gt; This is exactly the direction I have kept emphasizing in &lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism&lt;/a&gt;: the Agent is moving from &amp;ldquo;a class you import&amp;rdquo; to &amp;ldquo;a workload you must govern&amp;rdquo;.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="my-evaluation"&gt;My evaluation&lt;/h2&gt;
&lt;p&gt;Strengths:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proposes a fairly complete Agent reliability framework (the five-dimension model), turning a vague concept into a measurable engineering problem.&lt;/li&gt;
&lt;li&gt;Systematically ports traditional SRE thinking onto the Agent scenario, with clear mappings for fault tolerance, recovery, and monitoring.&lt;/li&gt;
&lt;li&gt;Offers strong engineering guidance on architecture selection (layered / microservices / event-driven / hybrid).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Weaknesses:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Lacks real production cases: the three cases in the paper (autonomous vehicle fleets, supply-chain optimization, customer-service bots) read more as illustrative scenarios than verifiable engineering practice.&lt;/li&gt;
&lt;li&gt;The math models (series/parallel reliability, Markov chains, MTBF/MTTR) are essentially standard reliability-engineering textbook material, not innovation.&lt;/li&gt;
&lt;li&gt;It never touches real Agent Runtime implementations like Kubernetes, Ray, Temporal, or LangGraph, so it feels light on engineering grounding.&lt;/li&gt;
&lt;li&gt;Its discussion of GPU and inference systems, the core AI Infra resource layer, is shallow, barely staying at the &amp;ldquo;compute/storage/network&amp;rdquo; level of abstraction.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
My score
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
&lt;p&gt;By &lt;strong&gt;academic novelty&lt;/strong&gt;: &lt;strong&gt;6.5 / 10&lt;/strong&gt;.
By &amp;ldquo;&lt;strong&gt;giving AI Infra practitioners a mental framework for Agent reliability&lt;/strong&gt;&amp;rdquo;: &lt;strong&gt;8 / 10&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Its biggest insight is not some technical detail but that opening line: &lt;em&gt;an Agent is not a single inference request but a long-running distributed state machine.&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The real contribution of this paper is that it &lt;strong&gt;re-categorizes&lt;/strong&gt; the reliability problem of Agentic AI. It is no longer a model problem, nor a prompt-engineering problem, but a distributed-systems problem, and specifically a distributed-systems problem that is stateful, long-running, and self-directing.&lt;/p&gt;
&lt;p&gt;For AI Infra practitioners, this means two things. First, the traditional SRE toolbox (checkpointing, redundancy, heartbeat, circuit breaker, RTO/RPO) can be ported over directly, but it must be extended to the &amp;ldquo;behavior layer&amp;rdquo;. Second, the focus of infrastructure will shift from &amp;ldquo;how to serve models faster&amp;rdquo; to &amp;ldquo;how to run Agents more reliably&amp;rdquo;, which is the Agent Runtime.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The AI platforms of the future must not only run fast, they must run stable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism: Insights from McKinsey Ark on 2026 Infrastructure Trends&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;GPU to Token Observability: An Eight-Layer Observation System from Hardware to Business&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ai-2026-infra-agentic-runtime"&gt;AI 2026: Infrastructure, Agents, and the Next Cloud-Native Shift&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>From GPU to Token: The 8-Layer Observability Stack for AI Infrastructure</title><link>https://jimmysong.io/blog/gpu-to-token-observability/</link><pubDate>Tue, 09 Jun 2026 03:31:16 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gpu-to-token-observability/</guid><description>From GPU hardware, Kubernetes scheduling, inference engines to token cost — understanding the 8-layer observability architecture for modern AI infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;GPU utilization is not the destination. Token cost is the true North Star metric for AI infrastructure.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For the past few years, one of the hottest topics in AI infrastructure has been GPU scheduling. Whether it&amp;rsquo;s Kubernetes, Volcano, Kueue, or &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;, they all fundamentally solve the same problem — how to make expensive and scarce GPUs more efficiently utilized.&lt;/p&gt;
&lt;p&gt;But as more enterprises begin running production-grade Large Language Model (LLM) services, a new phenomenon has emerged: GPU utilization is high, but users still complain about slow response times; GPU clusters are near full capacity, but business throughput hasn&amp;rsquo;t grown proportionally; VRAM and compute still have headroom, yet TTFT (Time To First Token) continues to degrade. This reveals a fundamental truth — GPU utilization alone is no longer sufficient to describe the real operational state of modern AI systems.&lt;/p&gt;
&lt;p&gt;For traditional cloud-native applications, we focus on CPU, memory, network, and disk. For AI systems, however, we also need to care about: whether GPUs are truly performing useful computation, whether NCCL (NVIDIA Collective Communications Library) communication has become a bottleneck, whether Kubernetes has correctly allocated resources, whether KV Cache is exhausted, whether Token latency meets user experience requirements, and whether the cost per Token is reasonable. In other words, the observability target for AI infrastructure has expanded from the GPU to the entire inference chain.&lt;/p&gt;
&lt;p&gt;Recently, while participating in the development of an industry standard — &amp;ldquo;Technical Capability Requirements for Computing Power Efficiency Enhancement: Heterogeneous Computing Services&amp;rdquo; organized by the China Academy of Information and Communications Technology (CAICT) — a core question came up repeatedly in discussions with experts: &lt;strong&gt;How do GPU and Token relate? How do we measure GPU output through Tokens?&lt;/strong&gt; The standard introduces the concept of &amp;ldquo;Token as a Service,&amp;rdquo; shifting the unit of measurement for computing services from traditional GPU hours to Tokens, yet a mature practical framework for building a complete observability chain from hardware to Token was still missing.&lt;/p&gt;
&lt;p&gt;Around the same time, I came across a technical article on GPU and LLM observability (&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;GPU &amp;amp; LLM Inference Observability — Layer-by-Layer Coverage&lt;/a&gt;) that proposed a constructive approach: decomposing the AI system into eight observability layers from GPU hardware to business cost, which neatly fills the observability gap between GPU and Token. This article reorganizes and interprets the original content, removes product-specific implementation details and vendor promotion, and extends the model by incorporating Kubernetes, HAMi, and modern LLM inference architectures.&lt;/p&gt;
&lt;p&gt;This article aims to answer one question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After we&amp;rsquo;ve solved GPU scheduling, what should we observe next?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="8-layer-observability-architecture-overview"&gt;8-Layer Observability Architecture Overview&lt;/h2&gt;
&lt;p&gt;The diagram below shows the eight observability layers from GPU hardware to business cost, each corresponding to different observability targets and areas of responsibility:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-to-token-observability/gpu-8-layer-observability.svg" data-img="https://assets.jimmysong.io/images/blog/gpu-to-token-observability/gpu-8-layer-observability.svg" alt="Figure 1: 8-Layer AI Infrastructure Observability Architecture" data-caption="Figure 1: 8-Layer AI Infrastructure Observability Architecture"
width="750"
height="940"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: 8-Layer AI Infrastructure Observability Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;Original image source GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Each layer collects data through the OTel Collector and routes it to metrics, tracing, and logging backends, forming a complete observability loop. Let&amp;rsquo;s examine each layer in detail.&lt;/p&gt;
&lt;h2 id="l1-gpu-hardware-layer"&gt;L1 GPU Hardware Layer&lt;/h2&gt;
&lt;p&gt;This layer focuses on the GPU itself. The core question is: &lt;strong&gt;Is the GPU running healthy?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The table below lists the key observability metrics for the GPU hardware layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Key Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Compute&lt;/td&gt;
&lt;td&gt;GPU Utilization, SM Occupancy, Tensor Core Activity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;VRAM Usage, HBM Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interconnect&lt;/td&gt;
&lt;td&gt;NVLink Throughput, PCIe Throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reliability&lt;/td&gt;
&lt;td&gt;ECC Error, XID Error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thermal&lt;/td&gt;
&lt;td&gt;Temperature, Power Draw, Throttle Reason&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: L1 GPU Hardware Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;Many teams only track GPU Utilization, but there&amp;rsquo;s a critical distinction — &lt;strong&gt;GPU Utilization ≠ GPU Efficiency&lt;/strong&gt;. Two GPUs might both show 90% Utilization: one executing Tensor Core operations, the other waiting for memory access. The actual performance can be vastly different. Therefore, SM (Streaming Multiprocessor) Occupancy and Tensor Core Activity are often more valuable than Utilization alone.&lt;/p&gt;
&lt;h2 id="l2-cuda-runtime-and-communication-layer"&gt;L2 CUDA Runtime and Communication Layer&lt;/h2&gt;
&lt;p&gt;For distributed training, GPU computation is usually not the bottleneck — communication is. The core question at this layer is: &lt;strong&gt;Is the GPU computing, or waiting for other GPUs?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The table below lists the key metrics for the CUDA runtime and communication layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NCCL&lt;/td&gt;
&lt;td&gt;AllReduce, AllGather, ReduceScatter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Communication&lt;/td&gt;
&lt;td&gt;Duration, Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernel&lt;/td&gt;
&lt;td&gt;Execution Time, P99 Duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Straggler&lt;/td&gt;
&lt;td&gt;Rank Skew&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: L2 CUDA Runtime and Communication Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;In large-scale training scenarios, a single slow node can drag down the entire training job. By monitoring NCCL communication Duration and Bandwidth, as well as Skew between Ranks, you can quickly pinpoint communication bottlenecks.&lt;/p&gt;
&lt;h2 id="l3-host--os-layer"&gt;L3 Host / OS Layer&lt;/h2&gt;
&lt;p&gt;Many GPU problems are ultimately not GPU problems. When GPU utilization is abnormal, the root cause may lie in the host&amp;rsquo;s CPU, memory, disk, or network.&lt;/p&gt;
&lt;p&gt;The table below lists the key metrics at the host level:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Utilization, IO Wait&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Usage, Swap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk&lt;/td&gt;
&lt;td&gt;Throughput, Latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;td&gt;Retransmit, Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Process&lt;/td&gt;
&lt;td&gt;RSS, Thread Count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: L3 Host / OS Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;A common misconception: when GPU utilization is low, the first reaction is that the GPU isn&amp;rsquo;t fast enough. In reality, the GPU is often &lt;strong&gt;waiting for data&lt;/strong&gt; — a slow DataLoader, insufficient CPU, network congestion, or inadequate storage performance can all leave the GPU idle.&lt;/p&gt;
&lt;h2 id="l4-kubernetes-and-scheduling-layer"&gt;L4 Kubernetes and Scheduling Layer&lt;/h2&gt;
&lt;p&gt;This layer is the core of cloud-native AI infrastructure. The question to answer is: &lt;strong&gt;Who owns the GPU?&lt;/strong&gt; You need to know which Pod is using the GPU, which Namespace consumes the most resources, which team has the highest cost, and which model occupies the most VRAM.&lt;/p&gt;
&lt;p&gt;The table below lists the key observability dimensions at the scheduling layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Key Attributes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ownership&lt;/td&gt;
&lt;td&gt;Pod, Container&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workload&lt;/td&gt;
&lt;td&gt;Deployment, Job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Organization&lt;/td&gt;
&lt;td&gt;Namespace, Team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling&lt;/td&gt;
&lt;td&gt;GPU Sharing, GPU Partition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Topology&lt;/td&gt;
&lt;td&gt;Node, AZ, Region&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 4: L4 Kubernetes and Scheduling Layer Key Dimensions
&lt;/figcaption&gt;
&lt;p&gt;Projects like HAMi, Volcano, and Kueue primarily operate at this layer. HAMi solves core problems like GPU Sharing, GPU Partitioning, and heterogeneous GPU scheduling, while observability answers the question: &lt;strong&gt;Has scheduling actually improved resource utilization?&lt;/strong&gt; The observability data at this layer is the foundation for resource auditing and cost allocation.&lt;/p&gt;
&lt;h2 id="l5-training-runtime-layer"&gt;L5 Training Runtime Layer&lt;/h2&gt;
&lt;p&gt;The training phase requires monitoring the model&amp;rsquo;s runtime state. The table below lists the key metrics for the training runtime:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Efficiency&lt;/td&gt;
&lt;td&gt;MFU, TFLOPS, Step Time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gradient&lt;/td&gt;
&lt;td&gt;Norm, NaN, Inf&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Loss&lt;/td&gt;
&lt;td&gt;Training Loss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;DataLoader Wait&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpoint&lt;/td&gt;
&lt;td&gt;Save Time, Restore Time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: L5 Training Runtime Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;The most important metric here is &lt;strong&gt;MFU (Model FLOPs Utilization)&lt;/strong&gt;. Many training jobs show high GPU utilization but only 30% MFU, meaning a significant amount of GPU time is not being converted into actual training progress. MFU is the golden metric for measuring training efficiency — it directly reflects the ratio of hardware compute converted into effective training computation.&lt;/p&gt;
&lt;h2 id="l6-inference-engine-layer"&gt;L6 Inference Engine Layer&lt;/h2&gt;
&lt;p&gt;This is the most critical layer in production environments and the one seeing the fastest growth in industry attention. The table below lists the core metrics for inference engines:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;Time To First Token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ITL&lt;/td&gt;
&lt;td&gt;Inter Token Latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Queue Wait&lt;/td&gt;
&lt;td&gt;Request queuing time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;Token throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch Size&lt;/td&gt;
&lt;td&gt;Batch processing efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV Cache Usage&lt;/td&gt;
&lt;td&gt;KV Cache utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: L6 Inference Engine Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;Many teams still focus solely on GPU Utilization, but the true capacity metric for inference systems is often &lt;strong&gt;KV Cache (Key-Value Cache) Utilization&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="why-kv-cache-matters-more-than-gpu-utilization"&gt;Why KV Cache Matters More Than GPU Utilization&lt;/h3&gt;
&lt;p&gt;In modern LLM inference systems, GPU compute usually has headroom, but KV Cache often runs out first. A typical scenario: GPU utilization is only 60%, KV Cache is at 95%, new requests start queuing, and TTFT spikes rapidly.&lt;/p&gt;
&lt;p&gt;For inference systems, KV Cache is more akin to a database&amp;rsquo;s Buffer Pool — it often determines the capacity ceiling of the entire system. When KV Cache approaches saturation, the system is forced to perform Eviction, leading to context loss, request retries, and ultimately latency spikes and throughput drops.&lt;/p&gt;
&lt;h2 id="l7-genai-api-layer"&gt;L7 GenAI API Layer&lt;/h2&gt;
&lt;p&gt;With the development of OpenTelemetry GenAI Semantic Conventions, the industry is converging on unified AI observability standards. The table below lists the key metrics at the API layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input Tokens&lt;/td&gt;
&lt;td&gt;Input token count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output Tokens&lt;/td&gt;
&lt;td&gt;Output token count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request Duration&lt;/td&gt;
&lt;td&gt;Request latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;First Token latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token Throughput&lt;/td&gt;
&lt;td&gt;Token throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 7: L7 GenAI API Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;This layer marks the shift from infrastructure-centric to application-centric observability. API-layer metrics directly face end users and business systems, serving as the core data source for SLA/SLO quality assessment.&lt;/p&gt;
&lt;h2 id="l8-business-and-cost-layer"&gt;L8 Business and Cost Layer&lt;/h2&gt;
&lt;p&gt;Ultimately, enterprises don&amp;rsquo;t care about GPU utilization — they care about: &lt;strong&gt;Is the GPU creating value?&lt;/strong&gt; The table below lists the core metrics for the business and cost layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cost per GPU Hour&lt;/td&gt;
&lt;td&gt;GPU hourly cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per Token&lt;/td&gt;
&lt;td&gt;Token cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per Request&lt;/td&gt;
&lt;td&gt;Request cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idle GPU Cost&lt;/td&gt;
&lt;td&gt;Idle resource cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per Watt&lt;/td&gt;
&lt;td&gt;Inference efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 8: L8 Business and Cost Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;The core competitive metric for future AI infrastructure may not be GPU Utilization, but rather &lt;strong&gt;Cost per Useful Token&lt;/strong&gt; — the comprehensive cost of producing one useful Token. This metric unifies hardware cost, energy consumption, inference efficiency, and business value into a single measurement framework.&lt;/p&gt;
&lt;h2 id="cross-layer-troubleshooting"&gt;Cross-Layer Troubleshooting&lt;/h2&gt;
&lt;p&gt;Real production issues often span multiple layers. The table below lists common cross-layer failure symptoms and their potential causes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Possible Causes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT Increase&lt;/td&gt;
&lt;td&gt;CPU IO Wait, Queue Depth, KV Cache Pressure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput Drop&lt;/td&gt;
&lt;td&gt;NCCL, Batch Size, Network&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency Spike&lt;/td&gt;
&lt;td&gt;KV Cache Eviction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OOM&lt;/td&gt;
&lt;td&gt;Insufficient VRAM, Oversized Batch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training Stall&lt;/td&gt;
&lt;td&gt;NCCL Straggler&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 9: Cross-Layer Troubleshooting Reference
&lt;/figcaption&gt;
&lt;p&gt;Modern AI operations are evolving from single-point monitoring to cross-layer observability. The root cause of a single-layer metric anomaly often lies in another layer. Only by building cross-layer correlation capabilities can teams achieve rapid root-cause identification and precise troubleshooting.&lt;/p&gt;
&lt;h2 id="from-gpu-control-plane-to-ai-observability-plane"&gt;From GPU Control Plane to AI Observability Plane&lt;/h2&gt;
&lt;p&gt;Over the past few years, the industry has focused primarily on projects like Kubernetes Scheduler, Volcano, Kueue, and HAMi, solving the problem of &lt;strong&gt;how to allocate GPUs&lt;/strong&gt;. In the coming years, the industry will start asking &lt;strong&gt;whether GPUs are truly generating value&lt;/strong&gt;, leading to the formation of a new AI infrastructure technology stack:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Hardware
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Kubernetes Control Plane
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Inference Runtime
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Observability Plane
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Optimization Plane&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;GPU scheduling determines how resources are allocated. Observability determines whether resources are being used correctly. Together, they form the next-generation AI Infrastructure Stack. The evolution from GPU Control Plane to AI Observability Plane marks a new era where AI infrastructure transitions from &amp;ldquo;resource management&amp;rdquo; to &amp;ldquo;value management.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;AI infrastructure observability is undergoing a fundamental paradigm shift. We used to look only at GPU utilization; now we need to build a complete observability system across eight layers — from GPU hardware, CUDA runtime, host OS, Kubernetes scheduling, training runtime, inference engine, GenAI API, to business cost.&lt;/p&gt;
&lt;p&gt;Key takeaways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;L1-L3 focus on hardware and systems&lt;/strong&gt;: GPU health, CUDA communication efficiency, and host resource adequacy&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L4 focuses on resource allocation&lt;/strong&gt;: Kubernetes scheduling and GPU sharing, with tools like HAMi solving allocation problems&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L5-L6 focus on model runtime&lt;/strong&gt;: Training efficiency (MFU) and inference capacity (KV Cache) are the core metrics&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L7-L8 focus on business value&lt;/strong&gt;: From Token throughput to cost per Token, ultimately measuring whether GPUs are creating value&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;GPU scheduling is the starting point for AI infrastructure, but not the end goal. The 8-layer observability stack from GPU to Token is the critical closed loop that ensures AI infrastructure truly delivers business value.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;GPU &amp;amp; LLM Inference Observability — Layer-by-Layer Coverage - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>My Personal AI Stack: Building a Continuously Running Personal AI Infrastructure for ~$100/Month</title><link>https://jimmysong.io/blog/personal-ai-stack/</link><pubDate>Sun, 07 Jun 2026 16:58:18 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/personal-ai-stack/</guid><description>How I built a personal AI infrastructure using ChatGPT, OpenClaw, Obsidian, GitHub, Lark, GLM-5.1, and a Mac mini M4.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;A personal AI infrastructure is not a single tool — it&amp;rsquo;s a system of long-term synergy.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/banner.webp" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/banner.webp" alt="Figure 17: My Personal AI Stack" data-caption="Figure 17: My Personal AI Stack"
width="1774"
height="887"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 17: My Personal AI Stack&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Many people talk about AI Agents, Second Brain, Personal Knowledge Management (PKM), and digital avatars.&lt;/p&gt;
&lt;p&gt;But over the past year, I&amp;rsquo;ve come to realize that what I&amp;rsquo;m actually building is not some AI assistant — it&amp;rsquo;s a continuously running Personal AI Infrastructure.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not a single product, nor a single model. It&amp;rsquo;s a set of systems that work together over the long term.&lt;/p&gt;
&lt;p&gt;This system helps me think, research, write, code, manage knowledge, process emails, maintain my website, and accumulate long-term memory — every single day.&lt;/p&gt;
&lt;p&gt;If I had to summarize it in one sentence:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ChatGPT handles thinking, OpenClaw handles execution, Obsidian handles memory, GitHub handles publishing.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="from-tools-to-infrastructure"&gt;From Tools to Infrastructure&lt;/h2&gt;
&lt;p&gt;Most AI workflow articles follow a similar structure: start with the model, then the plugins, then the editor.&lt;/p&gt;
&lt;p&gt;But I increasingly feel that tools are not the point.&lt;/p&gt;
&lt;p&gt;What matters is how these tools work together.&lt;/p&gt;
&lt;p&gt;My work spans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI infrastructure research&lt;/li&gt;
&lt;li&gt;Open source community operations&lt;/li&gt;
&lt;li&gt;Technical writing&lt;/li&gt;
&lt;li&gt;Developer Relations&lt;/li&gt;
&lt;li&gt;Product and ecosystem building&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every day produces a massive amount of information:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ChatGPT conversations&lt;/li&gt;
&lt;li&gt;Technical research&lt;/li&gt;
&lt;li&gt;GitHub activity&lt;/li&gt;
&lt;li&gt;Community discussions&lt;/li&gt;
&lt;li&gt;Email newsletters&lt;/li&gt;
&lt;li&gt;Hacker News&lt;/li&gt;
&lt;li&gt;Discord&lt;/li&gt;
&lt;li&gt;WeChat and Lark messages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The problem is never a lack of information — it&amp;rsquo;s how to organize it.&lt;/p&gt;
&lt;p&gt;So I gradually built a Personal AI Stack around my own work.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Workflow Perspective
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This is not a single tool, not a single model. It emphasizes continuous synergy across four layers: thinking, memory, execution, and publishing.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="overall-architecture"&gt;Overall Architecture&lt;/h2&gt;
&lt;p&gt;The architecture diagram below shows how my Personal AI Stack works in layers.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/ddbe89d6ae15e708024e83d0dc1bd038.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/ddbe89d6ae15e708024e83d0dc1bd038.svg" alt="Figure 18: Personal AI Stack Architecture Layers" data-caption="Figure 18: Personal AI Stack Architecture Layers"
width="4610"
height="2138"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 18: Personal AI Stack Architecture Layers&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The core components for each layer are listed below.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Components&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interface Layer&lt;/td&gt;
&lt;td&gt;ChatGPT, Telegram, Discord, Lark, WeChat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning Layer&lt;/td&gt;
&lt;td&gt;ChatGPT, GLM-5.1, Claude Code, Codex&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory Layer&lt;/td&gt;
&lt;td&gt;Obsidian, Markdown, iCloud&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Execution Layer&lt;/td&gt;
&lt;td&gt;OpenClaw, Gmail, Calendar, Lark CLI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Publishing Layer&lt;/td&gt;
&lt;td&gt;GitHub, Hugo, Cloudflare Pages&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: Personal AI Stack Layers and Components
&lt;/figcaption&gt;
&lt;h2 id="chatgpt-my-thinking-system"&gt;ChatGPT: My Thinking System&lt;/h2&gt;
&lt;p&gt;Although many workflows revolve around Agents, the tool I use most frequently is actually ChatGPT.&lt;/p&gt;
&lt;p&gt;I primarily use it for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deep Research&lt;/li&gt;
&lt;li&gt;Technical analysis&lt;/li&gt;
&lt;li&gt;Architecture discussions&lt;/li&gt;
&lt;li&gt;Content planning&lt;/li&gt;
&lt;li&gt;Writing assistance&lt;/li&gt;
&lt;li&gt;Career decisions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Rather than calling it an assistant, it&amp;rsquo;s more like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Research Partner&lt;/li&gt;
&lt;li&gt;Technical Advisor&lt;/li&gt;
&lt;li&gt;Thinking Companion&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Years of accumulated conversations have helped it gradually understand my background, projects, and long-term goals.&lt;/p&gt;
&lt;p&gt;Many articles, talks, and technical judgments actually originate from these ongoing conversations.&lt;/p&gt;
&lt;h2 id="openclaw-my-execution-system"&gt;OpenClaw: My Execution System&lt;/h2&gt;
&lt;p&gt;OpenClaw is the OpenClaw Agent I deployed on my Mac mini M4 at home.&lt;/p&gt;
&lt;p&gt;I mainly interact with OpenClaw through Telegram.&lt;/p&gt;
&lt;p&gt;To avoid context mixing, I use separate Telegram groups for different topics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;HAMi&lt;/li&gt;
&lt;li&gt;AI Handbook&lt;/li&gt;
&lt;li&gt;Personal&lt;/li&gt;
&lt;li&gt;Work&lt;/li&gt;
&lt;li&gt;Research&lt;/li&gt;
&lt;li&gt;Blog&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This naturally creates Context Isolation.&lt;/p&gt;
&lt;p&gt;OpenClaw handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Gmail email management&lt;/li&gt;
&lt;li&gt;Apple Calendar scheduling&lt;/li&gt;
&lt;li&gt;Scheduled tasks&lt;/li&gt;
&lt;li&gt;Obsidian operations&lt;/li&gt;
&lt;li&gt;GitHub operations&lt;/li&gt;
&lt;li&gt;Website maintenance&lt;/li&gt;
&lt;li&gt;Lark knowledge base operations&lt;/li&gt;
&lt;li&gt;Automated workflows&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It&amp;rsquo;s more like a Chief of Staff than a chatbot.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
OpenClaw&amp;rsquo;s Positioning
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
OpenClaw&amp;rsquo;s value lies in unifying scattered execution entry points into structured workflows, rather than being a simple chat interface.
&lt;/div&gt;
&lt;/div&gt;
&lt;h3 id="openclaws-workflow"&gt;OpenClaw&amp;rsquo;s Workflow&lt;/h3&gt;
&lt;p&gt;OpenClaw&amp;rsquo;s main path unfolds through the Telegram entry point for multi-platform execution.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/c89afbc470d0c515198eb344d8dc1e4e.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/c89afbc470d0c515198eb344d8dc1e4e.svg" alt="Figure 19: OpenClaw Execution Path" data-caption="Figure 19: OpenClaw Execution Path"
width="3119"
height="1058"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 19: OpenClaw Execution Path&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="lark-company-workflow-entry-point"&gt;Lark: Company Workflow Entry Point&lt;/h2&gt;
&lt;p&gt;I had barely used Lark before.&lt;/p&gt;
&lt;p&gt;After joining my current company, which heavily promotes Lark adoption, all daily workflows, knowledge bases, and collaborative communication happen in Lark.&lt;/p&gt;
&lt;p&gt;At first, I just treated it as an enterprise IM tool.&lt;/p&gt;
&lt;p&gt;But looking at it now, Lark is more of a company-level workflow entry point:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Group chats&lt;/li&gt;
&lt;li&gt;Documents&lt;/li&gt;
&lt;li&gt;Knowledge bases&lt;/li&gt;
&lt;li&gt;Approvals&lt;/li&gt;
&lt;li&gt;Tasks&lt;/li&gt;
&lt;li&gt;Automation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For me, what truly changed the experience was Lark CLI.&lt;/p&gt;
&lt;p&gt;Through Lark CLI, I can more easily manage the company&amp;rsquo;s knowledge bases, documents, and some process-driven information.&lt;/p&gt;
&lt;p&gt;This turns Lark from merely a chat tool into something that OpenClaw can incorporate into its automation system.&lt;/p&gt;
&lt;p&gt;From the perspective of a Personal AI Stack, Lark serves as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The organizational workflow layer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It connects information, knowledge, and tasks in a company context.&lt;/p&gt;
&lt;h2 id="obsidian-working-memory"&gt;Obsidian: Working Memory&lt;/h2&gt;
&lt;p&gt;Obsidian is my most frequently used knowledge tool.&lt;/p&gt;
&lt;p&gt;But I don&amp;rsquo;t consider it my final knowledge base.&lt;/p&gt;
&lt;p&gt;For me:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Obsidian
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Working Memory&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This is where I store:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Daily Notes&lt;/li&gt;
&lt;li&gt;Weekly Reports&lt;/li&gt;
&lt;li&gt;Research Notes&lt;/li&gt;
&lt;li&gt;Inbox&lt;/li&gt;
&lt;li&gt;Drafts&lt;/li&gt;
&lt;li&gt;Fleeting thoughts&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Starting three years ago, I developed a habit of consistently writing weekly reports.&lt;/p&gt;
&lt;p&gt;These reports record:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Work progress&lt;/li&gt;
&lt;li&gt;Learning content&lt;/li&gt;
&lt;li&gt;Community activities&lt;/li&gt;
&lt;li&gt;Project evolution&lt;/li&gt;
&lt;li&gt;Personal reflections&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In a sense, weekly reports constitute my daily memory.&lt;/p&gt;
&lt;p&gt;And Obsidian is the carrier of these memories.&lt;/p&gt;
&lt;h2 id="github-and-jimmysongio-long-term-memory"&gt;GitHub and jimmysong.io: Long-term Memory&lt;/h2&gt;
&lt;p&gt;Many people think of a blog as a content publishing platform.&lt;/p&gt;
&lt;p&gt;But for me, jimmysong.io is closer to a long-term memory system.&lt;/p&gt;
&lt;p&gt;This website has been continuously maintained for nearly ten years.&lt;/p&gt;
&lt;p&gt;What&amp;rsquo;s recorded here is not just technical articles, but more importantly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;My viewpoints&lt;/li&gt;
&lt;li&gt;My judgments&lt;/li&gt;
&lt;li&gt;My experiences&lt;/li&gt;
&lt;li&gt;My growth trajectory&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unlike Obsidian, content that makes it to the website typically goes through:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Research
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Thinking
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Validation
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Writing
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Revision
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Publishing&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Therefore:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Obsidian
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Working Memory
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;jimmysong.io
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Long-term Memory&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The following flow shows the basic path from Obsidian to website publishing.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/e43d18197c002d3343a02ba47e100ece.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/e43d18197c002d3343a02ba47e100ece.svg" alt="Figure 20: Long-term Memory Publishing Flow" data-caption="Figure 20: Long-term Memory Publishing Flow"
width="2627"
height="172"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 20: Long-term Memory Publishing Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="information-flow"&gt;Information Flow&lt;/h2&gt;
&lt;p&gt;There are clear boundaries between content sources, curation methods, and output channels. This diagram reveals my information flow logic.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/769ffd8bfc324f83203fa12b51796292.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/769ffd8bfc324f83203fa12b51796292.svg" alt="Figure 21: Personal AI Stack Information Flow" data-caption="Figure 21: Personal AI Stack Information Flow"
width="2518"
height="1349"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 21: Personal AI Stack Information Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="github-execution-and-publishing-layer"&gt;GitHub: Execution and Publishing Layer&lt;/h2&gt;
&lt;p&gt;For me, GitHub is no longer just a code repository.&lt;/p&gt;
&lt;p&gt;It also handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Blog content&lt;/li&gt;
&lt;li&gt;Website source code&lt;/li&gt;
&lt;li&gt;Documentation system&lt;/li&gt;
&lt;li&gt;AI Handbook&lt;/li&gt;
&lt;li&gt;AI Native Landscape&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All content eventually makes its way into GitHub.&lt;/p&gt;
&lt;p&gt;GitHub Actions automatically handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Building&lt;/li&gt;
&lt;li&gt;Testing&lt;/li&gt;
&lt;li&gt;Publishing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Finally, everything is served through Cloudflare Pages.&lt;/p&gt;
&lt;p&gt;The diagram below shows the publishing path from Markdown to static site.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/e6cc6be64df6f178466a37b70068a3b5.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/e6cc6be64df6f178466a37b70068a3b5.svg" alt="Figure 22: GitHub Publishing Pipeline" data-caption="Figure 22: GitHub Publishing Pipeline"
width="2352"
height="141"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 22: GitHub Publishing Pipeline&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="my-development-toolchain"&gt;My Development Toolchain&lt;/h2&gt;
&lt;p&gt;I currently use three main AI development tools.&lt;/p&gt;
&lt;h3 id="claude-code"&gt;Claude Code&lt;/h3&gt;
&lt;p&gt;My daily development workhorse.&lt;/p&gt;
&lt;p&gt;Despite the name Claude Code, I primarily use Zhipu&amp;rsquo;s GLM-5.1 model.&lt;/p&gt;
&lt;p&gt;It handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Coding&lt;/li&gt;
&lt;li&gt;Refactoring&lt;/li&gt;
&lt;li&gt;Debugging&lt;/li&gt;
&lt;li&gt;Documentation maintenance&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="codex"&gt;Codex&lt;/h3&gt;
&lt;p&gt;Mainly used for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Project initialization&lt;/li&gt;
&lt;li&gt;Large-scale code generation&lt;/li&gt;
&lt;li&gt;Automated execution of complex tasks&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="openclaw"&gt;OpenClaw&lt;/h3&gt;
&lt;p&gt;Handles automation work beyond development.&lt;/p&gt;
&lt;p&gt;The three form a clear division of labor.&lt;/p&gt;
&lt;p&gt;The diagram below shows how my toolchain collaborates.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/1d9a2ef112b60d5c6d3014e75e449be3.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/1d9a2ef112b60d5c6d3014e75e449be3.svg" alt="Figure 23: Development Toolchain Collaboration" data-caption="Figure 23: Development Toolchain Collaboration"
width="654"
height="1342"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 23: Development Toolchain Collaboration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-not-claude"&gt;Why Not Claude?&lt;/h2&gt;
&lt;p&gt;This is one of the questions I get asked most often.&lt;/p&gt;
&lt;p&gt;Objectively speaking, Claude is indeed very strong in code generation and code understanding.&lt;/p&gt;
&lt;p&gt;I seriously considered using Claude as my primary model.&lt;/p&gt;
&lt;p&gt;But ultimately, I chose not to.&lt;/p&gt;
&lt;p&gt;The reason is not about model capability — it&amp;rsquo;s about overall return on investment.&lt;/p&gt;
&lt;p&gt;First, there&amp;rsquo;s the account issue.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve registered Claude accounts multiple times in the past, and each time I ran into restrictions and bans.&lt;/p&gt;
&lt;p&gt;Second, there&amp;rsquo;s the cost issue.&lt;/p&gt;
&lt;p&gt;Because I pay for all my AI tools out of pocket.&lt;/p&gt;
&lt;p&gt;So I care more about:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Capability / Cost&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Rather than:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Absolute Capability&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For enterprise users, Claude Max might be a very reasonable choice.&lt;/p&gt;
&lt;p&gt;But for individual paying users, the conclusion may be different.&lt;/p&gt;
&lt;p&gt;This article discusses a Personal AI Stack that is entirely self-funded.&lt;/p&gt;
&lt;p&gt;If the company reimburses expenses, or if you have an enterprise budget, many choices would change.&lt;/p&gt;
&lt;h2 id="why-not-self-host-large-models"&gt;Why Not Self-host Large Models?&lt;/h2&gt;
&lt;p&gt;This is another frequently asked question.&lt;/p&gt;
&lt;p&gt;Many people believe:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Buy a GPU + Open-source Model = Free AI&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In reality, this is often not the case.&lt;/p&gt;
&lt;p&gt;If the goal is simply to get a stable, powerful AI assistant, I lean toward:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Subscription &amp;gt; Self-hosting&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The reasons include:&lt;/p&gt;
&lt;h3 id="time-cost"&gt;Time Cost&lt;/h3&gt;
&lt;p&gt;Maintaining a model is work in itself.&lt;/p&gt;
&lt;p&gt;Including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CUDA&lt;/li&gt;
&lt;li&gt;Drivers&lt;/li&gt;
&lt;li&gt;Inference frameworks&lt;/li&gt;
&lt;li&gt;Model upgrades&lt;/li&gt;
&lt;li&gt;Networking issues&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All of these require time.&lt;/p&gt;
&lt;p&gt;And I&amp;rsquo;d rather spend my time on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Writing&lt;/li&gt;
&lt;li&gt;Community&lt;/li&gt;
&lt;li&gt;Product&lt;/li&gt;
&lt;li&gt;Technical research&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="cost-issue"&gt;Cost Issue&lt;/h3&gt;
&lt;p&gt;Currently:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ChatGPT Plus&lt;/li&gt;
&lt;li&gt;GLM Coding Plan Max&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total monthly cost is under 600 CNY.&lt;/p&gt;
&lt;p&gt;While a high-end GPU:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;RTX 5090D&lt;/li&gt;
&lt;li&gt;RTX PRO&lt;/li&gt;
&lt;li&gt;Enterprise GPU&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Typically costs tens of thousands of CNY.&lt;/p&gt;
&lt;p&gt;Plus:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity&lt;/li&gt;
&lt;li&gt;Depreciation&lt;/li&gt;
&lt;li&gt;Maintenance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For my use case, it&amp;rsquo;s simply not worth it.&lt;/p&gt;
&lt;h3 id="model-upgrade-speed"&gt;Model Upgrade Speed&lt;/h3&gt;
&lt;p&gt;Cloud models upgrade every month.&lt;/p&gt;
&lt;p&gt;Local models require manual follow-up.&lt;/p&gt;
&lt;p&gt;For knowledge workers:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Using the latest model is more important than owning a model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Subscription vs. Self-hosting
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
For individual users, subscription services typically offer advantages in time cost, upgrade speed, and stability. They let you focus on results rather than infrastructure maintenance.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="daily-tools"&gt;Daily Tools&lt;/h2&gt;
&lt;p&gt;Beyond the core systems described above, I also rely on some daily tools to complete the workflow.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Atlas Browser&lt;/strong&gt;: My primary browser for web reading, research, and information gathering. Noteworthy content is saved to my knowledge base via Obsidian Clipper.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Warp&lt;/strong&gt;: My most frequently used terminal tool. Its modern interactive experience and AI capabilities make command-line work more efficient.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Typora&lt;/strong&gt;: A Markdown editor I&amp;rsquo;ve used for a long time, ideal for immersive writing and long-form editing. Many blog posts and documents are completed here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CodexBar&lt;/strong&gt;: Used to monitor usage of ChatGPT, Codex, Claude Code, and other tools. For heavy AI users, token consumption has become a resource metric worth tracking.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sogou Input Method&lt;/strong&gt;: My primary voice input tool. Compared to keyboard input, voice better matches my thinking habits, especially when working remotely, writing, and communicating with AI.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These tools are not the core of the system themselves, but they form the foundational experience layer of the entire Personal AI Stack, making information acquisition, content creation, and daily development smoother.&lt;/p&gt;
&lt;h2 id="my-ai-usage-scale"&gt;My AI Usage Scale&lt;/h2&gt;
&lt;p&gt;Current approximate consumption:&lt;/p&gt;
&lt;h3 id="glm-51"&gt;GLM-5.1&lt;/h3&gt;
&lt;p&gt;Weekly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;600 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Monthly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2.4 billion Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="codex-1"&gt;Codex&lt;/h3&gt;
&lt;p&gt;Weekly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;200 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Monthly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;800 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Approximately 3.2 billion Tokens per month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That&amp;rsquo;s roughly 100 million Tokens consumed per day. Through ChatGPT Plus and GLM Coding Plan subscriptions, this is much more cost-effective than paying per token — otherwise, these tokens would cost at least $500 per month.&lt;/p&gt;
&lt;h2 id="how-much-does-this-system-cost-per-month"&gt;How Much Does This System Cost Per Month?&lt;/h2&gt;
&lt;p&gt;The table below compares the fixed monthly costs.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th style="text-align: right"&gt;Monthly Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT Plus&lt;/td&gt;
&lt;td style="text-align: right"&gt;$19.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM Coding Plan Max&lt;/td&gt;
&lt;td style="text-align: right"&gt;¥422.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iCloud+&lt;/td&gt;
&lt;td style="text-align: right"&gt;$0.99&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: Monthly Cost Breakdown
&lt;/figcaption&gt;
&lt;p&gt;Approximately:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥573 / month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;OpenClaw runs on my Mac mini M4.&lt;/p&gt;
&lt;p&gt;Hardware includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mac mini M4&lt;/li&gt;
&lt;li&gt;Samsung 990 Pro 1TB&lt;/li&gt;
&lt;li&gt;HAGIBIS Dock&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total investment:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥4,469&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Amortized over four years:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Approximately ¥100 / month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="total-cost"&gt;Total Cost&lt;/h3&gt;
&lt;p&gt;The entire Personal AI Stack&amp;rsquo;s fixed cost is approximately:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥700 CNY / month (~$100)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This cost pie chart shows the monthly spending structure of the Personal AI Stack.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/b67330204f6917024f61be43bba7004c.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/b67330204f6917024f61be43bba7004c.svg" alt="Figure 24: Personal AI Stack Monthly Cost Breakdown" data-caption="Figure 24: Personal AI Stack Monthly Cost Breakdown"
width="618"
height="384"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 24: Personal AI Stack Monthly Cost Breakdown&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="faq"&gt;FAQ&lt;/h2&gt;
&lt;h3 id="is-this-the-most-powerful-setup"&gt;Is This the Most Powerful Setup?&lt;/h3&gt;
&lt;p&gt;No.&lt;/p&gt;
&lt;p&gt;This is not a &amp;ldquo;most powerful AI tool configuration guide.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This is a completely self-funded Personal AI Stack designed for long-term individual use.&lt;/p&gt;
&lt;p&gt;If your company reimburses expenses, or if you have a higher budget, you could choose Claude Max, Cursor, more API services, or even a local GPU workstation.&lt;/p&gt;
&lt;p&gt;But my goal is not to pursue the absolute best — it&amp;rsquo;s to achieve stable, sustainable, and cumulative productivity within a personal budget.&lt;/p&gt;
&lt;h3 id="why-not-sync-all-chatgpt-conversations-to-obsidian"&gt;Why Not Sync All ChatGPT Conversations to Obsidian?&lt;/h3&gt;
&lt;p&gt;Because I don&amp;rsquo;t want to turn Obsidian into a chat log repository.&lt;/p&gt;
&lt;p&gt;What I care about more is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Which content is worth preserving long-term?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Manual saving is itself a curation process.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s more important than automatically syncing everything.&lt;/p&gt;
&lt;h3 id="why-not-use-notion"&gt;Why Not Use Notion?&lt;/h3&gt;
&lt;p&gt;It&amp;rsquo;s not because Notion isn&amp;rsquo;t good.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s because I prefer Markdown First.&lt;/p&gt;
&lt;p&gt;The benefits of Markdown include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Local-first&lt;/li&gt;
&lt;li&gt;Version controllable&lt;/li&gt;
&lt;li&gt;Portable&lt;/li&gt;
&lt;li&gt;AI-friendly&lt;/li&gt;
&lt;li&gt;Suitable for long-term preservation&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="why-use-telegram-only-for-openclaw"&gt;Why Use Telegram Only for OpenClaw?&lt;/h3&gt;
&lt;p&gt;Because Telegram&amp;rsquo;s Bot ecosystem and API are better suited as an Agent entry point.&lt;/p&gt;
&lt;p&gt;WeChat, Lark, and Discord are more for person-to-person communication and community collaboration for me.&lt;/p&gt;
&lt;h3 id="why-emphasize-lark"&gt;Why Emphasize Lark?&lt;/h3&gt;
&lt;p&gt;Because after joining my current company, I truly started using Lark.&lt;/p&gt;
&lt;p&gt;The company&amp;rsquo;s entire workflow revolves around Lark.&lt;/p&gt;
&lt;p&gt;For me, Lark isn&amp;rsquo;t a personal knowledge base — it&amp;rsquo;s a company knowledge and collaboration system.&lt;/p&gt;
&lt;p&gt;Through Lark CLI, it can further become a workflow entry point that OpenClaw can operate.&lt;/p&gt;
&lt;h3 id="why-not-use-a-local-ai-workstation"&gt;Why Not Use a Local AI Workstation?&lt;/h3&gt;
&lt;p&gt;Because my main work is not training models or running inference services.&lt;/p&gt;
&lt;p&gt;My main work is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Thinking&lt;/li&gt;
&lt;li&gt;Research&lt;/li&gt;
&lt;li&gt;Writing&lt;/li&gt;
&lt;li&gt;Open source community&lt;/li&gt;
&lt;li&gt;Software development&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Subscription models + a Mac mini is more than enough.&lt;/p&gt;
&lt;p&gt;A local AI workstation requires significant investment, complex maintenance, and the model upgrade speed may not keep up with cloud services.&lt;/p&gt;
&lt;h3 id="why-not-buy-a-high-end-nvidia-gpu"&gt;Why Not Buy a High-end NVIDIA GPU?&lt;/h3&gt;
&lt;p&gt;If the primary purpose is learning CUDA, GPU scheduling, or AI infrastructure, buying a GPU has value.&lt;/p&gt;
&lt;p&gt;But if the primary purpose is daily productivity, a high-end GPU isn&amp;rsquo;t necessarily cost-effective.&lt;/p&gt;
&lt;p&gt;For me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Subscribing to models is more important than owning a GPU.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="workflow-matters-more-than-models"&gt;Workflow Matters More Than Models&lt;/h2&gt;
&lt;p&gt;My biggest takeaway from the past year is:&lt;/p&gt;
&lt;p&gt;Many people believe the core of AI is the model.&lt;/p&gt;
&lt;p&gt;But my experience suggests the opposite.&lt;/p&gt;
&lt;p&gt;What truly affects productivity is often not the model leaderboard, but workflow design.&lt;/p&gt;
&lt;p&gt;A 95-score model in an excellent workflow is usually more valuable than a 100-score model in a chaotic workflow.&lt;/p&gt;
&lt;p&gt;In a sense, this is very similar to the evolution of the cloud-native world.&lt;/p&gt;
&lt;p&gt;GPUs are important, but scheduling systems are equally important.&lt;/p&gt;
&lt;p&gt;Models are important, but workflows are equally important.&lt;/p&gt;
&lt;p&gt;Agents are important, but long-term memory and execution systems are equally important.&lt;/p&gt;
&lt;p&gt;For me, the ultimate goal of the Personal AI Stack is not to replace people.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s to connect thinking, memory, and execution — freeing up more time for what truly matters.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This article defines my Personal AI Stack as a long-term runnable infrastructure, rather than a single tool or a single model.&lt;/p&gt;
&lt;p&gt;The core conclusion is: in a self-funded scenario, stable workflows, continuous memory systems, and executable automation are more productive than chasing the most powerful model.&lt;/p&gt;
&lt;p&gt;If your goal is to accumulate long-term value, investing time in &amp;ldquo;how things work together&amp;rdquo; often yields better returns than investing time in &amp;ldquo;model rankings.&amp;rdquo;&lt;/p&gt;</content:encoded></item><item><title>Token Is More Than a Billing Unit, It's Becoming the Resource Unit of the AI Era</title><link>https://jimmysong.io/blog/tokenomics-foundation-new-resource-unit/</link><pubDate>Thu, 04 Jun 2026 06:42:40 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/tokenomics-foundation-new-resource-unit/</guid><description>The Linux Foundation&amp;#39;s Tokenomics Foundation signals a shift: tokens are becoming a core resource in the AI era, much like CPUs in the cloud era.</description><content:encoded>
&lt;p&gt;The &lt;a href="https://www.linuxfoundation.org/" target="_blank" rel="noopener"&gt;Linux Foundation&lt;/a&gt; recently announced plans to establish the &lt;a href="https://www.tokeneconomics.com/insights/launch-tokenomics-foundation/" target="_blank" rel="noopener"&gt;Tokenomics Foundation&lt;/a&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/tokenomics-foundation-new-resource-unit/banner.webp" data-img="https://assets.jimmysong.io/images/blog/tokenomics-foundation-new-resource-unit/banner.webp" alt="Figure 1: Tokenomics Foundation" data-caption="Figure 1: Tokenomics Foundation"
width="1774"
height="887"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Tokenomics Foundation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;When I first saw the news, my immediate reaction wasn&amp;rsquo;t:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Yet another foundation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI infrastructure is shifting from &amp;ldquo;managing GPUs&amp;rdquo; to &amp;ldquo;managing Tokens.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That shift may matter more than the foundation itself.&lt;/p&gt;
&lt;p&gt;The official rationale from the Linux Foundation is straightforward:&lt;/p&gt;
&lt;p&gt;As enterprises begin deploying generative AI and agents at scale, the Token has become the new unit of technology spend. The foundation will partner with the &lt;a href="https://www.finops.org/" target="_blank" rel="noopener"&gt;FinOps Foundation&lt;/a&gt; to establish Token cost management, benchmarks, open standards, and best practices.&lt;/p&gt;
&lt;p&gt;If you stop there, most people would reduce it to:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;FinOps for AI&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Or:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI cost management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;But I think there&amp;rsquo;s more to it than that.&lt;/p&gt;
&lt;h2 id="the-cloud-era-managed-resources"&gt;The Cloud Era Managed Resources&lt;/h2&gt;
&lt;p&gt;For the past two decades, the infrastructure industry has been managing resources.&lt;/p&gt;
&lt;p&gt;We discussed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CPU&lt;/li&gt;
&lt;li&gt;Memory&lt;/li&gt;
&lt;li&gt;Storage&lt;/li&gt;
&lt;li&gt;Network&lt;/li&gt;
&lt;li&gt;GPU&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://kubernetes.io/" target="_blank" rel="noopener"&gt;Kubernetes&lt;/a&gt; was no different.&lt;/p&gt;
&lt;p&gt;Whether it was the scheduler, autoscaling, or resource quotas, everything fundamentally answered one question:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How do we allocate resources?&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That was the central question of the cloud era.&lt;/p&gt;
&lt;h2 id="the-ai-era-starts-managing-outcomes"&gt;The AI Era Starts Managing Outcomes&lt;/h2&gt;
&lt;p&gt;With the rise of AI, an interesting shift began.&lt;/p&gt;
&lt;p&gt;Enterprises increasingly don&amp;rsquo;t care about:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How many GPUs were used&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;They care far more about:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How many Tokens were produced&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;For most enterprises:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPU is cost.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Token is output.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;GPU, CPU, and memory are merely means of production.&lt;/li&gt;
&lt;li&gt;Token is the final deliverable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/jrstorment_tokenomics-finops-share-7467935235771367424-XEy5/" target="_blank" rel="noopener"&gt;J.R. Storment&lt;/a&gt; (former Executive Director of the FinOps Foundation) said something on LinkedIn that stuck with me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Tokens are what all the hardware is being built to produce.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;All hardware is ultimately producing Tokens.&lt;/p&gt;
&lt;p&gt;From this perspective:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Inference
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;A new value chain is forming.&lt;/p&gt;
&lt;h2 id="token-could-become-the-new-resource-model"&gt;Token Could Become the New Resource Model&lt;/h2&gt;
&lt;p&gt;Over the past few years, I&amp;rsquo;ve been tracking &lt;a href="https://jimmysong.io/blog/ai-inference-on-kubernetes/"&gt;how Kubernetes is evolving in the AI era&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;More and more signals suggest:&lt;/p&gt;
&lt;p&gt;Kubernetes is evolving from a &lt;strong&gt;Compute Control Plane&lt;/strong&gt; into an &lt;strong&gt;AI Control Plane&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Once AI becomes the dominant workload, the resources we manage may no longer be just:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;cpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;8&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;32Gi&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;gpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;They may gradually become:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;token
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;context
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;latency
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;throughput&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Metrics much closer to business value.&lt;/p&gt;
&lt;p&gt;The most interesting thing about the Tokenomics Foundation isn&amp;rsquo;t whether it will define new standards.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s that it implicitly acknowledges one thing:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token has started to become a resource.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Just like CPU in the cloud era.&lt;/p&gt;
&lt;h2 id="implications-for-ai-infrastructure"&gt;Implications for AI Infrastructure&lt;/h2&gt;
&lt;p&gt;For teams building AI infrastructure, this shift deserves serious thought.&lt;/p&gt;
&lt;p&gt;Today, many projects (including &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;) focus on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPU utilization&lt;/li&gt;
&lt;li&gt;GPU sharing&lt;/li&gt;
&lt;li&gt;GPU scheduling&lt;/li&gt;
&lt;li&gt;GPU virtualization&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are all important.&lt;/p&gt;
&lt;p&gt;But from an enterprise perspective, they ultimately don&amp;rsquo;t buy GPU utilization.&lt;/p&gt;
&lt;p&gt;They buy:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Cost Per Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Or even further:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Value Per Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Given the same pool of GPUs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whoever can produce more Tokens,&lt;/li&gt;
&lt;li&gt;Whoever can reduce the cost per million Tokens,&lt;/li&gt;
&lt;li&gt;Whoever can increase the business value of each Token,&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Will capture greater commercial value.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why I believe the significance of the Tokenomics Foundation lies beyond the standards themselves.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s driving the entire industry from:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Resource management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Toward:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Value management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="my-take"&gt;My Take&lt;/h2&gt;
&lt;p&gt;I don&amp;rsquo;t think the Tokenomics Foundation will become the next Kubernetes.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s more like the &lt;a href="https://www.finops.org/" target="_blank" rel="noopener"&gt;FinOps Foundation&lt;/a&gt;, &lt;a href="https://www.opencost.io/" target="_blank" rel="noopener"&gt;OpenCost&lt;/a&gt;, or &lt;a href="https://opentelemetry.io/" target="_blank" rel="noopener"&gt;OpenTelemetry&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;What it&amp;rsquo;s trying to define isn&amp;rsquo;t software.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s the &lt;strong&gt;metering system for the AI era&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The past decade:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;CPU was the language of infrastructure&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The next decade:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token may become the language of AI infrastructure&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If this trend holds, then the GPUs, inference frameworks, scheduling systems, and Agent Runtimes we discuss today are ultimately just parts of a larger Token economy.&lt;/p&gt;
&lt;p&gt;And that may be the real signal the Linux Foundation is sending by launching the Tokenomics Foundation.&lt;/p&gt;</content:encoded></item><item><title>AI Native Landscape Launches as a Standalone Site</title><link>https://jimmysong.io/blog/ai-native-landscape-launch/</link><pubDate>Thu, 04 Jun 2026 06:12:10 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-native-landscape-launch/</guid><description>AI Native Landscape has moved to landscape.jimmysong.io with 600+ curated open-source projects, AI skill search support, and a call for community contributions.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;AI Native Landscape now has its own home. 600+ curated open-source projects, a brand-new standalone site, and direct search from your AI coding tools.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="how-it-started"&gt;How It Started&lt;/h2&gt;
&lt;p&gt;Last July, I started collecting and organizing AI open-source projects on my personal website. The AI space was moving incredibly fast, with new projects popping up every day. GitHub is a treasure trove, full of tools and frameworks worth studying and using. I wanted to organize them systematically so others (and myself) could find the truly useful ones.&lt;/p&gt;
&lt;p&gt;At first, everything lived on jimmysong.io. But as the catalog grew, maintaining it became a pain. The main site already had plenty of content, and AI-related pages mixed in meant that updating a single project required rebuilding the entire site. So this year I decided to spin it out as a standalone open-source project: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;ai-native-landscape&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;All pages previously at &lt;code&gt;jimmysong.io/en/ai/*&lt;/code&gt; now redirect to the new site at &lt;a href="https://landscape.jimmysong.io" target="_blank" rel="noopener"&gt;landscape.jimmysong.io&lt;/a&gt;. Existing links and bookmarks still work.&lt;/p&gt;
&lt;h2 id="not-just-a-list-but-a-scored-one"&gt;Not Just a List, But a Scored One&lt;/h2&gt;
&lt;p&gt;There are plenty of AI project directories out there, but most just list a name and a link and call it a day. That&amp;rsquo;s not enough.&lt;/p&gt;
&lt;p&gt;GitHub has thousands upon thousands of projects. You see a name, click through, and find out the last commit was six months ago and nobody&amp;rsquo;s responding to Issues. You just spent time researching something that&amp;rsquo;s essentially abandoned.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why AI Native Landscape does something different: &lt;strong&gt;every project gets a score&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The scoring is based on four dimensions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Activity&lt;/strong&gt;: commit frequency, release cadence, Issue response time&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community&lt;/strong&gt;: contributor count, PR merge speed, discussion engagement&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality&lt;/strong&gt;: Stars, Forks, dependency relationships&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sustainability&lt;/strong&gt;: maintenance history, team stability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These four dimensions combine into an overall health score. At a glance, you can tell whether a project is worth your time.&lt;/p&gt;
&lt;p&gt;And the scores &lt;strong&gt;update daily&lt;/strong&gt;, automatically. Today&amp;rsquo;s score might differ from tomorrow&amp;rsquo;s.&lt;/p&gt;
&lt;p&gt;The project list itself is also &lt;strong&gt;human-curated&lt;/strong&gt;. Projects that have gone inactive don&amp;rsquo;t make the cut. I want every project on this list to be &amp;ldquo;alive,&amp;rdquo; so the time you spend studying them won&amp;rsquo;t be wasted.&lt;/p&gt;
&lt;h2 id="whats-covered"&gt;What&amp;rsquo;s Covered&lt;/h2&gt;
&lt;p&gt;The catalog spans eight major areas of AI:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agent Systems&lt;/td&gt;
&lt;td&gt;Frameworks, orchestration, and workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge &amp;amp; Context&lt;/td&gt;
&lt;td&gt;RAG, vector databases, document processing, knowledge graphs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Models &amp;amp; Modalities&lt;/td&gt;
&lt;td&gt;Foundation models, toolkits, speech and vision generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference &amp;amp; Runtime&lt;/td&gt;
&lt;td&gt;Model serving, inference engines, sandboxes, GPU acceleration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training, Evaluation &amp;amp; Optimization&lt;/td&gt;
&lt;td&gt;Frameworks, fine-tuning, benchmarks, observability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Tooling&lt;/td&gt;
&lt;td&gt;MCP protocols, coding agents, IDE tools, SDKs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Applications &amp;amp; Experience&lt;/td&gt;
&lt;td&gt;Chat interfaces, workflow automation, low-code platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform &amp;amp; Infrastructure&lt;/td&gt;
&lt;td&gt;Cloud-native AI, data platforms, security and operations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Currently &lt;strong&gt;600+&lt;/strong&gt; open-source projects, each with bilingual descriptions in English and Chinese.&lt;/p&gt;
&lt;h2 id="ai-skill-search"&gt;AI Skill Search&lt;/h2&gt;
&lt;p&gt;The landscape supports direct search from AI coding tools. No browser needed. Install with one command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npx landscape-search&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Works with Claude Code, Cursor, GitHub Copilot, OpenAI Codex, Cline, Aider, and other popular AI coding tools. After installation, just search in natural language:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Find me an MCP-compatible agent framework&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;No API key required. No local data files. Your AI tool fetches the live project catalog and returns ranked results.&lt;/p&gt;
&lt;h2 id="tech-stack"&gt;Tech Stack&lt;/h2&gt;
&lt;p&gt;Built with &lt;a href="https://astro.build" target="_blank" rel="noopener"&gt;Astro&lt;/a&gt;, deployed on &lt;a href="https://pages.cloudflare.com" target="_blank" rel="noopener"&gt;Cloudflare Pages&lt;/a&gt;, with &lt;a href="https://www.typescriptlang.org/" target="_blank" rel="noopener"&gt;TypeScript&lt;/a&gt; for type safety. Project data lives in bilingual Markdown files. The build pipeline validates, indexes, and generates OG images in one pass.&lt;/p&gt;
&lt;h2 id="get-involved"&gt;Get Involved&lt;/h2&gt;
&lt;p&gt;GitHub is everyone&amp;rsquo;s treasure trove, and I hope this project can grow with the community&amp;rsquo;s help.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Star the repo&lt;/strong&gt;: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;github.com/rootsongjc/ai-native-landscape&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Submit a project&lt;/strong&gt;: If you know an AI open-source project worth listing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/rootsongjc/ai-native-landscape/issues/new?template=add-project.md" target="_blank" rel="noopener"&gt;Open an Issue&lt;/a&gt; on GitHub&lt;/li&gt;
&lt;li&gt;Or submit a PR directly, add bilingual Markdown files under &lt;code&gt;data/projects/&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;See the &lt;a href="https://github.com/rootsongjc/ai-native-landscape/blob/main/docs/contributing.md" target="_blank" rel="noopener"&gt;contributing guide&lt;/a&gt; for details.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Live site: &lt;a href="https://landscape.jimmysong.io" target="_blank" rel="noopener"&gt;landscape.jimmysong.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;GitHub repo: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;rootsongjc/ai-native-landscape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Skill install: &lt;code&gt;npx landscape-search&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>AI Infra Industry Trends: From Compute Bottlenecks to Ecosystem Evolution</title><link>https://jimmysong.io/slide/ai-infra-trend-2026-slide/</link><pubDate>Sun, 31 May 2026 01:18:27 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/slide/ai-infra-trend-2026-slide/</guid><description>A practitioner&amp;#39;s perspective on AI infrastructure trends: evolving bottlenecks, roles of CPU/GPU/scheduling, ecosystem shifts, and compute demand across training, inference, and Agent workloads.</description><content:encoded>
&lt;p&gt;This slide deck offers a practitioner&amp;rsquo;s perspective on the key trends in AI Infrastructure: the core bottleneck has shifted from &amp;ldquo;compute scarcity&amp;rdquo; to &amp;ldquo;efficiency governance&amp;rdquo;, GPU virtualization and intelligent scheduling deliver the highest ROI, the cloud-native ecosystem is upgrading from &amp;ldquo;application scheduling&amp;rdquo; to &amp;ldquo;compute scheduling&amp;rdquo;, and Agent workloads will drive a new scheduling paradigm.&lt;/p&gt;
&lt;p&gt;Use the controls or keyboard shortcuts below to navigate the embedded interactive slides.&lt;/p&gt;
&lt;div class="slide-embed-container" data-height="auto" data-style=""&gt;
&lt;iframe
loading="lazy"
src="https://jimmysong.io/slides/ai-infra-trend-2026/index-en.html"
allowfullscreen="allowfullscreen"
allow="fullscreen"
title="Slide Presentation"
class="slide-embed-iframe"&gt;
&lt;/iframe&gt;
&lt;/div&gt;
&lt;figcaption&gt;Slide: AI Infra Industry Trends&lt;/figcaption&gt;
&lt;h2 id="key-topics"&gt;Key Topics&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Evolving AI Infra Bottlenecks&lt;/strong&gt; — From compute scarcity to efficiency governance&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real Roles of CPU / GPU / Scheduling&lt;/strong&gt; — How each layer functions in production&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cloud Native × Open-Source Scheduling&lt;/strong&gt; — The shift from Cloud Native to AI Native&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Training / Inference / Agent Compute Demand&lt;/strong&gt; — Resource characteristics and trend analysis by scenario&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="about-the-speaker"&gt;About the Speaker&lt;/h2&gt;
&lt;p&gt;Jimmy Song, CNCF Ambassador and founder of the Cloud Native Community (China). Long-time practitioner in cloud-native infrastructure and AI Infrastructure, focusing on GPU compute scheduling, virtualization, and open-source ecosystem development.&lt;/p&gt;</content:encoded></item><item><title>Kubernetes as the GPU Control Plane: HAMi v2.9 and Next-Gen AI Infra</title><link>https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/</link><pubDate>Thu, 14 May 2026 06:34:19 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/</guid><description>Observations on the evolution of AI infrastructure control planes, focusing on HAMi v2.9, GPU scheduling, and Kubernetes resource models.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Recently, I&amp;rsquo;ve been following the progress of domestic GPU scheduling and Kubernetes AI resource models. With the release of &lt;a href="https://project-hami.io/blog/hami-v2-9-0-release" target="_blank" rel="noopener"&gt;HAMi v2.9&lt;/a&gt;, I want to share several observations on how the AI Infra control plane is evolving.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/banner.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/banner.webp" alt="Figure 1: Kubernetes as the GPU Control Plane for AI" data-caption="Figure 1: Kubernetes as the GPU Control Plane for AI"
width="1983"
height="793"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Kubernetes as the GPU Control Plane for AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-discuss-this-topic-now"&gt;Why Discuss This Topic Now&lt;/h2&gt;
&lt;p&gt;When DeepSeek R1 was released in early 2025, most people focused on the fact that it trained a model competitive with OpenAI o1 for just $5.6 million. What struck me, however, was that as inference costs plummeted, GPU utilization issues would quickly come to the forefront.&lt;/p&gt;
&lt;p&gt;As models became more useful and inference demand exploded, &amp;ldquo;one GPU per model&amp;rdquo; rapidly became a luxury. Meanwhile, the NVIDIA H200 export saga accelerated the adoption of domestic compute. First, a sales ban; then, at the end of 2025, a 25% tariff under Trump; and by January 2026, Chinese customs had cleared zero units. Policy now mandates that over 40% of data center chips must be domestically produced by 2026.&lt;/p&gt;
&lt;p&gt;The reality is harsh: not only are GPUs scarce, but you must also learn to use NVIDIA, Ascend, Cambricon, Hygon, and other very different platforms simultaneously.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why I believe HAMi v2.9 is more significant than it appears on the surface.&lt;/p&gt;
&lt;h2 id="gpus-are-no-longer-just-about-the-card"&gt;GPUs Are No Longer Just About the Card&lt;/h2&gt;
&lt;p&gt;Kubernetes has always managed GPUs in a rather crude way:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;resources&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;nvidia.com/gpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This was sufficient in 2019, when the main question was simply whether a Pod needed a GPU or not.&lt;/p&gt;
&lt;p&gt;But that&amp;rsquo;s no longer enough. An inference service might only need 4GB of VRAM, multiple small models can share a single card, training jobs care about GPU topology and interconnect bandwidth, and multi-tenancy requires fault domain isolation. Treating GPUs as integer resources is like using an abacus for statistical analysis—not impossible, but the mental model is out of sync with reality.&lt;/p&gt;
&lt;p&gt;The most notable feature in HAMi v2.9 is the HAMi-core mode for the Ascend 910C. Previously, sharing Ascend cards relied on SR-IOV hardware virtualization, which was coarse-grained and inflexible. HAMi-core takes a different approach: it uses &lt;code&gt;LD_PRELOAD&lt;/code&gt; to intercept ACL calls in user space, enabling memory isolation at the MB level and compute throttling by percentage.&lt;/p&gt;
&lt;p&gt;In short: it&amp;rsquo;s managed by software, not hardware slicing.&lt;/p&gt;
&lt;p&gt;This is reminiscent of how SDN abstracted the network control plane from hardware devices—GPU partitioning is shifting from a hardware capability to a cluster control plane capability. Considering that Huawei shipped 810,000 Ascend 910C cards last year—nearly half of all domestic chips—this capability has significant real-world impact.&lt;/p&gt;
&lt;h2 id="dra-kubernetes-finally-has-a-robust-device-resource-model"&gt;DRA: Kubernetes Finally Has a Robust Device Resource Model&lt;/h2&gt;
&lt;p&gt;Kubernetes v1.34 (September 2025) officially promoted DRA (Dynamic Resource Allocation) to GA, and Red Hat OpenShift 4.21 followed suit. This is a big deal.&lt;/p&gt;
&lt;p&gt;The Device Plugin solved &amp;ldquo;how to connect GPUs to K8s,&amp;rdquo; but not &amp;ldquo;how to express complex AI resource requirements.&amp;rdquo; Device Plugins only know how many cards are on a node, not how much VRAM you need, what topology, or what isolation level.&lt;/p&gt;
&lt;p&gt;DRA standardizes device resource declaration, allocation, and management via &lt;code&gt;ResourceClaim&lt;/code&gt; and &lt;code&gt;DeviceClass&lt;/code&gt;. HAMi-DRA takes a pragmatic approach: it doesn&amp;rsquo;t require users to change how they declare resources. Instead, it uses a Mutating Webhook to automatically convert existing Device Plugin-style declarations into the DRA model. Legacy systems don&amp;rsquo;t need to change, but can still leverage new capabilities.&lt;/p&gt;
&lt;p&gt;I liken this to what CSI did for storage: it didn&amp;rsquo;t eliminate vendor differences, but allowed Kubernetes to consume different storage capabilities in a unified way. DRA does the same for AI accelerators—NVIDIA, Ascend, AMD, Vastai cards will never be identical, but the scheduling layer should speak a common language.&lt;/p&gt;
&lt;h2 id="a-complete-control-plane-path"&gt;A Complete Control Plane Path&lt;/h2&gt;
&lt;p&gt;If we look beyond individual features and consider HAMi-core, DRA, CDI, and the scheduler together, they actually correspond to different layers of GPU resource management:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi-core&lt;/strong&gt;: How to partition and isolate devices internally&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DRA&lt;/strong&gt;: How to declare, allocate, and bind resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CDI&lt;/strong&gt;: How to standardize device injection into container runtimes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduler/Webhook&lt;/strong&gt;: How to schedule, admit, and observe&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Connecting these layers, from top to bottom, forms the complete Kubernetes GPU Control Plane:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/kubernetes-gpu-control-plane-en.svg" data-img="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/kubernetes-gpu-control-plane-en.svg" alt="Figure 2: Kubernetes as the GPU Control Plane for AI Workloads" data-caption="Figure 2: Kubernetes as the GPU Control Plane for AI Workloads"
width="1296"
height="1302"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Kubernetes as the GPU Control Plane for AI Workloads&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is a complete control plane path. In v2.9, Volcano vGPU was upgraded to v0.19 with enhanced CDI support. While this may seem like a minor improvement in device injection, it actually completes a critical link in this chain.&lt;/p&gt;
&lt;h2 id="heterogeneity-is-the-main-battlefield-for-domestic-ai-infra"&gt;Heterogeneity Is the Main Battlefield for Domestic AI Infra&lt;/h2&gt;
&lt;p&gt;The reality for domestic AI clusters: you can&amp;rsquo;t build infrastructure around just one type of GPU.&lt;/p&gt;
&lt;p&gt;Enterprise environments often have NVIDIA, Ascend, Biren, Cambricon, Hygon, Muxi, Kunlunxin, Vastai, and other devices coexisting. Each card has different drivers, runtimes, virtualization capabilities, and monitoring methods. HAMi v2.9 adds support for Vastai, covering more than ten types of heterogeneous compute devices. Mixed training and inference, online and offline workloads, domestic and overseas GPUs, multi-team and multi-tenant resource pools—in these scenarios, unified scheduling is far more important than single-card performance.&lt;/p&gt;
&lt;h2 id="key-judgments"&gt;Key Judgments&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;GPU sharing will shift from a cost-saving measure to a default requirement.&lt;/strong&gt; After the explosion of inference workloads, not every workload deserves exclusive access to an entire card. Exclusive allocation will increasingly become a luxury.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DRA is the way forward, but migration will be gradual.&lt;/strong&gt; The Device Plugin ecosystem is too large to disappear overnight. HAMi-DRA&amp;rsquo;s compatibility layer shows the project team understands this.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous scheduling will become the core challenge for AI Infra.&lt;/strong&gt; Whoever can abstract different vendor devices into a unified scheduling language will control the key position in the control plane.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes will be reshaped by AI workloads.&lt;/strong&gt; From scheduling semantics to resource models, AI requires much greater expressiveness than traditional web services. DRA, CDI, topology-aware scheduling—these are not isolated evolutions, but all point to one thing: Kubernetes is evolving from a container orchestrator to the control plane for AI computing.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The significance of HAMi v2.9 is not just in supporting a particular device or partitioning method, but in making one thing clear: the next generation of AI infrastructure competition is not just about model frameworks or GPU counts, but about the &lt;strong&gt;control plane&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;GPUs are shifting from external devices on nodes to native resources within the Kubernetes control plane. Whoever defines the resource model for the AI era will define the long-term boundaries of AI Infra.&lt;/p&gt;</content:encoded></item><item><title>Kubernetes's Anxiety and Rebirth in the AI Wave</title><link>https://jimmysong.io/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/</link><pubDate>Fri, 03 Apr 2026 05:20:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/</guid><description>At KubeCon EU 2026, I witnessed Kubernetes&amp;#39; anxiety and transformation in the AI era. This article explores the challenges and future opportunities for Kubernetes in the age of AI.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Kubernetes hasn&amp;rsquo;t been replaced by AI, but it&amp;rsquo;s being redefined by it. Anxiety is the prelude to rebirth.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After attending KubeCon EU 2026 in Amsterdam, I&amp;rsquo;ve been pondering a key question: Kubernetes isn&amp;rsquo;t obsolete, but it&amp;rsquo;s no longer &amp;ldquo;enough&amp;rdquo;; it hasn&amp;rsquo;t been replaced by AI, but it&amp;rsquo;s being redefined by AI.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/keep-cloud-native-moving.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/keep-cloud-native-moving.webp" alt="Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date." data-caption="Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date."
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This was my third time attending KubeCon in Europe. Over the past few years, you can actually see the community&amp;rsquo;s mindset shift through the event slogans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;2024 Paris: &lt;strong&gt;La vie en Cloud Native&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ Cloud Native has become a &amp;ldquo;way of life,&amp;rdquo; the default state&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;2025 London: &lt;strong&gt;No slogan, just the 10th anniversary&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ Kubernetes reached a milestone, focusing on retrospection rather than moving forward&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;2026 Amsterdam: &lt;strong&gt;Keep Cloud Native Moving&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ But the question is: where is it moving?&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The absence of a slogan in 2025 was a signal in itself:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;When an ecosystem starts commemorating the past instead of defining the future, it&amp;rsquo;s already at an inflection point.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This article doesn&amp;rsquo;t recap the talks, but instead distills my observations at KubeCon into insights about Kubernetes&amp;rsquo; anxiety and rebirth in the AI wave.&lt;/p&gt;
&lt;h2 id="the-root-of-anxiety-is-kubernetes-facing-a-crisis"&gt;The Root of Anxiety: Is Kubernetes Facing a &amp;ldquo;Crisis&amp;rdquo;?&lt;/h2&gt;
&lt;p&gt;The biggest change at KubeCon was that &lt;strong&gt;AI has completely replaced traditional cloud native topics&lt;/strong&gt;. The focus shifted from service optimization and microservices management to how to deploy and manage AI workloads on Kubernetes, especially inference tasks and GPU scheduling.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainer-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainer-summit.webp" alt="Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI." data-caption="Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI."
width="4000"
height="2668"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Kubernetes, as the foundational infrastructure, was once the core of the cloud native world. With the explosive growth of AI models, &lt;strong&gt;the question now is whether Kubernetes can still serve as a &amp;ldquo;universal&amp;rdquo; platform for everything&lt;/strong&gt;, which has become a new source of anxiety.&lt;/p&gt;
&lt;p&gt;The AI boom brings real challenges: &lt;strong&gt;Can Kubernetes&amp;rsquo; &amp;ldquo;universality&amp;rdquo; adapt to the complexity of AI workloads?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-focus-brought-by-the-ai-boom"&gt;The Focus Brought by the AI Boom&lt;/h2&gt;
&lt;p&gt;AI&amp;rsquo;s popularity has shifted the cloud native spotlight entirely to artificial intelligence. AI coding, OpenClaw, large language models, and generative models have all drawn widespread attention. AI has become the core computing demand in the real world.&lt;/p&gt;
&lt;p&gt;This surge in demand raises the question: Can Kubernetes continue to serve as the infrastructure platform for complex tasks? Especially with issues like GPU sharing, inference model scheduling, VRAM allocation, and device attribute selection, is the traditional Kubernetes resource model sufficient?&lt;/p&gt;
&lt;p&gt;In the past, Kubernetes handled compute, storage, and networking as foundational infrastructure. But with the rapid development of AI, its &amp;ldquo;universality&amp;rdquo; is being challenged. Particularly for inference tasks, Kubernetes&amp;rsquo; model appears thin.&lt;/p&gt;
&lt;h2 id="comparing-with-openstack-will-kubernetes-repeat-history"&gt;Comparing with OpenStack: Will Kubernetes Repeat History?&lt;/h2&gt;
&lt;p&gt;OpenStack once aimed to be a complete open-source cloud platform, but ultimately failed to sustain growth due to &lt;strong&gt;complexity&lt;/strong&gt; and a lack of &lt;strong&gt;flexibility&lt;/strong&gt; in adapting to new technologies.&lt;/p&gt;
&lt;p&gt;Will Kubernetes follow the same path? I believe Kubernetes has different strengths: as a container and microservices orchestration platform, it&amp;rsquo;s widely adopted and has strong community and vendor support. It doesn&amp;rsquo;t try to replace all cloud provider capabilities but serves as an infrastructure control plane to help users manage resources.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainers-summit-group-photo.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainers-summit-group-photo.webp" alt="Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality." data-caption="Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality."
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;However, as AI workloads become mainstream, Kubernetes must find a new position to avoid being replaced by &amp;ldquo;AI-optimized platforms.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="kubernetes-challenge-the-gpu-resource-management-gap"&gt;Kubernetes&amp;rsquo; Challenge: The GPU Resource Management Gap&lt;/h3&gt;
&lt;p&gt;At KubeCon, NVIDIA announced the donation of the &lt;strong&gt;&lt;a href="https://github.com/kubernetes-sigs/nvidia-dra-driver-gpu" target="_blank" rel="noopener"&gt;GPU DRA&lt;/a&gt; (Dynamic Resource Allocation) driver&lt;/strong&gt; to the CNCF, marking the upstreaming of GPU resource management. GPU sharing and scheduling have become urgent issues for Kubernetes.&lt;/p&gt;
&lt;p&gt;Traditionally, Kubernetes relied on the &lt;strong&gt;Device Plugin&lt;/strong&gt; model to schedule GPUs, only supporting allocation by device count (e.g., &lt;code&gt;nvidia.com/gpu: 1&lt;/code&gt;). But for AI inference tasks, more information is needed for resource scheduling, such as &lt;strong&gt;VRAM size&lt;/strong&gt;, &lt;strong&gt;GPU topology&lt;/strong&gt;, and &lt;strong&gt;sharing strategies&lt;/strong&gt;. NVIDIA DRA makes GPU resource management more flexible and intelligent, gradually easing the &amp;ldquo;GPU resource crunch&amp;rdquo; in AI workloads.&lt;/p&gt;
&lt;p&gt;This shift means Kubernetes is no longer just a &amp;ldquo;container orchestration platform,&amp;rdquo; but is becoming the &lt;strong&gt;infrastructure layer for AI-specific resource scheduling&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Against this backdrop, both the community and industry are exploring finer-grained GPU resource abstraction and scheduling mechanisms. For example, the open-source project &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; is building a GPU resource management layer for AI workloads on top of Kubernetes, supporting GPU sharing, VRAM-level allocation, and heterogeneous device scheduling.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/hami-kubecon-demo.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/hami-kubecon-demo.webp" alt="Figure 4: HAMi demo at KubeCon EU 2026 Keynote" data-caption="Figure 4: HAMi demo at KubeCon EU 2026 Keynote"
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi demo at KubeCon EU 2026 Keynote&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;These efforts are not about replacing Kubernetes, but about filling the resource model gaps for the AI era. In the long run, this layer may evolve into a &amp;ldquo;GPU Abstraction Layer&amp;rdquo; similar to CNI/CSI, becoming a key part of AI-native infrastructure.&lt;/p&gt;
&lt;h3 id="the-production-gap-many-ai-pocs-few-in-production"&gt;The Production &amp;ldquo;Gap&amp;rdquo;: Many AI PoCs, Few in Production&lt;/h3&gt;
&lt;p&gt;A common post-event summary was: &lt;strong&gt;Many PoCs, but &amp;ldquo;everyday production deployments&amp;rdquo; are still rare&lt;/strong&gt;. Pulumi summarized it as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;lots of working demos, very few production setups people trust&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This shows that while many AI workload solutions succeed in technical demos, the transition from &lt;strong&gt;experimentation to production&lt;/strong&gt; remains difficult. Whether it&amp;rsquo;s GPU resource sharing or inference request scheduling, &lt;strong&gt;whether Kubernetes as the foundation can support this transformation&lt;/strong&gt; is still an open question.&lt;/p&gt;
&lt;h2 id="the-rise-of-inference-systems-kubernetes-scheduling-boundaries-are-challenged"&gt;The Rise of Inference Systems: Kubernetes&amp;rsquo; Scheduling Boundaries Are Challenged&lt;/h2&gt;
&lt;p&gt;Another major event at this KubeCon was &lt;a href="https://github.com/llm-d/llm-d" target="_blank" rel="noopener"&gt;llm-d&lt;/a&gt; being contributed to the CNCF as a Sandbox project.&lt;/p&gt;
&lt;p&gt;If GPU DRA represents the upstreaming of device resource models, then llm-d represents another critical evolution: &lt;strong&gt;Distributed LLM inference capabilities are moving from proprietary engineering implementations to standardized, community-driven collaboration in cloud native.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is significant not just because it&amp;rsquo;s another open-source project, but because it shows that Kubernetes&amp;rsquo; challenges in the AI era are no longer just about &amp;ldquo;how to schedule GPUs,&amp;rdquo; but also &amp;ldquo;how to host inference systems themselves.&amp;rdquo; As prefill/decode separation, request routing, KV cache management, and throughput optimization move into the infrastructure layer, Kubernetes&amp;rsquo; boundaries are being redefined.&lt;/p&gt;
&lt;p&gt;Traditionally, the Kubernetes scheduler focused on Pod scheduling. But in AI inference scenarios, scheduling is not just about picking a node—it&amp;rsquo;s about &lt;strong&gt;selecting the most suitable inference instance based on request characteristics&lt;/strong&gt;. Factors like model state, request queue depth, and cache hit rate all need to be considered. This process is increasingly managed by inference runtimes, forming new &amp;ldquo;request-level scheduling&amp;rdquo; systems.&lt;/p&gt;
&lt;p&gt;This leads to an &lt;strong&gt;overlap between the Kubernetes scheduler and inference systems&lt;/strong&gt;, forcing Kubernetes to rethink its role: should it keep expanding, or collaborate with inference systems?&lt;/p&gt;
&lt;h2 id="ai-native-infrastructure-the-key-challenge-for-production"&gt;AI-Native Infrastructure: The Key Challenge for Production&lt;/h2&gt;
&lt;p&gt;At the &lt;strong&gt;AI Native Summit&lt;/strong&gt;, the real needs for AI-native infrastructure were especially clear. The focus was no longer &amp;ldquo;can it run on Kubernetes,&amp;rdquo; but how to make AI workloads routine, stable, and production-ready on Kubernetes.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/ai-native-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/ai-native-summit.webp" alt="Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era." data-caption="Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era."
width="2970"
height="1980"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The core challenge is &lt;strong&gt;delivery&lt;/strong&gt;. Unlike traditional apps, AI model weights are often huge—tens of GB or even TB—making model delivery and data management extremely complex. Traditional container delivery systems (like image layers) struggle with such massive data and complex versioning.&lt;/p&gt;
&lt;p&gt;A key direction for Kubernetes is to &lt;strong&gt;standardize model weight and data delivery&lt;/strong&gt;, using &lt;strong&gt;ImageVolume&lt;/strong&gt; and &lt;strong&gt;OCI artifacts&lt;/strong&gt; to solve AI model delivery and version management on Kubernetes. This not only reduces &amp;ldquo;cold start&amp;rdquo; times but also provides infrastructure support for multi-tenancy and compliance.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Kubernetes won&amp;rsquo;t be replaced by AI, but it&amp;rsquo;s being reshaped as the core of infrastructure. This anxiety is the force driving its evolution—it&amp;rsquo;s moving from a &lt;strong&gt;&amp;ldquo;general-purpose infrastructure platform&amp;rdquo;&lt;/strong&gt; to an &lt;strong&gt;&amp;ldquo;AI-powered multifunctional base&amp;rdquo;&lt;/strong&gt;. Some even call it the AI operating system.&lt;/p&gt;
&lt;p&gt;In the future, Kubernetes&amp;rsquo; core competitiveness will no longer be just container management, but &lt;strong&gt;how effectively it can schedule and manage AI workloads&lt;/strong&gt;, and how it can make AI a routine part of operations. This was my biggest takeaway from the AI Native Summit and KubeCon, and it&amp;rsquo;s what I look forward to in the Kubernetes ecosystem over the next few years.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blogs.nvidia.com/blog/nvidia-at-kubecon-2026/" target="_blank" rel="noopener"&gt;Advancing Open Source AI, NVIDIA Donates Dynamic Resource Allocation Driver for GPUs to Kubernetes Community - blog.nvidia.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.pulumi.com/blog/kubecon-eu-2026-recap/" target="_blank" rel="noopener"&gt;KubeCon EU 2026 Recap: The Year AI Moved Into Production on Kubernetes - pulumi.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Day One in Amsterdam: Kubernetes Is Rethinking AI</title><link>https://jimmysong.io/blog/kubecon-eu-2026-day1-ai-infra/</link><pubDate>Sun, 22 Mar 2026 20:41:19 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubecon-eu-2026-day1-ai-infra/</guid><description>KubeCon Europe 2026 Day One: How Kubernetes is adapting to the AI infrastructure wave and the evolution of the GPU resource layer.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Today marks my first day at &lt;strong&gt;KubeCon Europe 2026&lt;/strong&gt;. The most striking feeling is: the world is vast, but this community is truly small.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/jimmy-at-kubecon-eu.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/jimmy-at-kubecon-eu.webp" alt="Figure 11: Jimmy on the first day of KubeCon EU 2026" data-caption="Figure 11: Jimmy on the first day of KubeCon EU 2026"
width="2400"
height="2400"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 11: Jimmy on the first day of KubeCon EU 2026&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One strong impression stands out:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The world is big, but this circle is really small.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="old-friends-new-cycle"&gt;&lt;strong&gt;Old Friends, New Cycle&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;At the Maintainer Summit, I met many familiar faces—&lt;/p&gt;
&lt;p&gt;Colleagues from Ant Group, friends from Tetrate, and some people I&amp;rsquo;ve known for nearly a decade. Together, we&amp;rsquo;ve journeyed from the early days of Kubernetes, Service Mesh, and cloud native infrastructure to today.&lt;/p&gt;
&lt;p&gt;In a sense, this generation has fully experienced:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The rise of Kubernetes&lt;/li&gt;
&lt;li&gt;The standardization of Cloud Native&lt;/li&gt;
&lt;li&gt;The microservices and service mesh boom&lt;/li&gt;
&lt;li&gt;And now, the era of AI Infrastructure&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This isn&amp;rsquo;t about &amp;ldquo;new people entering the field,&amp;rdquo; but rather—&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The same group stepping into a new technology cycle.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="what-is-the-maintainer-summit-discussing"&gt;&lt;strong&gt;What Is the Maintainer Summit Discussing?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If you ask:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What is the Kubernetes community most concerned about right now?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Today&amp;rsquo;s answer is very clear:&lt;/p&gt;
&lt;p&gt;👉 &lt;strong&gt;How to run AI workloads better on Kubernetes&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit.webp" alt="Figure 12: The Maintainer Summit’s main topic is AI Infra" data-caption="Figure 12: The Maintainer Summit’s main topic is AI Infra"
width="1440"
height="960"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 12: The Maintainer Summit’s main topic is AI Infra&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Many topics at the Maintainer Summit revolved around:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Scheduling models for LLM / AI workloads&lt;/li&gt;
&lt;li&gt;GPU / accelerator resource management&lt;/li&gt;
&lt;li&gt;Integrating inference systems with Kubernetes&lt;/li&gt;
&lt;li&gt;Redefining the roles of data plane vs. control plane&lt;/li&gt;
&lt;li&gt;How observability tools like OTel monitor AI workloads&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes hasn&amp;rsquo;t been replaced by AI; it&amp;rsquo;s actively &amp;ldquo;absorbing&amp;rdquo; AI.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="key-signal-gpus-are-becoming-the"&gt;&lt;strong&gt;Key Signal: GPUs Are Becoming the &amp;ldquo;Infrastructure Layer&amp;rdquo;&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Today, I had an in-depth discussion with CNCF TOC, Red Hat, and the vLLM community.&lt;/p&gt;
&lt;p&gt;The core question was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How should GPUs be &amp;ldquo;platformized&amp;rdquo;?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Some consensus is already clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPUs are no longer just devices&lt;/li&gt;
&lt;li&gt;They are now &lt;strong&gt;a schedulable, partitionable, and shareable resource layer&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/toc-meeting.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/toc-meeting.webp" alt="Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration" data-caption="Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration"
width="2400"
height="1648"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;At the Maintainer Summit in Amsterdam, we had deep discussions with CNCF TOC, Red Hat, and the vLLM community about GPU resource management and LLM Serving integration in Kubernetes scenarios, and explored potential collaboration between vLLM and HAMi.&lt;/p&gt;
&lt;p&gt;Behind this is a major paradigm shift:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Past&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Now&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPU = Node resource&lt;/td&gt;
&lt;td&gt;GPU = Infrastructure layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exclusive use&lt;/td&gt;
&lt;td&gt;Multi-tenant sharing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Static binding&lt;/td&gt;
&lt;td&gt;Dynamic scheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed within frameworks&lt;/td&gt;
&lt;td&gt;Unified management at the platform layer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This is exactly what we&amp;rsquo;ve been working on in &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="hami-from"&gt;&lt;strong&gt;HAMi: From &amp;ldquo;Project&amp;rdquo; to &amp;ldquo;Reference Pattern&amp;rdquo;&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Another interesting change today:&lt;/p&gt;
&lt;p&gt;HAMi is no longer just a &amp;ldquo;community project&amp;rdquo;—it&amp;rsquo;s becoming:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A reference implementation (reference pattern) for AI Infra&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit-hami.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit-hami.webp" alt="Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit" data-caption="Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit"
width="1440"
height="960"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is reflected in several ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Invited to present at the Maintainer Summit&lt;/li&gt;
&lt;li&gt;Participating in CNCF TOC discussions&lt;/li&gt;
&lt;li&gt;Involved in incubating review demos&lt;/li&gt;
&lt;li&gt;Exploring joint content with the vLLM community (even discussing a joint blog 👀)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Especially in conversations with Red Hat and vLLM, a clear trend emerged:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GPU resource management and LLM serving are becoming coupled&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Upper layer: vLLM / inference frameworks&lt;/li&gt;
&lt;li&gt;Lower layer: GPU scheduling / sharing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A new &amp;ldquo;interface layer&amp;rdquo; is gradually forming.&lt;/p&gt;
&lt;p&gt;This is a direction worth betting on.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/incubating-review.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/incubating-review.webp" alt="Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo" data-caption="Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo"
width="2400"
height="1489"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="a-caution-the-ai-infra-startup-boom-hasn"&gt;&lt;strong&gt;A Caution: The AI Infra Startup Boom Hasn&amp;rsquo;t Really Begun&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;At the same time, I have a somewhat &amp;ldquo;counterintuitive&amp;rdquo; observation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;We haven&amp;rsquo;t yet seen a large wave of AI Infra (K8s-focused) startups.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Most companies I saw today:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many are pivoting from CI/CD, Service Mesh, or Gateway&lt;/li&gt;
&lt;li&gt;Many are traditional cloud vendors extending into AI&lt;/li&gt;
&lt;li&gt;Many are working on models, agents, or even lower-level tech&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But those truly focused on:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Making AI workloads run better on Kubernetes&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There are actually not many startups at this layer.&lt;/p&gt;
&lt;p&gt;This could mean two things:&lt;/p&gt;
&lt;h3 id="1-this-layer-isn"&gt;&lt;strong&gt;1) This Layer Isn&amp;rsquo;t Fully Formed Yet&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Currently, most activity is at:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The model layer (LLM / foundation models)&lt;/li&gt;
&lt;li&gt;The application layer (Agent / Copilot)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But not at:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The scheduling layer&lt;/li&gt;
&lt;li&gt;The resource layer&lt;/li&gt;
&lt;li&gt;The runtime layer&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="2-or-the-barrier-to-entry-is-very-high"&gt;&lt;strong&gt;2) Or, the Barrier to Entry Is Very High&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Because at its core, this is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The intersection of Cloud Native × GPU × AI workload&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It&amp;rsquo;s not just &amp;ldquo;wrapping AI,&amp;rdquo; but a fundamental re-architecture at the infrastructure level.&lt;/p&gt;
&lt;h2 id="my-take"&gt;&lt;strong&gt;My Take&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If we break down the AI technology stack:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Agent / Application
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;LLM Serving (vLLM, etc.)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI Runtime / Scheduling
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Resource Layer
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Hardware&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Most innovation today is concentrated in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The top two layers (Agent / LLM)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But the real long-term moat lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The middle two layers (Runtime + Resource Layer)&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And Kubernetes is very likely to remain:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The default platform for this middle layer&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Today&amp;rsquo;s takeaway:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes is not obsolete; it&amp;rsquo;s being redefined.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And our generation is shifting from:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Cloud Native Builders&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;AI Infrastructure Builders&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;More to come tomorrow.&lt;/p&gt;</content:encoded></item><item><title>HAMi Website Refactor: Why HAMi Docs and Website Underwent a Complete Redesign</title><link>https://jimmysong.io/blog/hami-website-redesign/</link><pubDate>Tue, 17 Mar 2026 08:55:52 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/hami-website-redesign/</guid><description>A systematic upgrade to HAMi’s website and docs, improving community visibility, content structure, search, and usability.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;This redesign is more than a style update—it&amp;rsquo;s a step toward clearer technical communication and better user experience. Try the new HAMi website at &lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;https://project-hami.io&lt;/a&gt; and submit issues &lt;a href="https://github.com/Project-HAMi/website/issues" target="_blank" rel="noopener"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Over the past two months, I conducted a thorough refactor of the documentation website (see &lt;a href="https://github.com/Project-HAMi/website/pulls?q=is%3Apr&amp;#43;is%3Aclosed&amp;#43;author%3Arootsongjc" target="_blank" rel="noopener"&gt;GitHub&lt;/a&gt;). Externally, it looks like a &amp;ldquo;visual redesign&amp;rdquo;, but from the perspective of community maintainers and content builders, it&amp;rsquo;s a comprehensive upgrade of information architecture, content system, and frontend experience.&lt;/p&gt;
&lt;p&gt;This article aims to systematically explain three things: why we did this refactor, what exactly changed, and what these changes mean for the HAMi community.&lt;/p&gt;
&lt;h2 id="why-refactor-the-website-and-documentation"&gt;Why Refactor the Website and Documentation&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; is a CNCF-hosted open source project initiated and contributed by &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt;, with growing influence in GPU virtualization, heterogeneous compute scheduling, and AI infrastructure. The community content is expanding, and user types are becoming more diverse: from first-time visitors to engineers and enterprise users seeking deployment docs, architecture diagrams, case studies, and ecosystem information.&lt;/p&gt;
&lt;p&gt;The original site was functional, but as content grew, several issues became apparent:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The homepage lacked information density, making it hard to quickly grasp the project&amp;rsquo;s overall value.&lt;/li&gt;
&lt;li&gt;Connections between docs, blogs, and community info were not smooth; content entry points were scattered.&lt;/li&gt;
&lt;li&gt;Search experience was unstable; external solutions were not ideal in practice.&lt;/li&gt;
&lt;li&gt;Mobile experience had many details needing improvement, especially navigation, card layouts, and footer areas.&lt;/li&gt;
&lt;li&gt;Visual style was inconsistent, making it hard to convey community influence and engineering maturity.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For a fast-evolving open source community, the website is not just a &amp;ldquo;place for docs&amp;rdquo;, but the public interface of the community. It needs to serve as project introduction, knowledge gateway, adoption proof, community connector, and brand expression.&lt;/p&gt;
&lt;p&gt;So the goal of this refactor was clear: not just superficial beautification, but to truly upgrade the website into HAMi&amp;rsquo;s systematic community entry point.&lt;/p&gt;
&lt;h2 id="what-was-done-in-this-refactor"&gt;What Was Done in This Refactor&lt;/h2&gt;
&lt;p&gt;This update was not a single-point change, but a series of systematic improvements.&lt;/p&gt;
&lt;h3 id="homepage-redesign-and-complete-information-architecture-overhaul"&gt;Homepage Redesign and Complete Information Architecture Overhaul&lt;/h3&gt;
&lt;p&gt;The most obvious change is the homepage.&lt;/p&gt;
&lt;p&gt;We redesigned the homepage structure, moving away from simply stacking content blocks, and instead organizing the page around the main narrative: &amp;ldquo;Project Positioning → Core Capabilities → Ecosystem Entry → Content Accumulation → Community Trust&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Specifically, the homepage received several key upgrades:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Rebuilt the Hero section to strengthen first-screen information delivery and action entry.&lt;/li&gt;
&lt;li&gt;Optimized CTA design so users can quickly access docs, blogs, and resources.&lt;/li&gt;
&lt;li&gt;Added and enhanced multiple homepage sections to showcase project value and community reach in a more structured way.&lt;/li&gt;
&lt;li&gt;Adjusted visual hierarchy, background atmosphere, and scroll rhythm, transforming the homepage from a &amp;ldquo;content list&amp;rdquo; into a &amp;ldquo;narrative page&amp;rdquo;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These changes include Hero animations and atmosphere layers, research/story sections, new resource entry sections, refreshed CTAs, unified background design, and ongoing reduction of visual noise. Together, they solve a core problem: enabling visitors to understand what HAMi is and why it&amp;rsquo;s worth exploring further within seconds.&lt;/p&gt;
&lt;h3 id="architecture-diagrams"&gt;Architecture Diagrams&lt;/h3&gt;
&lt;p&gt;Key diagrams were redrawn for clearer technical communication. This helps users grasp HAMi&amp;rsquo;s role in AI infrastructure.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-hero-diagram.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-hero-diagram.webp" alt="Figure 1: HAMi website homepage architecture diagram" data-caption="Figure 1: HAMi website homepage architecture diagram"
width="3160"
height="1714"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: HAMi website homepage architecture diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For HAMi, this change is critical. The community faces not just a single feature, but a set of system-level challenges involving Kubernetes, schedulers, GPU Operators, heterogeneous devices, and enterprise platforms. Improved diagrams make the website a better technical entry point.&lt;/p&gt;
&lt;h3 id="added-case-studies-community-and-ecosystem-sections-to-make-impact-visible"&gt;Added Case Studies, Community, and Ecosystem Sections to Make Impact Visible&lt;/h3&gt;
&lt;p&gt;Another important direction was strengthening the &amp;ldquo;community proof&amp;rdquo; layer.&lt;/p&gt;
&lt;p&gt;Many open source project sites fall into the trap of having complete docs, but users can&amp;rsquo;t tell if the project is truly adopted, if the community is active, or if the ecosystem is expanding. The HAMi website redesign consciously addresses this.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/ecosystem.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/ecosystem.webp" alt="Figure 2: HAMi ecosystem and device support" data-caption="Figure 2: HAMi ecosystem and device support"
width="2200"
height="454"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: HAMi ecosystem and device support&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/adopters.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/adopters.webp" alt="Figure 3: HAMi adopters" data-caption="Figure 3: HAMi adopters"
width="3688"
height="1534"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: HAMi adopters&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/contributors.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/contributors.webp" alt="Figure 4: HAMi contributor organizations" data-caption="Figure 4: HAMi contributor organizations"
width="3662"
height="674"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi contributor organizations&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="blog--reading-experience"&gt;Blog &amp;amp; Reading Experience&lt;/h3&gt;
&lt;p&gt;Blog cards, lists, and metadata were unified for easier reading and sharing. Blogs are now a core communication layer.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-blog.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-blog.webp" alt="Figure 5: HAMi website blog list page" data-caption="Figure 5: HAMi website blog list page"
width="2318"
height="1088"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: HAMi website blog list page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="mobile-optimization"&gt;Mobile Optimization&lt;/h3&gt;
&lt;p&gt;Navigation, card layouts, footer, and search were improved for smoother mobile browsing.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/mobile.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/mobile.webp" alt="Figure 6: HAMi website mobile view" data-caption="Figure 6: HAMi website mobile view"
width="654"
height="1418"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: HAMi website mobile view&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="footer--search"&gt;Footer &amp;amp; Search&lt;/h3&gt;
&lt;p&gt;Footer layout was enhanced for better navigation and credibility. Built-in search replaced unreliable external solutions, improving content accessibility.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/footer.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/footer.webp" alt="Figure 7: HAMi website footer" data-caption="Figure 7: HAMi website footer"
width="2472"
height="718"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: HAMi website footer&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/search.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/search.webp" alt="Figure 8: HAMi website built-in search" data-caption="Figure 8: HAMi website built-in search"
width="1330"
height="1214"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 8: HAMi website built-in search&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="what-this-redesign-means-for-the-hami-community"&gt;What This Redesign Means for the HAMi Community&lt;/h2&gt;
&lt;p&gt;From screenshots, it looks like &amp;ldquo;the website looks better&amp;rdquo;. But from a community-building perspective, its significance is deeper.&lt;/p&gt;
&lt;p&gt;First, HAMi&amp;rsquo;s external expression is more systematic.&lt;/p&gt;
&lt;p&gt;The website is no longer just a collection of scattered pages, but is forming a complete narrative chain: users can understand project value from the homepage, capability details from docs, practical paths from blogs, and community impact from ecosystem modules.&lt;/p&gt;
&lt;p&gt;Second, community content assets are reorganized.&lt;/p&gt;
&lt;p&gt;Previously, valuable articles, diagrams, and explanations existed but were hard to find. Now, through homepage sections, navigation, and search refactor, these contents are more effectively connected.&lt;/p&gt;
&lt;p&gt;Third, HAMi&amp;rsquo;s community image is more mature.&lt;/p&gt;
&lt;p&gt;A mature open source project needs not just an active code repository, but clear, stable, and sustainable website expression. Structure, style, and usability are part of the community&amp;rsquo;s engineering capability.&lt;/p&gt;
&lt;p&gt;Fourth, this lays the foundation for expanding case studies, adopters, contributors, and ecosystem content.&lt;/p&gt;
&lt;p&gt;With the framework sorted, adding more case studies, collaboration entry points, or showcasing more adopters and partners will be more natural and easier for users to understand.&lt;/p&gt;
&lt;h2 id="as-a-community-contributor-my-top-three-takeaways-from-this-redesign"&gt;As a Community Contributor, My Top Three Takeaways from This Redesign&lt;/h2&gt;
&lt;p&gt;In summary, I believe this refactor got three things right:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Upgraded the website from a &amp;ldquo;content dump&amp;rdquo; to a &amp;ldquo;community gateway&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Combined visual optimization with information architecture adjustment, not just a skin change.&lt;/li&gt;
&lt;li&gt;Improved basic experiences like search, mobile, navigation, and footer.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These may not be as flashy as launching a new feature, but they directly impact content dissemination, user comprehension, and the project&amp;rsquo;s long-term image.&lt;/p&gt;
&lt;p&gt;For infrastructure projects like HAMi, technical capability is fundamental, but clearly communicating, organizing, and continuously presenting that capability is also a form of infrastructure.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;This HAMi documentation and website refactor is essentially an upgrade to the community&amp;rsquo;s &amp;ldquo;expression layer&amp;rdquo; infrastructure.&lt;/p&gt;
&lt;p&gt;It improves visual and reading experience, reorganizes content, homepage narrative, search paths, mobile access, and community signal display. Homepage redesign, architecture diagram redraw, unified blog style, mobile optimization, enhanced footer, and switching from external to built-in search together constitute a true &amp;ldquo;refactor&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Externally, it helps more people quickly understand HAMi; internally, it provides a stable platform for the community to accumulate case studies, expand the ecosystem, and serve adopters and contributors.&lt;/p&gt;
&lt;p&gt;The website is not an accessory to the open source community, but part of its long-term influence. HAMi&amp;rsquo;s redesign is about taking this seriously.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re interested in Kubernetes GPU virtualization, add me on WeChat &lt;code&gt;jimmysong&lt;/code&gt; or scan the QR code below.&lt;/p&gt;
&lt;div class="cta-group"&gt;
&lt;a href="https://github.com/project-hami/hami" class="btn btn-sm btn-primary"&gt;Check out the HAMi project on GitHub&lt;/a&gt;
&lt;/div&gt;</content:encoded></item><item><title>GTC 2026 Eve: AI is Becoming the New Infrastructure</title><link>https://jimmysong.io/blog/gtc-2026-ai-native-infrastructure/</link><pubDate>Sun, 15 Mar 2026 11:34:06 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gtc-2026-ai-native-infrastructure/</guid><description>On the eve of GTC 2026, rethinking whether AI is becoming the new infrastructure from NVIDIA&amp;#39;s AI Five-Layer Cake, the rise of agent runtime, to AI-native infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;AI is quietly reshaping the infrastructure landscape, and GTC 2026 may become a key node in this transformation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Next week, one of the most important technology conferences in the AI industry, &lt;a href="https://www.nvidia.com/gtc/" target="_blank" rel="noopener"&gt;&lt;strong&gt;NVIDIA GTC 2026&lt;/strong&gt;&lt;/a&gt;, will be held in San Jose, USA.&lt;/p&gt;
&lt;p&gt;For many people, GTC is just a GPU technology conference. But if you follow the development of the AI industry over the past few years, you&amp;rsquo;ll find an interesting phenomenon:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Many important narratives about AI infrastructure are gradually taking shape at GTC.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From CUDA, DGX, to AI Factory, and most recently Jensen Huang&amp;rsquo;s proposed &lt;strong&gt;AI Five-Layer Cake&lt;/strong&gt;, NVIDIA is constantly attempting to redefine the computing infrastructure of the AI era.&lt;/p&gt;
&lt;p&gt;This is why many people call GTC:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI&amp;rsquo;s &amp;ldquo;Woodstock.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-gtc.webp" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-gtc.webp" alt="Figure 1: NVIDIA GTC Conference" data-caption="Figure 1: NVIDIA GTC Conference"
width="2212"
height="1152"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: NVIDIA GTC Conference&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This year&amp;rsquo;s GTC (March 16-19) is expected to cover various levels of the AI stack, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Chips&lt;/li&gt;
&lt;li&gt;AI Data Centers&lt;/li&gt;
&lt;li&gt;AI Agents&lt;/li&gt;
&lt;li&gt;Robotics&lt;/li&gt;
&lt;li&gt;Inference Computing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;According to &lt;a href="https://blogs.nvidia.com/blog/gtc-2026-news/" target="_blank" rel="noopener"&gt;NVIDIA&amp;rsquo;s official blog&lt;/a&gt;, this year&amp;rsquo;s keynote will focus on &lt;strong&gt;the complete AI stack from chips to applications&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;If we put these signals together, we can actually see a larger trend:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is transforming from an &amp;ldquo;applied technology&amp;rdquo; into &amp;ldquo;infrastructure.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-perspective-of-industrial-revolutions"&gt;The Perspective of Industrial Revolutions&lt;/h2&gt;
&lt;p&gt;From a longer time scale, the technological revolutions in human history are essentially infrastructure revolutions.&lt;/p&gt;
&lt;p&gt;We usually divide industrial revolutions into four times.&lt;/p&gt;
&lt;p&gt;In the table below, you can see the infrastructure corresponding to each industrial revolution:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Industrial Revolution&lt;/th&gt;
&lt;th&gt;Infrastructure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Steam Revolution&lt;/td&gt;
&lt;td&gt;Steam Engine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Electrical Revolution&lt;/td&gt;
&lt;td&gt;Power Grid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Digital Revolution&lt;/td&gt;
&lt;td&gt;Computer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internet Era&lt;/td&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Industrial Revolutions and Corresponding Infrastructure
&lt;/figcaption&gt;
&lt;h3 id="first-industrial-revolution-steam"&gt;First Industrial Revolution: Steam&lt;/h3&gt;
&lt;p&gt;The steam engine allowed humans to utilize mechanical power on a large scale for the first time. Production no longer relied on human or animal power, but on machines.&lt;/p&gt;
&lt;h3 id="second-industrial-revolution-electricity"&gt;Second Industrial Revolution: Electricity&lt;/h3&gt;
&lt;p&gt;Electricity changed not only the source of power, but also the organization of production. Assembly lines, large-scale manufacturing, and modern industrial systems are all built on the foundation of the power grid.&lt;/p&gt;
&lt;h3 id="third-industrial-revolution-computers"&gt;Third Industrial Revolution: Computers&lt;/h3&gt;
&lt;p&gt;Computers allowed information to be processed digitally. Software became a production tool.&lt;/p&gt;
&lt;h3 id="fourth-industrial-revolution-internet-and-intelligence"&gt;Fourth Industrial Revolution: Internet and Intelligence&lt;/h3&gt;
&lt;p&gt;The internet connects all computers together. Cloud computing transforms computing resources into infrastructure. And AI gives machines a certain degree of &amp;ldquo;cognitive ability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-true-significance-of-ai"&gt;The True Significance of AI&lt;/h2&gt;
&lt;p&gt;If we observe these industrial revolutions, we discover a pattern:&lt;/p&gt;
&lt;p&gt;Each industrial revolution produces a new &lt;strong&gt;General Purpose Infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;And AI is likely to become the next-generation infrastructure.&lt;/p&gt;
&lt;p&gt;NVIDIA even directly stated in a &lt;a href="https://blogs.nvidia.com/blog/ai-5-layer-cake/" target="_blank" rel="noopener"&gt;recent article&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI is essential infrastructure, like electricity and the internet.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;p&gt;AI is no longer just an applied technology, but a &lt;strong&gt;new factor of production&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="nvidias-five-layer-cake"&gt;NVIDIA&amp;rsquo;s Five-Layer Cake&lt;/h2&gt;
&lt;p&gt;Recently, Jensen Huang proposed a very interesting concept: &lt;strong&gt;AI Five-Layer Cake&lt;/strong&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-five-layer-cake.webp" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-five-layer-cake.webp" alt="Figure 2: AI Five Layer Cake (Image source: &amp;lt;a href=&amp;#34;https://blogs.nvidia.com/blog/ai-5-layer-cake/&amp;#34; target=&amp;#34;_blank&amp;#34; rel=&amp;#34;noopener&amp;#34;&amp;gt;NVIDIA&amp;lt;/a&amp;gt;)" data-caption="Figure 2: AI Five Layer Cake (Image source: &amp;lt;a href=&amp;#34;https://blogs.nvidia.com/blog/ai-5-layer-cake/&amp;#34; target=&amp;#34;_blank&amp;#34; rel=&amp;#34;noopener&amp;#34;&amp;gt;NVIDIA&amp;lt;/a&amp;gt;)"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: AI Five Layer Cake (Image source: &lt;a href="https://blogs.nvidia.com/blog/ai-5-layer-cake/" target="_blank" rel="noopener"&gt;NVIDIA&lt;/a&gt;)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;AI is broken down into five layers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Energy&lt;/li&gt;
&lt;li&gt;Chips&lt;/li&gt;
&lt;li&gt;AI Infrastructure&lt;/li&gt;
&lt;li&gt;Models&lt;/li&gt;
&lt;li&gt;Applications&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This model actually illustrates one thing:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is a complete industrial system.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Jensen Huang even described AI at Davos as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;One of the largest-scale infrastructure constructions in human history.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="signals-gtc-2026-may-release"&gt;Signals GTC 2026 May Release&lt;/h2&gt;
&lt;p&gt;This year&amp;rsquo;s GTC is expected to release several important directions.&lt;/p&gt;
&lt;h3 id="inference-computing"&gt;Inference Computing&lt;/h3&gt;
&lt;p&gt;The focus of AI in the past was training. But the main load of AI in the future is likely to be &lt;strong&gt;Inference&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Analysts expect that by 2030, &lt;strong&gt;75% of computing demand in the AI data center market will come from inference&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="agentic-ai"&gt;Agentic AI&lt;/h3&gt;
&lt;p&gt;The past AI model was:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;User → Model → Answer&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The Agent model is more complex:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;User → Agent → Tools → Model → Action&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The flowchart below shows the main interaction paths in the Agent model:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agentic-ai-interaction-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agentic-ai-interaction-en.svg" alt="Figure 3: Agentic AI Interaction Flow" data-caption="Figure 3: Agentic AI Interaction Flow"
width="936"
height="536"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Agentic AI Interaction Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;AI is no longer just answering questions, but &lt;strong&gt;executing tasks&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="agent-platform"&gt;Agent Platform&lt;/h3&gt;
&lt;p&gt;Recent media reports suggest that NVIDIA may launch a new Agent platform: &lt;strong&gt;NemoClaw&lt;/strong&gt;, aimed at helping enterprises deploy AI Agents.&lt;/p&gt;
&lt;p&gt;If this project is truly released, it means NVIDIA&amp;rsquo;s stack will become the following structure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-agent-platform-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-agent-platform-en.svg" alt="Figure 4: NVIDIA Agent Platform Architecture" data-caption="Figure 4: NVIDIA Agent Platform Architecture"
width="416"
height="816"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: NVIDIA Agent Platform Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is actually a complete AI stack.&lt;/p&gt;
&lt;h2 id="agents-change-computing-workloads"&gt;Agents Change Computing Workloads&lt;/h2&gt;
&lt;p&gt;The emergence of Agents brings new computing workload issues.&lt;/p&gt;
&lt;p&gt;Past AI workloads were mainly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training&lt;/li&gt;
&lt;li&gt;Inference&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But Agents bring a third type of workload:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent Workloads&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The figure below shows the diverse workload types related to Agents:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agent-workloads-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agent-workloads-en.svg" alt="Figure 5: Agent Workloads Structure" data-caption="Figure 5: Agent Workloads Structure"
width="1376"
height="316"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Agent Workloads Structure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The characteristic of this workload is &lt;strong&gt;highly fragmented&lt;/strong&gt;. GPUs are no longer occupied for long periods, but rather face many small requests. This poses new challenges for infrastructure.&lt;/p&gt;
&lt;h2 id="ai-native-infrastructure"&gt;AI-Native Infrastructure&lt;/h2&gt;
&lt;p&gt;For the past few years, I&amp;rsquo;ve been thinking about a question:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is AI-native infrastructure?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It is clearly not just &amp;ldquo;Kubernetes with GPUs.&amp;rdquo; I&amp;rsquo;m more inclined to believe it needs to possess several characteristics.&lt;/p&gt;
&lt;h3 id="gpu-as-a-first-class-resource"&gt;GPU as a First-Class Resource&lt;/h3&gt;
&lt;p&gt;In the cloud computing era, CPU is the core resource. In the AI era, &lt;strong&gt;GPU is the core resource&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="heterogeneous-computing"&gt;Heterogeneous Computing&lt;/h3&gt;
&lt;p&gt;Real-world AI chips are not limited to NVIDIA:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;NVIDIA&lt;/li&gt;
&lt;li&gt;Ascend&lt;/li&gt;
&lt;li&gt;Cambricon&lt;/li&gt;
&lt;li&gt;Metax&lt;/li&gt;
&lt;li&gt;Moore Threads&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Future AI infrastructure must be able to manage &lt;strong&gt;heterogeneous computing&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="gpu-sharing"&gt;GPU Sharing&lt;/h3&gt;
&lt;p&gt;GPU is a very expensive resource. If it cannot be shared, utilization will be very low. This is why GPU virtualization and slicing are becoming increasingly important.&lt;/p&gt;
&lt;h3 id="ai-scheduling"&gt;AI Scheduling&lt;/h3&gt;
&lt;p&gt;AI scheduling includes not only traditional CPU and Memory, but also:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;VRAM
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Topology
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Bandwidth&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="a-possible-ai-tech-stack"&gt;A Possible AI Tech Stack&lt;/h2&gt;
&lt;p&gt;Combining the above trends, the future AI stack may present the following structure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-tech-stack-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-tech-stack-en.svg" alt="Figure 6: AI Tech Stack Evolution" data-caption="Figure 6: AI Tech Stack Evolution"
width="416"
height="956"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: AI Tech Stack Evolution&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This structure is very close to NVIDIA&amp;rsquo;s Five-Layer Cake.&lt;/p&gt;
&lt;h2 id="my-judgment"&gt;My Judgment&lt;/h2&gt;
&lt;p&gt;Combining signals from GTC, AI Factory, Agents, and AI Five-Layer Cake, we can see a very obvious trend:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is rewriting computing infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Future competition may not just be &amp;ldquo;who has the best model,&amp;rdquo; but:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Who has the best AI Infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Just like the past few decades:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity determines industrial capability&lt;/li&gt;
&lt;li&gt;Internet determines information capability&lt;/li&gt;
&lt;li&gt;Cloud computing determines software capability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The future may be:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Infrastructure determines intelligence capability.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;If we stretch the time scale a bit longer, we may be in a new historical stage.&lt;/p&gt;
&lt;p&gt;AI is no longer just a technological tool. It is becoming &lt;strong&gt;new infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Just like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity&lt;/li&gt;
&lt;li&gt;Internet&lt;/li&gt;
&lt;li&gt;Cloud computing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And AI-native infrastructure is likely to become one of the most important technology directions for the next decade.&lt;/p&gt;</content:encoded></item><item><title>When GPUs Move Toward Open Scheduling: Structural Shifts in AI Native Infrastructure</title><link>https://jimmysong.io/blog/gpu-open-scheduling-hami-2025/</link><pubDate>Fri, 13 Feb 2026 14:32:46 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gpu-open-scheduling-hami-2025/</guid><description>A CTO/VP view on open GPU scheduling: CDI, Kubernetes DRA, virtualization data planes, ecosystem governance, and lock-in risk.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The future of GPU scheduling isn&amp;rsquo;t about whose implementation is more &amp;ldquo;black-box&amp;rdquo;—it&amp;rsquo;s about who can standardize device resource contracts into something governable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/banner.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/banner.webp" alt="Figure 1: GPU Open Scheduling" data-caption="Figure 1: GPU Open Scheduling"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: GPU Open Scheduling&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Have you ever wondered: why are GPUs so expensive, yet overall utilization often hovers around 10–20%?&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/underutilization.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/underutilization.webp" alt="Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization" data-caption="Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This isn&amp;rsquo;t a problem you solve with &amp;ldquo;better scheduling algorithms.&amp;rdquo; It&amp;rsquo;s a &lt;strong&gt;structural problem&lt;/strong&gt; - GPU scheduling is undergoing a shift from &amp;ldquo;proprietary implementation&amp;rdquo; to &amp;ldquo;open scheduling,&amp;rdquo; similar to how networking converged on CNI and storage converged on CSI.&lt;/p&gt;
&lt;p&gt;In the &lt;a href="https://dynamia.ai/blog/hami-2025-recap" target="_blank" rel="noopener"&gt;HAMi 2025 Annual Review&lt;/a&gt;, we noted: &amp;ldquo;HAMi 2025 is no longer just about GPU sharing tools—it&amp;rsquo;s a more structural signal: GPUs are moving toward open scheduling.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;By 2025, the signals of this shift became visible: Kubernetes Dynamic Resource Allocation (DRA) graduated to GA and became enabled by default, NVIDIA GPU Operator started defaulting to &lt;a href="https://github.com/cncf-tags/container-device-interface" target="_blank" rel="noopener"&gt;CDI&lt;/a&gt; (Container Device Interface), and HAMi&amp;rsquo;s production-grade case studies under CNCF are moving &amp;ldquo;GPU sharing&amp;rdquo; from experimental capability to operational excellence.&lt;/p&gt;
&lt;p&gt;This post analyzes this structural shift from an AI Native Infrastructure perspective, and what it means for &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; and the industry.&lt;/p&gt;
&lt;h2 id="why-open-scheduling-matters"&gt;Why &amp;ldquo;Open Scheduling&amp;rdquo; Matters&lt;/h2&gt;
&lt;p&gt;In multi-cloud and hybrid cloud environments, GPU model diversity significantly amplifies operational costs. One large internet company&amp;rsquo;s platform spans H200/H100/A100/V100/4090 GPUs across five clusters. If you can only allocate &amp;ldquo;whole GPUs,&amp;rdquo; resource misalignment becomes inevitable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Open scheduling&amp;rdquo; isn&amp;rsquo;t a slogan—it&amp;rsquo;s a set of engineering contracts being solidified into the mainstream stack.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="standardized-resource-expression"&gt;Standardized Resource Expression&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt; GPUs were extended resources. The scheduler didn&amp;rsquo;t understand if they represented memory, compute, or device types.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dra-evolution.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dra-evolution.webp" alt="Figure 3: Open Scheduling Standardization Evolution" data-caption="Figure 3: Open Scheduling Standardization Evolution"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Open Scheduling Standardization Evolution&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Now:&lt;/strong&gt; Kubernetes DRA provides objects like DeviceClass, ResourceClaim, and ResourceSlice. This lets drivers and cluster administrators define device categories and selection logic (including CEL-based selectors), while Kubernetes handles the full loop: match devices → bind claims → place Pods onto nodes with access to allocated devices.&lt;/p&gt;
&lt;p&gt;Even more importantly, Kubernetes 1.34 stated that core APIs in the &lt;code&gt;resource.k8s.io&lt;/code&gt; group graduated to GA, DRA became stable and enabled by default, and the community committed to avoiding breaking changes going forward. This means the ecosystem can invest with confidence in a stable, standard API.&lt;/p&gt;
&lt;h3 id="standardized-device-injection"&gt;Standardized Device Injection&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt; Device injection relied on vendor-specific hooks and runtime class patterns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Now:&lt;/strong&gt; The Container Device Interface (CDI) abstracts device injection into an open specification. NVIDIA&amp;rsquo;s Container Toolkit explicitly describes CDI as an open specification for container runtimes, and NVIDIA GPU Operator 25.10.0 defaults to enabling CDI on install/upgrade—directly leveraging runtime-native CDI support (containerd, CRI-O, etc.) for GPU injection.&lt;/p&gt;
&lt;p&gt;This means &amp;ldquo;devices into containers&amp;rdquo; is also moving toward replaceable, standardized interfaces.&lt;/p&gt;
&lt;h2 id="hami-from-sharing-tool-to-governable-data-plane"&gt;HAMi: From &amp;ldquo;Sharing Tool&amp;rdquo; to &amp;ldquo;Governable Data Plane&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;On this standardization path, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;&amp;rsquo;s role needs redefinition: &lt;strong&gt;it&amp;rsquo;s not about replacing Kubernetes—it&amp;rsquo;s about turning GPU virtualization and slicing into a declarative, schedulable, governable data plane.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="data-plane-perspective"&gt;Data Plane Perspective&lt;/h3&gt;
&lt;p&gt;HAMi&amp;rsquo;s core contribution expands the allocatable unit from &amp;ldquo;whole GPU integers&amp;rdquo; to finer-grained shares (memory and compute), forming a complete allocation chain:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Device discovery:&lt;/strong&gt; Identify available GPU devices and models&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling placement:&lt;/strong&gt; Use Scheduler Extender to make native schedulers &amp;ldquo;understand&amp;rdquo; vGPU resource models (Filter/Score/Bind phases)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;In-container enforcement:&lt;/strong&gt; Inject share constraints into container runtime&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metric export:&lt;/strong&gt; Provide observable metrics for utilization, isolation, and more&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This transforms &amp;ldquo;sharing&amp;rdquo; from ad-hoc &amp;ldquo;it runs&amp;rdquo; experimentation into engineering capability that can be declared in YAML, scheduled by policy, and validated by metrics.&lt;/p&gt;
&lt;h3 id="scheduling-mechanism-enhancement-not-replacement"&gt;Scheduling Mechanism: Enhancement, Not Replacement&lt;/h3&gt;
&lt;p&gt;HAMi&amp;rsquo;s scheduling doesn&amp;rsquo;t replace Kubernetes—it uses a &lt;strong&gt;Scheduler Extender&lt;/strong&gt; pattern to let the native scheduler understand vGPU resource models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Filter:&lt;/strong&gt; Filter nodes based on memory, compute, device type, topology, and other constraints&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Score:&lt;/strong&gt; Apply configurable policies like binpack, spread, topology-aware scoring&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bind:&lt;/strong&gt; Complete final device-to-Pod binding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This architecture positions HAMi naturally as an execution layer under higher-level &amp;ldquo;AI control planes&amp;rdquo; (queuing, quotas, priorities)—working alongside Volcano, Kueue, Koordinator, and others.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/hami-scheduler-extender.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/hami-scheduler-extender.webp" alt="Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)" data-caption="Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="production-evidence-from-can-we-share-to-can-we-operate"&gt;Production Evidence: From &amp;ldquo;Can We Share?&amp;rdquo; to &amp;ldquo;Can We Operate?&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.cncf.io/case-studies/?_sft_lf_project=hami" target="_blank" rel="noopener"&gt;CNCF public case studies&lt;/a&gt; provide concrete answers: &lt;strong&gt;in a hybrid, multi-cloud platform built on Kubernetes and HAMi, 10,000+ Pods run concurrently, and GPU utilization improves from 13% to 37% (nearly 3×).&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/case-studies.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/case-studies.webp" alt="Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%&amp;#43; utilization, SF Technology 57% savings" data-caption="Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%&amp;#43; utilization, SF Technology 57% savings"
width="2466"
height="1508"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%+ utilization, SF Technology 57% savings&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Here are highlights from several cases:&lt;/p&gt;
&lt;h3 id="case-study-1-ke-holdings-february-5-2026"&gt;Case Study 1: Ke Holdings (February 5, 2026)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Environment:&lt;/strong&gt; 5 clusters spanning public and private clouds&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU models:&lt;/strong&gt; H200/H100/A100/V100/4090 and more&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Architecture:&lt;/strong&gt; Separate &amp;ldquo;GPU clusters&amp;rdquo; for large training tasks (dedicated allocation) vs &amp;ldquo;vGPU clusters&amp;rdquo; with HAMi fine-grained memory slicing for high-density inference&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Concurrent scale:&lt;/strong&gt; 10,000+ Pods&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcome:&lt;/strong&gt; Overall GPU utilization improved from 13% to 37% (nearly 3×)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-2-daocloud-december-2-2025"&gt;Case Study 2: DaoCloud (December 2, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hard constraints:&lt;/strong&gt; Must remain cloud-native, vendor-agnostic, and compatible with CNCF toolchain&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adoption outcomes:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Average GPU utilization: 80%+&lt;/li&gt;
&lt;li&gt;GPU-related operating cost reduction: 20–30%&lt;/li&gt;
&lt;li&gt;Coverage: 10+ data centers, 10,000+ GPUs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit benefit:&lt;/strong&gt; Unified abstraction layer across NVIDIA and domestic GPUs, reducing vendor dependency&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-3-prep-edu-august-20-2025"&gt;Case Study 3: Prep EDU (August 20, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Negative experience:&lt;/strong&gt; Isolation failures in other GPU-sharing approaches caused memory conflicts and instability&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Positive outcome:&lt;/strong&gt; HAMi&amp;rsquo;s vGPU scheduling, GPU type/UUID targeting, and compatibility with NVIDIA GPU Operator and RKE2 became decisive factors for production adoption&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environment:&lt;/strong&gt; Heterogeneous RTX 4070/4090 cluster&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-4-sf-technology-september-18-2025"&gt;Case Study 4: SF Technology (September 18, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Project:&lt;/strong&gt; EffectiveGPU (built on HAMi)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use cases:&lt;/strong&gt; Large model inference, test services, speech recognition, domestic AI hardware (Huawei Ascend, Baidu Kunlun, etc.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcomes:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;GPU savings: Large model inference runs 65 services on 28 GPUs (37 saved); test cluster runs 19 services on 6 GPUs (13 saved)&lt;/li&gt;
&lt;li&gt;Overall savings: Up to 57% GPU savings for production and test clusters&lt;/li&gt;
&lt;li&gt;Utilization improvement: Up to 100% GPU utilization improvement with GPU virtualization&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Highlights:&lt;/strong&gt; Cross-node collaborative scheduling, priority-based preemption, memory over-subscription&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These cases demonstrate a consistent pattern: &lt;strong&gt;GPU virtualization becomes economically meaningful only when it participates in a governable contract—where utilization, isolation, and policy can be expressed, measured, and improved over time.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="strategic-implications-for-dynamia"&gt;Strategic Implications for Dynamia&lt;/h2&gt;
&lt;p&gt;From Dynamia&amp;rsquo;s perspective (and as VP of Open Source Ecosystem), the strategic value of HAMi becomes clear:&lt;/p&gt;
&lt;h3 id="two-layer-architecture-open-source-vs-commercial"&gt;Two-Layer Architecture: Open Source vs Commercial&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi (CNCF open source project):&lt;/strong&gt; Responsible for &amp;ldquo;adoption and trust,&amp;rdquo; focused on GPU virtualization and compute efficiency&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamia enterprise products and services:&lt;/strong&gt; Responsible for &amp;ldquo;production and scale,&amp;rdquo; providing commercial distributions and enterprise services built on HAMi&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dynamia-hami-dual-mechanism.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dynamia-hami-dual-mechanism.webp" alt="Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial" data-caption="Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This boundary is the foundation for long-term trust—project and company offerings remain separate, with commercial distributions and services built on the open source project.&lt;/p&gt;
&lt;h3 id="global-narrative-strategy"&gt;Global Narrative Strategy&lt;/h3&gt;
&lt;p&gt;The internal alignment memo recommends a bilingual approach:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First layer:&lt;/strong&gt; Lead globally with &amp;ldquo;GPU virtualization / sharing / utilization&amp;rdquo; (Chinese can directly use &amp;ldquo;GPU virtualization and heterogeneous scheduling,&amp;rdquo; but English first layer should avoid &amp;ldquo;heterogeneous&amp;rdquo; as a headline)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second layer:&lt;/strong&gt; When users discuss mixed GPUs or workload diversity, introduce &amp;ldquo;heterogeneous&amp;rdquo; to confirm capability boundaries—never as the opening hook&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core anchor:&lt;/strong&gt; Maintain &amp;ldquo;HAMi (project and community) ≠ company products&amp;rdquo; as the non-negotiable baseline for long-term positioning&lt;/p&gt;
&lt;h3 id="the-right-commercialization-landing"&gt;The Right Commercialization Landing&lt;/h3&gt;
&lt;p&gt;DaoCloud&amp;rsquo;s case study already set vendor-agnostic and CNCF toolchain compatibility as hard constraints, framing vendor dependency reduction as a business and operational benefit—not just a technical detail. Project-HAMi&amp;rsquo;s official documentation lists &amp;ldquo;avoid vendor lock&amp;rdquo; as a core value proposition.&lt;/p&gt;
&lt;p&gt;In this context, &lt;strong&gt;the right commercialization landing isn&amp;rsquo;t &amp;ldquo;closed-source scheduling&amp;rdquo;—it&amp;rsquo;s productizing capabilities around real enterprise complexity:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Systematic compatibility matrix&lt;/li&gt;
&lt;li&gt;SLO and tail-latency governance&lt;/li&gt;
&lt;li&gt;Metering for billing&lt;/li&gt;
&lt;li&gt;RBAC, quotas, multi-cluster governance&lt;/li&gt;
&lt;li&gt;Upgrade and rollback safety&lt;/li&gt;
&lt;li&gt;Faster path-to-production for DRA/CDI and other standardization efforts&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="forward-view-the-next-23-years"&gt;Forward View: The Next 2–3 Years&lt;/h2&gt;
&lt;p&gt;My strong judgment: &lt;strong&gt;over the next 2–3 years, GPU scheduling competition will shift from &amp;ldquo;whose implementation is more black-box&amp;rdquo; to &amp;ldquo;whose contract is more open.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The reasons are practical:&lt;/p&gt;
&lt;h3 id="hardware-form-factors-and-supply-chains-are-diversifying"&gt;Hardware Form Factors and Supply Chains Are Diversifying&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI&amp;rsquo;s February 12, 2026 &amp;ldquo;GPT‑5.3‑Codex‑Spark&amp;rdquo; release emphasizes ultra-low latency serving, including persistent WebSockets and a dedicated serving tier on Cerebras hardware&lt;/li&gt;
&lt;li&gt;Large-scale GPU-backed financing announcements (for pan-European deployments) illustrate the infrastructure scale and financial engineering surrounding accelerator fleets&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These signals suggest that heterogeneity will grow: mixed accelerators, mixed clouds, mixed workload types.&lt;/p&gt;
&lt;h3 id="low-latency-inference-tiers-will-force-systematic-scheduling"&gt;Low-Latency Inference Tiers Will Force Systematic Scheduling&lt;/h3&gt;
&lt;p&gt;Low-latency inference tiers (beyond just GPUs) will force resource scheduling toward &amp;ldquo;multi-accelerator, multi-layer cache, multi-class node&amp;rdquo; architectural design—scheduling must inherently be heterogeneous.&lt;/p&gt;
&lt;h3 id="open-scheduling-is-risk-management-not-idealism"&gt;Open Scheduling Is Risk Management, Not Idealism&lt;/h3&gt;
&lt;p&gt;In this world, &amp;ldquo;open scheduling&amp;rdquo; isn&amp;rsquo;t idealism—it&amp;rsquo;s risk management. Building schedulable governable &amp;ldquo;control plane + data plane&amp;rdquo; combinations around DRA/CDI and other solidifying open interfaces, ones that are pluggable, multi-tenant governable, and co-evolvable with the ecosystem—this looks like the truly sustainable path for AI Native Infrastructure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The next battleground isn&amp;rsquo;t &amp;ldquo;whose scheduling is smarter&amp;rdquo;—it&amp;rsquo;s &amp;ldquo;who can standardize device resource contracts into something governable.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;When you place HAMi 2025 back in the broader AI Native Infrastructure context, it&amp;rsquo;s no longer just the year of &amp;ldquo;GPU sharing tools&amp;rdquo;—it&amp;rsquo;s a more structural signal: &lt;strong&gt;GPUs are moving toward open scheduling.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/future-vision-open-scheduling.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/future-vision-open-scheduling.webp" alt="Figure 7: Open Scheduling Future Vision" data-caption="Figure 7: Open Scheduling Future Vision"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: Open Scheduling Future Vision&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The driving forces come from both ends:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Upstream:&lt;/strong&gt; Standards like DRA/CDI continue to solidify&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downstream:&lt;/strong&gt; Scale and diversity (multi-cloud, multi-model, even accelerators beyond GPUs)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For Dynamia, HAMi&amp;rsquo;s significance has transcended &amp;ldquo;GPU sharing tool&amp;rdquo;: it turns GPU virtualization and slicing into declarative, schedulable, measurable data planes—letting queues, quotas, priorities, and multi-tenancy actually close the governance loop.&lt;/p&gt;</content:encoded></item><item><title>AI Learning Resources: 44 Curated Collections from Our Cleanup</title><link>https://jimmysong.io/blog/ultimate-ai-learning-resources/</link><pubDate>Sun, 08 Feb 2026 12:20:05 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ultimate-ai-learning-resources/</guid><description>A curated collection of AI learning resources we removed from the AI Resources list: awesome lists, courses, tutorials, and cookbooks. These educational materials deserve their own spotlight.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;The best way to learn AI is to start building. These resources will guide your journey.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ultimate-ai-learning-resources/banner.webp" data-img="https://assets.jimmysong.io/images/blog/ultimate-ai-learning-resources/banner.webp" alt="Figure 1: AI Learning Resources Collection" data-caption="Figure 1: AI Learning Resources Collection"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: AI Learning Resources Collection&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In my ongoing effort to keep the AI Resources list focused on &lt;strong&gt;production-ready tools and frameworks&lt;/strong&gt;, I&amp;rsquo;ve removed &lt;strong&gt;44 collection-type projects&lt;/strong&gt;—courses, tutorials, awesome lists, and cookbooks.&lt;/p&gt;
&lt;p&gt;These resources aren&amp;rsquo;t gone—they&amp;rsquo;ve been moved here. This post is a &lt;strong&gt;curated collection&lt;/strong&gt; of those educational materials, organized by type and topic. Whether you&amp;rsquo;re a complete beginner or an experienced practitioner, you&amp;rsquo;ll find something valuable here.&lt;/p&gt;
&lt;h2 id="why-remove-collections-from-ai-resources"&gt;Why Remove Collections from AI Resources?&lt;/h2&gt;
&lt;p&gt;My AI Resources list now focuses on &lt;strong&gt;concrete tools and frameworks&lt;/strong&gt;—projects you can directly use in production. Collections, while valuable, serve a different purpose: &lt;strong&gt;education and discovery&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;By separating them, I:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Keep the resources list actionable and focused&lt;/li&gt;
&lt;li&gt;Create a dedicated space for learning materials&lt;/li&gt;
&lt;li&gt;Make it easier to find what you need&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-awesome-lists-14-collections"&gt;📚 Awesome Lists (14 Collections)&lt;/h2&gt;
&lt;p&gt;Awesome lists are community-curated collections of the best resources. They&amp;rsquo;re perfect for discovering new tools and staying updated.&lt;/p&gt;
&lt;h3 id="must-explore-awesome-lists"&gt;Must-Explore Awesome Lists&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/filipecalegario/awesome-generative-ai" target="_blank" rel="noopener"&gt;Awesome Generative AI&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Models, tools, tutorials, and research papers&lt;/li&gt;
&lt;li&gt;Great for: Comprehensive overview of generative AI landscape&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/hannibal046/awesome-llm" target="_blank" rel="noopener"&gt;Awesome LLM&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM resources: papers, tools, datasets, applications&lt;/li&gt;
&lt;li&gt;Great for: Deep dive into large language models&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/arindam200/awesome-ai-apps" target="_blank" rel="noopener"&gt;Awesome AI Apps&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Practical LLM applications, RAG examples, agent implementations&lt;/li&gt;
&lt;li&gt;Great for: Real-world implementation examples&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/hesreallyhim/awesome-claude-code" target="_blank" rel="noopener"&gt;Awesome Claude Code&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Claude Code commands, files, and workflows&lt;/li&gt;
&lt;li&gt;Great for: Maximizing Claude Code productivity&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/punkpeye/awesome-mcp-servers" target="_blank" rel="noopener"&gt;Awesome MCP Servers&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP servers for modular AI backend systems&lt;/li&gt;
&lt;li&gt;Great for: Building with Model Context Protocol&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="specialized-awesome-lists"&gt;Specialized Awesome Lists&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/f/awesome-chatgpt-prompts" target="_blank" rel="noopener"&gt;Awesome ChatGPT Prompts&lt;/a&gt;&lt;/strong&gt; - Prompt examples for various scenarios&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/shubhamsaboo/awesome-llm-apps" target="_blank" rel="noopener"&gt;Awesome LLM Apps&lt;/a&gt;&lt;/strong&gt; - LLM applications with code examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/bradyfu/awesome-multimodal-large-language-models" target="_blank" rel="noopener"&gt;Awesome Multimodal LLM&lt;/a&gt;&lt;/strong&gt; - Multimodal model resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/punkpeye/awesome-mcp-clients" target="_blank" rel="noopener"&gt;Awesome MCP Clients&lt;/a&gt;&lt;/strong&gt; - MCP client tools and SDKs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/composiohq/awesome-claude-skills" target="_blank" rel="noopener"&gt;Awesome Claude Skills&lt;/a&gt;&lt;/strong&gt; - Claude Skills and workflows&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/github/awesome-copilot" target="_blank" rel="noopener"&gt;Awesome GitHub Copilot&lt;/a&gt;&lt;/strong&gt; - Copilot customizations&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/zerolu/awesome-nanobanana-pro" target="_blank" rel="noopener"&gt;Awesome Nano Banana Pro&lt;/a&gt;&lt;/strong&gt; - Image model prompts and examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/alchemyst-ai/awesome-saas" target="_blank" rel="noopener"&gt;Awesome SaaS&lt;/a&gt;&lt;/strong&gt; - AI platform templates&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/voltagent/awesome-claude-code-subagents" target="_blank" rel="noopener"&gt;Awesome Claude Code Subagents&lt;/a&gt;&lt;/strong&gt; - Claude Code subagents&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-courses--tutorials-9-curricula"&gt;🎓 Courses &amp;amp; Tutorials (9 Curricula)&lt;/h2&gt;
&lt;p&gt;Structured learning paths from universities and tech companies.&lt;/p&gt;
&lt;h3 id="microsofts-ai-curriculum"&gt;Microsoft&amp;rsquo;s AI Curriculum&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ai-for-beginners" target="_blank" rel="noopener"&gt;AI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;12 weeks, 24 lessons covering neural networks, deep learning, CV, NLP&lt;/li&gt;
&lt;li&gt;Great for: Complete AI foundation&lt;/li&gt;
&lt;li&gt;Format: Lessons, quizzes, projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ml-for-beginners" target="_blank" rel="noopener"&gt;Machine Learning for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;12-week, 26-lesson curriculum on classic ML&lt;/li&gt;
&lt;li&gt;Great for: ML fundamentals without deep math&lt;/li&gt;
&lt;li&gt;Format: Project-based exercises&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/generative-ai-for-beginners" target="_blank" rel="noopener"&gt;Generative AI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;18 lessons on building GenAI applications&lt;/li&gt;
&lt;li&gt;Great for: Practical GenAI development&lt;/li&gt;
&lt;li&gt;Format: Hands-on projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ai-agents-for-beginners" target="_blank" rel="noopener"&gt;AI Agents for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;11 lessons on agent systems&lt;/li&gt;
&lt;li&gt;Great for: Understanding autonomous agents&lt;/li&gt;
&lt;li&gt;Format: Project-driven learning&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/edgeai-for-beginners" target="_blank" rel="noopener"&gt;EdgeAI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Optimization, deployment, and real-world Edge AI&lt;/li&gt;
&lt;li&gt;Great for: On-device AI applications&lt;/li&gt;
&lt;li&gt;Format: Practical tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/mcp-for-beginners" target="_blank" rel="noopener"&gt;MCP for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model Context Protocol curriculum&lt;/li&gt;
&lt;li&gt;Great for: Building with MCP&lt;/li&gt;
&lt;li&gt;Format: Cross-language examples and labs&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="official-platform-courses"&gt;Official Platform Courses&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/course" target="_blank" rel="noopener"&gt;Hugging Face Learn Center&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Free courses on LLMs, deep RL, CV, audio&lt;/li&gt;
&lt;li&gt;Great for: Hands-on Hugging Face ecosystem&lt;/li&gt;
&lt;li&gt;Format: Interactive notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/openai/openai-cookbook" target="_blank" rel="noopener"&gt;OpenAI Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Runnable examples using OpenAI API&lt;/li&gt;
&lt;li&gt;Great for: OpenAI API best practices&lt;/li&gt;
&lt;li&gt;Format: Code examples and guides&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/pytorch/tutorials" target="_blank" rel="noopener"&gt;PyTorch Tutorials&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Basics to advanced deep learning&lt;/li&gt;
&lt;li&gt;Great for: PyTorch mastery&lt;/li&gt;
&lt;li&gt;Format: Comprehensive tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-cookbooks--example-collections-5-collections"&gt;🍳 Cookbooks &amp;amp; Example Collections (5 Collections)&lt;/h2&gt;
&lt;p&gt;Practical code examples and recipes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/anthropics/claude-cookbooks" target="_blank" rel="noopener"&gt;Claude Cookbooks&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Notebooks and examples for building with Claude&lt;/li&gt;
&lt;li&gt;Great for: Anthropic Claude integration&lt;/li&gt;
&lt;li&gt;Format: Jupyter notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/cookbook" target="_blank" rel="noopener"&gt;Hugging Face Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Practical AI cookbook with Jupyter notebooks&lt;/li&gt;
&lt;li&gt;Great for: Open models and tools&lt;/li&gt;
&lt;li&gt;Format: Hands-on examples&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/RationaleInstitute/tinker-cookbook" target="_blank" rel="noopener"&gt;Tinker Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training and fine-tuning examples&lt;/li&gt;
&lt;li&gt;Great for: Fine-tuning workflows&lt;/li&gt;
&lt;li&gt;Format: Platform-specific recipes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/e2b-dev/e2b-cookbook" target="_blank" rel="noopener"&gt;E2B Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Examples for building LLM apps&lt;/li&gt;
&lt;li&gt;Great for: LLM application development&lt;/li&gt;
&lt;li&gt;Format: Recipes and tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/jamwithai/arxiv-paper-curator" target="_blank" rel="noopener"&gt;arXiv Paper Curator&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;6-week course on RAG systems&lt;/li&gt;
&lt;li&gt;Great for: Production-ready RAG&lt;/li&gt;
&lt;li&gt;Format: Project-based learning&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-guides--handbooks-5-resources"&gt;📖 Guides &amp;amp; Handbooks (5 Resources)&lt;/h2&gt;
&lt;p&gt;In-depth guides on specific topics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dair-ai/prompt-engineering-guide" target="_blank" rel="noopener"&gt;Prompt Engineering Guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Comprehensive prompt engineering resources&lt;/li&gt;
&lt;li&gt;Great for: Mastering prompt design&lt;/li&gt;
&lt;li&gt;Format: Guides, papers, lectures, notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/evaluation-guidebook" target="_blank" rel="noopener"&gt;Evaluation Guidebook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM evaluation best practices from Hugging Face&lt;/li&gt;
&lt;li&gt;Great for: Assessing LLM performance&lt;/li&gt;
&lt;li&gt;Format: Practical guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/davidkimai/context-engineering" target="_blank" rel="noopener"&gt;Context Engineering&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Design and optimize context beyond prompt engineering&lt;/li&gt;
&lt;li&gt;Great for: Advanced context management&lt;/li&gt;
&lt;li&gt;Format: Practical handbook&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/coleam00/context-engineering-intro" target="_blank" rel="noopener"&gt;Context Engineering Intro&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Template and guide for context engineering&lt;/li&gt;
&lt;li&gt;Great for: Providing project context to AI assistants&lt;/li&gt;
&lt;li&gt;Format: Template + guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/IIETER/IIETER" target="_blank" rel="noopener"&gt;Vibe-Coding Workflow&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;5-step prompt template for building MVPs with LLMs&lt;/li&gt;
&lt;li&gt;Great for: Rapid prototyping with AI&lt;/li&gt;
&lt;li&gt;Format: Workflow template&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-template--workflow-collections"&gt;🗂️ Template &amp;amp; Workflow Collections&lt;/h2&gt;
&lt;p&gt;Reusable templates and workflows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/davila7/claude-code-templates" target="_blank" rel="noopener"&gt;Claude Code Templates&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Code templates for various programming scenarios&lt;/li&gt;
&lt;li&gt;Great for: Claude AI development&lt;/li&gt;
&lt;li&gt;Format: Template collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/zie619/n8n-workflows" target="_blank" rel="noopener"&gt;n8n Workflows&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2,000+ professionally organized n8n workflows&lt;/li&gt;
&lt;li&gt;Great for: Workflow automation&lt;/li&gt;
&lt;li&gt;Format: Searchable catalog&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/nusquama/n8nworkflows.xyz" target="_blank" rel="noopener"&gt;N8N Workflows Catalog&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Community-driven reusable workflow templates&lt;/li&gt;
&lt;li&gt;Great for: Workflow import and versioning&lt;/li&gt;
&lt;li&gt;Format: Template catalog&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-research--evaluation"&gt;📊 Research &amp;amp; Evaluation&lt;/h2&gt;
&lt;p&gt;Academic and evaluation resources.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/amberljc/llmsys-paperlist" target="_blank" rel="noopener"&gt;LLMSys PaperList&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Curated list of LLM systems papers&lt;/li&gt;
&lt;li&gt;Great for: Research on training, inference, serving&lt;/li&gt;
&lt;li&gt;Format: Paper collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/cheahjs/free-llm-api-resources" target="_blank" rel="noopener"&gt;Free LLM API Resources&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM providers with free/trial API access&lt;/li&gt;
&lt;li&gt;Great for: Experimentation without cost&lt;/li&gt;
&lt;li&gt;Format: Provider list&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-other-notable-resources"&gt;🎨 Other Notable Resources&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/x1xhlol/system-prompts-and-models-of-ai-tools" target="_blank" rel="noopener"&gt;System Prompts and Models of AI Tools&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Community-curated collection of system prompts and AI tool examples&lt;/li&gt;
&lt;li&gt;Great for: Prompt and agent engineering&lt;/li&gt;
&lt;li&gt;Format: Resource collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/epfml/ml_course" target="_blank" rel="noopener"&gt;ML Course CS-433&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;EPFL Machine Learning Course&lt;/li&gt;
&lt;li&gt;Great for: Academic ML foundation&lt;/li&gt;
&lt;li&gt;Format: Lectures, labs, projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/stas00/ml-engineering" target="_blank" rel="noopener"&gt;Machine Learning Engineering&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ML engineering open-book: compute, storage, networking&lt;/li&gt;
&lt;li&gt;Great for: Production ML systems&lt;/li&gt;
&lt;li&gt;Format: Comprehensive guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/neural-maze/realtime-phone-agents-course" target="_blank" rel="noopener"&gt;Realtime Phone Agents Course&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Build low-latency voice agents&lt;/li&gt;
&lt;li&gt;Great for: Voice AI applications&lt;/li&gt;
&lt;li&gt;Format: Hands-on course&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/johnma2006/m3-workshop" target="_blank" rel="noopener"&gt;LLMs from Scratch&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Build a working LLM from first principles&lt;/li&gt;
&lt;li&gt;Great for: Understanding LLM internals&lt;/li&gt;
&lt;li&gt;Format: Repository + book materials&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-how-to-use-this-collection"&gt;💡 How to Use This Collection&lt;/h2&gt;
&lt;h3 id="for-complete-beginners"&gt;For Complete Beginners&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Start with&lt;/strong&gt;: Microsoft&amp;rsquo;s AI for Beginners&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Practice with&lt;/strong&gt;: PyTorch Tutorials&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explore&lt;/strong&gt;: Awesome AI Apps for inspiration&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-developers"&gt;For Developers&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Build skills&lt;/strong&gt;: OpenAI Cookbook + Claude Cookbooks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find tools&lt;/strong&gt;: Awesome Generative AI + Awesome LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn workflows&lt;/strong&gt;: n8n Workflows Catalog&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-researchers"&gt;For Researchers&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Stay updated&lt;/strong&gt;: Awesome Generative AI + LLMSys PaperList&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deep dive&lt;/strong&gt;: Awesome LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Implement&lt;/strong&gt;: Hugging Face Cookbook&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-product-builders"&gt;For Product Builders&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Find examples&lt;/strong&gt;: Awesome AI Apps&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn workflows&lt;/strong&gt;: n8n Workflows Catalog&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Study patterns&lt;/strong&gt;: Awesome LLM Apps&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="-what-was-not-removed"&gt;🔄 What Was NOT Removed&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Agent frameworks and production tools remain in the AI Resources list&lt;/strong&gt;, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AutoGen&lt;/strong&gt; - Microsoft&amp;rsquo;s multi-agent framework&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CrewAI&lt;/strong&gt; - High-performance multi-agent orchestration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LangGraph&lt;/strong&gt; - Stateful multi-agent applications&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flowise&lt;/strong&gt; - Visual agent platform&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Langflow&lt;/strong&gt; - Visual workflow builder&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;And 80+ more agent frameworks&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are &lt;strong&gt;functional tools&lt;/strong&gt; you can use to build applications, not educational collections. They belong in the AI Resources list.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="-summary"&gt;📝 Summary&lt;/h2&gt;
&lt;p&gt;I removed &lt;strong&gt;44 collection-type projects&lt;/strong&gt; from the AI Resources list to keep it focused on production tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;14 Awesome Lists&lt;/strong&gt; - Discover new tools and stay updated&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;9 Courses &amp;amp; Tutorials&lt;/strong&gt; - Structured learning paths&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;5 Cookbooks&lt;/strong&gt; - Practical code examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;5 Guides &amp;amp; Handbooks&lt;/strong&gt; - In-depth resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;4 Template Collections&lt;/strong&gt; - Reusable workflows&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;7 Other Resources&lt;/strong&gt; - Research and evaluation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These resources remain &lt;strong&gt;incredibly valuable&lt;/strong&gt; for learning and discovery. They just serve a different purpose than the production-focused tools in my AI Resources list.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Next Steps&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Bookmark this post for future reference&lt;/li&gt;
&lt;li&gt;Explore the AI Resources list for production tools (agent frameworks, databases, etc.)&lt;/li&gt;
&lt;li&gt;Check out my blog for more AI engineering insights&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Acknowledgments&lt;/strong&gt;: This collection was compiled during my AI Resources cleanup initiative. Special thanks to all the maintainers of these awesome lists, courses, and collections for their invaluable contributions to the AI community.&lt;/p&gt;</content:encoded></item><item><title>Standing on Giants' Shoulders: The Traditional Infrastructure Powering Modern AI</title><link>https://jimmysong.io/blog/giants-beneath-ai-feet/</link><pubDate>Sun, 08 Feb 2026 08:00:00 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/giants-beneath-ai-feet/</guid><description>Before ChatGPT and TensorFlow, there was Hadoop, Kafka, and Kubernetes. This post honors the traditional open source infrastructure that became the foundation of today&amp;#39;s AI revolution.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;If I have seen further, it is by standing on the shoulders of giants.&amp;rdquo; — Isaac Newton&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/giants-beneath-ai-feet/banner.webp" data-img="https://assets.jimmysong.io/images/blog/giants-beneath-ai-feet/banner.webp" alt="Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI" data-caption="Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the excitement surrounding LLMs, vector databases, and AI agents, it&amp;rsquo;s easy to forget that modern AI didn&amp;rsquo;t emerge from a vacuum. Today&amp;rsquo;s AI revolution stands upon decades of infrastructure work—distributed systems, data pipelines, search engines, and orchestration platforms that were built long before &amp;ldquo;AI Native&amp;rdquo; became a buzzword.&lt;/p&gt;
&lt;p&gt;This post is a tribute to those traditional open source projects that became the invisible foundation of AI infrastructure. They&amp;rsquo;re not &amp;ldquo;AI projects&amp;rdquo; per se, but without them, the AI revolution as we know it wouldn&amp;rsquo;t exist.&lt;/p&gt;
&lt;h2 id="the-evolution-from-big-data-to-ai"&gt;The Evolution: From Big Data to AI&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Era&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;Core Technologies&lt;/th&gt;
&lt;th&gt;AI Connection&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2000s&lt;/td&gt;
&lt;td&gt;Web Search &amp;amp; Indexing&lt;/td&gt;
&lt;td&gt;Lucene, Elasticsearch&lt;/td&gt;
&lt;td&gt;Semantic search foundations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Big Data &amp;amp; Distributed Computing&lt;/td&gt;
&lt;td&gt;Hadoop, Spark, Kafka&lt;/td&gt;
&lt;td&gt;Data processing at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Cloud Native&lt;/td&gt;
&lt;td&gt;Docker, Kubernetes&lt;/td&gt;
&lt;td&gt;Model deployment platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Stream Processing&lt;/td&gt;
&lt;td&gt;Flink, Storm, Pulsar&lt;/td&gt;
&lt;td&gt;Real-time ML inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2020s&lt;/td&gt;
&lt;td&gt;AI Native&lt;/td&gt;
&lt;td&gt;Transformers, Vector DBs&lt;/td&gt;
&lt;td&gt;Built on everything above&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Evolution of Data Infrastructure
&lt;/figcaption&gt;
&lt;h2 id="big-data-frameworks-the-data-engines"&gt;Big Data Frameworks: The Data Engines&lt;/h2&gt;
&lt;p&gt;Before we could train models on petabytes of data, we needed ways to store, process, and move that data.&lt;/p&gt;
&lt;h3 id="apache-hadoop-2006"&gt;Apache Hadoop (2006)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/hadoop" target="_blank" rel="noopener"&gt;https://github.com/apache/hadoop&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Hadoop democratized big data by making distributed computing accessible. Its HDFS filesystem and MapReduce paradigm proved that commodity hardware could process web-scale datasets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Modern ML training datasets live in HDFS-compatible storage&lt;/li&gt;
&lt;li&gt;Data lakes built on Hadoop became training data reservoirs&lt;/li&gt;
&lt;li&gt;Proved that distributed computing could scale horizontally&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="apache-kafka-2011"&gt;Apache Kafka (2011)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/kafka" target="_blank" rel="noopener"&gt;https://github.com/apache/kafka&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Kafka redefined data streaming with its log-based architecture. It became the nervous system for real-time data flows in enterprises worldwide.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Real-time feature pipelines for ML models&lt;/li&gt;
&lt;li&gt;Event-driven architectures for AI agent systems&lt;/li&gt;
&lt;li&gt;Streaming inference pipelines&lt;/li&gt;
&lt;li&gt;Model telemetry and monitoring backbones&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="apache-spark-2014"&gt;Apache Spark (2014)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/spark" target="_blank" rel="noopener"&gt;https://github.com/apache/spark&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Spark brought in-memory computing to big data, making iterative algorithms (like ML training) practical at scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MLlib made ML accessible to data engineers&lt;/li&gt;
&lt;li&gt;Distributed data processing for model training&lt;/li&gt;
&lt;li&gt;Spark ML became the de facto standard for big data ML&lt;/li&gt;
&lt;li&gt;Proved that in-memory computing could accelerate ML workloads&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="search-engines-the-retrieval-foundation"&gt;Search Engines: The Retrieval Foundation&lt;/h2&gt;
&lt;p&gt;Before RAG (Retrieval-Augmented Generation) became a buzzword, search engines were solving retrieval at scale.&lt;/p&gt;
&lt;h3 id="elasticsearch-2010"&gt;Elasticsearch (2010)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/elastic/elasticsearch" target="_blank" rel="noopener"&gt;https://github.com/elastic/elasticsearch&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Elasticsearch made full-text search accessible and scalable. Its distributed architecture and RESTful API became the standard for search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;pioneered distributed inverted index structures&lt;/li&gt;
&lt;li&gt;Proved that horizontal scaling was possible for search workloads&lt;/li&gt;
&lt;li&gt;Many &amp;ldquo;AI search&amp;rdquo; systems actually use Elasticsearch under the hood&lt;/li&gt;
&lt;li&gt;Query DSL influenced modern vector database query languages&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="opensearch-2021"&gt;OpenSearch (2021)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/opensearch-project/opensearch" target="_blank" rel="noopener"&gt;https://github.com/opensearch-project/opensearch&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;When AWS forked Elasticsearch, it ensured search infrastructure remained truly open. OpenSearch continues the mission of accessible, scalable search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maintains open source innovation in search&lt;/li&gt;
&lt;li&gt;Vector search capabilities added in 2023&lt;/li&gt;
&lt;li&gt;Demonstrates community fork resilience&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="databases-from-sql-to-vectors"&gt;Databases: From SQL to Vectors&lt;/h2&gt;
&lt;p&gt;The evolution from relational databases to vector databases represents a paradigm shift—but both have AI relevance.&lt;/p&gt;
&lt;h3 id="traditional-databases-that-paved-the-way"&gt;Traditional Databases That Paved the Way&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dgraph&lt;/strong&gt; (2015) - Graph database proving that specialized data structures enable new use cases&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TDengine&lt;/strong&gt; (2019) - Time-series database for IoT ML workloads&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OceanBase&lt;/strong&gt; (2021) - Distributed database showing ACID transactions could scale&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proved that specialized database engines could outperform general-purpose ones&lt;/li&gt;
&lt;li&gt;Database internals (indexing, sharding, replication) are now applied to vector databases&lt;/li&gt;
&lt;li&gt;Multi-model databases (graph + vector + relational) are becoming the norm for AI apps&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="cloud-native-the-runtime-foundation"&gt;Cloud Native: The Runtime Foundation&lt;/h2&gt;
&lt;p&gt;When Docker and Kubernetes emerged, they weren&amp;rsquo;t built for AI—but AI couldn&amp;rsquo;t scale without them.&lt;/p&gt;
&lt;h3 id="docker-2013--kubernetes-2014"&gt;Docker (2013) &amp;amp; Kubernetes (2014)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/kubernetes/kubernetes" target="_blank" rel="noopener"&gt;https://github.com/kubernetes/kubernetes&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes became the operating system for cloud-native applications. Its declarative API and controller pattern made it perfect for AI workloads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model deployment platforms (KServe, Seldon Core) run on K8s&lt;/li&gt;
&lt;li&gt;GPU orchestration (NVIDIA GPU Operator, Volcano, HAMi) extends K8s&lt;/li&gt;
&lt;li&gt;Kubeflow made K8s the standard for ML pipelines&lt;/li&gt;
&lt;li&gt;Microservice patterns enable modular AI agent architectures&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="service-mesh--serverless"&gt;Service Mesh &amp;amp; Serverless&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Istio&lt;/strong&gt; (2016), &lt;strong&gt;Knative&lt;/strong&gt; (2018) - Service mesh and serverless platforms that proved:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Network-level observability applies to AI model calls&lt;/li&gt;
&lt;li&gt;Scale-to-zero is essential for cost-effective inference&lt;/li&gt;
&lt;li&gt;Traffic splitting enables A/B testing of ML models&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Gateway patterns evolved from API gateways + service mesh&lt;/li&gt;
&lt;li&gt;Serverless inference platforms use Knative-style autoscaling&lt;/li&gt;
&lt;li&gt;Observability patterns (tracing, metrics) are now standard for ML systems&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="api-gateways-from-rest-to-llm"&gt;API Gateways: From REST to LLM&lt;/h2&gt;
&lt;p&gt;API gateways weren&amp;rsquo;t designed for AI, but they became the foundation of AI Gateway patterns.&lt;/p&gt;
&lt;h3 id="kong-apisix-kgateway"&gt;Kong, APISIX, KGateway&lt;/h3&gt;
&lt;p&gt;These API gateways solved rate limiting, auth, and routing at scale. When LLMs emerged, the same patterns applied:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Gateway Evolution&lt;/strong&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Traditional API Gateway (2010s)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Rate Limiting → Token Bucket Rate Limiting
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Auth → API Key + Organization Management
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Routing → Model Routing (GPT-4 → Claude → Local Models)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Observability → LLM-specific Telemetry (token usage, cost)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI Gateway (2024)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proved that centralized API management scales&lt;/li&gt;
&lt;li&gt;Plugin architectures enable LLM-specific features&lt;/li&gt;
&lt;li&gt;Traffic management patterns apply to prompt routing&lt;/li&gt;
&lt;li&gt;Security patterns (mTLS, JWT) now protect AI endpoints&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="workflow-orchestration-the-pipeline-backbone"&gt;Workflow Orchestration: The Pipeline Backbone&lt;/h2&gt;
&lt;p&gt;Data engineering needs pipelines. ML engineering needs pipelines. AI agents need workflows.&lt;/p&gt;
&lt;h3 id="apache-airflow-2015"&gt;Apache Airflow (2015)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/airflow" target="_blank" rel="noopener"&gt;https://github.com/apache/airflow&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Airflow made pipeline orchestration accessible with its DAG-based approach. It became the standard for ETL and data engineering.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ML pipeline orchestration (feature engineering, training, evaluation)&lt;/li&gt;
&lt;li&gt;Proved that DAG-based workflow definition works at scale&lt;/li&gt;
&lt;li&gt;Prompt engineering pipelines use Airflow-style orchestration&lt;/li&gt;
&lt;li&gt;Scheduler patterns are now applied to AI agent workflows&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="n8n-prefect-flyte"&gt;n8n, Prefect, Flyte&lt;/h3&gt;
&lt;p&gt;Modern workflow platforms that evolved from Airflow&amp;rsquo;s foundations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;n8n&lt;/strong&gt; (2019) - Visual workflow automation with AI capabilities&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prefect&lt;/strong&gt; (2018) - Python-native workflow orchestration for ML&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flyte&lt;/strong&gt; (2019) - Kubernetes-native workflow orchestration for ML/data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-modal agents need workflow orchestration&lt;/li&gt;
&lt;li&gt;RAG pipelines are essentially ETL pipelines for embeddings&lt;/li&gt;
&lt;li&gt;Prompt chaining is DAG-based orchestration&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="data-formats-the-lakehouse-foundation"&gt;Data Formats: The Lakehouse Foundation&lt;/h2&gt;
&lt;p&gt;Before we could train on massive datasets, we needed formats that supported ACID transactions and schema evolution.&lt;/p&gt;
&lt;h3 id="delta-lake-apache-iceberg-apache-hudi"&gt;Delta Lake, Apache Iceberg, Apache Hudi&lt;/h3&gt;
&lt;p&gt;These table formats brought reliability to data lakes:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training datasets need versioning and reproducibility&lt;/li&gt;
&lt;li&gt;Feature stores use Delta/Iceberg as storage formats&lt;/li&gt;
&lt;li&gt;Proved that &amp;ldquo;big data&amp;rdquo; could have transactional semantics&lt;/li&gt;
&lt;li&gt;Schema evolution handles ML feature drift&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-invisible-thread-why-these-projects-matter"&gt;The Invisible Thread: Why These Projects Matter&lt;/h2&gt;
&lt;p&gt;What do all these projects have in common?&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;They solved scaling first&lt;/strong&gt; - AI training/inference needs horizontal scaling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They proved distributed systems work&lt;/strong&gt; - Modern AI is fundamentally distributed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They created ecosystem patterns&lt;/strong&gt; - Plugin systems, extension points, APIs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They established best practices&lt;/strong&gt; - Observability, security, CI/CD&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They built developer habits&lt;/strong&gt; - YAML configs, declarative APIs, CLI tools&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="the-ai-native-continuum"&gt;The AI Native Continuum&lt;/h2&gt;
&lt;p&gt;Modern &amp;ldquo;AI Native&amp;rdquo; infrastructure didn&amp;rsquo;t replace these projects—it builds on them:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Traditional Project&lt;/th&gt;
&lt;th&gt;AI Native Evolution&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hadoop HDFS&lt;/td&gt;
&lt;td&gt;Distributed model storage&lt;/td&gt;
&lt;td&gt;HDFS for datasets, S3 for checkpoints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kafka&lt;/td&gt;
&lt;td&gt;Real-time feature pipelines&lt;/td&gt;
&lt;td&gt;Kafka → Feature Store → Model Serving&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spark ML&lt;/td&gt;
&lt;td&gt;Distributed ML training&lt;/td&gt;
&lt;td&gt;MLlib → PyTorch Distributed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Elasticsearch&lt;/td&gt;
&lt;td&gt;Vector search&lt;/td&gt;
&lt;td&gt;ES → Weaviate/Qdrant/Milvus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubernetes&lt;/td&gt;
&lt;td&gt;ML orchestration&lt;/td&gt;
&lt;td&gt;K8s → Kubeflow/KServe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Istio&lt;/td&gt;
&lt;td&gt;AI Gateway service mesh&lt;/td&gt;
&lt;td&gt;Istio → LLM Gateway with mTLS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Airflow&lt;/td&gt;
&lt;td&gt;ML pipeline orchestration&lt;/td&gt;
&lt;td&gt;Airflow → Prefect/Flyte for ML&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: From Traditional to AI Native
&lt;/figcaption&gt;
&lt;h2 id="why-were-removing-them-from-ai-resources-list"&gt;Why We&amp;rsquo;re Removing Them from AI Resources List&lt;/h2&gt;
&lt;p&gt;This post honors these projects, but we&amp;rsquo;re also removing them from our AI Resources list. Here&amp;rsquo;s why:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;They&amp;rsquo;re not &amp;ldquo;AI Projects&amp;rdquo;—they&amp;rsquo;re foundational infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hadoop, Kafka, Spark&lt;/strong&gt; are data engineering tools, not ML frameworks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Elasticsearch&lt;/strong&gt; is search, not semantic search&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kubernetes&lt;/strong&gt; is general-purpose orchestration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API gateways&lt;/strong&gt; serve REST/GraphQL, not just LLMs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;But their absence doesn&amp;rsquo;t diminish their importance.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;By removing them, we acknowledge that:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;AI has its own ecosystem&lt;/strong&gt; - Transformers, vector DBs, LLM ops&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Traditional infra has its own domain&lt;/strong&gt; - Data engineering, cloud native&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The intersection is where innovation happens&lt;/strong&gt; - AI-native data platforms, LLM ops on K8s&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="the-giants-we-stand-on"&gt;The Giants We Stand On&lt;/h2&gt;
&lt;p&gt;The next time you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deploy a model on Kubernetes&lt;/li&gt;
&lt;li&gt;Stream features through Kafka&lt;/li&gt;
&lt;li&gt;Search embeddings with a vector database&lt;/li&gt;
&lt;li&gt;Orchestrate a RAG pipeline with Prefect&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Remember: You&amp;rsquo;re standing on the shoulders of Hadoop, Kafka, Elasticsearch, Kubernetes, and countless others. They built the roads we now drive on.&lt;/p&gt;
&lt;h2 id="the-future-building-new-giants"&gt;The Future: Building New Giants&lt;/h2&gt;
&lt;p&gt;Just as Hadoop and Kafka enabled modern AI, today&amp;rsquo;s AI infrastructure will become tomorrow&amp;rsquo;s foundation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Vector databases&lt;/strong&gt; may become the new standard for all search&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM observability&lt;/strong&gt; may evolve into general distributed tracing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI agent orchestration&lt;/strong&gt; may reinvent workflow automation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU scheduling&lt;/strong&gt; may influence general-purpose resource management&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The cycle continues. The giants of today will be the foundations of tomorrow.&lt;/p&gt;
&lt;h2 id="conclusion-gratitude-and-continuity"&gt;Conclusion: Gratitude and Continuity&lt;/h2&gt;
&lt;p&gt;As we clean up our AI Resources list to focus on AI-native projects, we don&amp;rsquo;t forget where we came from. Traditional big data and cloud native infrastructure made the AI revolution possible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;To the Hadoop committers, Kafka maintainers, Kubernetes contributors, and all who built the foundation: Thank you.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Your work enabled ChatGPT, enabled Transformers, enabled everything we now call &amp;ldquo;AI.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Standing on your shoulders, we see further.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Acknowledgments&lt;/strong&gt;: This post was inspired by the need to refactor our AI Resources list. The 27 projects mentioned here are being removed—not because they&amp;rsquo;re unimportant, but because they deserve their own category: &lt;strong&gt;The Foundation&lt;/strong&gt;.&lt;/p&gt;</content:encoded></item><item><title>My First Month at Dynamia: Why AI Native Infra Is Worth the Investment</title><link>https://jimmysong.io/blog/why-i-join-dynamia-ai-native-infra/</link><pubDate>Fri, 06 Feb 2026 12:56:35 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/why-i-join-dynamia-ai-native-infra/</guid><description>Observations from my first month at Dynamia: From cloud native to AI Native Infra, why this direction is worth investing in, and the key issues and opportunities in compute governance.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Time flies—it&amp;rsquo;s already been a month since I joined Dynamia. In this article, I want to share my observations from this past month: why AI Native Infra is a direction worth investing in, and some considerations for those thinking about their own career or technical direction.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;After nearly five years of remote work, I officially joined &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; last month as VP of Open Source Ecosystem. This decision was not sudden, but a natural extension of my journey from cloud native to AI Native Infra.&lt;/p&gt;
&lt;p&gt;But this article is not just about my personal choice. I want to answer a more universal question: &lt;strong&gt;In the wave of AI infrastructure startups, why is compute governance a direction worth investing in?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For the past decade, I have worked continuously in the infrastructure space: from Kubernetes to Service Mesh, and now to AI Infra. I am increasingly convinced that the core challenge in the AI era is not &amp;ldquo;can the model run,&amp;rdquo; but &amp;ldquo;can compute resources be run efficiently, reliably, and in a controlled manner.&amp;rdquo; This conviction has only grown stronger through my observations and reflections during this first month at Dynamia.&lt;/p&gt;
&lt;p&gt;This article answers three questions: What is AI Native Infra? Why is GPU virtualization a necessity? Why did I choose Dynamia and HAMi?&lt;/p&gt;
&lt;h2 id="what-is-ai-native-infra"&gt;What Is AI Native Infra&lt;/h2&gt;
&lt;p&gt;The core of &lt;a href="https://jimmysong.io/book/ai-native-infra/"&gt;AI Native Infrastructure&lt;/a&gt; is not about adding another platform layer, but about redefining the governance target: expanding from &amp;ldquo;services and containers&amp;rdquo; to &amp;ldquo;model behaviors and compute assets.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I summarize it as three key shifts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Models as execution entities&lt;/strong&gt;: Governance now includes not just processes, but also model behaviors.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compute as a scarce asset&lt;/strong&gt;: GPU, memory, and bandwidth must be scheduled and metered precisely.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Uncertainty as the default&lt;/strong&gt;: Systems must remain observable and recoverable amid fluctuations.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In essence, AI Native Infra is about upgrading compute governance from &amp;ldquo;resource allocation&amp;rdquo; to &amp;ldquo;sustainable business capability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="why-gpu-virtualization-is-essential"&gt;Why GPU Virtualization Is Essential&lt;/h2&gt;
&lt;p&gt;Many teams focus on model inference optimization, but in production, enterprises first encounter the problem of &amp;ldquo;underutilized GPUs.&amp;rdquo; This is where GPU virtualization delivers value.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Structural idleness&lt;/strong&gt;: Small tasks monopolize large GPUs, leaving them idle for long periods.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pseudo-isolation risks&lt;/strong&gt;: Native sharing lacks hard boundaries, so a single task OOM can cause cascading failures.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling failures&lt;/strong&gt;: Some users queue for GPUs while others occupy but do not use them, leading to both shortages and idleness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fragmentation waste&lt;/strong&gt;: There may be enough total GPU, but not enough full cards, making efficient packing impossible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vendor lock-in anxiety&lt;/strong&gt;: Proprietary, tightly coupled solutions make migration costs uncontrollable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In short: GPUs must not only be allocatable, but also splittable, isolatable, schedulable, and governable.&lt;/p&gt;
&lt;h2 id="the-relationship-between-hami-and-dynamia"&gt;The Relationship Between HAMi and Dynamia&lt;/h2&gt;
&lt;p&gt;This is the most frequently asked question. Here is the shortest answer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi&lt;/strong&gt;: A CNCF-hosted open source project and community focused on GPU virtualization and heterogeneous compute scheduling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamia&lt;/strong&gt;: The founding and leading company behind HAMi, providing enterprise-grade products and services based on HAMi.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open source projects are not the same as company products, but the two evolve together. HAMi drives industry adoption and technical trust, while Dynamia brings these capabilities into enterprise production environments at scale. This &amp;ldquo;dual engine&amp;rdquo; approach is what makes Dynamia unique.&lt;/p&gt;
&lt;h2 id="what-hami-provides"&gt;What HAMi Provides&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; (&lt;em&gt;Heterogeneous AI Computing Virtualization Middleware&lt;/em&gt;) delivers three key capabilities on Kubernetes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Virtualization and partitioning&lt;/strong&gt;: Split physical GPUs into logical resources on demand to improve utilization.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling and topology awareness&lt;/strong&gt;: Place workloads optimally based on topology to reduce communication bottlenecks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Isolation and observability&lt;/strong&gt;: Support quotas, policies, and monitoring to reduce production risks.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Currently, HAMi has attracted over 360 contributors from 16 countries, with more than 200 enterprise end users, and its international influence continues to grow.&lt;/p&gt;
&lt;h2 id="market-trends-the-ai-infrastructure-startup-wave"&gt;Market Trends: The AI Infrastructure Startup Wave&lt;/h2&gt;
&lt;p&gt;AI infrastructure is experiencing a new wave of startups. The vLLM team&amp;rsquo;s company raised $150 million, SGLang&amp;rsquo;s commercial spin-off RadixArk is valued at $4 billion, and Databricks acquired MosaicML for $1.3 billion—all pointing to a consensus: &lt;strong&gt;Whoever helps enterprises run large models more efficiently and cost-effectively will hold the keys to next-generation AI infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Against this backdrop, &lt;strong&gt;the positioning of Dynamia and HAMi&lt;/strong&gt; is even clearer. Many teams focus on &amp;ldquo;model performance acceleration&amp;rdquo; and &amp;ldquo;inference optimization&amp;rdquo; (like vLLM, SGLang), while we focus on &lt;strong&gt;&amp;ldquo;resource scheduling and virtualization&amp;rdquo;&lt;/strong&gt;—enabling better orchestration of existing accelerated hardware resources.&lt;/p&gt;
&lt;p&gt;The two are complementary: the former makes individual models run faster and cheaper, while the latter ensures that compute allocation at the cluster level is efficient, fair, and controllable. This is similar to extending Kubernetes&amp;rsquo; CPU/memory scheduling philosophy to GPU and heterogeneous compute management in the AI era.&lt;/p&gt;
&lt;h2 id="why-ai-native-infra-is-worth-the-investment"&gt;Why AI Native Infra Is Worth the Investment&lt;/h2&gt;
&lt;p&gt;My observations this month have convinced me that &lt;strong&gt;compute governance is the most undervalued yet most promising area in AI infrastructure&lt;/strong&gt;. If you are considering a career or technical investment, here is my assessment:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, this is a real and urgent pain point&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Model training and inference optimization attract a lot of attention, but in production, enterprises first encounter the problem of &amp;ldquo;underutilized GPUs&amp;rdquo;—structural idleness, scheduling failures, fragmentation waste, and vendor lock-in anxiety. Without solving these problems, even the fastest models cannot scale in production. GPU virtualization and heterogeneous compute scheduling are the &amp;ldquo;infrastructure below infrastructure&amp;rdquo; for enterprise AI transformation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, this is a clear long-term track&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Frameworks like vLLM and SGLang emerge constantly, making individual models run faster. But who ensures that compute allocation at the cluster level is efficient, fair, and controllable? This is similar to extending Kubernetes&amp;rsquo; success in CPU/memory scheduling to GPU and heterogeneous compute management in the AI era. This is not something that can be finished in a year or two, but a direction for continuous construction over the next five to ten years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Third, this is an open and verifiable path&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Dynamia chose to build on HAMi as an open source foundation, first solving general capabilities, then supporting enterprise adoption. This means the technical direction is transparent and verifiable in the community. You can form your own judgment by participating in open source, observing adoption, and evaluating the ecosystem—rather than relying on the black-box promises of proprietary solutions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fourth, this is a window of opportunity that is opening now&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;AI infrastructure is being redefined. Investing in its construction today will continue to yield value in the coming years. The vLLM team&amp;rsquo;s company raised $150 million, SGLang&amp;rsquo;s commercial spin-off RadixArk is valued at $4 billion, Databricks acquired MosaicML for $1.3 billion—all validating the same trend: &lt;strong&gt;Whoever helps enterprises run large models more efficiently will hold the keys to next-generation AI infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I hope to bring my experience in cloud native and open source communities to the next stage of HAMi and Dynamia: turning GPU resources from a &amp;ldquo;cost center&amp;rdquo; into an &amp;ldquo;operational asset.&amp;rdquo; This is not just my career choice, but my judgment and investment in the direction of next-generation infrastructure.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Join the HAMi Community
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Add me on WeChat (&lt;code&gt;jimmysong&lt;/code&gt;) to join the &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; community focused on GPU virtualization and heterogeneous compute scheduling.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;If you are also interested in HAMi, GPU virtualization, AI Native Infra, or Dynamia, feel free to reach out.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;From cloud native to AI Native Infra, my observations this month have only strengthened my conviction: &lt;strong&gt;The true upper limit of AI applications is determined by the infrastructure&amp;rsquo;s ability to govern compute resources.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;HAMi addresses the fundamental issues of GPU virtualization and heterogeneous compute scheduling, while Dynamia is driving these capabilities into large-scale production. If you are also looking for a technical direction worth long-term investment, AI Native Infra—especially compute governance and scheduling—is a track with real pain points, a clear path, an open ecosystem, and an opening window of opportunity.&lt;/p&gt;
&lt;p&gt;Joining Dynamia is not just a career choice, but a commitment to building the next generation of infrastructure. I hope the observations and reflections in this article can provide some reference for you as you evaluate technical directions and career opportunities.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;If you are also interested in HAMi, GPU virtualization, AI Native Infra, or Dynamia, feel free to reach out.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>The True Inflection Point of ADD: When Spec Becomes the Core Asset of AI-Era Software</title><link>https://jimmysong.io/blog/add-inflection-point-spec-as-core-asset/</link><pubDate>Tue, 20 Jan 2026 07:51:36 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/add-inflection-point-spec-as-core-asset/</guid><description>Exploring how Spec becomes the governable core asset in Agent-Driven Development (ADD) and the trend toward control-plane engineering systems.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The role of Spec is undergoing a fundamental transformation, becoming the governance anchor of engineering systems in the AI era.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="the-essence-of-software-engineering-and-the-cost-structure-shift-brought-by-ai"&gt;The Essence of Software Engineering and the Cost Structure Shift Brought by AI&lt;/h2&gt;
&lt;p&gt;From first principles, software engineering has always been about one thing: &lt;strong&gt;stably, controllably, and reproducibly transforming human intent into executable systems.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Artificial Intelligence (AI) does not change this engineering essence, but it dramatically alters the cost structure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Implementation costs plummet:&lt;/strong&gt; Code, tests, and boilerplate logic are rapidly commoditized.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consistency costs rise sharply:&lt;/strong&gt; Intent drift, hidden conflicts, and cross-module inconsistencies become more frequent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance costs are amplified:&lt;/strong&gt; As agents can act directly, auditability, accountability, and explainability become hard constraints.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, in the era of Agent-Driven Development (ADD), the core issue is not &amp;ldquo;can agents do the work,&amp;rdquo; but how to maintain controllability and intent preservation in engineering systems under highly autonomous agents.&lt;/p&gt;
&lt;h2 id="the-add-era-inflection-point-three-structural-preconditions"&gt;The ADD Era Inflection Point: Three Structural Preconditions&lt;/h2&gt;
&lt;p&gt;Many attribute the &amp;ldquo;explosion&amp;rdquo; of ADD to more mature multi-agent systems, stronger models, or more automated tools. In reality, the true structural inflection point arises only when these three conditions are met:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agents have acquired multi-step execution capabilities&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;With frameworks like LangChain, LangGraph, and CrewAI, agents are no longer just prompt invocations, but long-lived entities capable of planning, decomposition, execution, and rollback.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agents are entering real enterprise delivery pipelines&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Once in enterprise R&amp;amp;D, the question shifts from &amp;ldquo;can it generate&amp;rdquo; to &amp;ldquo;who approved it, is it compliant, can it be rolled back.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traditional engineering tools lack a control plane for the agent era&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Tools like Git, CI, and Issue Trackers were designed for &amp;ldquo;human developer collaboration,&amp;rdquo; not for &amp;ldquo;agent execution.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;When these three factors converge, ADD inevitably shifts from an &amp;ldquo;efficiency tool&amp;rdquo; to a &amp;ldquo;governance system.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-changing-role-of-spec-from-documentation-to-system-constraint"&gt;The Changing Role of Spec: From Documentation to System Constraint&lt;/h2&gt;
&lt;p&gt;In the context of ADD, Spec is undergoing a fundamental shift:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Spec is no longer &amp;ldquo;documentation for humans,&amp;rdquo; but &amp;ldquo;the source of constraints and facts for systems and agents to execute.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Spec now serves at least three roles:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Verifiable expression of intent and boundaries&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Requirements, acceptance criteria, and design principles are no longer just text, but objects that can be checked, aligned, and traced.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stable contracts for organizational collaboration&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When agents participate in delivery, verbal consensus and tacit knowledge quickly fail. Versioned, auditable artifacts become the foundation of collaboration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy surface for agent execution&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Agents can write code, modify configurations, and trigger pipelines. Spec must become the constraint on &amp;ldquo;what can and cannot be done.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;From this perspective, the status of Spec is approaching that of the &lt;strong&gt;Control Plane&lt;/strong&gt; in AI-native infrastructure.&lt;/p&gt;
&lt;h2 id="the-reality-of-multi-agent-workflows-orchestration-and-governance-first"&gt;The Reality of Multi-Agent Workflows: Orchestration and Governance First&lt;/h2&gt;
&lt;p&gt;In recent systems (such as &lt;a href="https://apoxai.com" target="_blank" rel="noopener"&gt;APOX&lt;/a&gt; and other enterprise products), an industry consensus is emerging:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-agent collaboration no longer pursues &amp;ldquo;full automation,&amp;rdquo; but is staged and gated.&lt;/li&gt;
&lt;li&gt;Frameworks like LangGraph are used to build persistent, debuggable agent workflows.&lt;/li&gt;
&lt;li&gt;RAG (e.g., based on Milvus) is used to accumulate historical Specs, decisions, and context as long-term memory.&lt;/li&gt;
&lt;li&gt;The IDE mainly focuses on execution efficiency, not engineering governance.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/add-inflection-point-spec-as-core-asset/apox.webp" data-img="https://assets.jimmysong.io/images/blog/add-inflection-point-spec-as-core-asset/apox.webp" alt="Figure 1: APOX user interface" data-caption="Figure 1: APOX user interface"
width="1400"
height="1045"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: APOX user interface&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;APOX (AI Product Orchestration eXtended) is a multi-agent collaboration workflow platform for enterprise software delivery. Its core goals are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;To connect the entire process from product requirements to executable code with a governable Agentflow and explicit engineering artifact chain.&lt;/li&gt;
&lt;li&gt;To assign dedicated AI agents to each delivery stage (such as PRD, PO, Architecture, Developer, Implementation, Coding, etc.).&lt;/li&gt;
&lt;li&gt;To embed manual approval gates and full audit trails at every step, solving the &amp;ldquo;intent drift and consistency&amp;rdquo; governance problem that traditional AI coding tools cannot address.&lt;/li&gt;
&lt;li&gt;The platform provides a VS Code plugin for real-time sync between local IDE and web artifacts, allowing Specs, code, tasks, and approval statuses to coexist in the repository.&lt;/li&gt;
&lt;li&gt;Supports assigning different base models to different agents according to enterprise needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;APOX is not about simply speeding up code generation, but about elevating &amp;ldquo;Spec&amp;rdquo; from auxiliary documentation to a verifiable, constrainable, and traceable core asset in engineering—building a control plane and workflow governance system suitable for Agent-Driven Development.&lt;/p&gt;
&lt;p&gt;Such systems emphasize:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An explicit artifact chain from PRD → Spec → Task → Implementation.&lt;/li&gt;
&lt;li&gt;Manual confirmation and audit points at every stage.&lt;/li&gt;
&lt;li&gt;Bidirectional sync between Spec, code, repository, and IDE.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is not about &amp;ldquo;smarter AI,&amp;rdquo; but about engineering systems adapting to the agent era.&lt;/p&gt;
&lt;h2 id="the-long-term-value-of-spec-the-core-anchor-of-engineering-assets"&gt;The Long-Term Value of Spec: The Core Anchor of Engineering Assets&lt;/h2&gt;
&lt;p&gt;This is not to devalue code, but to acknowledge reality:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;There will always be long-term differentiation in algorithms and model capabilities.&lt;/li&gt;
&lt;li&gt;General engineering implementation is rapidly homogenizing.&lt;/li&gt;
&lt;li&gt;What is hard to replicate is: how to define problems, constrain systems, and govern change.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the ADD era, the value of Spec is reflected in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Determining what agents can and cannot do.&lt;/li&gt;
&lt;li&gt;Carrying the organization&amp;rsquo;s long-term understanding of the system.&lt;/li&gt;
&lt;li&gt;Serving as the anchor for audit, compliance, and accountability.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Code will be rewritten again and again; Spec is the long-term asset.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="risks-and-challenges-of-add-living-spec-and-governance-constraints"&gt;Risks and Challenges of ADD: Living Spec and Governance Constraints&lt;/h2&gt;
&lt;p&gt;ADD also faces significant risks:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can Spec become a Living Spec&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That is, when key implementation changes occur, can the system detect &amp;ldquo;intent changes&amp;rdquo; and prompt Spec updates, rather than allowing silent drift?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can governance achieve low friction but strong constraints&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If gates are too strict, teams will bypass them; if too loose, the system loses control.&lt;/p&gt;
&lt;p&gt;These two factors determine whether ADD is &amp;ldquo;the next engineering paradigm&amp;rdquo; or &amp;ldquo;just another tool bubble.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-trend-toward-control-planes-in-engineering-systems"&gt;The Trend Toward Control Planes in Engineering Systems&lt;/h2&gt;
&lt;p&gt;From a broader perspective, ADD is the inevitable result of engineering systems becoming &amp;ldquo;control planes&amp;rdquo;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Engineering systems are evolving from &amp;ldquo;human collaboration tools&amp;rdquo; to &amp;ldquo;control systems for agent execution.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In this structure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Agent / IDE is the &lt;strong&gt;execution plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;RAG / Memory is the &lt;strong&gt;state and memory plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Spec is the intent and policy plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Gates, audit, and traceability form the governance loop.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This closely aligns with the evolution path of AI-native infrastructure.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The winners of the ADD era will not be the systems with &amp;ldquo;the most agents or the fastest generation,&amp;rdquo; but those that first upgrade Spec from documentation to a governable, auditable, and executable asset. As automation advances, the true scarcity is the long-term control of intent.&lt;/p&gt;</content:encoded></item><item><title>AI Voice Dictation Input Methods Are Becoming the New Shortcut Key for the Programming Era</title><link>https://jimmysong.io/blog/ai-voice-dictation-input-method-comparison/</link><pubDate>Sun, 18 Jan 2026 06:53:08 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-voice-dictation-input-method-comparison/</guid><description>Comparing Miaoyan, Zhipu, and Shandianshuo voice input methods for developers: speed, stability, command capabilities, and cost models.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Voice input methods are not just about being &amp;ldquo;fast&amp;rdquo;—they are becoming a brand new gateway for developers to collaborate with AI.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
Warning
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
On January 12, 2026, due to financial difficulties encountered during operations, the Miaoyan project announced the cessation of operations and the team was disbanded. The application will no longer be updated or maintained, but existing versions can continue to be used on the current device and system, and do not store any audio or transcription content.
&lt;/div&gt;
&lt;/div&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/banner.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/banner.webp" alt="Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience." data-caption="Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience."
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="ai-voice-input-methods-are-becoming-the-new-shortcut-key-in-the-programming-era"&gt;AI Voice Input Methods Are Becoming the &amp;ldquo;New Shortcut Key&amp;rdquo; in the Programming Era&lt;/h2&gt;
&lt;p&gt;I am increasingly convinced of one thing: &lt;strong&gt;PC-based AI voice input methods are evolving from mere &amp;ldquo;input tools&amp;rdquo; into the foundational interaction layer for the era of programming and AI collaboration.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not just about typing faster—it determines how you deliver your &lt;strong&gt;intent&lt;/strong&gt; to the system, whether you&amp;rsquo;re writing documentation, code, or collaborating with AI in IDEs, terminals, or chat windows.&lt;/p&gt;
&lt;p&gt;Because of this, the differences in voice input method experiences are far more significant than they appear on the surface.&lt;/p&gt;
&lt;h2 id="my-six-evaluation-criteria-for-ai-voice-input-methods"&gt;My Six Evaluation Criteria for AI Voice Input Methods&lt;/h2&gt;
&lt;p&gt;After long-term, high-frequency use, I have developed a set of criteria to assess the real-world performance of AI voice input methods:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Response speed&lt;/strong&gt;: Does text appear quickly enough after pressing the shortcut to keep up with your thoughts?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Continuous input stability&lt;/strong&gt;: Does it remain reliable during extended use, or does it suddenly fail or miss recognition?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mixed Chinese-English and technical terms&lt;/strong&gt;: Can it reliably handle code, paths, abbreviations, and product names?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Developer friendliness&lt;/strong&gt;: Is it truly designed for command line, IDE, and automation scenarios?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interaction restraint&lt;/strong&gt;: Does it avoid introducing distracting features that interfere with input itself?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Subscription and cost structure&lt;/strong&gt;: Is it a standalone paid product, or can it be bundled with existing tool subscriptions?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Based on these criteria, I focused on comparing &lt;strong&gt;Miaoyan&lt;/strong&gt;, &lt;strong&gt;Shandianshuo&lt;/strong&gt;, and &lt;strong&gt;Zhipu AI Voice Input Method&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="miaoyan-currently-the-most-developer-oriented-domestic-product"&gt;Miaoyan: Currently the Most &amp;ldquo;Developer-Oriented&amp;rdquo; Domestic Product&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://miaoyan.cn" target="_blank" rel="noopener"&gt;Miaoyan&lt;/a&gt; was the first domestic AI voice input method I used extensively, and it remains the one I am most willing to use continuously.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/miaoyan.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/miaoyan.webp" alt="Figure 2: Miaoyan is currently my most-used Mac voice input method." data-caption="Figure 2: Miaoyan is currently my most-used Mac voice input method."
width="2272"
height="1624"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Miaoyan is currently my most-used Mac voice input method.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="command-mode-the-key-differentiator-for-developer-productivity"&gt;Command Mode: The Key Differentiator for Developer Productivity&lt;/h3&gt;
&lt;p&gt;It&amp;rsquo;s important to clarify that &lt;strong&gt;Miaoyan&amp;rsquo;s command mode is not about editing text via voice&lt;/strong&gt;. Instead:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You describe your need in natural language, and the system directly generates an &lt;strong&gt;executable command-line command&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is crucial for developers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It&amp;rsquo;s not just about input&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s about turning voice into an automation entry point&lt;/li&gt;
&lt;li&gt;Essentially, it connects voice to the CLI or toolchain&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This design is clearly focused on &lt;strong&gt;engineering efficiency&lt;/strong&gt;, not office document polishing.&lt;/p&gt;
&lt;h3 id="usage-experience-summary"&gt;Usage Experience Summary&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Fast response, nearly instant&lt;/li&gt;
&lt;li&gt;Output is relatively clean, with minimal guessing&lt;/li&gt;
&lt;li&gt;Interaction design is restrained, with no unnecessary concepts&lt;/li&gt;
&lt;li&gt;Developer-friendly mindset&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But there are some practical limitations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It is a &lt;strong&gt;completely standalone product&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Requires a separate subscription&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Still in relatively small-scale use&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a product strategy perspective, it feels more like a &amp;ldquo;pure tool&amp;rdquo; than part of an ecosystem.&lt;/p&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
Note
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
On January 12, 2026, due to financial difficulties encountered during operations, the Miaoyan project announced the cessation of operations and the team was disbanded. The application will no longer be updated or maintained, but existing versions can continue to be used on the current device and system, and do not store any audio or transcription content.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="shandianshuo-local-first-approach-developer-experience-depends-on-your-setup"&gt;Shandianshuo: Local-First Approach, Developer Experience Depends on Your Setup&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://shandianshuo.cn" target="_blank" rel="noopener"&gt;Shandianshuo&lt;/a&gt; takes a different approach: it treats voice input as a &amp;ldquo;local-first foundational capability,&amp;rdquo; emphasizing low latency and privacy (at least in its product narrative). The natural advantages of this approach are speed and controllable marginal costs, making it suitable as a &amp;ldquo;system capability&amp;rdquo; that&amp;rsquo;s always available, rather than a cloud service.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/shandianshuo.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/shandianshuo.webp" alt="Figure 3: Shandianshuo settings page" data-caption="Figure 3: Shandianshuo settings page"
width="2556"
height="2080"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Shandianshuo settings page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;However, from a developer&amp;rsquo;s perspective, its upper limit often depends on &amp;ldquo;how you implement enhanced capabilities&amp;rdquo;:&lt;/p&gt;
&lt;p&gt;If you only use it for basic transcription, the experience is more like a high-quality local input tool. But if you want better mixed Chinese-English input, technical term correction, symbol and formatting handling, the common approach is to add optional AI correction/enhancement capabilities, which usually requires extra configuration (such as providing your own API key or subscribing to enhanced features). The key trade-off here is not &amp;ldquo;can it be used,&amp;rdquo; but &amp;ldquo;how much configuration cost are you willing to pay for enhanced capabilities.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;If you want voice input to be a &amp;ldquo;lightweight, stable, non-intrusive&amp;rdquo; foundation, Shandianshuo is worth considering. But if your goal is to make voice input part of your developer workflow (such as command generation or executable actions), it needs to offer stronger productized design at the &amp;ldquo;command layer&amp;rdquo; and in terms of controllability.&lt;/p&gt;
&lt;h2 id="zhipu-ai-voice-input-method-stable-but-with-friction"&gt;Zhipu AI Voice Input Method: Stable but with Friction&lt;/h2&gt;
&lt;p&gt;I also thoroughly tested the &lt;a href="https://autoglm.zhipuai.cn/autotyper/" target="_blank" rel="noopener"&gt;Zhipu AI Voice Input Method&lt;/a&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/autoglm.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/autoglm.webp" alt="Figure 4: Zhipu Voice Input Method settings interface" data-caption="Figure 4: Zhipu Voice Input Method settings interface"
width="2430"
height="1824"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Zhipu Voice Input Method settings interface&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Its strengths include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;More stable for long-term continuous input&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Rarely becomes completely unresponsive&lt;/li&gt;
&lt;li&gt;Good tolerance for longer Chinese input&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But with frequent use, some issues stand out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Idle misrecognition&lt;/strong&gt;: If you press the shortcut but don&amp;rsquo;t speak, it may output random characters, disrupting your input flow&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Occasionally messy output&lt;/strong&gt;: Sometimes adds irrelevant words, making it less controllable than Miaoyan&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Basic recognition errors&lt;/strong&gt;: For example, &amp;ldquo;Zhipu&amp;rdquo; being recognized as &amp;ldquo;Zhipu&amp;rdquo; (with a different character), which is a trust issue for professional users&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Feature-heavy design&lt;/strong&gt;: Various tone and style features increase cognitive load&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="subscription-bundling-zhipus-practical-advantage"&gt;Subscription Bundling: Zhipu&amp;rsquo;s Practical Advantage&lt;/h2&gt;
&lt;p&gt;Although I prefer Miaoyan in terms of experience, &lt;strong&gt;Zhipu has a very practical advantage&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;If you already subscribe to Zhipu&amp;rsquo;s programming package, &lt;strong&gt;the voice input method is included for free&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No need to pay separately for the input method&lt;/li&gt;
&lt;li&gt;Lower psychological and decision-making cost&lt;/li&gt;
&lt;li&gt;More likely to become the &amp;ldquo;default tool&amp;rdquo; that stays&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a business perspective, this is a very smart strategy.&lt;/p&gt;
&lt;h2 id="main-comparison-table"&gt;Main Comparison Table&lt;/h2&gt;
&lt;p&gt;The following table compares the three products across key dimensions for quick reference.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Miaoyan&lt;/th&gt;
&lt;th&gt;Shandianshuo&lt;/th&gt;
&lt;th&gt;Zhipu AI Voice Input Method&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Response Speed&lt;/td&gt;
&lt;td&gt;Fast, nearly instant&lt;/td&gt;
&lt;td&gt;Usually fast (local-first)&lt;/td&gt;
&lt;td&gt;Slightly slower than Miaoyan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous Stability&lt;/td&gt;
&lt;td&gt;Stable&lt;/td&gt;
&lt;td&gt;Depends on setup and environment&lt;/td&gt;
&lt;td&gt;Very stable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idle Misrecognition&lt;/td&gt;
&lt;td&gt;Rare&lt;/td&gt;
&lt;td&gt;Generally restrained (varies by version)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Obvious: outputs characters even if silent&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output Cleanliness/Control&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;More like an &amp;ldquo;input tool&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Occasionally messy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Differentiator&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Natural language → executable command&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Local-first / optional enhancements&lt;/td&gt;
&lt;td&gt;Ecosystem-attached capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subscription &amp;amp; Cost&lt;/td&gt;
&lt;td&gt;Standalone, separate purchase&lt;/td&gt;
&lt;td&gt;Basic usable; enhancements often require setup/subscription&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Bundled free with programming package&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My Current Preference&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Best experience&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;More like a &amp;ldquo;foundation approach&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Easy to keep but not clean enough&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Core Comparison of Miaoyan, Shandianshuo, and Zhipu AI Voice Input Methods
&lt;/figcaption&gt;
&lt;h2 id="user-loyalty-to-ai-voice-input-methods"&gt;User Loyalty to AI Voice Input Methods&lt;/h2&gt;
&lt;p&gt;The switching cost for voice input methods is actually low: just a shortcut key and a habit of output.&lt;/p&gt;
&lt;p&gt;What really determines whether users stick around is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether the output is controllable&lt;/li&gt;
&lt;li&gt;Whether it keeps causing annoying minor issues&lt;/li&gt;
&lt;li&gt;Whether it integrates into your existing workflow and payment structure&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For me personally:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The best and smoothest experience is still Miaoyan&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The one most likely to stick around is probably Zhipu&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shandianshuo is more of a &amp;ldquo;foundation approach&amp;rdquo; and worth watching for how its enhancements evolve&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These points are not contradictory.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Miaoyan is more mature in &lt;strong&gt;engineering orientation, command capabilities, and input control&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Zhipu has practical advantages in &lt;strong&gt;stability and subscription bundling&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Shandianshuo takes a &lt;strong&gt;local-first + optional enhancement&lt;/strong&gt; approach, with the key being how it balances &amp;ldquo;basic capability&amp;rdquo; and &amp;ldquo;enhancement cost&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Who truly becomes the &amp;ldquo;default gateway&amp;rdquo; depends on reducing distractions, fixing frequent minor issues, and treating voice input as true &amp;ldquo;infrastructure&amp;rdquo; rather than an add-on feature&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The competition among AI voice input methods is no longer about recognition accuracy, but about who can own the shortcut key you press every day.&lt;/strong&gt;&lt;/p&gt;</content:encoded></item><item><title>From Spatial Data to AI Open Source: Technical Standards, Data Sovereignty, and the Global Divide</title><link>https://jimmysong.io/blog/spatial-data-ai-open-source-standards-sovereignty/</link><pubDate>Sun, 11 Jan 2026 03:29:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/spatial-data-ai-open-source-standards-sovereignty/</guid><description>How technical standards and data sovereignty shape AI open source paths and infrastructure competition in the global AI era.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The divide in technical standards and data sovereignty determines the global competitive landscape of infrastructure open source in the AI era.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In this article, I will use the differences in air quality data presentation in Apple Maps and Weather as a starting point to explore how technical standards and data sovereignty influence the open source paths of AI in different countries. I will further analyze why, in the AI era, infrastructure-level open source has become the key battleground for ecosystem dominance.&lt;/p&gt;
&lt;h2 id="authors-note"&gt;Author&amp;rsquo;s Note&lt;/h2&gt;
&lt;p&gt;This article originates from a very everyday observation: Why is air quality data in China shown as &amp;ldquo;points&amp;rdquo; in Apple Maps and Weather, while in other countries it is often displayed as &amp;ldquo;areas&amp;rdquo;?&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/aqi-map.webp" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/aqi-map.webp" alt="Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries" data-caption="Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries"
width="1650"
height="1864"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;At first glance, it seems like a product experience difference. But when I reconsidered this issue in the context of engineering, standards, and system design, I realized it actually points to a much bigger question: how different countries understand the relationship between technology, standards, openness, and sovereignty.&lt;/p&gt;
&lt;p&gt;As an engineer who has long worked in cloud native, AI infrastructure, and open source ecosystems, I gradually realized that this difference is not limited to air quality or map data. In the AI era, it is further amplified, directly affecting how we open source models, build infrastructure, and whether we can participate in the formulation of global rules.&lt;/p&gt;
&lt;p&gt;Writing this article is not about judging right or wrong, but about using a concrete example to explain a structural difference and discuss the long-term impact and real opportunities this difference may bring in the AI era.&lt;/p&gt;
&lt;p&gt;What is especially important: at the level of AI infrastructure and infra-level open source, the competition has just begun. China is not without opportunities, but the choice of path will become more critical than ever.&lt;/p&gt;
&lt;h2 id="differences-in-air-quality-data-presentation-a-microcosm-of-technical-standards-and-sovereignty"&gt;Differences in Air Quality Data Presentation: A Microcosm of Technical Standards and Sovereignty&lt;/h2&gt;
&lt;p&gt;The following image illustrates the divide between spatial data, AI open source, and technical standards. By comparing how air quality data is presented in Apple Maps and Weather in different countries, you can intuitively feel the differences in technical standards and sovereignty strategies behind the scenes.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/banner.webp" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/banner.webp" alt="Figure 2: The divide between spatial data, AI open source, and technical standards" data-caption="Figure 2: The divide between spatial data, AI open source, and technical standards"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The divide between spatial data, AI open source, and technical standards&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;If you regularly use global products such as maps, weather, traffic, or various data services, you may notice a recurring phenomenon that is rarely discussed seriously: the way data is presented in China often differs significantly from global mainstream standards.&lt;/p&gt;
&lt;p&gt;A very intuitive example comes from the air quality display in Apple Maps or Weather. In China, air quality is usually shown as discrete points; in the US, Europe, Japan, and other countries, it is often rendered as continuous coverage areas.&lt;/p&gt;
&lt;p&gt;At first glance, this seems like a product experience difference, and may even lead people to mistakenly believe that &amp;ldquo;China&amp;rsquo;s data is incomplete.&amp;rdquo; But if you treat it as an engineering or system design issue, you will find: this is not a matter of data capability, but a different choice in technical standards, data sovereignty, and openness strategies.&lt;/p&gt;
&lt;p&gt;And this choice is not limited to air quality.&lt;/p&gt;
&lt;h2 id="air-quality-is-just-a-slice-greater-differences-in-spatial-public-data"&gt;Air Quality Is Just a Slice: Greater Differences in Spatial Public Data&lt;/h2&gt;
&lt;p&gt;Air quality is just a highly visible and relatively low-risk example. Similar differences have long existed in broader spatial and public data domains.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maps and coordinate systems&lt;/li&gt;
&lt;li&gt;Surveying and high-precision spatial data&lt;/li&gt;
&lt;li&gt;Real-time traffic and population movement&lt;/li&gt;
&lt;li&gt;Remote sensing, environmental, and urban operation data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In global mainstream systems, such data is usually regarded as public information infrastructure. It is standardized, gridded, API-ified, allows interpolation, modeling, and redistribution, and is widely used in research, business, and product innovation.&lt;/p&gt;
&lt;p&gt;In China, this data often takes another form: hierarchical, discrete, strictly defined, and with centralized interpretation authority.&lt;/p&gt;
&lt;p&gt;This is not a technical preference in a single field, but a systemic logic of technology and governance.&lt;/p&gt;
&lt;h2 id="three-global-paths"&gt;Three Global Paths&lt;/h2&gt;
&lt;p&gt;Placing China in a global context, we can see that there are roughly three different paths worldwide regarding &amp;ldquo;how public data and technical standards are opened.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Engineering-Open Type: Standards and Ecosystem First&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Represented by the US and some European countries, the core features of this system are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Public data prioritized as infrastructure&lt;/li&gt;
&lt;li&gt;Standards and interfaces come first&lt;/li&gt;
&lt;li&gt;Encourages engineering autonomy and ecosystem evolution&lt;/li&gt;
&lt;li&gt;Tolerates model inference and uncertainty&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This path directly shaped the global landscape of foundational software and infrastructure-level open source. Linux, Kubernetes, and the cloud native system are essentially products of openness at the rules layer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Governance-Sovereignty Type: Control and Auditability First&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Represented by China, this path emphasizes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sensitivity of spatial and public data&lt;/li&gt;
&lt;li&gt;Data as part of governance capability&lt;/li&gt;
&lt;li&gt;Standards, definitions, and release methods are highly bound&lt;/li&gt;
&lt;li&gt;Emphasizes traceability, accountability, and controllability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this system, &amp;ldquo;point data&amp;rdquo; is not a sign of technological backwardness, but a governable technical form. When a technical system is designed as a governance system, its primary goal is not reusability, but controllability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compromise-Coordinated Type: Cautious Openness, Engineering Internationalization&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Some countries try to find a balance between the two, maintaining caution in spatial data while being highly internationalized in engineering and industry. This shows that the difference is not about being advanced or backward, but about different objective functions.&lt;/p&gt;
&lt;p&gt;The following diagram compares the core characteristics, typical cases, and advantages/challenges of these three paths from a global perspective. The &amp;ldquo;Engineering-Open Type&amp;rdquo; on the left shapes the global infrastructure software landscape through standards and ecosystems; the &amp;ldquo;Governance-Sovereignty Type&amp;rdquo; in the middle emphasizes data sovereignty and security controllability but has limitations in influence at the rules layer; the &amp;ldquo;Compromise-Coordinated Type&amp;rdquo; on the right attempts to find a balance between security and openness. The divide between these three paths directly affects the infrastructure competition landscape of various countries in the AI era.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/global-three-paths-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/global-three-paths-en.svg" alt="Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards" data-caption="Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards"
width="2663"
height="1862"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-essence-of-point-vs-area-in-air-quality"&gt;The Essence of &amp;ldquo;Point&amp;rdquo; vs. &amp;ldquo;Area&amp;rdquo; in Air Quality&lt;/h2&gt;
&lt;p&gt;Among all spatial public data, air quality is an ideal observation window:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Does not directly involve military or core economic security&lt;/li&gt;
&lt;li&gt;Highly visible, updated daily, and perceptible to everyone&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;China does not lack air quality data; on the contrary, the density of monitoring stations is among the highest in the world. The real difference lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether interpolation is allowed&lt;/li&gt;
&lt;li&gt;Whether model inference is allowed&lt;/li&gt;
&lt;li&gt;Whether platforms are allowed to reinterpret the data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;ldquo;Point&amp;rdquo; means authenticity and traceability; &amp;ldquo;area&amp;rdquo; means models, inference, and redistribution of interpretive authority. This is precisely the watershed between technical standards and data sovereignty.&lt;/p&gt;
&lt;p&gt;The following diagram compares two different technical paths. The left side, &amp;ldquo;Governance-Sovereignty Type,&amp;rdquo; emphasizes data traceability and controllability, using discrete point-based data presentation. The right side, &amp;ldquo;Engineering-Open Type,&amp;rdquo; allows model interpolation and inference, providing more user-friendly experience through continuous area-based coverage. The essence of this difference lies not in the level of technical capability, but in the different choices made between data sovereignty, governance capability, and open ecosystems.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/data-sovereignty-comparison-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/data-sovereignty-comparison-en.svg" alt="Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation" data-caption="Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation"
width="2263"
height="1562"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-amplification-effect-in-the-ai-era"&gt;The Amplification Effect in the AI Era&lt;/h2&gt;
&lt;p&gt;With the above logic in mind, many phenomena in the AI era become less confusing.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Why are Chinese AI companies more willing to open source large language model (LLM) weights, while American companies have clearly shifted toward closed source in recent years?&lt;/li&gt;
&lt;li&gt;Why is foundational software and infrastructure-level open source still mainly led by the US?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key is not &amp;ldquo;whether to open source,&amp;rdquo; but &amp;ldquo;which layer is open sourced.&amp;rdquo;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model weights are static, declarable assets&lt;/li&gt;
&lt;li&gt;Infrastructure, runtimes, protocols, and standards are dynamic, evolving system rules&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open sourcing weights is essentially openness at the asset layer; infrastructure-level open source means relinquishing control over operating rules and interpretive authority.&lt;/p&gt;
&lt;p&gt;The following diagram compares two different layers of AI open source. The left side shows &amp;ldquo;Model Weight Layer Open Source,&amp;rdquo; which is a typical feature of Chinese path—opening static digital assets with low cost and controllable risk, but not involving rule-making. The right side shows &amp;ldquo;Infrastructure Layer Open Source,&amp;rdquo; which is a core strategy of US path—by open sourcing development tools, protocol standards, runtimes, and compute scheduling and other infrastructure, defining how AI is used, thereby mastering ecosystem rules and interpretive authority. Key insight: Open sourcing model weights does not equal mastering AI ecosystem, and the real competitive focus is shifting to the infrastructure layer of &amp;ldquo;how AI runs.&amp;rdquo;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/ai-opensource-layers-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/ai-opensource-layers-en.svg" alt="Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure" data-caption="Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure"
width="2363"
height="1862"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-us-approach-focusing-on-rules-and-runtime-layers"&gt;The US Approach: Focusing on Rules and Runtime Layers&lt;/h2&gt;
&lt;p&gt;In the past year or two, US-led AI open source and ecosystem initiatives have shown a highly consistent direction: not rushing to open source the strongest models, but focusing on defining &amp;ldquo;how AI is used.&amp;rdquo;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Linux Foundation established &lt;a href="https://aaif.io" target="_blank" rel="noopener"&gt;AAIF&lt;/a&gt; (Agentic AI Foundation), focusing on AI infrastructure, standards, and toolchain collaboration&lt;/li&gt;
&lt;li&gt;Protocols like MCP (Model Context Protocol) aim to define common interaction methods between agents and tools/systems&lt;/li&gt;
&lt;li&gt;Major tech companies are generally focusing on APIs, platforms, runtimes, and ecosystem binding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The commonality of these actions: competing in model capability, but controlling the usage rules.&lt;/p&gt;
&lt;h2 id="chinas-shift-from-model-oriented-to-infrastructure-oriented"&gt;China&amp;rsquo;s Shift: From Model-Oriented to Infrastructure-Oriented&lt;/h2&gt;
&lt;p&gt;It is important to emphasize that this difference does not mean China is unaware of the issue.&lt;/p&gt;
&lt;p&gt;Whether in policy discussions or within industry and research institutions, the risk of &amp;ldquo;only open sourcing models without controlling infrastructure and standard dominance&amp;rdquo; has been repeatedly discussed.&lt;/p&gt;
&lt;p&gt;The real challenge lies in how to achieve a directional shift within the existing governance logic and risk framework. This shift has already appeared in some concrete practices.&lt;/p&gt;
&lt;h2 id="exploration-and-practice-at-the-infrastructure-layer"&gt;Exploration and Practice at the Infrastructure Layer&lt;/h2&gt;
&lt;p&gt;In the AI era, infrastructure often starts with the most engineering-driven problems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HAMi Project&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Projects like &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; do not focus on model capability, but on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Abstraction, allocation, and isolation of GPU resources&lt;/li&gt;
&lt;li&gt;How multi-tenant AI workloads are run&lt;/li&gt;
&lt;li&gt;How computing power transitions from hardware assets to governable system resources&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The significance of such projects is not about being &amp;ldquo;SOTA,&amp;rdquo; but about entering the domain of &amp;ldquo;how AI runs.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Runtime Reconstruction from a System Software Review&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Exploration at the research institution level is also noteworthy. The &lt;a href="https://www.flagos.io" target="_blank" rel="noopener"&gt;FlagOS&lt;/a&gt; initiative by the Beijing Academy of Artificial Intelligence is a clear signal: AI is being redefined as a system software issue, not just a model or algorithm problem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-Term Tech Stack Investment by Industry Players&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the industry, Huawei&amp;rsquo;s strategy reflects a similar direction: not simply open sourcing models, but attempting to build a complete, controllable AI tech stack, from computing power to frameworks, platforms, and ecosystems. This is a slower, heavier, but more infrastructure-competitive path.&lt;/p&gt;
&lt;h2 id="realistic-assessment-the-starting-point-of-ai-infrastructure-competition"&gt;Realistic Assessment: The Starting Point of AI Infrastructure Competition&lt;/h2&gt;
&lt;p&gt;Taking a longer view, we find an easily overlooked fact:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;At the level of AI infrastructure and infra-level open source, there is no settled pattern between China and the US.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The US advantage lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mature engineering culture&lt;/li&gt;
&lt;li&gt;Standard organizations and foundation mechanisms&lt;/li&gt;
&lt;li&gt;High proficiency in openness at the rules layer&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;China&amp;rsquo;s variables include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Huge AI application scenarios&lt;/li&gt;
&lt;li&gt;Extreme demand for computing power and system efficiency&lt;/li&gt;
&lt;li&gt;Ongoing directional adjustments&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The real uncertainty is not &amp;ldquo;whether we can catch up,&amp;rdquo; but whether it is possible to gradually open up space for engineering autonomy and standard co-construction while maintaining governance bottom lines.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The &amp;ldquo;points&amp;rdquo; and &amp;ldquo;areas&amp;rdquo; of air quality, model weights and the world of operations—behind these appearances lies not a simple technical route dispute, but how a country finds its own balance between openness, standards, and sovereignty.&lt;/p&gt;
&lt;p&gt;In the AI era, this issue will not disappear, but will become more concrete and more engineering-driven. And this is precisely where there are still opportunities for China&amp;rsquo;s AI infrastructure open source.&lt;/p&gt;</content:encoded></item><item><title>Joining Dynamia: Embarking on a New Journey in AI Native Infrastructure</title><link>https://jimmysong.io/blog/joining-dynamia/</link><pubDate>Wed, 07 Jan 2026 07:49:21 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/joining-dynamia/</guid><description>Joining Dynamia as Open Source Ecosystem VP to drive AI-native infrastructure ecosystem development, transforming compute from hardware consumption to core asset.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Compute governance is the critical bottleneck for AI scaling. From hardware consumption to core asset, this long-undervalued path needs to be redefined.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/joining-dynamia/banner.webp" data-img="https://assets.jimmysong.io/images/blog/joining-dynamia/banner.webp" alt="Figure 1: Dynamia.ai" data-caption="Figure 1: Dynamia.ai"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Dynamia.ai&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="a-new-beginning"&gt;A New Beginning&lt;/h2&gt;
&lt;p&gt;I have officially joined &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; as &lt;strong&gt;Open Source Ecosystem VP&lt;/strong&gt;, responsible for the long-term development of the company in open source, technical narrative, and &lt;strong&gt;AI Native Infrastructure&lt;/strong&gt; ecosystem directions.&lt;/p&gt;
&lt;h2 id="why-i-chose-dynamia"&gt;Why I Chose Dynamia&lt;/h2&gt;
&lt;p&gt;I chose to join Dynamia not because it&amp;rsquo;s a company trying to &amp;ldquo;solve all AI problems,&amp;rdquo; but precisely the opposite—it&amp;rsquo;s because Dynamia &lt;strong&gt;focuses intensely on one unavoidable, yet long-undervalued core issue in AI Native Infrastructure&lt;/strong&gt;: compute, especially &lt;strong&gt;Graphics Processing Units&lt;/strong&gt; (GPU), are evolving from &amp;ldquo;technical resources&amp;rdquo; into infrastructure elements that require refined governance and economic management.&lt;/p&gt;
&lt;p&gt;Through years of practice in cloud native, distributed systems, and AI infrastructure (AI Infra), I&amp;rsquo;ve formed a clear judgment: as Large Language Models (LLM) and &lt;strong&gt;AI Agents&lt;/strong&gt; enter the stage of large-scale deployment, the real bottleneck limiting system scalability and sustainability is no longer just model capability itself, but how compute is measured, allocated, isolated, and scheduled, and how a governable, accountable, and optimizable operational mechanism is formed at the system level. From this perspective, the core challenge of AI infrastructure is essentially evolving into a &amp;ldquo;resource governance and Token economy&amp;rdquo; problem.&lt;/p&gt;
&lt;h2 id="about-dynamia-and-hami"&gt;About Dynamia and HAMi&lt;/h2&gt;
&lt;p&gt;Dynamia is an AI-native infrastructure technology company rooted in open source DNA, driving efficiency leaps in heterogeneous compute through technological innovation. Its leading open source project, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; (Heterogeneous AI Computing Virtualization Middleware), is a &lt;strong&gt;Cloud Native Computing Foundation&lt;/strong&gt; (CNCF) sandbox project providing GPU, NPU and other heterogeneous device virtualization, sharing, isolation, and topology-aware scheduling capabilities, widely adopted by 50+ enterprises and institutions.&lt;/p&gt;
&lt;h2 id="dynamias-technical-approach"&gt;Dynamia&amp;rsquo;s Technical Approach&lt;/h2&gt;
&lt;p&gt;In this context, Dynamia&amp;rsquo;s technical approach—starting from &lt;strong&gt;the GPU layer, which is the most expensive, scarcest, and least unified abstraction layer in AI systems&lt;/strong&gt;, treating compute as a foundational resource that can be measured, partitioned, scheduled, governed, and even &amp;ldquo;tokenized&amp;rdquo; for refined accounting and optimization—aligns highly with my long-term judgment on AI-native infrastructure.&lt;/p&gt;
&lt;p&gt;This path doesn&amp;rsquo;t use &amp;ldquo;model capabilities&amp;rdquo; or &amp;ldquo;application innovation&amp;rdquo; as selling points in the short term, nor is it easily packaged into simple stories. However, with rising compute costs, heterogeneous accelerators becoming the norm, and AI systems moving toward multi-tenant and large-scale operations, these infrastructure-level capabilities are gradually becoming prerequisites for the establishment and expansion of AI systems.&lt;/p&gt;
&lt;h2 id="future-focus"&gt;Future Focus&lt;/h2&gt;
&lt;p&gt;As Dynamia&amp;rsquo;s Open Source Ecosystem VP, I will focus on &lt;strong&gt;technical narrative of AI-native infrastructure, open source ecosystem building, and global developer collaboration&lt;/strong&gt;, promoting compute from &amp;ldquo;hardware resource being consumed&amp;rdquo; to &lt;strong&gt;governable, measurable, and optimizable AI infrastructure core asset&lt;/strong&gt;, laying the foundation for the scaling and sustainable evolution of AI systems in the next stage.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Joining Dynamia is an important milestone in my career and a concrete action demonstrating my long-term optimism about AI-native infrastructure. Compute governance is not a short-term trend that yields quick results, but an infrastructure proposition that cannot be bypassed for AI large-scale deployment. I look forward to exploring, building, and landing solutions on this long-undervalued path with global developers.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia Official Website&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi - Heterogeneous AI Computing Virtualization Middleware (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Running Parallel AI Agents on My Mac: Hands-On with Verdent's Standalone App</title><link>https://jimmysong.io/blog/verdent-standalone-app-parallel-agents/</link><pubDate>Sun, 04 Jan 2026 02:25:48 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/verdent-standalone-app-parallel-agents/</guid><description>A hands-on experience with Verdent&amp;#39;s standalone Mac app, exploring how parallel AI agents, isolated workspaces, and task-oriented workflows change real-world development.</description><content:encoded>
&lt;p&gt;I&amp;rsquo;ve been spending more time recently experimenting with vibe coding tools on real projects, not demos. One of those projects is my own website, where I constantly tweak content structure, navigation, and layout.&lt;/p&gt;
&lt;p&gt;During this process, I started using &lt;a href="https://verdent.ai" target="_blank" rel="noopener"&gt;Verdent&amp;rsquo;s standalone Mac app&lt;/a&gt; more seriously. What stood out was not any single feature, but how different the experience felt compared to traditional AI coding tools.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-standalone-app-ui.webp" data-img="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-standalone-app-ui.webp" alt="Figure 1: Verdent Standalone App UI" data-caption="Figure 1: Verdent Standalone App UI"
width="3836"
height="2240"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Verdent Standalone App UI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Verdent doesn&amp;rsquo;t behave like an assistant waiting for instructions. It behaves more like an environment where work happens in parallel.&lt;/p&gt;
&lt;h2 id="a-different-starting-point-tasks-not-chats"&gt;A Different Starting Point: Tasks, Not Chats&lt;/h2&gt;
&lt;p&gt;Most AI coding tools begin with a conversation. Verdent begins with &lt;strong&gt;tasks&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When I opened my website repository in the Verdent app, I didn&amp;rsquo;t start with a long prompt. I created multiple tasks directly: one to rethink navigation and SEO structure, another to explore homepage layout improvements, and a third to review existing content organization.&lt;/p&gt;
&lt;p&gt;Each task immediately spun up its own agent and workspace. From the beginning, the app encouraged me to think in parallel, the same way I normally would when sketching ideas on paper or jumping between files.&lt;/p&gt;
&lt;p&gt;This framing alone changes how you work.&lt;/p&gt;
&lt;h2 id="built-for-multitasking-without-losing-context"&gt;Built for Multitasking, Without Losing Context&lt;/h2&gt;
&lt;p&gt;Switching contexts is unavoidable in real development work. What usually breaks is continuity.&lt;/p&gt;
&lt;p&gt;Verdent handles this well. Each task preserves its full context independently. I could stop one task mid-way, switch to another, and come back later without re-explaining the problem or reloading files.&lt;/p&gt;
&lt;p&gt;For example, while one agent was analyzing my site&amp;rsquo;s navigation structure, another was exploring layout options. I moved between them freely. Nothing was lost. Each agent remembered exactly what it was doing.&lt;/p&gt;
&lt;p&gt;This feels closer to how developers think than how chat-based tools operate.&lt;/p&gt;
&lt;h2 id="safe-parallel-coding-with-workspaces"&gt;Safe Parallel Coding with Workspaces&lt;/h2&gt;
&lt;p&gt;Parallel work only becomes truly safe when code changes are isolated. When parallelism moves from discussion to actual code modification, risk management becomes essential.&lt;/p&gt;
&lt;p&gt;Verdent solves this with &lt;strong&gt;Workspaces&lt;/strong&gt;. Each workspace is an isolated, independent code environment with its own change history, commit log, and branches. This isn&amp;rsquo;t just about separation—it&amp;rsquo;s about making concurrent code changes manageable.&lt;/p&gt;
&lt;p&gt;What this means in practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple tasks can write code simultaneously&lt;/li&gt;
&lt;li&gt;Changes remain isolated from each other&lt;/li&gt;
&lt;li&gt;If conflicts arise, they&amp;rsquo;re visible and cleanly resolvable&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I intentionally let different agents operate on overlapping parts of my project: one modifying Markdown content and links, another adjusting CSS and layout logic. Both ran in parallel. No conflicts emerged. Later, I reviewed the diffs from each workspace and merged only what made sense.&lt;/p&gt;
&lt;p&gt;This kind of isolation removes significant anxiety from AI-assisted coding. You stop worrying about breaking things and start experimenting more freely, knowing that each change exists in its own contained environment.&lt;/p&gt;
&lt;h2 id="parallel-agent-execution-feels-like-delegation"&gt;Parallel Agent Execution Feels Like Delegation&lt;/h2&gt;
&lt;p&gt;Parallelism doesn’t mean that all agents complete the same phase of work at the same time—instead, by isolating and overlapping phases, what was once a strictly sequential process is compressed into a more efficient, collaborative mode.&lt;/p&gt;
&lt;p&gt;In Verdent, each agent runs in its own workspace, essentially an automatically managed branch or worktree. In practice, I often create multiple tasks with different responsibilities for the same requirement, such as planning, implementation, and review. But this doesn’t mean they all complete the same phase simultaneously.&lt;/p&gt;
&lt;p&gt;These tasks are triggered as needed, each running for a period and producing clear artifacts as boundaries for collaboration. The planning task generates planning documents or constraint specifications; the implementation task advances code changes based on those documents and produces diffs; the review task, according to the established planning goals and audit criteria, performs staged reviews of the generated changes. By overlapping phases around artifacts, the originally strict sequential process is compressed into a workflow that more closely resembles team collaboration.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The value of splitting into multiple tasks is not parallel execution, but parallel cognition and clear collaboration boundaries.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;While it’s technically possible to put multiple roles into a single task, this causes planning, implementation, and review to share the same context, which weakens role isolation and the auditability of results.&lt;/p&gt;
&lt;h2 id="configurability-and-design-trade-offs"&gt;Configurability and Design Trade-offs&lt;/h2&gt;
&lt;p&gt;Beyond the workflow model itself, Verdent exposes a surprisingly rich set of configurable capabilities.&lt;/p&gt;
&lt;p&gt;It allows users to customize MCP settings, define subagents with configurable prompts, and create reusable commands via slash (&lt;code&gt;/&lt;/code&gt;) shortcuts. Personal rules can be written to influence agent behavior and response style, and command-level permissions can be configured to enforce basic security boundaries. Verdent also supports multiple mainstream foundation models, including GPT, Claude, Gemini, and K2. For users who prefer a lightweight coding experience without a full IDE, Verdent offers DiffLens as an alternative review-oriented interface. Both &lt;a href="https://www.verdent.ai/pricing" target="_blank" rel="noopener"&gt;subscription-based and credit-based pricing models&lt;/a&gt; are supported.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-settings.webp" data-img="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-settings.webp" alt="Figure 2: Verdent Settings" data-caption="Figure 2: Verdent Settings"
width="2780"
height="1648"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Verdent Settings&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;That said, Verdent makes a clear set of trade-offs. It is not built around tab-based code completion, nor does it offer a plugin system. If it did, it would start to resemble a traditional IDE - which does not seem to be its goal. Verdent is not designed for direct, fine-grained code manipulation; most changes are mediated through conversational tasks and agent-driven edits. This makes the experience clean and focused, but it also means that for large, highly complex codebases, Verdent may function better as a complementary orchestration layer rather than a full-time development environment.&lt;/p&gt;
&lt;h2 id="where-verdent-fits-today"&gt;Where Verdent Fits Today&lt;/h2&gt;
&lt;p&gt;There are many AI-assisted coding tools emerging right now. Some focus on smarter editors, others on faster generation.&lt;/p&gt;
&lt;p&gt;Verdent feels different because it focuses on &lt;strong&gt;orchestration&lt;/strong&gt;, not just assistance.&lt;/p&gt;
&lt;p&gt;It doesn&amp;rsquo;t try to replace your editor. It sits one level above, coordinating planning, execution, and review across multiple agents.&lt;/p&gt;
&lt;p&gt;That makes it particularly suitable for exploratory work, refactoring, and early-stage design - exactly the kind of work I was doing on my website.&lt;/p&gt;
&lt;h2 id="final-thoughts"&gt;Final Thoughts&lt;/h2&gt;
&lt;p&gt;Using Verdent&amp;rsquo;s standalone app didn&amp;rsquo;t just speed things up. It changed how I structured work.&lt;/p&gt;
&lt;p&gt;Instead of doing everything sequentially, I started thinking in parallel again - and letting the system support that way of thinking.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://verdent.ai" target="_blank" rel="noopener"&gt;Verdent&lt;/a&gt; feels less like an AI feature and more like an environment that assumes AI is already part of how development happens.&lt;/p&gt;
&lt;p&gt;For developers experimenting with AI-native workflows, that shift is worth paying attention to.&lt;/p&gt;</content:encoded></item><item><title>2025 Annual Review: The Transformation Journey from Cloud Native to AI Native</title><link>https://jimmysong.io/blog/2025-annual-review/</link><pubDate>Wed, 31 Dec 2025 10:02:01 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/2025-annual-review/</guid><description>A look back at the major changes in 2025: shifting from Cloud Native to AI Native Infrastructure, AI tool ecosystem, and major website improvements.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The waves of technology keep evolving; only by actively embracing change can we continue to create value. In 2025, I chose to move from Cloud Native to AI Native—this year marked a key turning point for personal growth and system reinvention.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;2025 was a turning point for me. This year, I not only changed my technical direction but also the way I approach problems. Moving from Cloud Native infrastructure to AI Native Infrastructure was not just a migration of content, but an upgrade in mindset.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/2025-annual-review/banner.webp" data-img="https://assets.jimmysong.io/images/blog/2025-annual-review/banner.webp" alt="Figure 1: Farewell 2025!" data-caption="Figure 1: Farewell 2025!"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Farewell 2025!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This year, I conducted a large-scale refactoring of the website and systematically organized the content. Beyond the technical improvements, I want to share my thoughts and changes throughout the year.&lt;/p&gt;
&lt;h2 id="a-bold-shift-embracing-the-ai-native-era"&gt;A Bold Shift: Embracing the AI Native Era&lt;/h2&gt;
&lt;p&gt;At the beginning of 2025, I made an important decision: to reposition myself from a Cloud Native Evangelist to an AI Infrastructure Architect. This was not just a change in title, but a strategic transformation after careful consideration.&lt;/p&gt;
&lt;p&gt;As I witnessed the surge of AI technologies and the rise of Agent-based applications reshaping software, I realized that clinging to the boundaries of Cloud Native might mean missing an era. So, I systematically adjusted the website’s content structure, shifting the focus toward AI Native Infrastructure.&lt;/p&gt;
&lt;p&gt;This transformation was not about abandoning the past, but extending forward from the foundation of Cloud Native. Classic content like Kubernetes and Istio remains and is continuously updated, but new topics such as AI Agent and the AI Native Landscape have been added, forming a more complete knowledge map.&lt;/p&gt;
&lt;h2 id="content-creation-from-technical-details-to-ecosystem-perspective"&gt;Content Creation: From Technical Details to Ecosystem Perspective&lt;/h2&gt;
&lt;h3 id="ai-agent-building-systematic-knowledge"&gt;AI Agent: Building Systematic Knowledge&lt;/h3&gt;
&lt;p&gt;Agents represent a major evolution in software for the AI era. When I tried to understand Agent design principles, I found fragmented information everywhere but lacked a systematic knowledge base.&lt;/p&gt;
&lt;p&gt;So I created content that analyzes the Agent context lifecycle and control loop mechanisms, summarizing several proven architectural patterns. To make complex knowledge easier to digest, I organized it into logical sections so readers can learn step by step.&lt;/p&gt;
&lt;h3 id="ai-tool-ecosystem-mapping-the-open-source-landscape"&gt;AI Tool Ecosystem: Mapping the Open Source Landscape&lt;/h3&gt;
&lt;p&gt;AI tools and frameworks are emerging rapidly, with new projects appearing daily. To help readers quickly grasp the ecosystem, I built a comprehensive AI OSS database.&lt;/p&gt;
&lt;p&gt;This database covers everything from Agent frameworks to development tools and deployment services. I not only included active projects but also established an archive mechanism, preserving detailed information on over 150 historical projects. More importantly, I developed a scoring system to objectively evaluate projects across dimensions like quality and sustainability, helping readers decide which tools are worth investing time in.&lt;/p&gt;
&lt;h3 id="blogging-capturing-technology-trends-faster"&gt;Blogging: Capturing Technology Trends Faster&lt;/h3&gt;
&lt;p&gt;In 2025, I wrote over 120 blog posts. Compared to previous years, these articles focused more on observing and reflecting on technology trends, rather than just technical tutorials.&lt;/p&gt;
&lt;p&gt;I started paying attention to deeper questions: How will AI infrastructure evolve? What does Beijing’s open source initiative mean for the AI industry? What ripple effects might a tech acquisition trigger? These articles allowed me and my readers to not only see &amp;ldquo;what&amp;rdquo; technology is, but also &amp;ldquo;why&amp;rdquo; and &amp;ldquo;what’s next.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="user-experience-making-knowledge-easier-to-discover-and-consume"&gt;User Experience: Making Knowledge Easier to Discover and Consume&lt;/h2&gt;
&lt;p&gt;No matter how good the content is, if it can’t be easily found and read, its value is greatly diminished. In 2025, I invested significant effort into website functionality, with one goal: to provide readers with a smoother reading experience.&lt;/p&gt;
&lt;h3 id="comprehensive-search-upgrade"&gt;Comprehensive Search Upgrade&lt;/h3&gt;
&lt;p&gt;As the volume of content grew, the original search function could no longer meet demand. I redesigned the search system to support fuzzy search and result scoring, and optimized index loading performance. More importantly, the new search interface is more user-friendly, supporting keyboard navigation and category filtering so users can find what they want faster.&lt;/p&gt;
&lt;h3 id="multi-device-experience-optimization"&gt;Multi-Device Experience Optimization&lt;/h3&gt;
&lt;p&gt;Mobile reading experience has improved significantly. I refactored the mobile navigation and table of contents, making reading on phones much smoother. Dark mode is now more refined, fixing several display issues and ensuring images and diagrams look good on dark backgrounds.&lt;/p&gt;
&lt;h3 id="efficiency-revolution-in-content-distribution"&gt;Efficiency Revolution in Content Distribution&lt;/h3&gt;
&lt;p&gt;A major change was optimizing the WeChat Official Account publishing workflow. Previously, publishing website content to WeChat required manual handling of many details; now, it’s almost one-click export. This workflow automatically processes images, metadata, styles, and all details, reducing a half-hour task to just a few minutes.&lt;/p&gt;
&lt;p&gt;Additionally, I added a glossary feature for technical term highlighting and tooltips; improved SEO and social sharing metadata; and cleaned up outdated content. These seemingly minor improvements quietly enhance the user experience.&lt;/p&gt;
&lt;h2 id="content-evolution-more-dimensional-knowledge-expression"&gt;Content Evolution: More Dimensional Knowledge Expression&lt;/h2&gt;
&lt;p&gt;Looking back at content creation in 2025, I found clear changes in several dimensions.&lt;/p&gt;
&lt;h3 id="from-tutorials-to-observations"&gt;From Tutorials to Observations&lt;/h3&gt;
&lt;p&gt;Early content leaned toward technical tutorials and practical guides, showing &amp;ldquo;how to do.&amp;rdquo; This year, I focused more on &amp;ldquo;why&amp;rdquo; and &amp;ldquo;what are the trends.&amp;rdquo; I wrote more technology trend analyses, ecosystem maps, and in-depth case studies. These may not directly teach you how to use an API, but they help you understand the direction of technological evolution.&lt;/p&gt;
&lt;h3 id="from-chinese-to-bilingual"&gt;From Chinese to Bilingual&lt;/h3&gt;
&lt;p&gt;AI is a global wave and cannot be limited to the Chinese-speaking world. In 2025, I wrote bilingual documentation for almost all new AI tools, and important blog posts also have English versions. This increased the workload, but allowed the content to reach a broader audience.&lt;/p&gt;
&lt;h3 id="from-text-to-multimedia"&gt;From Text to Multimedia&lt;/h3&gt;
&lt;p&gt;Text is efficient, but not all knowledge is best expressed in words. This year, I used many architecture and schematic diagrams to explain complex concepts, adding 59 new charts. These visual elements lower the barrier to understanding, making abstract concepts more intuitive. I also optimized image display in dark mode to ensure consistent visual experience.&lt;/p&gt;
&lt;h2 id="development-approach-embracing-ai-assisted-programming"&gt;Development Approach: Embracing AI-Assisted Programming&lt;/h2&gt;
&lt;p&gt;2025 was not only a year of shifting content themes toward AI, but also a year of deep practice in AI-assisted programming.&lt;/p&gt;
&lt;p&gt;I developed a VS Code plugin and created many prompts to automate repetitive tasks. I experimented with various AI programming tools and settled on a toolchain that suits me. I even migrated the website to Cloudflare Pages and used its edge computing services to develop a chatbot. These practices greatly improved development efficiency, giving me more time to focus on thinking and creating rather than mechanical coding.&lt;/p&gt;
&lt;p&gt;This made me realize: AI will not replace developers, but developers who use AI well will replace those who do not. I also shared more insights to help others master AI-assisted programming.&lt;/p&gt;
&lt;h2 id="looking-ahead-to-2026-keep-moving-forward"&gt;Looking Ahead to 2026: Keep Moving Forward&lt;/h2&gt;
&lt;p&gt;Looking back at 2025, the site underwent a profound transformation—from a Cloud Native tech blog to an AI infrastructure knowledge base. But this is just the beginning, not the end.&lt;/p&gt;
&lt;p&gt;Looking forward to 2026, I plan to continue deepening in several areas:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Enhancing the knowledge system&lt;/strong&gt;: Continue to supplement GPU infrastructure and AI Agent content, especially practical cases and performance tuning knowledge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tracking ecosystem evolution&lt;/strong&gt;: AI tools and frameworks iterate rapidly; I need to keep up with this fast-changing ecosystem and update content in a timely manner.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deepening engineering practice&lt;/strong&gt;: Share more practical AI engineering experience to help readers turn theory into practice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exploring knowledge connections&lt;/strong&gt;: Consider building a knowledge graph to connect different content sections, providing smarter navigation and recommendations.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;2025 was a year of change and growth. From Cloud Native to AI Native, from technical practice to ecosystem observation, both the content and functionality of the site have made qualitative leaps.&lt;/p&gt;
&lt;p&gt;What makes me happiest is that this transformation allowed me and my readers to stand at the forefront of the technology wave. We are not just learning new technologies, but thinking about how technology changes the world and the way we write software.&lt;/p&gt;
&lt;p&gt;The waves of technology keep evolving; only by actively embracing change can we continue to create value. Thank you to every reader for your companionship and support. I look forward to sharing more insights and practices in 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Further Reading&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Native Landscape&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/"&gt;2025 Blog Posts&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>The Butterfly Effect After Manus Was Acquired by Meta</title><link>https://jimmysong.io/blog/manus-meta-acquisition-butterfly-effect/</link><pubDate>Tue, 30 Dec 2025 03:30:51 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/manus-meta-acquisition-butterfly-effect/</guid><description>Manus&amp;#39;s acquisition by Meta sparked polarized opinions. This article explores the butterfly effect in AI applications and key lessons for entrepreneurs on growth strategies.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The success or failure of AI applications often lies not in the technology itself, but in the ability to scale delivery and create a closed loop.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/manus-meta-acquisition-butterfly-effect/banner.webp" data-img="https://assets.jimmysong.io/images/blog/manus-meta-acquisition-butterfly-effect/banner.webp" alt="Figure 1: The Butterfly Effect After Manus Was Acquired by Meta" data-caption="Figure 1: The Butterfly Effect After Manus Was Acquired by Meta"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: The Butterfly Effect After Manus Was Acquired by Meta&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="when-those-who-discuss-it-are-not-those-who-pay-for-it"&gt;When &amp;ldquo;Those Who Discuss It&amp;rdquo; Are Not &amp;ldquo;Those Who Pay for It&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;On December 30, 2025, a piece of news went viral: Manus was acquired by Meta for billions of dollars (&lt;a href="https://manus.im/blog/manus-joins-meta-for-next-era-of-innovation" target="_blank" rel="noopener"&gt;Manus Joins Meta for Next Era of Innovation&lt;/a&gt;). This startup, founded in China and under pressure from tech giants since its inception, completed a whirlwind journey in less than a year—from explosive growth, relocating to Singapore, to being acquired by a global giant.&lt;/p&gt;
&lt;p&gt;According to Manus&amp;rsquo;s official statement, its products and subscriptions will continue to be available via the app and website, and the company will remain operational in Singapore. The team will join Meta to provide general Agent capabilities for Meta&amp;rsquo;s consumer and enterprise products (including Meta AI).&lt;/p&gt;
&lt;p&gt;Rather than focusing on &amp;ldquo;who won,&amp;rdquo; I&amp;rsquo;m more interested in the chain reaction this event triggered: it activated completely opposite judgment systems among different groups, and this split is reshaping the growth paths and strategies for AI applications and startups.&lt;/p&gt;
&lt;h2 id="two-public-opinion-arenas-blessings-and-doubts-coexist"&gt;Two Public Opinion Arenas: Blessings and Doubts Coexist&lt;/h2&gt;
&lt;p&gt;After Manus was acquired, the mainstream sentiment in social circles was one of congratulations and excitement. Many saw it as a stellar example of a Chinese team going global—achieving remarkable results in the most competitive field in a very short time.&lt;/p&gt;
&lt;p&gt;Meanwhile, the comment sections of public accounts became &amp;ldquo;venting valves for counter-narratives,&amp;rdquo; with skepticism centering on three main points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether the technology has real barriers (e.g., &amp;ldquo;there are countless similar products,&amp;rdquo; &amp;ldquo;it&amp;rsquo;s not hard for big companies to build their own&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;Valuation and bubble concerns (e.g., &amp;ldquo;another case of the AI bubble&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;Distrust in the buyer&amp;rsquo;s judgment (e.g., &amp;ldquo;giants making desperate bets,&amp;rdquo; &amp;ldquo;history repeating itself&amp;rdquo;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This divergence isn&amp;rsquo;t about who understands AI better, but about different evaluation frameworks: social circles focus on &amp;ldquo;trajectory and outcome,&amp;rdquo; while comment sections focus on &amp;ldquo;legitimacy and worthiness.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="where-does-the-100m-arr-come-from-the-target-users-arent-in-our-social-circles"&gt;Where Does the $100M ARR Come From: The Target Users Aren&amp;rsquo;t in Our Social Circles&lt;/h2&gt;
&lt;p&gt;Many people are impressed by Manus&amp;rsquo;s marketing buzz and controversies, which can lead to skepticism. But if it achieved a &amp;ldquo;strict $100M ARR&amp;rdquo; in 10 months, one fact is clear: &lt;strong&gt;its revenue doesn&amp;rsquo;t depend on broad consensus, but comes from a highly concentrated group of global users with strong willingness to pay.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Manus&amp;rsquo;s core user profile is closer to &amp;ldquo;individuals as production units,&amp;rdquo; including freelancers, indie developers, independent researchers, and key deliverers in small and medium businesses. They don&amp;rsquo;t care about debates over &amp;ldquo;wrapping&amp;rdquo; or not; they care about &amp;ldquo;can I deliver end-to-end tasks,&amp;rdquo; and &amp;ldquo;can this help me hire one less person, work fewer late nights, or avoid juggling ten tools.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This leads to a counterintuitive phenomenon: &lt;strong&gt;those who discuss the most may not pay, while those who pay steadily are often silent.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For these users, tools are not identity badges—they are profit levers.&lt;/p&gt;
&lt;h2 id="three-lessons-for-entrepreneurs-the-growth-paradigm-in-the-ai-application-era-has-changed"&gt;Three Lessons for Entrepreneurs: The Growth Paradigm in the AI Application Era Has Changed&lt;/h2&gt;
&lt;p&gt;Based on the above, the Manus case offers three lessons for entrepreneurs:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Growth No Longer Equals Positive Reviews&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;AI applications can commercialize first and build consensus later. Public opinion can remain divided for a long time, but cash flow doesn&amp;rsquo;t wait for unified recognition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Heavy Marketing&amp;rdquo; Is Becoming a Capability, Not a Stigma&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;As foundational models and capabilities spread rapidly, differentiation is quickly erased. Being seen, understood, and paid for is itself part of the moat. Not all marketing deserves respect, but &amp;ldquo;distribution and mindshare&amp;rdquo; have become unavoidable battlegrounds for AI applications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization Is No Longer a Bonus, but May Be a Survival Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From payment willingness, compliance boundaries, talent density to valuation systems, market structure means many teams &amp;ldquo;can only complete the loop overseas.&amp;rdquo; It&amp;rsquo;s not romantic, but it&amp;rsquo;s reality.&lt;/p&gt;
&lt;h2 id="a-personal-reflection"&gt;A Personal Reflection&lt;/h2&gt;
&lt;p&gt;As someone long engaged in cloud native and AI infrastructure, I&amp;rsquo;m used to evaluating products by their &amp;ldquo;technical barriers.&amp;rdquo; But cases like Manus remind me: at the AI application layer, barriers may not first appear in models or code, but often in &lt;strong&gt;organizational speed, productization capability, delivery loop, and distribution efficiency&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When a system can reliably turn &amp;ldquo;capability&amp;rdquo; into &amp;ldquo;results,&amp;rdquo; it has built a commercial moat—even if its tech stack doesn&amp;rsquo;t meet outsiders&amp;rsquo; ideals of &amp;ldquo;purity.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The biggest butterfly effect of Manus being acquired by Meta may not be the deal itself, but making more entrepreneurs realize: &lt;strong&gt;in the AI era, the winning move is shifting from &amp;ldquo;what model you use&amp;rdquo; to &amp;ldquo;whether you can deliver results at scale.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The acquisition of Manus by Meta is not just a convergence of capital and technology, but also a microcosm of the changing growth paradigm in the AI application era. For entrepreneurs, understanding and mastering &amp;ldquo;user structure,&amp;rdquo; &amp;ldquo;distribution capability,&amp;rdquo; and &amp;ldquo;global closed loops&amp;rdquo; will be key to future competition.&lt;/p&gt;</content:encoded></item><item><title>AI Infra Open Source in China: Analysis of Beijing and Shanghai's Plans</title><link>https://jimmysong.io/blog/beijing-open-source-plan-ai-infra-analysis/</link><pubDate>Thu, 25 Dec 2025 10:01:13 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/beijing-open-source-plan-ai-infra-analysis/</guid><description>Beijing and Shanghai&amp;#39;s open source plans reveal opportunities and challenges for China&amp;#39;s AI infrastructure, balancing technology and governance.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Institutionalized open source marks a new starting point for China&amp;rsquo;s AI Infra, but true breakthroughs and risks lie in the engineering and governance details.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="perspective-on-beijing-and-shanghais-open-source-plans"&gt;Perspective on Beijing and Shanghai&amp;rsquo;s Open Source Plans&lt;/h2&gt;
&lt;p&gt;Using the simultaneous release of open source ecosystem plans by Beijing and Shanghai as a lens, and drawing on China&amp;rsquo;s past foundation practices and international open source governance experience, this article explores the real opportunities, structural constraints, and potential risks as AI Infrastructure (AI Infra, Artificial Intelligence Infrastructure) enters a new phase of institutionalized open source.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beijing-open-source-plan-ai-infra-analysis/banner.webp" data-img="https://assets.jimmysong.io/images/blog/beijing-open-source-plan-ai-infra-analysis/banner.webp" alt="Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans" data-caption="Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-compare-beijing-and-shanghai-together"&gt;Why Compare Beijing and Shanghai Together&lt;/h2&gt;
&lt;p&gt;It is rare for me to write an article solely because of a local policy document. However, during Christmas, both Beijing and Shanghai&amp;rsquo;s Bureaus of Economy and Information Technology released their respective open source ecosystem construction plans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/9YEL1HORWatsol3nRT596w" target="_blank" rel="noopener"&gt;Building an Open Source Innovation Highland! Beijing Releases Open Source Ecosystem Construction Implementation Plan&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/QZl66fUllKiePwQ7euhiGQ" target="_blank" rel="noopener"&gt;Shanghai&amp;rsquo;s Implementation Plan for Strengthening the Open Source System | Infographic&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This time, the fact that both cities released their plans on the same day sends a signal worth serious attention: China is attempting to advance open source in a more systematic and institutionalized way, especially regarding open source capabilities related to AI Infra.&lt;/p&gt;
&lt;p&gt;If you only look at Beijing&amp;rsquo;s plan, it is easy to interpret it as a local industrial policy upgrade. But when you consider both Beijing and Shanghai&amp;rsquo;s plans together, it looks more like a clearly defined &amp;ldquo;dual-center structure.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The question is no longer whether to develop open source, but:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the AI era, what institutional forms, engineering paths, and governance models will open source take?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="open-source-as-industrial-infrastructure-engineering"&gt;Open Source as &amp;ldquo;Industrial Infrastructure Engineering&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Both Beijing and Shanghai&amp;rsquo;s plans reflect a highly consistent judgment:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Open source is no longer seen as a spontaneous community activity, but as an industrial infrastructure capability that requires systematic construction.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is especially evident in the field of AI Infra.&lt;/p&gt;
&lt;p&gt;Issues such as computing power scheduling, model evaluation, toolchains, data elements, license compliance, and supply chain security—previously hidden in &amp;ldquo;engineering details&amp;rdquo;—are now systematically incorporated into policy language for the first time. This at least shows that decision-makers have realized:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI competition is not only about model parameter scale&lt;/li&gt;
&lt;li&gt;It is even more about toolchains, infrastructure, evaluation systems, and engineering capabilities&lt;/li&gt;
&lt;li&gt;These capabilities are naturally more suitable for building public foundations through open source&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this respect, Beijing and Shanghai are highly aligned.&lt;/p&gt;
&lt;h2 id="two-open-source-paths-infra-vs-platform"&gt;Two Open Source Paths: Infra vs. Platform&lt;/h2&gt;
&lt;p&gt;When we zoom in, the differences between the two plans become clear.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Beijing: &amp;ldquo;Foundation-Oriented&amp;rdquo; Open Source Path for AI Infra&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Beijing&amp;rsquo;s plan focuses on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Heterogeneous computing power scheduling&lt;/li&gt;
&lt;li&gt;Model evaluation toolchains&lt;/li&gt;
&lt;li&gt;Data elements and data governance&lt;/li&gt;
&lt;li&gt;RISC-V software-hardware collaboration&lt;/li&gt;
&lt;li&gt;SBOM, license compatibility, open source compliance&lt;/li&gt;
&lt;li&gt;Supply chain security and industrial resilience&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a typical perspective of &amp;ldquo;treating AI as an infrastructure problem.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;It is less concerned with the number of projects or community size, and more with:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether reusable engineering capabilities can be formed&lt;/li&gt;
&lt;li&gt;Whether these can be trusted by industry and government over the long term&lt;/li&gt;
&lt;li&gt;Whether they can stand up to scrutiny in terms of security, compliance, and governance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To some extent, Beijing is answering the question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;How can open source become a &amp;ldquo;governable, auditable, and scalable public capability&amp;rdquo;?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Shanghai: &amp;ldquo;Scale and Internationalization&amp;rdquo; Path for AI Platform&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In contrast, Shanghai&amp;rsquo;s plan has a different focus:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Building an international open source community for artificial intelligence&lt;/li&gt;
&lt;li&gt;Covering the entire platform chain from development, training, testing, hosting, to operation&lt;/li&gt;
&lt;li&gt;Overseas sites, multilingual support, international activities&lt;/li&gt;
&lt;li&gt;Resource linkage through computing vouchers and model vouchers&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Open source platform first release / global simultaneous release&amp;rdquo; dual-release mechanism&lt;/li&gt;
&lt;li&gt;Clear targets for community, enterprise, and developer scale&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Shanghai cares more about:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How open source can achieve scale effects&lt;/li&gt;
&lt;li&gt;How it can support the growth of commercial enterprises&lt;/li&gt;
&lt;li&gt;How it can be seen and adopted globally&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a path of &amp;ldquo;treating open source as a global digital product and platform capability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="together-a-complete-but-tension-filled-structure"&gt;Together: A Complete but Tension-Filled Structure&lt;/h2&gt;
&lt;p&gt;When viewed together, Beijing and Shanghai&amp;rsquo;s plans form a more complete picture:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Beijing is responsible for &amp;ldquo;making open source solid,&amp;rdquo; while Shanghai is responsible for &amp;ldquo;taking open source global.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Structurally, this is a clear division of labor:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Beijing focuses on institutions, governance, and foundational capabilities&lt;/li&gt;
&lt;li&gt;Shanghai focuses on community, commercialization, and international communication&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These two paths are not in conflict; in theory, they are even complementary. The real question is whether they can form positive feedback in practice, rather than operating in silos.&lt;/p&gt;
&lt;h2 id="cautious-attitude-toward-institutionalized-platformized-open-source"&gt;Cautious Attitude Toward &amp;ldquo;Institutionalized, Platformized Open Source&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Precisely because both plans are so &amp;ldquo;systematic,&amp;rdquo; I am even more cautious.&lt;/p&gt;
&lt;p&gt;The reason is simple: this is not China&amp;rsquo;s first attempt to promote open source through foundations, associations, or platforms.&lt;/p&gt;
&lt;p&gt;Over the past decade, we have seen similar paths repeatedly, and recurring structural problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The difficulty of establishing neutrality and multi-party trust is extremely high&lt;/li&gt;
&lt;li&gt;There is a huge gap between showcase metrics (quantity, activities, certifications) and ecosystem strength&lt;/li&gt;
&lt;li&gt;Commercialization and long-term maintenance mechanisms are hard to sustain&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These problems will not disappear just because the plans are more comprehensive.&lt;/p&gt;
&lt;h2 id="four-risks-to-watch-under-the-dual-plans"&gt;Four Risks to Watch Under the Dual Plans&lt;/h2&gt;
&lt;p&gt;If we are to &amp;ldquo;listen to their words and watch their actions,&amp;rdquo; I would focus on the following four risks:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will Metrics Hijack Engineering Reality&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When &amp;ldquo;internationally influential projects,&amp;rdquo; &amp;ldquo;star projects,&amp;rdquo; and &amp;ldquo;first-release projects&amp;rdquo; become hard metrics, will this induce packaging, migration, and short-term hype, rather than truly solving engineering problems?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will It Slide Toward Platform Centralism&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The long-term pattern of AI Infra is closer to a model that prioritizes protocols, standards, and interoperability. If it eventually evolves into &amp;ldquo;a few platforms concentrating resources and discourse power,&amp;rdquo; it may be efficient in the short term but will suppress external participation and international collaboration in the long run.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is Internationalization Underestimated as an &amp;ldquo;Operational Issue&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;True international collaboration is never just about language, sites, or events; it also involves governance structures, compliance boundaries, and supply chain trust.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will Application Demonstrations Become One-Off Projects&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If &amp;ldquo;first plans&amp;rdquo; and &amp;ldquo;computing vouchers&amp;rdquo; are just procurement tactics without continuous iteration and community feedback mechanisms, the long-term benefit to the ecosystem will be very limited.&lt;/p&gt;
&lt;h2 id="what-are-the-hard-results-of-ai-infra-open-source-after-three-years"&gt;What Are the &amp;ldquo;Hard Results&amp;rdquo; of AI Infra Open Source After Three Years&lt;/h2&gt;
&lt;p&gt;If we review the success of this round of institutionalized open source after three years, I would look for three types of results:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether de facto standards and interoperable ecosystems have emerged, including scheduling interfaces, evaluation benchmarks, Agent tool invocation protocols, and observability semantics.&lt;/li&gt;
&lt;li&gt;Whether compliance and supply chain security have become public capabilities—SBOM, license compatibility, vulnerability monitoring—truly productized and service-oriented.&lt;/li&gt;
&lt;li&gt;Whether a sustainable maintenance business mechanism has been established, allowing core maintainers to stay long-term, rather than relying on passion and subsidies.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;If I were to use a North Star metric to measure the success of these plans, it would be the emergence of several outstanding open source commercial companies rooted in China and serving the world.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The open source ecosystem plans of Beijing and Shanghai mark a new phase of institutionalization and engineering for AI Infra open source in China. Over the next three years, the real achievements will not be about meeting targets, but about forming sustainable engineering capabilities, de facto standards, and maintenance mechanisms. Only through continuous participation and practice can open source become the public foundation of AI infrastructure.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jxj.beijing.gov.cn/zwgk/2024zcwj/202512/t20251224_4360437.html" target="_blank" rel="noopener"&gt;Beijing Open Source Ecosystem Construction Implementation Plan (2026–2028) - jxj.beijing.gov.cn&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/QZl66fUllKiePwQ7euhiGQ" target="_blank" rel="noopener"&gt;Shanghai&amp;rsquo;s Implementation Plan for Strengthening the Open Source System | Infographic - mp.weixin.qq.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>From 2025 Onwards, Software Engineering Shifts from Code-Centric to Runtime and Cost-Centric</title><link>https://jimmysong.io/blog/software-engineering-shift-runtime-cost-2025/</link><pubDate>Wed, 24 Dec 2025 14:59:11 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/software-engineering-shift-runtime-cost-2025/</guid><description>In 2025, software engineering shifts from code-centric to runtime and cost governance. AI and Agents move complexity to runtime, compute, and budget layers, reshaping engineering value.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;In 2025, the core of software engineering is no longer just about code itself, but about runtime controllability and cost governance. This shift is fundamentally reshaping the industry&amp;rsquo;s underlying logic.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Looking back at 2025, I became increasingly aware that this year was not about &amp;ldquo;code becoming unimportant,&amp;rdquo; but rather that &lt;strong&gt;the value coordinates of engineering have shifted as a whole&lt;/strong&gt;. For more than a decade, software engineering has focused on code quality, architectural evolution, and delivery efficiency. But starting in 2025, the key to system success is shifting—&lt;strong&gt;towards whether the runtime is controllable and whether costs are governable&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This is not just a slogan, but a conclusion repeatedly validated by my real-world experiences throughout the year.&lt;/p&gt;
&lt;h2 id="my-2025-from-platform-engineering-to-runtime-challenges"&gt;My 2025: From &amp;ldquo;Platform Engineering&amp;rdquo; to &amp;ldquo;Runtime Challenges&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;In my annual review, I noted a clear change: I spent less time on &amp;ldquo;how to write a good system,&amp;rdquo; and more time on &amp;ldquo;how to keep the system running stably, reliably, and affordably.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This shift in focus is a natural extension of a decade of cloud native evolution.&lt;/p&gt;
&lt;p&gt;The following timeline diagram illustrates how my focus has changed over recent years:
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/focus-shift-timeline-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/focus-shift-timeline-en.svg" alt="Figure 1: My Focus Shift Timeline" data-caption="Figure 1: My Focus Shift Timeline"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: My Focus Shift Timeline&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;My focus shifted from cloud native platform engineering to LLM application engineering, then to AI infrastructure, and finally to Agentic Runtime with governance and cost control.&lt;/p&gt;
&lt;p&gt;When AI workloads truly enter business scenarios, the core challenges engineers face also change:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Are inference, training, and evaluation competing for the same compute pool?&lt;/li&gt;
&lt;li&gt;Is GPU utilization consistently below expectations?&lt;/li&gt;
&lt;li&gt;Does cost scale linearly and uncontrollably with concurrency?&lt;/li&gt;
&lt;li&gt;Does the system have failure isolation and replay capabilities?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These issues go far beyond the code level.&lt;/p&gt;
&lt;h2 id="industry-consensus-ai-is-shifting-the-focus-of-engineering"&gt;Industry Consensus: AI Is Shifting the Focus of Engineering&lt;/h2&gt;
&lt;p&gt;By 2025, an industry consensus is emerging: AI is rewriting software engineering. But the real change is not happening in the IDE or code completion speed—it is reflected in &lt;strong&gt;the migration of engineering complexity&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Previously, complexity was concentrated in code and interfaces, and problems were solved through abstraction, refactoring, and testing.&lt;/p&gt;
&lt;p&gt;Now, complexity has shifted to the runtime, resource, and cost layers, and must be addressed through scheduling, isolation, observability, and governance.&lt;/p&gt;
&lt;p&gt;This is why the same AI tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Serve as &amp;ldquo;accelerators&amp;rdquo; for junior engineers&lt;/li&gt;
&lt;li&gt;But act as &amp;ldquo;magnifiers&amp;rdquo; for senior engineers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;AI tools amplify whether you truly understand how systems run in production.&lt;/p&gt;
&lt;h2 id="why-cost-becomes-a-first-principle"&gt;Why &amp;ldquo;Cost&amp;rdquo; Becomes a First Principle&lt;/h2&gt;
&lt;p&gt;In traditional cloud native systems, low CPU utilization is often just an efficiency issue; but in AI systems, &lt;strong&gt;low GPU utilization is often a cash flow problem&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In 2025, I repeatedly encountered scenarios like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Resources &amp;ldquo;seem insufficient,&amp;rdquo; but utilization is not actually high&lt;/li&gt;
&lt;li&gt;Scaling up to solve queuing issues ends up increasing unit costs&lt;/li&gt;
&lt;li&gt;The system lacks clear budget and quota boundaries, so throttling becomes the only way to stop the bleeding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The root cause of these phenomena is not model selection, but &lt;strong&gt;the lack of a runtime and cost control plane tailored for AI workloads&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The following flowchart visually illustrates the cyclical relationship between GPU resources and cost pressures:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/gpu-cost-cycle-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/gpu-cost-cycle-en.svg" alt="Figure 2: GPU Resource and Cost Cycle in AI Systems" data-caption="Figure 2: GPU Resource and Cost Cycle in AI Systems"
width="2263"
height="320"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: GPU Resource and Cost Cycle in AI Systems&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In AI systems, limited GPU supply leads to queuing and waiting, which causes throughput to drop. Attempts to solve this through blind scaling only increase unit costs and create budget pressure, ultimately forcing the adoption of finer scheduling and governance strategies.&lt;/p&gt;
&lt;p&gt;Engineering problems ultimately manifest as cost issues.&lt;/p&gt;
&lt;h2 id="the-rise-of-agents-the-real-challenge-is-at-runtime"&gt;The Rise of Agents: The Real Challenge Is at Runtime&lt;/h2&gt;
&lt;p&gt;In 2025, Agent (Intelligent Agent, Agent, Intelligent Agent) became a hot topic; by 2026, it will enter the &amp;ldquo;can it actually run&amp;rdquo; stage.&lt;/p&gt;
&lt;p&gt;The challenge for Agents has never been about &amp;ldquo;how smart they are,&amp;rdquo; but rather:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether there are clear permission and data boundaries&lt;/li&gt;
&lt;li&gt;Whether they run in an isolated execution environment&lt;/li&gt;
&lt;li&gt;Whether they can be observed, evaluated, and replayed&lt;/li&gt;
&lt;li&gt;Whether they are subject to explicit cost and budget constraints&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These capabilities form the outline of &lt;strong&gt;Agentic Runtime (Agentic Runtime, Intelligent Agent Runtime)&lt;/strong&gt; that I have been trying to clarify throughout the year.&lt;/p&gt;
&lt;p&gt;The following flowchart shows the core capability layers of Agentic Runtime:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/agentic-runtime-layers-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/agentic-runtime-layers-en.svg" alt="Figure 3: Agentic Runtime Capability Layers" data-caption="Figure 3: Agentic Runtime Capability Layers"
width="463"
height="983"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Agentic Runtime Capability Layers&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Agentic Runtime builds from the foundation of Agents and workflows, connecting through orchestration and tool protocols, with the runtime managing state, memory, and evaluation. It provides secure execution environments (Sandbox and Policy), and ultimately implements a resource and cost control plane that unifies GPU, quota, and billing management.&lt;/p&gt;
&lt;p&gt;Without a runtime, an Agent is just a demo; without cost constraints, an Agent is just a risk amplifier.&lt;/p&gt;
&lt;h2 id="outlook-for-2026-the-foundation-of-engineering-matters-again"&gt;Outlook for 2026: The &amp;ldquo;Foundation&amp;rdquo; of Engineering Matters Again&lt;/h2&gt;
&lt;p&gt;Looking ahead to 2026, I remain cautiously optimistic.&lt;/p&gt;
&lt;p&gt;I do not believe the future belongs to &amp;ldquo;those who write the best prompts,&amp;rdquo; but more likely to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Those who understand runtime boundaries&lt;/li&gt;
&lt;li&gt;Those who can govern compute as a constrained resource&lt;/li&gt;
&lt;li&gt;Those who design AI systems as long-running systems&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From 2025 onwards, software engineering is no longer code-centric, but &lt;strong&gt;runtime and cost-centric&lt;/strong&gt;. This is not a regression, but a return: a return to being responsible for the whole system and for real-world constraints.&lt;/p&gt;
&lt;p&gt;For me personally, this is both a year-end summary and the direction I will continue to invest in for the coming years.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;In 2025, the focus of software engineering has shifted from code itself to runtime and cost governance. The rise of AI and Agents has not diminished the value of engineering, but has pushed complexity to a higher level. In the future, understanding runtime, managing compute and cost will become the new core competencies for engineers. I hope this year-end review provides some inspiration and reflection for fellow professionals.&lt;/p&gt;</content:encoded></item><item><title>From Cloud Native to AI Native: Why Kubernetes Is the Foundation for Next-Gen AI Agents</title><link>https://jimmysong.io/blog/ai-native-from-cloud-native/</link><pubDate>Wed, 24 Dec 2025 12:25:52 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-native-from-cloud-native/</guid><description>Explores why AI Agents need Kubernetes infrastructure and how Agent orchestration, MCP services, and AI gateways enable production-ready AI architectures.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;As a long-time practitioner in the cloud native field, I am increasingly convinced of one thing: &lt;strong&gt;AI Agents are not just a change in application form, but a migration of infrastructure paradigms.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;As artificial intelligence evolves from demos and copilots to systems that truly take on tasks and responsibilities, &lt;strong&gt;AI Agents&lt;/strong&gt; are becoming the new execution units in enterprise IT architectures. They not only &amp;ldquo;think,&amp;rdquo; but also &lt;strong&gt;act&lt;/strong&gt;: they can invoke tools, access systems, and collaborate to achieve goals.&lt;/p&gt;
&lt;p&gt;This raises an important question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What kind of infrastructure should such systems run on?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In my view, Kubernetes remains a solid choice for large-scale scenarios—but only if we &lt;strong&gt;reimagine Kubernetes in an AI-native way&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="cloud-native-challenges-for-production-grade-ai-agents"&gt;Cloud Native Challenges for Production-Grade AI Agents&lt;/h2&gt;
&lt;p&gt;In real production environments, AI Agents expose infrastructure needs that are fundamentally different from traditional microservices. Agents are not &amp;ldquo;just another HTTP service&amp;rdquo;; they have three distinct characteristics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Behavior is non-deterministic&lt;/strong&gt; (driven by model inference)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Execution paths are dynamic&lt;/strong&gt; (tool invocation cannot be fully enumerated in advance)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decisions must be auditable, constrained, and reviewable&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If we simply apply existing cloud native infrastructure, we quickly hit bottlenecks.&lt;/p&gt;
&lt;p&gt;The following table summarizes the main challenges and risks AI Agents face in cloud native environments:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Challenge Category&lt;/th&gt;
&lt;th&gt;Real Needs of Agents&lt;/th&gt;
&lt;th&gt;What Happens If Missing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Policy &amp;amp; Security&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dynamic control of tool and data access based on context, identity, and task&lt;/td&gt;
&lt;td&gt;Agents have &amp;ldquo;superuser&amp;rdquo; privileges, risks are uncontrollable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not just &amp;ldquo;did it succeed,&amp;rdquo; but also &lt;strong&gt;why was this decision made&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hard to debug, hard to review, hard to hold accountable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance &amp;amp; Consistency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Platform-level guardrails enforce organizational policies&lt;/td&gt;
&lt;td&gt;Each Agent could become a &amp;ldquo;shadow AI&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Challenges and Risks for AI Agents in Cloud Native Environments
&lt;/figcaption&gt;
&lt;p&gt;All these issues point to one conclusion:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI Agents must be treated as first-class citizens in Kubernetes, not just ordinary workloads.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="core-architecture-making-agents-native-kubernetes-objects"&gt;Core Architecture: Making Agents Native Kubernetes Objects&lt;/h2&gt;
&lt;p&gt;Looking back at the evolution of cloud native technologies, we&amp;rsquo;ve gone through similar stages:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Physical machines → Virtual machines&lt;/li&gt;
&lt;li&gt;Virtual machines → Containers&lt;/li&gt;
&lt;li&gt;Containers → Microservices&lt;/li&gt;
&lt;li&gt;Microservices → Declarative, governable platforms&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;AI Agents are simply the next step.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A production-ready AI Agent architecture requires at least three layers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Agent Orchestration Layer&lt;/strong&gt;: Declaratively define Agents&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool Service-ization Layer (MCP Services)&lt;/strong&gt;: Turn capabilities into governable services&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Native Data Plane / Gateway&lt;/strong&gt;: Unify policy, security, and protocols&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="agent-orchestration-layer-declarative-agent-management"&gt;Agent Orchestration Layer: Declarative Agent Management&lt;/h2&gt;
&lt;p&gt;Agents should no longer be &amp;ldquo;runtime objects&amp;rdquo; inside an SDK—they should be managed like Pods or Deployments.&lt;/p&gt;
&lt;p&gt;Key concepts:&lt;/p&gt;
&lt;h3 id="agents-as-kubernetes-resources"&gt;Agents as Kubernetes Resources&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Agents are defined using &lt;strong&gt;CRD (CustomResourceDefinition)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Lifecycle managed via &lt;code&gt;kubectl&lt;/code&gt; or GitOps&lt;/li&gt;
&lt;li&gt;Agent &lt;strong&gt;models, tools, and policies&lt;/strong&gt; are all explicitly declared&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A typical Agent definition includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agent logic&lt;/strong&gt; (inference loop)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model configuration&lt;/strong&gt; (specifying which large language model to use)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Callable toolset&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;This closely mirrors how we once decomposed &amp;ldquo;applications&amp;rdquo; into Deployments, Services, and ConfigMaps.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="tool-service-ization-layer-mcp-services-are-essential"&gt;Tool Service-ization Layer: MCP Services Are Essential&lt;/h2&gt;
&lt;p&gt;In Agent architectures, &lt;strong&gt;tools&lt;/strong&gt; are where real &amp;ldquo;actions&amp;rdquo; happen.&lt;/p&gt;
&lt;p&gt;Early MCP tools were often:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Local processes&lt;/li&gt;
&lt;li&gt;Tightly coupled to a single Agent&lt;/li&gt;
&lt;li&gt;Lacking versioning, permissions, and auditing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is unsustainable in enterprise environments.&lt;/p&gt;
&lt;h3 id="the-essence-of-mcp-service-ization"&gt;The Essence of MCP Service-ization&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Tools → &lt;strong&gt;Remote services&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Services → &lt;strong&gt;Kubernetes native workloads&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Capabilities → &lt;strong&gt;Reusable, governable, auditable&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This step is fundamentally similar to how we once turned scripts into microservices.&lt;/p&gt;
&lt;h2 id="ai-native-gateway-the-control-plane-entry-for-the-agent-world"&gt;AI Native Gateway: The &amp;ldquo;Control Plane Entry&amp;rdquo; for the Agent World&lt;/h2&gt;
&lt;p&gt;As the number of Agents grows and tools/models diversify, &lt;strong&gt;connectivity itself becomes a system risk&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Traditional API Gateways do not understand scenarios like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP&lt;/li&gt;
&lt;li&gt;Agent-to-Agent (A2A) communication&lt;/li&gt;
&lt;li&gt;Model invocation context&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thus, we need an &lt;strong&gt;AI native gateway&lt;/strong&gt; dedicated to mediation and governance.&lt;/p&gt;
&lt;p&gt;It must understand at least three types of traffic:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A2T&lt;/strong&gt;: Agent → Tool&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A2L&lt;/strong&gt;: Agent → LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A2A&lt;/strong&gt;: Agent ↔ Agent&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And enforce, across these paths:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Identity and authorization&lt;/li&gt;
&lt;li&gt;Policy and guardrails&lt;/li&gt;
&lt;li&gt;Auditing and rate limiting&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="architecture-overview"&gt;Architecture Overview&lt;/h2&gt;
&lt;p&gt;The diagram below illustrates the core layers and traffic paths of an AI-native system on Kubernetes:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-native-from-cloud-native/5be5cb784d4b228006abdf024bb99d6f.svg" data-img="https://assets.jimmysong.io/images/blog/ai-native-from-cloud-native/5be5cb784d4b228006abdf024bb99d6f.svg" alt="Figure 3: AI Native Architecture Layers and Traffic Paths" data-caption="Figure 3: AI Native Architecture Layers and Traffic Paths"
width="1311"
height="1642"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: AI Native Architecture Layers and Traffic Paths&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;AI Agents do not negate cloud native; on the contrary:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI Agents are the natural extension of cloud native in the era of intelligence.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Declarative → Agent definitions&lt;/li&gt;
&lt;li&gt;Service → MCP Services&lt;/li&gt;
&lt;li&gt;Service Mesh → AI Native Gateway&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If Kubernetes is the &amp;ldquo;automated factory,&amp;rdquo; then AI Agents are the &lt;strong&gt;intelligent workers who actually get things done&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;And the AI native gateway is the &lt;strong&gt;security and governance system tailored for these intelligent workers&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This is not an optional architecture—it is &lt;strong&gt;the only path for AI to reach production&lt;/strong&gt;.&lt;/p&gt;</content:encoded></item><item><title>AI Open Source Landscape: A One-Stop Guide to AI Project Navigation and Scoring System</title><link>https://jimmysong.io/blog/ai-oss-landscape-intro/</link><pubDate>Tue, 23 Dec 2025 08:34:05 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-oss-landscape-intro/</guid><description>Comprehensive introduction to the AI Open Source Landscape&amp;#39;s positioning, interface, scoring model, and data mechanisms to help developers efficiently discover quality AI projects.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The AI Open Source Landscape is not just a project directory, but an innovative attempt to bring transparency and quantifiability to the AI open source ecosystem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note: This article is intended for general readers and focuses on platform features and usage scenarios. If you want to see the technical details and formulas behind the scoring, please refer to:&lt;/strong&gt; AI Project Scoring and Inclusion Criteria.&lt;/p&gt;
&lt;h2 id="project-background-and-positioning"&gt;Project Background and Positioning&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape aims to provide developers, researchers, and enterprise users with a one-stop navigation and evaluation platform for AI open source projects. With the rapid development of large language models (LLM, Large Language Model), multimodal models (Multimodal Model), and other AI technologies, the open source community has seen a surge of innovative projects. However, information is scattered and quality varies, making it difficult for users to filter and make decisions.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/ai-oss-landscape.webp" data-img="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/ai-oss-landscape.webp" alt="Figure 1: AI Open Source Landscape" data-caption="Figure 1: AI Open Source Landscape"
width="3653"
height="2494"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: AI Open Source Landscape&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The AI Open Source Landscape systematically collects mainstream AI open source projects. As of the time of writing, it has included 851 open source projects. This landscape combines a multi-dimensional scoring system to help users efficiently discover, compare, and select the most suitable AI tools and frameworks for their needs. The platform not only focuses on models themselves, but also covers datasets, inference engines, evaluation tools, application frameworks, and the entire ecosystem chain, striving to promote transparency, quantifiability, and sustainable development in the AI open source ecosystem.&lt;/p&gt;
&lt;h2 id="main-interface-and-feature-highlights"&gt;Main Interface and Feature Highlights&lt;/h2&gt;
&lt;p&gt;The platform homepage presents project distribution in both landscape and list views, supporting category filtering, keyword search, and tag navigation to help users quickly locate target projects.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/project-details.webp" data-img="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/project-details.webp" alt="Figure 2: Open Source Project Detail Page" data-caption="Figure 2: Open Source Project Detail Page"
width="2780"
height="2915"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Open Source Project Detail Page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For general readers, the main experience points include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Card view: One-sentence overview, star rating, and overall score for quick browsing and comparison.&lt;/li&gt;
&lt;li&gt;Health card: Displays overall health and key dimensions (activity, community, influence, sustainability) on the project page or sidebar, with the latest update marked for easy assessment of maintenance status.&lt;/li&gt;
&lt;li&gt;Detail page: Provides more background information, project links, and application scenarios to help you evaluate suitability for your needs.&lt;/li&gt;
&lt;li&gt;Smart badges: Visually display labels such as &amp;ldquo;Active&amp;rdquo;, &amp;ldquo;New Project&amp;rdquo;, &amp;ldquo;Popular&amp;rdquo;, &amp;ldquo;Archived&amp;rdquo; on cards, helping you quickly capture key project features.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you are interested in the specific rules for badge determination or scoring, detailed explanations are available on the Scoring Rules Page.&lt;/p&gt;
&lt;h2 id="scoring-and-ranking-mechanism"&gt;Scoring and Ranking Mechanism&lt;/h2&gt;
&lt;p&gt;The platform uses multi-dimensional scores to reflect the overall health and popularity of projects. The main dimensions include: &lt;strong&gt;Activity&lt;/strong&gt;, &lt;strong&gt;Community&lt;/strong&gt;, &lt;strong&gt;Quality&lt;/strong&gt;, &lt;strong&gt;Sustainability&lt;/strong&gt;, and the comprehensive &lt;strong&gt;Health&lt;/strong&gt; score. These scores help you quickly judge whether a project is suitable for production or experimentation.&lt;/p&gt;
&lt;h2 id="data-sources-and-update-mechanism"&gt;Data Sources and Update Mechanism&lt;/h2&gt;
&lt;p&gt;The platform&amp;rsquo;s data mainly comes from GitHub, project lists, official documentation, and community recommendations. We regularly and automatically synchronize and update metrics to ensure that the &amp;ldquo;last updated&amp;rdquo; and scores displayed on the interface reflect the current maintenance status of projects. Projects that have not been updated for a long time or are determined to be &amp;ldquo;inactive&amp;rdquo; are moved to the Archived Page. Archived projects remain searchable and retain historical scores, but will not appear in the default view of active rankings, making it easier for readers to focus on projects that are still maintained and active.&lt;/p&gt;
&lt;p&gt;For general readers, the key points are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The page displays key metrics and &amp;ldquo;last updated&amp;rdquo; time, helping you quickly judge whether a project is still maintained.&lt;/li&gt;
&lt;li&gt;The AI Open Source Landscape continuously iterates on the scoring model to improve fairness and differentiation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="how-to-contribute-and-correct-data"&gt;How to Contribute and Correct Data&lt;/h2&gt;
&lt;p&gt;If you want a project to be included or its data updated, you can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/rootsongjc/rootsongjc.github.io/issues/new?template=ai-resource.md" target="_blank" rel="noopener"&gt;Submit an AI Open Source Project Inclusion Request&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Keep the project&amp;rsquo;s README, License, documentation, and other information complete in the repository to facilitate our data collection and assessment.&lt;/li&gt;
&lt;li&gt;For faster synchronization or if you encounter data issues, contact the maintainers via project issues or raise a request in the site discussion area.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="typical-use-cases-or-user-feedback"&gt;Typical Use Cases or User Feedback&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape has been widely used in various scenarios such as AI developer selection, enterprise technology research, and academic studies. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Developers can quickly filter models or tools that meet their needs through the platform, saving significant research time.&lt;/li&gt;
&lt;li&gt;Enterprise technical teams use the ranking lists for competitor analysis and technology planning.&lt;/li&gt;
&lt;li&gt;Educational and research institutions refer to the landscape to understand trends in the AI open source ecosystem, supporting course design and topic selection.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Some users have commented that the platform is &amp;ldquo;comprehensive, well-structured, and fair in scoring,&amp;rdquo; greatly improving the efficiency of AI project selection and learning. Community suggestions continue to drive ongoing improvements in platform features and content.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape systematically and quantitatively organizes the AI open source ecosystem: the backend worker is responsible for reliable data collection and scoring calculations (supporting backfill and migration), while frontend components handle fast rendering and visualization (including smart badges, health cards, and metric explanations).&lt;/p&gt;
&lt;p&gt;If you want to learn more about the scoring details or participate in improvements:&lt;/p&gt;
&lt;p&gt;The community is welcome to join in evaluation, backfilling historical data, and refining scoring rules, working together to make the AI open source ecosystem more transparent and sustainable.&lt;/p&gt;</content:encoded></item></channel></rss>