<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Jimmy Song – Blog</title><link>https://jimmysong.io/blog/</link><description>Recent content in Blog on Jimmy Song</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>Jimmy Song</managingEditor><webMaster>Jimmy Song</webMaster><follow_challenge><feedId>51621818828612637</feedId><userId>59800919738273792</userId></follow_challenge><lastBuildDate>Tue, 26 Aug 2025 10:14:34 +0800</lastBuildDate><atom:link href="https://jimmysong.io/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>After HAMi Became a CNCF Incubating Project: Open Source Is Moving from Code to Consensus</title><link>https://jimmysong.io/blog/hami-cncf-incubating-consensus/</link><pubDate>Wed, 08 Jul 2026 13:51:44 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/hami-cncf-incubating-consensus/</guid><description>Code is cheap; consensus is the new scarce good.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;After HAMi became a CNCF Incubating project, I want to talk about something overlooked: AI is shifting the scarce resource of open source communities from code to consensus.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;On July 2, 2026, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; officially became a CNCF Incubating project (see &lt;a href="https://project-hami.io/blog/hami-cncf-incubating" target="_blank" rel="noopener"&gt;the announcement&lt;/a&gt;). For an open source project, this means more than recognition of its technical capability; it means that community governance, ecosystem building, and real-world adoption have all entered a new phase.&lt;/p&gt;
&lt;p&gt;But if you only read HAMi&amp;rsquo;s growth as &amp;ldquo;a GPU virtualization project succeeded,&amp;rdquo; you might miss the more important shift.&lt;/p&gt;
&lt;p&gt;My time building the HAMi community has left me with one increasingly strong feeling: AI is changing how open source communities produce. As AI coding drives the cost of producing code lower and lower, the core of competition in open source will no longer be who wrote the most code, but who can build stronger technical consensus, attract more contributors, and form a sustainable ecosystem network.&lt;/p&gt;
&lt;p&gt;This article is my attempt to lay out that argument clearly.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-incubating.webp" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-incubating.webp" alt="Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project" data-caption="Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project"
width="1024"
height="1536"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 9: Congratulations to HAMi for becoming a CNCF Incubating project&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="hamis-growth-shows-that-a-projects-real-asset-is-its-community"&gt;HAMi&amp;rsquo;s Growth Shows That a Project&amp;rsquo;s Real Asset Is Its Community&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s start with the data. Here is where HAMi stands today:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Stars&lt;/td&gt;
&lt;td&gt;3,700+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contributors&lt;/td&gt;
&lt;td&gt;Nearly 500, from 27 countries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Participating organizations&lt;/td&gt;
&lt;td&gt;Multiple, and growing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release cadence&lt;/td&gt;
&lt;td&gt;Once every three months&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: HAMi community key metrics (as of July 2026)
&lt;/figcaption&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-contributors-map.webp" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/hami-contributors-map.webp" alt="Figure 10: HAMi community contributors map, contributors from 27 countries" data-caption="Figure 10: HAMi community contributors map, contributors from 27 countries"
width="2050"
height="870"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 10: HAMi community contributors map, contributors from 27 countries&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I&amp;rsquo;m not listing these numbers to show off &amp;ldquo;growth metrics.&amp;rdquo; I&amp;rsquo;m making a different point: an open source project is shifting from &amp;ldquo;software maintained by a team&amp;rdquo; into &amp;ldquo;a technical community of people gathered around a shared goal.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;That distinction matters. Software can be forked, rewritten, or generated overnight by AI. But a community with a shared goal, trust, and rhythm cannot be forked. That is the irreplaceable asset of an open source project.&lt;/p&gt;
&lt;p&gt;From a governance-maturity perspective, HAMi has passed through three milestones:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/a221c9ebe476891231ef44e3de11b783.svg" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/a221c9ebe476891231ef44e3de11b783.svg" alt="Figure 11: Evolution of HAMi’s governance maturity" data-caption="Figure 11: Evolution of HAMi’s governance maturity"
width="1372"
height="168"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 11: Evolution of HAMi’s governance maturity&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Note the middle stretch: from entering Sandbox in August 2024 to reaching Incubating in July 2026, roughly two years. During those two years the code certainly grew, but what actually convinced the CNCF Technical Oversight Committee (TOC) was the diversification of the community, the formalization of governance, and real production adoption. None of that is something you produce by writing code.&lt;/p&gt;
&lt;h2 id="when-hami-open-sourced-in-2021-there-was-no-ai-coding"&gt;When HAMi Open-Sourced in 2021, There Was No AI Coding&lt;/h2&gt;
&lt;p&gt;HAMi was first open-sourced in 2021. Back then, most developers did not see AI coding the way we do today. Whether an open source project survived depended on developers genuinely investing their time, discussing problems in issues, submitting code through pull requests, and building trust through code review.&lt;/p&gt;
&lt;p&gt;Today, the environment has changed.&lt;/p&gt;
&lt;p&gt;In a recent HAMi community livestream (&lt;a href="https://www.bilibili.com/video/BV1r6EC6zEWS/" target="_blank" rel="noopener"&gt;Mastering HAMi DRA, Yang Shouren, HAMi Community Livestream Episode 2&lt;/a&gt;), someone asked the maintainers a question: &amp;ldquo;How much of HAMi&amp;rsquo;s code now comes from AI assistance?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The answer from &lt;a href="https://github.com/shouren" target="_blank" rel="noopener"&gt;Yang Shouren&lt;/a&gt; stuck with me: &lt;strong&gt;about half of the code in the HAMi community today is already AI-assisted.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Half. And that share is still rising.&lt;/p&gt;
&lt;p&gt;This immediately raises a sharp question: if AI can write more and more code, what is the value of an open source community? Could one person plus a few AI agents just fork a &amp;ldquo;new HAMi&amp;rdquo;?&lt;/p&gt;
&lt;p&gt;My answer is no, because what AI lowers is the cost of producing code, not the cost of building technical consensus.&lt;/p&gt;
&lt;h2 id="in-the-ai-era-code-is-no-longer-scarce-consensus-is"&gt;In the AI Era, Code Is No Longer Scarce; Consensus Is&lt;/h2&gt;
&lt;p&gt;Let me sharpen that point.&lt;/p&gt;
&lt;p&gt;An AI agent can already do a lot today: write code, fix bugs, add tests, generate docs, produce migration scripts. These capabilities are getting stronger fast. But there are a few things in a community that AI cannot replace today, and in my judgment will not replace soon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, deciding which problems are worth solving.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Take HAMi. Why does GPU sharing matter? Not because &amp;ldquo;slicing cards finely&amp;rdquo; is a cool technique. It matters because once AI infrastructure scales, reality looks like this: GPU costs are enormous, heterogeneous hardware keeps multiplying, and Kubernetes&amp;rsquo; native resource model is no longer enough.&lt;/p&gt;
&lt;p&gt;The community has to first agree that &amp;ldquo;this problem is worth investing in&amp;rdquo; before anyone writes any code. That agreement is a human-to-human matter, supported by real scenarios, real costs, and real pain. AI can solve a problem you have already defined, but &amp;ldquo;which problem is worth defining&amp;rdquo; is decided by community consensus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, choosing a technical path.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In GPU virtualization there are many paths: MIG, MPS, time-slicing, vGPU, DRA. Each has trade-offs, and choosing wrong can cost you two years of detours.&lt;/p&gt;
&lt;p&gt;Code can be generated, but architectural choice is fundamentally a value judgment. HAMi&amp;rsquo;s decision on Ascend 910C to move from hardware SR-IOV to userspace HAMi-core was not about someone writing a better piece of code; it was about the maintainers holding to a judgment that &amp;ldquo;hardware partitioning is too coarse, software partitioning is more flexible.&amp;rdquo; That kind of judgment is ground out through repeated discussion, failure, and validation in the community, not prompted out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Third, trust.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Users don&amp;rsquo;t choose HAMi because of &amp;ldquo;how much AI-generated code is in this repo.&amp;rdquo; They care about: who maintains it? Who reviews it? Are there real production cases? Does the community respond when something breaks?&lt;/p&gt;
&lt;p&gt;Each of these is a relationship between people, a product of community organization, not a product of code quality.&lt;/p&gt;
&lt;p&gt;Put these three together, and the production model of open source communities in the AI era is shifting:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/3cc15ad018c9b11909feea79769e166c.svg" data-img="https://assets.jimmysong.io/images/blog/hami-cncf-incubating-consensus/3cc15ad018c9b11909feea79769e166c.svg" alt="Figure 12: How the open source production model is changing in the AI era" data-caption="Figure 12: How the open source production model is changing in the AI era"
width="2333"
height="357"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 12: How the open source production model is changing in the AI era&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the past, code came first and the community sedimented out of the code; in the future, consensus comes first, AI rapidly turns consensus into code, and code flows back to test the consensus. The center of gravity of the scarce resource moves from &amp;ldquo;code&amp;rdquo; on the left to &amp;ldquo;consensus&amp;rdquo; on the right.&lt;/p&gt;
&lt;h2 id="in-the-ai-era-open-source-governance-itself-has-to-level-up"&gt;In the AI Era, Open Source Governance Itself Has to Level Up&lt;/h2&gt;
&lt;p&gt;Since AI has become a new category of contributor, a community&amp;rsquo;s governance rules have to keep up.&lt;/p&gt;
&lt;p&gt;My advice is: don&amp;rsquo;t treat AI merely as a tool, treat it as a new type of contributor. It used to be &amp;ldquo;developers write code, humans review&amp;rdquo;; in the future it will be &amp;ldquo;humans set intent, AI generates code, the community reviews, and shared knowledge is distilled.&amp;rdquo; There is an extra layer in the middle, and an extra layer of governance complexity.&lt;/p&gt;
&lt;p&gt;HAMi is already responding to this. Its &lt;a href="https://github.com/Project-HAMi/HAMi/blob/master/CONTRIBUTING.md" target="_blank" rel="noopener"&gt;CONTRIBUTING.md&lt;/a&gt; is explicit:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you are using any kind of AI assistance to contribute to HAMi, it must be disclosed in the pull request.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In other words, if you use AI to help with a contribution, you must declare it in the PR. But the community also knows that a norm without a gate is not enough (see the discussion in &lt;a href="https://github.com/Project-HAMi/HAMi/issues/1998" target="_blank" rel="noopener"&gt;Issue #1998&lt;/a&gt;), and there is already ongoing discussion about how to give that norm real enforcement.&lt;/p&gt;
&lt;p&gt;This is actually a problem every AI-era open source project will run into. I&amp;rsquo;d break it into a few questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Must AI-generated code always be declared?&lt;/li&gt;
&lt;li&gt;Which model, and what context, did the contributor use?&lt;/li&gt;
&lt;li&gt;How do you ensure the security of AI code, avoiding injection and licensing risks?&lt;/li&gt;
&lt;li&gt;What process should maintainers use to review an AI diff they may not be able to fully trace themselves?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Whoever figures out and operationalizes these rules first will keep contribution quality stable in the AI era. HAMi&amp;rsquo;s exploration here is worth a look for every open source project.&lt;/p&gt;
&lt;h2 id="what-cncf-incubating-really-means"&gt;What CNCF Incubating Really Means&lt;/h2&gt;
&lt;p&gt;Back to the promotion itself.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d lean against treating it as an &amp;ldquo;honor.&amp;rdquo; What CNCF Incubating really validates is not code quality, but whether a project has the capacity to become infrastructure. It examines a whole package: technical maturity, community governance, production adoption, and ecosystem building.&lt;/p&gt;
&lt;p&gt;HAMi&amp;rsquo;s case, in one sentence, is not &amp;ldquo;a Chinese team built a GPU project.&amp;rdquo; It is this:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI Infrastructure community, jointly shaped by developers from around the world, is taking shape.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The first is a product story; the second is an ecosystem story. The Incubating recognition from the CNCF is recognizing the latter, because competition over infrastructure is never competition between individual products; it is competition between ecosystem networks.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The open source competition of the next decade will not be just a competition of code, but a competition of communities.&lt;/p&gt;
&lt;p&gt;Once AI gives everyone near-infinite capacity to produce code, the truly scarce capabilities will be three: finding the right problem, building technical consensus, and organizing developers worldwide to solve a problem together.&lt;/p&gt;
&lt;p&gt;HAMi&amp;rsquo;s path to CNCF Incubating is just one snapshot of how open source communities are evolving in the AI era. Code will keep getting cheaper, and consensus will keep getting more expensive. Whoever understands this inversion will be the one who can build open source communities with real depth in the AI era.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Join the HAMi Community
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Add me on WeChat (&lt;code&gt;jimmysong&lt;/code&gt;) or follow &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi on GitHub&lt;/a&gt; to join the community focused on GPU virtualization and heterogeneous compute scheduling. Let&amp;rsquo;s talk about open source governance in the AI era.
&lt;/div&gt;
&lt;/div&gt;</content:encoded></item><item><title>Olares and HAMi: A New Inflection Point for Desktop AI Workstations</title><link>https://jimmysong.io/blog/olares-hami-edge-ai-cloud/</link><pubDate>Wed, 24 Jun 2026 22:30:00 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/olares-hami-edge-ai-cloud/</guid><description>HAMi moves from cluster to desktop with Olares.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;HAMi used to save cards in the cluster. Now it decides how good a desktop AI workstation feels.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/banner.webp" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/banner.webp" alt="Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations" data-caption="Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations"
width="1983"
height="793"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Olares and HAMi: a new inflection point for desktop AI workstations&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="an-old-friend-mentions-a-name"&gt;An Old Friend Mentions a Name&lt;/h2&gt;
&lt;p&gt;A few days ago an old friend came by to chat. We used to run the cloud-native community scene together back home, so we go way back. She recently joined a company called Olares, and somewhere in the conversation she dropped this: their project had integrated HAMi.&lt;/p&gt;
&lt;p&gt;I run the HAMi community day to day, so whenever I hear someone using it for something, I want to take a closer look. I went and dug through the &lt;a href="https://github.com/beclab/Olares" target="_blank" rel="noopener"&gt;Olares&lt;/a&gt; repo and website, and my first reaction was, huh, this is actually interesting.&lt;/p&gt;
&lt;p&gt;Isn&amp;rsquo;t this exactly the kind of local AI workstation I&amp;rsquo;d been eyeing forever but never pulled the trigger on? I wrote in &lt;a href="https://jimmysong.io/blog/personal-ai-stack"&gt;My Personal AI Stack&lt;/a&gt; that for someone like me who mainly works with my head, subscribing to models beats maintaining a high-end GPU. But how far can &amp;ldquo;one machine running the full AI stack&amp;rdquo; really go, and who has actually built it, I&amp;rsquo;ve wanted to see with my own eyes.&lt;/p&gt;
&lt;h2 id="what-is-this-thing-really"&gt;What Is This Thing, Really&lt;/h2&gt;
&lt;p&gt;Olares isn&amp;rsquo;t a &amp;ldquo;NAS with a GPU bolted on,&amp;rdquo; and it isn&amp;rsquo;t a &amp;ldquo;mini PC with Ollama installed.&amp;rdquo; It&amp;rsquo;s more like a personal cloud built on Kubernetes: local models, apps, identity, remote access, dev environment, storage, even GPU governance, all packed into one machine as a desktop cloud OS.&lt;/p&gt;
&lt;p&gt;HAMi isn&amp;rsquo;t filler here. It&amp;rsquo;s the layer that turns &amp;ldquo;one card&amp;rdquo; into &amp;ldquo;a resource pool you can share, isolate, and schedule.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="how-it-differs-from-the-ai-mini-pc-crowd"&gt;How It Differs from the AI Mini-PC Crowd&lt;/h2&gt;
&lt;p&gt;There&amp;rsquo;s no shortage of things calling themselves AI workstations, but most are just beefier mini PCs. Olares is different. It&amp;rsquo;s genuinely designed as a cloud.&lt;/p&gt;
&lt;p&gt;The company positions this machine outright as a 24/7 personal AI cloud, not a PC you sit in front of. That single framing explains almost everything that follows: why it has to use Kubernetes, why it cares so much about network traversal and remote access, why GPU governance suddenly matters here.&lt;/p&gt;
&lt;p&gt;Two other details tell you a lot: it supports Thunderbolt 5 external eGPUs, and two machines can cluster up. In other words, it was never a sealed box. It starts as a single node and grows upward.&lt;/p&gt;
&lt;h2 id="stuffing-a-whole-cloud-runtime-into-one-box"&gt;Stuffing a Whole Cloud Runtime into One Box&lt;/h2&gt;
&lt;p&gt;Architecturally, what Olares does is compress an entire cloud runtime into a single machine: auth and authorization, app lifecycle, tunnels and traversal, secrets, observability middleware, not a layer skipped.&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t &amp;ldquo;a few AI apps preinstalled.&amp;rdquo; It&amp;rsquo;s a complete cloud runtime stuffed into a desktop device.&lt;/p&gt;
&lt;p&gt;The diagram below is my own re-drawn layering based on its public docs, with the HAMi layer pulled out explicitly because it&amp;rsquo;s the point of this whole piece.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/olares-architecture-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/olares-architecture-en.svg" alt="Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane" data-caption="Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane"
width="1382"
height="1392"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Olares system layered architecture, with HAMi as the GPU resource plane&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="one-detail-that-caught-my-eye"&gt;One Detail That Caught My Eye&lt;/h2&gt;
&lt;p&gt;One thing jumped out while I was reading: &lt;strong&gt;Olares&amp;rsquo;s docs haven&amp;rsquo;t kept up with its own product.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Its 2025 architecture page still says &lt;code&gt;nvshare&lt;/code&gt;, noting that GPUs are limited to one card per node. But flip through the release notes and 1.12 already integrates the HAMi scheduler, with exclusive, time-slicing, and memory-slicing modes all there; 1.12.2 adds multi-GPU; 1.12.5 supports DGX Spark outright and rolls automatic scheduling across all three modes.&lt;/p&gt;
&lt;p&gt;The product is outrunning its docs. That alone tells you something: it&amp;rsquo;s shifting from a &amp;ldquo;personal cloud OS&amp;rdquo; toward a &amp;ldquo;local AI cloud OS.&amp;rdquo; That&amp;rsquo;s also why I wanted to write a whole piece on it.&lt;/p&gt;
&lt;h2 id="what-hami-actually-does-in-there"&gt;What HAMi Actually Does in There&lt;/h2&gt;
&lt;p&gt;HAMi is a CNCF Sandbox project, positioned as a heterogeneous AI compute virtualization middleware. On the path there&amp;rsquo;s a Webhook, a scheduler, a Device Plugin, and HAMi-core handling in-container resource control. I summed it up in one line in &lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes Is Becoming the GPU Control Plane of the AI Era&lt;/a&gt;: it turns GPU slicing from &amp;ldquo;a hardware capability&amp;rdquo; into &amp;ldquo;a control-plane capability.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;On Olares, that&amp;rsquo;s the real thing behind the GPU mode toggle in settings. On the surface it&amp;rsquo;s a UI option; underneath, it&amp;rsquo;s a policy switch on the resource plane.&lt;/p&gt;
&lt;p&gt;The most critical piece is almost certainly HAMi-core, which in one sentence: intercepts CUDA calls inside the container and does memory virtualization, compute throttling, and utilization monitoring. It doesn&amp;rsquo;t slice with hardware. It manages with software.&lt;/p&gt;
&lt;p&gt;The diagram below puts that injection path and the three GPU modes side by side.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-core-gpu-modes-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-core-gpu-modes-en.svg" alt="Figure 3: HAMi-core injection path and Olares’s three GPU modes" data-caption="Figure 3: HAMi-core injection path and Olares’s three GPU modes"
width="1421"
height="1022"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: HAMi-core injection path and Olares’s three GPU modes&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The evidence lines up too. In Olares&amp;rsquo;s GitHub issues you can see HAMi&amp;rsquo;s &lt;code&gt;libvgpu.so&lt;/code&gt; crashing under WSL2, the discussion mentions it getting injected into every process via &lt;code&gt;/etc/ld.so.preload&lt;/code&gt;, and the Olares team itself says this implementation &amp;ldquo;drew heavy inspiration&amp;rdquo; from the HAMi project.&lt;/p&gt;
&lt;p&gt;One thing I should be clear about, though: whether Olares uses upstream HAMi-core directly or maintains a forked version, there&amp;rsquo;s no single official answer in public. I won&amp;rsquo;t fill that in for them.&lt;/p&gt;
&lt;p&gt;The three modes make sense in order. Full-card exclusive is for heavy loads; time-slicing lets lightweight services take turns, and Olares even swaps inactive models into memory first, with roughly 5% switching overhead; memory-slicing cuts the memory into fixed quotas so multiple apps run together. On something like DGX Spark, where CPU and GPU share memory, it defaults to memory-slicing because there&amp;rsquo;s no traditional memory paging in and out to begin with.&lt;/p&gt;
&lt;h2 id="why-this-matters-for-hami"&gt;Why This Matters for HAMi&lt;/h2&gt;
&lt;p&gt;HAMi used to live mostly in big clusters: multi-tenant, inference serving, mixed training-and-inference, heterogeneous cards. NIO running it for sharing across 80 nodes and 600 cards is the textbook example. That line is well-trodden.&lt;/p&gt;
&lt;p&gt;What feels new about Olares is that it moves the same problem onto a single machine.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-scene-shift-en.svg" data-img="https://assets.jimmysong.io/images/blog/olares-hami-edge-ai-cloud/hami-scene-shift-en.svg" alt="Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience" data-caption="Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience"
width="1382"
height="1042"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi’s value narrative shifts from cluster efficiency to edge product experience&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What you actually run at once on Olares is never just one Ollama: a local model, a chat UI, a research agent, an image or video pipeline, plus a pile of OCR, speech-to-text, and automation tools. In a scenario like that, without a GPU scheduling layer, the GPU always degenerates into &amp;ldquo;whoever starts first grabs it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;So for HAMi this matters. It goes from &amp;ldquo;a tool that saves cards for platform engineers&amp;rdquo; to &amp;ldquo;something that directly decides whether an ordinary user has a good time.&amp;rdquo; In the cluster it manages utilization and cost; on Olares it manages concurrent experience, model switching, and whether this machine is actually pleasant to use. HAMi doesn&amp;rsquo;t only talk to platform engineers anymore.&lt;/p&gt;
&lt;h2 id="where-edge-ai-goes-from-here"&gt;Where Edge AI Goes from Here&lt;/h2&gt;
&lt;p&gt;My sense is, edge AI won&amp;rsquo;t settle at &amp;ldquo;a stronger local model box.&amp;rdquo; It&amp;rsquo;ll grow into a full stack of &amp;ldquo;control plane plus model plane plus application plane.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;NVIDIA pushing DGX Spark and RTX Spark onto the desktop has already moved the story from &amp;ldquo;run models locally&amp;rdquo; to &amp;ldquo;run agents locally.&amp;rdquo; Olares&amp;rsquo;s CUDA-plus-x86 is one path, DGX Spark&amp;rsquo;s unified memory is another, and below that there&amp;rsquo;s the DIY mini-PC-plus-eGPU route you assemble yourself. Interestingly, Olares lists eGPU and dual-node as supported paths itself, so the line between appliance and DIY isn&amp;rsquo;t that sharp.&lt;/p&gt;
&lt;p&gt;But my guess is, the one that breaks out won&amp;rsquo;t be the one with the most ferocious specs. It&amp;rsquo;ll be the one that fuses resource governance, dev experience, app ecosystem, security, and remote access into one closed loop. If desktop AI workstations really become a category, what people compete on isn&amp;rsquo;t GPU and memory, it&amp;rsquo;s the control plane.&lt;/p&gt;
&lt;p&gt;I went ahead and added Olares to the &lt;a href="https://landscape.jimmysong.io/projects/olares/" target="_blank" rel="noopener"&gt;AI Native Landscape&lt;/a&gt; I maintain, so I can keep watching how it grows.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In one line: Olares is a desktop AI OS with Kubernetes at its core, and HAMi is the layer that turns its single-machine GPU into a shareable, isolatable, schedulable resource pool.&lt;/p&gt;
&lt;p&gt;From cluster to desktop, HAMi&amp;rsquo;s story is expanding from &amp;ldquo;saving cards&amp;rdquo; to &amp;ldquo;making edge AI usable.&amp;rdquo; If desktop AI workstations become a real category, what people compete on is the control plane, not single-card performance.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/beclab/Olares" target="_blank" rel="noopener"&gt;Olares - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://landscape.jimmysong.io/projects/olares/" target="_blank" rel="noopener"&gt;Olares - AI Native Landscape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;HAMi official site - project-hami.io&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Every Nation Begins with Textiles</title><link>https://jimmysong.io/blog/every-nation-starts-with-textiles/</link><pubDate>Sat, 20 Jun 2026 05:43:26 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/every-nation-starts-with-textiles/</guid><description>From Anji bamboo weaving to industrialization</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;In a bamboo-weaving workshop in Anji, holding a strip of bamboo split hair-thin, I realized for the first time: a bolt of cloth, a machine, a country, all begin with this single thread in the hand.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/banner.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/banner.webp" alt="Figure 1: Bamboo strip" data-caption="Figure 1: Bamboo strip"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Bamboo strip&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;This is something I&amp;rsquo;d been thinking about for a long time, without ever quite sorting it out.&lt;/p&gt;
&lt;p&gt;In February this year, our company offsite took us to Anji, in Zhejiang province. It&amp;rsquo;s famous bamboo country in China, the mountains covered in moso bamboo, and the locals have been weaving with it for generations. Bamboo weaving is a national intangible cultural heritage. The offsite included a hands-on session for us to try it ourselves.&lt;/p&gt;
&lt;p&gt;I hadn&amp;rsquo;t expected much from this kind of &amp;ldquo;experiential activity.&amp;rdquo; But once I actually sat down, picked up a bamboo strip, and listened to the old master explain how to lift, press, and raise the threads, my mind started to wander. Up and down, lift and press, the warp and weft interlacing in his hands until a pattern slowly emerged. The motion was deeply repetitive, a fixed rhythm, a clear logic.&lt;/p&gt;
&lt;p&gt;And in that moment an almost absurd thought popped into my head: this is just 0 and 1.&lt;/p&gt;
&lt;p&gt;A warp thread raised is 1, pressed down is 0. Unfold a piece of cloth and it&amp;rsquo;s a two-dimensional structure made of 0s and 1s. And the thing the old master kept referring to as the &amp;ldquo;flower program&amp;rdquo; (花本), the pre-designed sequence of warp-lifting, is essentially a program: follow it, and you weave exactly the pattern you intended.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/red-horse.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/red-horse.webp" alt="Figure 2: A woven bamboo horse I finished in Anji, mounted for framing" data-caption="Figure 2: A woven bamboo horse I finished in Anji, mounted for framing"
width="1567"
height="1498"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: A woven bamboo horse I finished in Anji, mounted for framing&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;That thought flashed by and I didn&amp;rsquo;t make much of it. But then in March I went to the Netherlands and spent a week in Amsterdam. The two experiences sat together, and some feelings that had been vague began to sharpen: how could a single strip of bamboo, a single motion, connect to the looms of two centuries ago, to today&amp;rsquo;s computers, even to a country&amp;rsquo;s industrialization?&lt;/p&gt;
&lt;p&gt;I looked it up later and found I wasn&amp;rsquo;t the first to make the connection. Many historians of technology argue that weaving is one of the earliest forms of large-scale information encoding. Now, in June, I&amp;rsquo;m sitting down to try to lay out this thread. This isn&amp;rsquo;t a technical piece. It&amp;rsquo;s a personal, cultural reflection.&lt;/p&gt;
&lt;h2 id="a-cultural-contrast"&gt;A Cultural Contrast&lt;/h2&gt;
&lt;p&gt;Let me start with the Netherlands.&lt;/p&gt;
&lt;p&gt;What struck me most about that country wasn&amp;rsquo;t the windmills or the tulips, but a quality that the whole nation seems to exude: restraint, order, and an almost engineering-like way of being. The canals are dug, the land reclaimed from the sea is called a &amp;ldquo;polder,&amp;rdquo; and the windmills were never there to look pretty. They were built to pump water and drain the land. A country that sits below sea level literally engineered itself out of the sea. Though the windmills aren&amp;rsquo;t only about drainage. The wind in the Netherlands is genuinely strong. In March I got battered by it. It&amp;rsquo;s a wind that pushes at you constantly. So a below-sea-level country hauled itself out of the sea with engineering, and along the way put that fierce wind to work too.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/canal.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/canal.webp" alt="Figure 3: A canal and church in central Amsterdam" data-caption="Figure 3: A canal and church in central Amsterdam"
width="1800"
height="2699"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: A canal and church in central Amsterdam&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/windmill.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/windmill.webp" alt="Figure 4: Windmills at Zaanse Schans" data-caption="Figure 4: Windmills at Zaanse Schans"
width="1800"
height="2699"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Windmills at Zaanse Schans&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This quality seeps into every corner of daily life. Shops close around six in the evening, the streets slowly quiet down, and you almost never hear a car horn. Cars stop well ahead of time to let you cross. One evening I wandered a little too close to the bike lane, and a passing car actually rolled down its window to remind me to use the sidewalk. It surprised me at first, and then I understood: the rules there are clear, and everyone assumes you&amp;rsquo;ll follow them too.&lt;/p&gt;
&lt;p&gt;This temperament even shows up in how people get around. In the city center, not many people drive. Bicycles are the default. For one thing, the Netherlands is flat, so cycling takes no effort. For another, dedicated bike lanes are everywhere, and the streets are narrow, which makes them ill-suited to cars. Add that commutes are generally short, and cycling is just right. What&amp;rsquo;s interesting is that even in the southern suburbs, where the roads are wide and every household has a car, you still see a lot of people on bikes. For them, a bicycle isn&amp;rsquo;t exercise. It&amp;rsquo;s the default way to commute.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/street.webp" data-img="https://assets.jimmysong.io/images/blog/every-nation-starts-with-textiles/street.webp" alt="Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists" data-caption="Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists"
width="1800"
height="2400"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: A street in the southern suburbs of Amsterdam: wide roads, cars in every driveway, and still plenty of cyclists&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I set that against the noise and speed of China, the loud, garish signboards back home. These two temperaments are really two paths of modernization. The Netherlands is an early, slow-and-steady kind of country, from the Age of Discovery to the world&amp;rsquo;s first stock exchange, from water engineering to today&amp;rsquo;s ASML and Philips. It has always moved forward in a &amp;ldquo;long-term, restrained, systematic&amp;rdquo; way. China is a late developer chasing from behind, relying on speed, scale, and a generation&amp;rsquo;s sheer exertion to cover in a few decades what took others centuries.&lt;/p&gt;
&lt;p&gt;Neither path is higher or lower than the other, but they settle into completely different urban textures and national characters. The ease of &amp;ldquo;closing at six&amp;rdquo; that the Dutch have is something China won&amp;rsquo;t reach for many years; and the Chinese speed of &amp;ldquo;just do it,&amp;rdquo; of laying a high-speed rail line overnight, is something the Netherlands probably couldn&amp;rsquo;t match either.&lt;/p&gt;
&lt;p&gt;But here&amp;rsquo;s the interesting part. If you trace both civilizations back through history, these two seemingly different cultures share a common starting point. Textiles.&lt;/p&gt;
&lt;h2 id="why-textiles-were-the-starting-point-of-industrialization"&gt;Why Textiles Were the Starting Point of Industrialization&lt;/h2&gt;
&lt;p&gt;Our generation is so familiar with the phrase &amp;ldquo;Industrial Revolution&amp;rdquo; that it&amp;rsquo;s almost gone numb. But have you ever asked: where did the first Industrial Revolution actually break out? Not in steel, not in coal, not in the railways. In textiles.&lt;/p&gt;
&lt;p&gt;The British Industrial Revolution is practically a history of textile machinery. The flying shuttle in 1733 made weaving faster. The spinning jenny in 1764 let one person spin many threads at once. The water frame in 1769 began replacing human power with water power. The power loom in 1785 turned weaving into an automated process. Every one of these inventions happened inside the textile industry.&lt;/p&gt;
&lt;p&gt;Why textiles of all things? At first it feels counterintuitive, since textiles seem too &amp;ldquo;light,&amp;rdquo; lacking the heft of making steel or guns. But think about it for a moment and it becomes obvious.&lt;/p&gt;
&lt;p&gt;Textiles are a basic need. Everyone has to wear clothes, clothes wear out, and they have to be replaced constantly. That&amp;rsquo;s an enormous market that already existed back in the agrarian age. Once the demand is there, any gain in efficiency turns immediately into profit, into capital that can be reinvested.&lt;/p&gt;
&lt;p&gt;More importantly, the processes of textile production are especially easy to mechanize. Spinning and weaving are deeply repetitive motions with a fixed rhythm and a clear logic, so machines can directly replace the human hand. Steelmaking, by contrast, has a far higher technical barrier, the science of chemistry wasn&amp;rsquo;t yet mature, and machine manufacturing itself needed an existing industrial base before it could even start.&lt;/p&gt;
&lt;p&gt;And textile production has a moderate investment threshold. Building a textile mill was far cheaper than building a steel plant, so merchant capital and money from colonial trade flowed in easily. A large share of early British industrial capital came from the cotton trade and colonial trade, and that money went first into textile mills.&lt;/p&gt;
&lt;p&gt;The most crucial point is that textiles don&amp;rsquo;t exist in isolation. They pull an entire supply chain along with them. To spin and weave, you first have to grow and transport cotton. To run machines, you have to build textile machinery. To power the machines, you need steam engines. To fuel steam engines, you have to mine coal. To move cotton and cloth, you have to build railways. So the steam engine was first deployed at scale in textile mills, the earliest railways carried cotton and cloth, and the earliest machine manufacturing served textile equipment. A supposedly &amp;ldquo;light&amp;rdquo; industry ended up dragging the entire mechanical and energy industries into being.&lt;/p&gt;
&lt;p&gt;So textiles became the progenitor of industry not because they were the most complex, but because they were the first to satisfy four conditions at once: large demand, mechanizable, low threshold, and able to pull others along. Textiles were humanity&amp;rsquo;s first trial plot on the way from handicraft into an industrial system.&lt;/p&gt;
&lt;h2 id="late-developers-all-come-up-this-same-road"&gt;Late Developers All Come Up This Same Road&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s a pattern I&amp;rsquo;d never noticed before: the industrialization of late developers, almost without exception, begins with textiles.&lt;/p&gt;
&lt;p&gt;Take Japan. Today when you think of Japanese industry, you think of Toyota cars. But how did Toyota begin? Its founder, Sakichi Toyoda, didn&amp;rsquo;t start by making cars. He invented automatic looms. In 1924 he invented the Type G automatic loom, and in 1926 he founded Toyoda Automatic Loom Works, the origin of the entire Toyota Group. It was his son Kiichiro who later moved the business into automobiles.&lt;/p&gt;
&lt;p&gt;So Toyota cars quite literally &amp;ldquo;grew out of a loom.&amp;rdquo; Its whole system of lean production, kanban management, total quality control, carries in its bones the obsession with precision, process, and zero defects that came from making looms.&lt;/p&gt;
&lt;p&gt;Then look at South Korea. Samsung today is a global electronics and semiconductor giant, but when it was founded it exported dried fish, vegetables, and fruit, only later moving into textiles and sugar. As South Korea industrialized after the war, textiles and garments were among its earliest foreign-exchange earners, and Samsung built its first fortune on trade and textiles in that period before pivoting to electronics in the 1970s. The Hyundai Group took a similar path, starting in engineering and heavy industry before expanding into automobiles, shipbuilding, and heavy industry.&lt;/p&gt;
&lt;p&gt;Britain was first; Japan and Korea are the late-developer cases. You&amp;rsquo;ll notice the sequence is strikingly consistent across all three: textiles first, then light industry, then machinery, then heavy industry, and finally high-tech. Nobody designed this path. It was forced by cost and by the market. Because textiles are the lowest-threshold, largest-market form of manufacturing, every nation that wants to turn from an agrarian country into an industrial one has to start from the easiest, most necessary step.&lt;/p&gt;
&lt;p&gt;Behind this lies a very plain truth: modernization doesn&amp;rsquo;t happen in one leap. It needs a starting point that can earn money, accumulate experience, and train the first generation of industrial workers. And textiles happen to be exactly that kind of starting point.&lt;/p&gt;
&lt;h2 id="where-china-sits-on-this-road"&gt;Where China Sits on This Road&lt;/h2&gt;
&lt;p&gt;Writing this far, I can&amp;rsquo;t help but turn and look back at China itself.&lt;/p&gt;
&lt;p&gt;China&amp;rsquo;s modern industrialization also began with textiles. The most famous Chinese-owned enterprises of the late Qing and early Republic were almost all cotton mills. Zhang Jian founded the Dasheng Cotton Mill in Nantong. The Rong brothers, Zongjing and Desheng, founded the Shenxin Cotton Mills in Wuxi. They were China&amp;rsquo;s first generation of national industrial capitalists, rising out of textiles and holding up half of modern Chinese industry. The thinking of Zhang Jian&amp;rsquo;s generation was plain and direct: foreigners were using machines to weave cloth and taking our money, so we had to set up our own mills, weave our own cloth, and keep that money at home.&lt;/p&gt;
&lt;p&gt;That was China&amp;rsquo;s first stretch. Back then, China was walking the very road that Britain, Japan, and Korea had all walked.&lt;/p&gt;
&lt;p&gt;Then the road broke. War, turmoil, the planned economy. Chinese industrialization took many detours. It wasn&amp;rsquo;t until reform and opening up that it reconnected. In the 1980s the coast was blanketed with processing factories for garments, shoes, and toys. In essence it was still that same old road of &amp;ldquo;starting with textiles and light industry,&amp;rdquo; only this time China relaunched it through exports, cheap labor, and the role of factory to the world.&lt;/p&gt;
&lt;p&gt;Further on, through the 1990s and 2000s, China pushed into heavy industry: steel, cement, shipbuilding, chemicals, catching up to the world&amp;rsquo;s front ranks one after another. Then electronics, home appliances, mobile phones. Then the internet, high-speed rail, new-energy vehicles, solar, semiconductors, and the AI that everyone talks about today.&lt;/p&gt;
&lt;p&gt;String this line together and you realize China is actually completing the same road that every late developer has walked, only faster, fiercer, and at a far greater scale. From Zhang Jian&amp;rsquo;s cotton mills to today&amp;rsquo;s new-energy vehicles and large AI models is barely over a hundred years. In just over a century we&amp;rsquo;ve run the entire course of industrialization that took Britain more than two hundred and Japan more than a hundred, and we&amp;rsquo;re still going.&lt;/p&gt;
&lt;p&gt;This often leaves me with a complicated feeling. On one hand, admiration: generation after generation, from Zhang Jian to today&amp;rsquo;s engineers, really did turn an agrarian country into the factory of the world and then into a country that now leads in quite a few high-tech fields. On the other hand, a faint unease: this road was run so fast that a lot of things got left behind.&lt;/p&gt;
&lt;h2 id="a-thread-that-never-broke"&gt;A Thread That Never Broke&lt;/h2&gt;
&lt;p&gt;Inside this long thread of industrialization, there&amp;rsquo;s one detail I never forgot: that moment in the bamboo-weaving workshop in Anji back in February.&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;flower program&amp;rdquo; the old master described actually has a very old tradition in China. The drawloom of the Han dynasty used a pre-designed system of cords to control which warp threads rose and fell, so a weaver following it could reproduce complex patterns. Then in the early nineteenth century, the Frenchman Joseph Marie Jacquard pushed the idea a big step forward, using punched cards to control the weaving pattern: where the card had a hole, a hook passed through and lifted the warp; where there was no hole, the hook was blocked and the warp stayed down. Hole or no hole, that&amp;rsquo;s 1 and 0.&lt;/p&gt;
&lt;p&gt;Jacquard&amp;rsquo;s punched cards were later borrowed by the computing pioneer Charles Babbage for his Analytical Engine, and Ada Lovelace used them to write the first computer program in history. Further on, IBM rose on punched-card machines, and if you trace today&amp;rsquo;s entire computing industry back to its source, it leads, improbably, to a loom.&lt;/p&gt;
&lt;p&gt;And that&amp;rsquo;s not the end of it. The Transformer, matrix operations, the weights of neural networks in today&amp;rsquo;s AI, at bottom are all about computing relationships across vast fields of &amp;ldquo;warp and weft.&amp;rdquo; A loom decides which warp threads to lift and when; a neural network decides which parameters relate to which others. The logic is the same.&lt;/p&gt;
&lt;p&gt;So the word &amp;ldquo;weaving&amp;rdquo; is more than the starting point of an industry. It&amp;rsquo;s a metaphor: for thousands of years, human beings have been learning how to encode complex things into structure, from a bolt of cloth, to a program, to a neural network. Every nation does this. They&amp;rsquo;re just at different stages.&lt;/p&gt;
&lt;p&gt;I plan to write the technical bloodline of this story separately, in another piece. Here I&amp;rsquo;ll only point to it, to make one thing clear: the thread that began with that strip of bamboo has never broken.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;From a bamboo-weaving session in Anji in February, to the canals and windmills of Amsterdam in March, and on to the industrialization of Britain, Japan, Korea, and China, what I want to say really comes down to one thing.&lt;/p&gt;
&lt;p&gt;Every nation&amp;rsquo;s modernization is a road from a single thread woven into a net. Textiles are the common starting point of that road, not because textiles are so lofty, but because they are the most necessary, the easiest, the quickest to earn that first money and train that first generation of workers. From that starting point, some walk with restraint, like the Netherlands; some move with ferocity, like China. Some take two hundred years, some a hundred, some only a few decades. But the starting point is the same, and so is the direction: from that thread in the hand, step by step, woven into the whole of industrial civilization as we know it today.&lt;/p&gt;
&lt;p&gt;The Dutch ease of closing at six, of no honking, of swans gliding through the canals, is a place China hasn&amp;rsquo;t reached yet. The Chinese speed of just doing it, of rolling something out overnight, is something the Netherlands couldn&amp;rsquo;t pull off either. Between these two temperaments of modernization there is no higher or lower, only trade-offs. And behind the trade-offs lies the question of where a country sits in its industrialization, and what rhythm it&amp;rsquo;s willing to pay for that net.&lt;/p&gt;
&lt;p&gt;That day in Anji, holding a strip of bamboo, I clumsily followed the old master, lifting and pressing, weaving a small, lopsided patch of pattern. I wasn&amp;rsquo;t thinking about any of this then. But looking back, from that single lift and press, you can trace upward to the Han dynasty drawloom, outward to a country&amp;rsquo;s industrialization, and forward, faintly, to today&amp;rsquo;s computers and AI. A single strip of bamboo, connected to so much.&lt;/p&gt;
&lt;p&gt;Perhaps the evolution of civilization, in the end, is just humanity learning, over and over, how to weave one thing into another.&lt;/p&gt;
&lt;p&gt;Starting from a single thread.&lt;/p&gt;</content:encoded></item><item><title>Why GPUs Became the Foundation of AI: A GPU Primer for K8s Veterans</title><link>https://jimmysong.io/blog/why-gpu-foundation-of-ai/</link><pubDate>Wed, 17 Jun 2026 14:04:42 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/why-gpu-foundation-of-ai/</guid><description>A GPU explainer for Kubernetes veterans new to AI. Maps token, model, training, inference, Transformer, Tensor Core, HBM, and KV cache to concepts you already know.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Note
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This post builds on Ameen Alam&amp;rsquo;s three-part GPU Architecture series and draws on TrendForce, NVIDIA (Slinky/Slurm, the GPUDirect data path) and Mirantis material on GPU infrastructure and agentic AI. It&amp;rsquo;s written for Kubernetes veterans who have never touched a GPU and never trained or served a model.
&lt;/div&gt;
&lt;/div&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/banner.webp" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/banner.webp" alt="Figure 1: GPU: the foundation of AI" data-caption="Figure 1: GPU: the foundation of AI"
width="1717"
height="916"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: GPU: the foundation of AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-a-cloud-native-veteran-is-writing-about-gpus"&gt;Why a cloud-native veteran is writing about GPUs&lt;/h2&gt;
&lt;p&gt;For the past decade my home turf has been containers and Kubernetes. From service mesh to the whole cloud-native ecosystem, I&amp;rsquo;ve spent almost every day on scheduling, networking, storage and observability, but always on the CPU side of cloud native. Last year I formally moved into AI infrastructure (AI Infra).&lt;/p&gt;
&lt;p&gt;Once I dove in, I found the concept density absurd. Token, Transformer, Tensor Core, HBM, KV Cache come at you one after another, and almost every doc and article assumes you already know them, which is deeply unfriendly to engineers who have never trained a model or run inference.&lt;/p&gt;
&lt;p&gt;I quickly hit on a trick: &lt;strong&gt;don&amp;rsquo;t learn from scratch, use what you already know by heart as a translator.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The mental model a Kubernetes veteran already carries, scheduling, Jobs, Deployments, controllers, caches, utilization, multi-tenancy, and then microservices, service mesh, distributed systems, event-driven, high-availability, maps onto AI Infra almost one to one. Once you build that mapping, the intimidating concepts become &amp;ldquo;oh, it&amp;rsquo;s just the GPU version of X&amp;rdquo;. This &amp;ldquo;translate AI through cloud-native eyes&amp;rdquo; approach got me up to speed fast, and I&amp;rsquo;m writing it down to help friends with the same background skip the detour.&lt;/p&gt;
&lt;p&gt;This is the first post in that translation series, tackling the most fundamental question: &lt;strong&gt;why does AI absolutely need GPUs?&lt;/strong&gt; Later posts will cover GPU resource management, scheduling and observability, topics a K8s veteran knows well. If you&amp;rsquo;re also crossing over from cloud native, I hope this saves you a few days of digging.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You type a sentence and the AI replies word by word. For someone who has used K8s, the most natural mental model is: AI is a &amp;ldquo;workload&amp;rdquo; that runs on GPUs, and a GPU is a special kind of &amp;ldquo;node&amp;rdquo; built for it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="translating-the-jargon-into-k8s"&gt;Translating the jargon into K8s&lt;/h2&gt;
&lt;p&gt;The AI world throws around terms that read like scripture to anyone who hasn&amp;rsquo;t done training or inference. Here&amp;rsquo;s a cheat sheet; I&amp;rsquo;ll re-explain each term in plain words below.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;AI concept&lt;/th&gt;
&lt;th&gt;In plain words&lt;/th&gt;
&lt;th&gt;K8s veteran&amp;rsquo;s analogy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Token&lt;/td&gt;
&lt;td&gt;A small unit text is chopped into; AI generates them one at a time&lt;/td&gt;
&lt;td&gt;A log line, a text chunk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;A &amp;ldquo;scoring program&amp;rdquo; holding billions of numbers (weights)&lt;/td&gt;
&lt;td&gt;A giant image whose &amp;ldquo;weights&amp;rdquo; aren&amp;rsquo;t code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training&lt;/td&gt;
&lt;td&gt;Feed it huge data, tune params repeatedly, build the model&lt;/td&gt;
&lt;td&gt;Run a Job to build an image; offline, batch, care about throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;Model is done, serve user requests and emit answers&lt;/td&gt;
&lt;td&gt;Run a Deployment taking traffic; online, care about latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformer&lt;/td&gt;
&lt;td&gt;The architecture shared by nearly all LLMs (GPT, Claude, LLaMA)&lt;/td&gt;
&lt;td&gt;A controller design pattern&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tensor Core&lt;/td&gt;
&lt;td&gt;A GPU unit that only does &amp;ldquo;matrix multiply&amp;rdquo; but insanely fast&lt;/td&gt;
&lt;td&gt;A sidecar worker that only does batched multiply&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HBM&lt;/td&gt;
&lt;td&gt;The GPU&amp;rsquo;s on-board high-bandwidth memory; model and cache live here&lt;/td&gt;
&lt;td&gt;Node-local RAM, but with bandwidth that crushes normal RAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV Cache&lt;/td&gt;
&lt;td&gt;Attention state stored at inference time, grows with the conversation&lt;/td&gt;
&lt;td&gt;A pod-local session notebook that gets thicker the longer you talk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: AI jargon → K8s veteran cheat sheet
&lt;/figcaption&gt;
&lt;p&gt;Memorize this table; the rest follows.&lt;/p&gt;
&lt;h2 id="the-question-most-people-never-ask"&gt;The question most people never ask&lt;/h2&gt;
&lt;p&gt;GPUs are everywhere in AI now, so accepted that most people skip past it. We care about which card to rent, which framework to use, which model to deploy, but rarely stop to ask the deeper question:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why GPUs?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Why not the more general CPU? Why did a chip originally built to render game graphics become the foundation of the most important technology shift in a generation?&lt;/p&gt;
&lt;p&gt;The answer isn&amp;rsquo;t just &amp;ldquo;it&amp;rsquo;s fast&amp;rdquo;, it&amp;rsquo;s about &lt;strong&gt;how computation itself is organized&lt;/strong&gt;, and why AI workloads demand an architecture fundamentally different from fifty years of software.&lt;/p&gt;
&lt;h2 id="two-philosophies-of-computing-the-craftsman-and-the-factory"&gt;Two philosophies of computing: the craftsman and the factory&lt;/h2&gt;
&lt;p&gt;You know CPUs well. K8s&amp;rsquo;s control plane, etcd, the scheduler all run on CPU; it excels at executing complex instructions one after another, with few cores (8 to 128) but each highly capable. &lt;strong&gt;A CPU optimizes for latency: how fast can I finish one complex task?&lt;/strong&gt; Like a master craftsman doing one intricate job at a time.&lt;/p&gt;
&lt;p&gt;A GPU takes the opposite path: &lt;strong&gt;execute simple instructions across massive amounts of data at once.&lt;/strong&gt; A modern data-center GPU has thousands of small cores, and &lt;strong&gt;it optimizes for throughput: how many simple tasks can I finish at the same moment?&lt;/strong&gt; Like a factory floor of thousands of workers, each doing one identical step simultaneously.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/cpu-vs-gpu-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/cpu-vs-gpu-en.svg" alt="Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing" data-caption="Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing"
width="1157"
height="340"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The CPU craftsman vs the GPU factory: latency-first vs throughput-first computing&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For any single complex job, the craftsman (CPU) is faster; but as long as the work is uniform and parallelizable, the factory (GPU) crushes the craftsman in total output per second. For decades CPU dominated because most software (web servers, databases, operating systems) is inherently serial. Then deep learning arrived.&lt;/p&gt;
&lt;h2 id="why-ai-broke-the-cpu"&gt;Why AI broke the CPU&lt;/h2&gt;
&lt;p&gt;The core of a neural network is, essentially, a giant pile of &lt;strong&gt;matrix multiplications&lt;/strong&gt; (an operation that batch-multiplies-and-adds two sets of numbers). When a model processes a token, it runs thousands or tens of thousands of these multiplies, and they &lt;strong&gt;don&amp;rsquo;t depend on each other&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The key insight in one line: &lt;strong&gt;these operations are naturally parallel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A 128-core CPU looks at this workload and sweeps through it sequentially; a many-thousand-core GPU chops the matrix into blocks and hands them to its small cores to run at once. Same work, orders of magnitude faster on GPU, not because a single GPU core is faster (it&amp;rsquo;s slower), but because &lt;strong&gt;the problem is parallel and the GPU is built for exactly that shape of computation.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the real reason AI runs on GPUs: not marketing, not legacy, architectural fit.&lt;/p&gt;
&lt;h2 id="the-three-things-inside-a-gpu-in-k8s-terms"&gt;The three things inside a GPU, in K8s terms&lt;/h2&gt;
&lt;h3 id="1-tensor-core-the-sidecar-that-only-does-batched-multiply"&gt;1. Tensor Core: the sidecar that only does batched multiply&lt;/h3&gt;
&lt;p&gt;A GPU has two kinds of cores. Regular CUDA cores are general-purpose workers that can compute anything; &lt;strong&gt;a Tensor Core is a dedicated unit that does exactly one thing: multiply two small matrices in a single shot.&lt;/strong&gt; But that one thing it does insanely fast, finishing in one cycle what would take a regular core thousands.&lt;/p&gt;
&lt;p&gt;AI&amp;rsquo;s core operation is matrix multiplication, so Tensor Cores are tailor-made for it. In K8s terms: regular CUDA cores are like general pods in a Deployment that do everything; a Tensor Core is like a highly specialized sidecar that only does &amp;ldquo;batched multiply&amp;rdquo;, single-function but with crushing throughput.&lt;/p&gt;
&lt;h3 id="2-hbm-the-nodes-high-speed-local-memory"&gt;2. HBM: the node&amp;rsquo;s high-speed local memory&lt;/h3&gt;
&lt;p&gt;Computing fast isn&amp;rsquo;t enough; you have to get the data. A GPU&amp;rsquo;s memory is layered just like a CPU node&amp;rsquo;s, and if you understand a K8s node&amp;rsquo;s L1/L2 cache, RAM and local disk, you understand the GPU&amp;rsquo;s:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/gpu-memory-hierarchy-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/gpu-memory-hierarchy-en.svg" alt="Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI" data-caption="Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI"
width="1037"
height="393"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: GPU memory hierarchy: faster and smaller up top, slower and bigger below; HBM bandwidth is the lifeblood of AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The bottom layer, &lt;strong&gt;HBM (High Bandwidth Memory)&lt;/strong&gt;, is the GPU&amp;rsquo;s &amp;ldquo;main memory&amp;rdquo;; model weights, intermediate results and KV cache all live here. It&amp;rsquo;s characterized by large capacity (hundreds of GB per card) and extremely high bandwidth.&lt;/p&gt;
&lt;p&gt;Why build HBM at all? Because traditional VRAM (GDDR) lies flat on the circuit board, you run out of routing space and hit a bandwidth wall. HBM&amp;rsquo;s answer is to &lt;strong&gt;stack memory chips vertically&lt;/strong&gt; (connected by Through-Silicon Vias, TSVs), packed right against the GPU die with an interface thousands of lines wide running in parallel. By analogy: normal memory is like a warehouse spread across a parking lot where the movers can&amp;rsquo;t keep up; HBM is like building that warehouse into a dozens-story tower right next to the workshop, with TSV elevators running up and down at full speed.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a counter-intuitive fact: &lt;strong&gt;modern AI is memory-bound more than compute-bound.&lt;/strong&gt; Tensor Cores do matrix multiplication faster than HBM can feed them data, so the GPU is often waiting. That&amp;rsquo;s why each new GPU generation (H200 → Blackwell → Vera Rubin) sees its most important upgrade in HBM bandwidth, not compute (4.8 → 8 → 22 TB/s, &lt;a href="https://www.linkedin.com/pulse/hidden-technology-behind-modern-ai-gpus-ameen-alam-tzkre" target="_blank" rel="noopener"&gt;source&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Since data movement is the bottleneck, NVIDIA&amp;rsquo;s other move is to &lt;strong&gt;let data bypass the CPU and reach the GPU directly&lt;/strong&gt;, the three &amp;ldquo;expressways&amp;rdquo; collectively called &lt;a href="https://www.linkedin.com/pulse/why-nvidia-gpu-architecture-perfect-ai-insights-data-path-kawonise-a821e/" target="_blank" rel="noopener"&gt;GPUDirect&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPUDirect Storage&lt;/strong&gt;: data goes straight from NVMe to GPU memory, without detouring through host memory and the CPU.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPUDirect RDMA&lt;/strong&gt;: GPU memory talks directly to the NIC (InfiniBand), so cross-node gradient exchange no longer hops through the CPU.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NVLink&lt;/strong&gt;: GPUs inside one machine connect directly and share memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In one line: &lt;strong&gt;the CPU is no longer the traffic cop for every data movement.&lt;/strong&gt; In K8s terms, it&amp;rsquo;s like sending data over a direct data plane instead of routing every packet through the apiserver. NVIDIA&amp;rsquo;s edge isn&amp;rsquo;t just fast compute; it&amp;rsquo;s that the entire data highway from storage to GPU, GPU to GPU, and GPU to network has been straightened out.&lt;/p&gt;
&lt;h3 id="3-transformer--token-the-program-that-emits-answers-word-by-word"&gt;3. Transformer + Token: the program that emits answers word by word&lt;/h3&gt;
&lt;p&gt;A Transformer isn&amp;rsquo;t hardware, it&amp;rsquo;s a &lt;strong&gt;model architecture&lt;/strong&gt; (a program structure). GPT, Claude, LLaMA and other large models are all built on it. What it does is actually plain:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Read a piece of text, predict the next token.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A token is a small unit text is chopped into (a character, half a word). A Transformer feeds the input token sequence in, emits &amp;ldquo;the most likely next token&amp;rdquo;, appends it, feeds the new sequence back in, predicts the next-next, and so on, generating the whole answer one piece at a time.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/token-transformer-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/token-transformer-en.svg" alt="Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next" data-caption="Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next"
width="836"
height="285"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Token and Transformer: AI plays relay, reading the tokens so far at each step and predicting the next&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In K8s terms: a Transformer is a controller whose reconcile logic is &amp;ldquo;look at the current state (tokens so far) → emit the next action (the next token)&amp;rdquo;, looping continuously. That&amp;rsquo;s where the &amp;ldquo;word by word&amp;rdquo; effect of AI assistants comes from.&lt;/p&gt;
&lt;h2 id="training-vs-inference-writing-the-recipe-vs-serving-the-dish"&gt;Training vs Inference: writing the recipe vs serving the dish&lt;/h2&gt;
&lt;p&gt;This is the easiest pair to confuse when starting out, yet the most critical. &lt;strong&gt;Training and inference are two different things with completely different hardware needs.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Training&lt;/strong&gt;: feed huge data, repeatedly tune those billions of parameters, &amp;ldquo;teach&amp;rdquo; the model into existence. It runs for weeks to months as an offline batch job, &lt;strong&gt;cares about throughput, not per-request latency.&lt;/strong&gt; In K8s terms, it&amp;rsquo;s a long-running &lt;strong&gt;Job&lt;/strong&gt; whose goal is to build a model image.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference&lt;/strong&gt;: once the model is trained and online, it takes each user&amp;rsquo;s input, computes an answer and emits it back. The user is waiting, &lt;strong&gt;so it cares about latency&lt;/strong&gt;; it must serve thousands of users at once, &lt;strong&gt;so it cares about concurrency.&lt;/strong&gt; It&amp;rsquo;s like a &lt;strong&gt;Deployment&lt;/strong&gt; taking traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/training-vs-inference-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/training-vs-inference-en.svg" alt="Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)" data-caption="Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)"
width="1038"
height="347"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Training vs Inference: training is a Job writing the recipe (offline, throughput), inference is a Deployment serving the dish (online, latency)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This directly affects which GPU metrics you should care about. The headline numbers (Tensor Core FLOPS, NVLink bandwidth, cluster size) are mostly &lt;strong&gt;training&lt;/strong&gt; metrics; what actually determines whether your AI assistant feels snappy (HBM capacity, memory bandwidth, KV cache management) are &lt;strong&gt;inference&lt;/strong&gt; metrics, and they get little airtime. Yet inference is where most production AI actually runs.&lt;/p&gt;
&lt;h2 id="multi-gpu-collaboration-an-ai-cluster-is-just-a-distributed-system"&gt;Multi-GPU collaboration: an AI cluster is just a distributed system&lt;/h2&gt;
&lt;p&gt;With single-card covered, back to reality: the biggest models don&amp;rsquo;t fit on one card, and training routinely needs hundreds or thousands of cards working together. At that point the GPU cluster is essentially a &lt;strong&gt;distributed system&lt;/strong&gt;, and almost everything you know from cloud native applies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Many GPUs = a microservice cluster.&lt;/strong&gt; A card is like a service instance; the model is sharded across cards, each computes a slice and the results are stitched back. That&amp;rsquo;s exactly the microservices playbook of splitting by responsibility and sharing load.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inter-card communication = the service-mesh data plane.&lt;/strong&gt; Cards constantly exchange data (gradients especially during training). Within a rack it&amp;rsquo;s NVLink (several TB/s), across racks it&amp;rsquo;s InfiniBand. It&amp;rsquo;s just the east-west traffic between microservices: in-node pod-to-pod direct is fastest (NVLink), cross-node cross-cluster goes over the network (IB). NVIDIA packaging GPUs, switches, DPUs and RDMA into a whole rack (&lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;Vera Rubin NVL72&lt;/a&gt;) is exactly a service mesh unifying the data plane, control plane and observability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distributed training&amp;rsquo;s all-reduce = distributed consensus.&lt;/strong&gt; Each training step synchronizes and averages the gradients across all cards, a step called all-reduce. Anyone who has done etcd or Raft gets it instantly: it&amp;rsquo;s distributed-system consensus and consistency, except here you&amp;rsquo;re syncing gradients, not a state-machine log.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Continuous batching = event-driven.&lt;/strong&gt; During inference, new requests are dropped into the running batch as they arrive, no waiting for the whole batch to finish. Requests are events, the batcher is a consumer, batch-then-process: that&amp;rsquo;s the event-driven / message-queue mindset, all to keep the GPU busy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disaggregated inference = microservices.&lt;/strong&gt; Split prefill (process input, compute-heavy) and decode (emit tokens, memory-bandwidth-heavy) into separate GPU pools that scale independently, with KV cache passed between them over NVLink/RDMA. That&amp;rsquo;s entirely the microservices pattern of splitting by responsibility and scaling each part; KV cache is the session context passed between services.&lt;/p&gt;
&lt;p&gt;See the pattern? An AI cluster isn&amp;rsquo;t a new species, it&amp;rsquo;s &lt;strong&gt;your familiar distributed-systems, microservices and service-mesh toolkit replayed on GPU hardware.&lt;/strong&gt; The instincts you built tuning traffic on Istio and Envoy, or consensus on etcd, are worth exactly as much here.&lt;/p&gt;
&lt;h2 id="running-gpus-on-k8s-slurm-for-training-k8s-for-inference"&gt;Running GPUs on K8s: Slurm for training, K8s for inference&lt;/h2&gt;
&lt;p&gt;By now, as a K8s veteran, you&amp;rsquo;re bound to ask: so what actually schedules a GPU cluster? The interesting answer: &lt;strong&gt;training and inference often run on two different schedulers.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On the training side, the HPC world&amp;rsquo;s heavyweight is Slurm.&lt;/strong&gt; It manages 65% of the world&amp;rsquo;s TOP500 supercomputers (&lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;NVIDIA&lt;/a&gt;), and large AI training teams have years invested in Slurm scripts, fair-share policies and accounting. Training is a long-running batch job that wants topology awareness (place chatty GPUs close together), long exclusive holds and high throughput, areas Slurm has refined for over a decade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On the inference side, the home field is Kubernetes.&lt;/strong&gt; Inference is an online service that wants fast autoscaling, per-request scheduling and integration with service mesh and observability, exactly K8s&amp;rsquo;s strengths.&lt;/p&gt;
&lt;p&gt;So what if a team needs both, maintain two environments? NVIDIA open-sourced the &lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;Slinky&lt;/a&gt; project to solve exactly this, in a very K8s-native way:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/SlinkyProject/slurm-operator" target="_blank" rel="noopener"&gt;slurm-operator&lt;/a&gt;&lt;/strong&gt;: turns each Slurm daemon (slurmctld for scheduling, slurmd for compute, slurmdbd for accounting) into a K8s CRD and Pod, with the control plane made highly available through Pod regeneration instead of Slurm&amp;rsquo;s native HA. Config changes sync automatically via ConfigMap/Secret, workers autoscale with HPA, scale-in drains running jobs first, and upgrades use PodDisruptionBudget to protect in-flight work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU Operator + DCGM Exporter&lt;/strong&gt;: auto-installs drivers and the device plugin, and can label metrics by Slurm job ID, giving you &lt;strong&gt;per-job&lt;/strong&gt; GPU metrics (scraped by Prometheus).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ComputeDomains + DRA&lt;/strong&gt;: for cross-node-NVLink machines like GB200 NVL72, K8s uses DRA (Dynamic Resource Allocation) to dynamically manage the cross-node GPU interconnect domain, so distributed training hits full NVLink bandwidth across nodes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;NVIDIA itself runs Slinky in production up to 8,000+ GPUs, with NCCL all-reduce/all-gather performance matching bare Slurm, the K8s layer adding almost no overhead (&lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;NVIDIA production data&lt;/a&gt;). &lt;strong&gt;K8s is becoming the substrate for GPU computing, with Slurm as the scheduling layer on top.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Training and inference have fundamentally different GPU-scheduling needs, so they can&amp;rsquo;t be managed the same way. That&amp;rsquo;s also why GPU resource management (covered later) has to be scenario-specific, and why solutions like &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; aim to be a unified GPU resource-management layer across Slurm + K8s hybrid environments.&lt;/p&gt;
&lt;h2 id="the-real-trap-of-inference-kv-cache"&gt;The real trap of inference: KV Cache&lt;/h2&gt;
&lt;p&gt;If there&amp;rsquo;s one concept that separates &amp;ldquo;GPU theory&amp;rdquo; from &amp;ldquo;AI inference reality&amp;rdquo;, it&amp;rsquo;s &lt;strong&gt;KV Cache&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When a Transformer generates each token, it has to reference the relationships among all previous tokens (this is &amp;ldquo;attention&amp;rdquo;). To avoid recomputing from scratch every time, it stores each token&amp;rsquo;s attention state, and that store is the KV Cache.&lt;/p&gt;
&lt;p&gt;The problem is it &lt;strong&gt;grows without bound&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user sends a 10,000-token conversation; the model stores a state entry for each.&lt;/li&gt;
&lt;li&gt;To generate the next token, &lt;strong&gt;it reads the entire KV Cache.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;A 70B model at 32K context can produce a KV Cache of 8 to 16 GB per request, larger than the model itself.&lt;/li&gt;
&lt;li&gt;50 concurrent users land, and KV Cache eats hundreds of GB of memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/kv-cache-growth-en.svg" data-img="https://assets.jimmysong.io/images/blog/why-gpu-foundation-of-ai/kv-cache-growth-en.svg" alt="Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook" data-caption="Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook"
width="980"
height="353"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: KV Cache is a session notebook that thickens the longer you talk: every new token forces a full read of the whole notebook&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In K8s terms: KV Cache is like a &lt;strong&gt;session notebook&lt;/strong&gt; local to a Pod. The longer the conversation, the thicker the notebook, and every new word forces a full read from cover to cover. So inference memory is often eaten not by the model but by this &amp;ldquo;notebook&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s also why we got paged attention (managing the notebook like virtual-memory pages to cut fragmentation, popularized by &lt;a href="https://github.com/vllm-project/vllm" target="_blank" rel="noopener"&gt;vLLM&lt;/a&gt;) and NVIDIA &lt;a href="https://github.com/ai-dynamo/dynamo" target="_blank" rel="noopener"&gt;Dynamo&lt;/a&gt;&amp;rsquo;s multi-tier cache (hot pages in HBM, warm offloaded to CPU memory, cold spilled to NVMe). &lt;strong&gt;Half an inference engineer&amp;rsquo;s job is managing this notebook.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="gpu-utilization-lies"&gt;GPU utilization lies&lt;/h2&gt;
&lt;p&gt;K8s folks know &amp;ldquo;utilization lies&amp;rdquo; best: a pod reporting Running isn&amp;rsquo;t necessarily doing work, it might be in CPU steal or waiting on IO. GPUs are worse.&lt;/p&gt;
&lt;p&gt;That &amp;ldquo;GPU utilization %&amp;rdquo; in monitoring usually only measures &amp;ldquo;is the GPU executing any kernel&amp;rdquo;, not how efficiently. A card can report 90% utilization while its Tensor Cores are actually busy only 30% of the time, with the rest spent on memory ops, kernel-launch overhead, or just waiting for data.&lt;/p&gt;
&lt;p&gt;The more honest metric is &lt;strong&gt;SM Efficiency (SM activity rate)&lt;/strong&gt;: it looks at how many SMs are doing useful work each clock cycle. A card showing 100% utilization in nvidia-smi may have an SM Efficiency of only 20-30%. Many companies think their &amp;ldquo;GPUs are maxed out&amp;rdquo; when in fact huge amounts of compute are spinning idle. So to judge whether a GPU is truly working, don&amp;rsquo;t look at utilization, look at SM Efficiency.&lt;/p&gt;
&lt;p&gt;The metrics that actually matter (mapping to the QPS, P99 latency and resource levels you watch in K8s):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Token throughput&lt;/strong&gt;: tokens generated per second per card, how many users you can serve (like QPS).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TTFT (Time To First Token)&lt;/strong&gt;: time from receiving a request to emitting the first token, sets the &amp;ldquo;responsiveness feel&amp;rdquo; (like cold-start time).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TPOT (Time Per Output Token, also called ITL / inter-token latency)&lt;/strong&gt;: average time to produce each token, sets streaming smoothness (like P99). Serving frameworks like vLLM and TGI generally use TPOT; NVIDIA more often calls it ITL, same thing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory breakdown&lt;/strong&gt;: how much is weights vs KV cache vs transient activations, tells you whether you&amp;rsquo;re memory-bound (like breaking down a pod&amp;rsquo;s memory usage).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="why-nvidia-is-so-hard-to-dislodge"&gt;Why NVIDIA is so hard to dislodge&lt;/h2&gt;
&lt;p&gt;You can&amp;rsquo;t talk GPUs without NVIDIA&amp;rsquo;s dominance. The hardware is excellent, but hardware alone can&amp;rsquo;t explain why AMD, Intel and a crowd of startups have failed to gain ground. The answer is &lt;strong&gt;ecosystem depth&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;CUDA isn&amp;rsquo;t just a parallel-computing platform, it&amp;rsquo;s a programming model, compiler, runtime and a stack of libraries, refined for 18 years. Nearly every AI framework (PyTorch, TensorFlow, JAX) grew up on CUDA first and was ported elsewhere as an afterthought. In K8s terms: CUDA is to GPUs roughly what the Linux kernel + containerd + the whole CNCF toolchain are to the container ecosystem, &lt;strong&gt;not something you replace by swapping a kernel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Switching to another card means re-validating every layer of your inference stack (kernels, libraries, frameworks, serving, monitoring, ops). The switching cost isn&amp;rsquo;t buying different hardware, it&amp;rsquo;s &lt;strong&gt;rebuilding an entire software ecosystem.&lt;/strong&gt; That&amp;rsquo;s the &amp;ldquo;CUDA moat&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="the-future-ai-factories"&gt;The future: AI factories&lt;/h2&gt;
&lt;p&gt;NVIDIA no longer describes its business in terms of &amp;ldquo;GPUs&amp;rdquo; or even &amp;ldquo;data centers&amp;rdquo;, but as &amp;ldquo;AI factories&amp;rdquo;: facilities that continuously convert electricity, silicon and data into intelligence. Beneath the language is a real architectural shift:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Power is now the primary constraint&lt;/strong&gt;: a single Vera Rubin rack draws over 100 kW (&lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;source&lt;/a&gt;), and a mid-size training cluster needs 10 to 50 MW, comparable to a small town. The GPU is no longer the hard part; securing reliable power is.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cooling moves off air&lt;/strong&gt;: liquid cooling is now standard for high-density GPUs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Networking decides GPU placement&lt;/strong&gt;: AI cluster design now starts with &lt;strong&gt;network topology and works backward to where GPUs go&lt;/strong&gt;, the reverse of traditional data centers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory bandwidth remains the scaling frontier&lt;/strong&gt;: each generation adds compute, but what actually moves the experience is HBM bandwidth.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s fundamentally HA + datacenter ops&lt;/strong&gt;: power redundancy, cooling failover, latency determined by network topology, single-point failure and DR are all old problems for anyone who has done distributed-systems HA. The AI factory isn&amp;rsquo;t new magic; it&amp;rsquo;s your HA-architecture skillset moved into a data center an order of magnitude denser in power and compute.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dont-retire-the-cpu-just-yet-agentic-ai-is-rebalancing-the-cpugpu-ratio"&gt;Don&amp;rsquo;t retire the CPU just yet: agentic AI is rebalancing the CPU:GPU ratio&lt;/h2&gt;
&lt;p&gt;After all this GPU praise, you might think the CPU is sidelined in the AI era. The opposite is true: &lt;strong&gt;agentic AI is putting the CPU back at center stage.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In traditional LLM inference the CPU mostly just compresses and routes data for the GPU, so AI data centers ran CPU:GPU ratios as low as 1:4 or even 1:8. Agents are different: they plan tasks autonomously, call tools, route data between sub-agents and decide whether a task is complete, and all of that &lt;strong&gt;orchestration logic lands squarely on the CPU.&lt;/strong&gt; Add that agents are often trained with reinforcement learning, where every action has to be evaluated by the CPU, and the CPU load gets heavier still.&lt;/p&gt;
&lt;p&gt;The signal for K8s veterans is clear: &lt;strong&gt;the future AI node is a mixed-workload node where CPU and GPU are billed together&lt;/strong&gt;, and scheduling and resource management must handle both, not just stare at the GPU.&lt;/p&gt;
&lt;h2 id="what-this-means-for-me-a-k8s-veteran"&gt;What this means for me (a K8s veteran)&lt;/h2&gt;
&lt;p&gt;After all this hardware, it lands on what I work on: &lt;strong&gt;only by first understanding why the GPU is the foundation of AI can you understand why &amp;ldquo;GPU resource management&amp;rdquo; is a real problem.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When a card costs hundreds of thousands of yuan, a whole rack draws over a hundred kilowatts, and production GPUs still run below capacity because memory bandwidth, KV cache or batching aren&amp;rsquo;t tuned, just &amp;ldquo;handing out cards&amp;rdquo; is nowhere near enough. How to run multiple tenants safely on one card (like K8s scheduling many pods onto one node), how to partition resources between the very different workloads of training and inference, how to push utilization from &amp;ldquo;looks full&amp;rdquo; to &amp;ldquo;actually full&amp;rdquo;, that&amp;rsquo;s exactly what GPU virtualization and sharing (the work I do) solves.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a sobering number: per the &lt;a href="https://www.mirantis.com/blog/gpu-infrastructure-automation-and-strategy/" target="_blank" rel="noopener"&gt;ClearML AI infrastructure survey&lt;/a&gt;, only about &lt;strong&gt;7%&lt;/strong&gt; of enterprises hit over 85% GPU utilization at peak, more than half sit at 51-70%, and 15% are below 50%. In other words, a big chunk of the GPUs companies pay dearly for are spinning idle. That&amp;rsquo;s usually not a hardware shortage, it&amp;rsquo;s &lt;strong&gt;scheduling and management not keeping up.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;My own take is that GPU resource management is moving through three stages: &lt;strong&gt;Allocation → Utilization → Efficiency.&lt;/strong&gt; Stage one only answers &amp;ldquo;who gets this card&amp;rdquo; (device plugin, exclusive or time-slicing); stage two asks &amp;ldquo;is this card full&amp;rdquo; (dynamic batching, elastic autoscaling, where most enterprises are stuck); stage three asks &amp;ldquo;is this card&amp;rsquo;s compute producing maximum value&amp;rdquo; (SM Efficiency, tokens per watt, bandwidth utilization). The real challenge is stage three, and the future GPU scheduler shouldn&amp;rsquo;t be a resource allocator but a &lt;strong&gt;fine-grained orchestrator of the GPU&amp;rsquo;s internals&lt;/strong&gt;, reading how many SMs are active, the Tensor Core utilization, how much memory KV cache holds, and only then deciding whether one more request fits.&lt;/p&gt;
&lt;p&gt;Plainly put, &lt;strong&gt;the AI era is replaying the cloud-native story: from &amp;ldquo;single-machine exclusive&amp;rdquo; toward &amp;ldquo;multi-tenant sharing + scheduling + observability.&amp;rdquo;&lt;/strong&gt; And the GPU is the main stage of that play.&lt;/p&gt;
&lt;h2 id="this-is-only-the-first-half-who-else-is-at-the-table-besides-nvidia"&gt;This is only the first half: who else is at the table besides NVIDIA&lt;/h2&gt;
&lt;p&gt;By now you&amp;rsquo;ve probably noticed &lt;strong&gt;this post barely mentions anyone but NVIDIA.&lt;/strong&gt; Unavoidable, it&amp;rsquo;s the absolute protagonist today. But if you think the AI accelerator world begins and ends with NVIDIA, you&amp;rsquo;re very wrong.&lt;/p&gt;
&lt;p&gt;In fact, an &amp;ldquo;anti-NVIDIA alliance&amp;rdquo; is gathering from all sides:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cloud-vendor silicon&lt;/strong&gt;: Google&amp;rsquo;s TPU is on its sixth generation and backs almost all of its own AI; AWS&amp;rsquo;s Trainium (training) and Inferentia (inference) keep spreading; Meta and Microsoft aren&amp;rsquo;t sitting still.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Traditional chip giants&lt;/strong&gt;: AMD presses hard with the Instinct line and ROCm, Intel holds ground with Gaudi and oneAPI.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;China&amp;rsquo;s heterogeneous accelerator ecosystem&lt;/strong&gt;: Huawei Ascend, Hygon DCU, Cambricon, Moore Threads, Enflame, Kunlunxin, Metax, Biren&amp;hellip;, sprinting through the domestic-substitution window, with their software stacks climbing from &amp;ldquo;works&amp;rdquo; toward &amp;ldquo;works well&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The open ecosystem fights back&lt;/strong&gt;: the UXL Alliance, OpenAI Triton, and PyTorch&amp;rsquo;s native AMD/TPU backends are all gnawing at the walls of the CUDA moat.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What does that mean for a K8s veteran? It means &lt;strong&gt;the future AI cluster is almost certainly heterogeneous&lt;/strong&gt;: a rack might hold NVIDIA, AMD, TPU and domestic cards side by side, and one schedule has to manage several completely different kinds of hardware.&lt;/p&gt;
&lt;p&gt;So the really interesting questions follow:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How do you mix multiple accelerator types in one cluster and still schedule and observe them uniformly?&lt;/li&gt;
&lt;li&gt;Where do K8s device plugins and DRA (Dynamic Resource Allocation) evolve to, so they can elegantly describe this menagerie of hardware?&lt;/li&gt;
&lt;li&gt;Will the CUDA moat be breached by the open ecosystem, or will NVIDIA rule long-term like x86 did?&lt;/li&gt;
&lt;li&gt;What pieces are still missing before domestic heterogeneous accelerators are truly production-ready?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&amp;rsquo;ll dig into those in later posts. &lt;strong&gt;This one only nails &amp;ldquo;why GPUs&amp;rdquo;; the next will tackle the big chess game of how AI infrastructure should schedule things once &amp;ldquo;GPUs aren&amp;rsquo;t just one kind&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/why-gpus-became-foundation-modern-ai-ameen-alam-1oiee" target="_blank" rel="noopener"&gt;Why GPUs Became the Foundation of Modern AI&lt;/a&gt; (Part 1 of the trilogy)&lt;/li&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/hidden-technology-behind-modern-ai-gpus-ameen-alam-tzkre" target="_blank" rel="noopener"&gt;The Hidden Technology Behind Modern AI GPUs&lt;/a&gt; (Part 2)&lt;/li&gt;
&lt;li&gt;Ameen Alam, &lt;a href="https://www.linkedin.com/pulse/real-ai-infrastructure-future-gpus-ameen-alam-os7ff" target="_blank" rel="noopener"&gt;Real AI Infrastructure and the Future of GPUs&lt;/a&gt; (Part 3)&lt;/li&gt;
&lt;li&gt;TrendForce, &lt;a href="https://insights.trendforce.com/p/agentic-ai-cpu-gpu" target="_blank" rel="noopener"&gt;The Great Rebalance: How Agentic AI Is Reshaping the CPU/GPU Ratio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Anton Polyakov (NVIDIA), &lt;a href="https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/" target="_blank" rel="noopener"&gt;Running Large-Scale GPU Workloads on Kubernetes with Slurm&lt;/a&gt; (Slinky / Slurm on K8s)&lt;/li&gt;
&lt;li&gt;Kawonise, &lt;a href="https://www.linkedin.com/pulse/why-nvidia-gpu-architecture-perfect-ai-insights-data-path-kawonise-a821e/" target="_blank" rel="noopener"&gt;Why NVIDIA GPU Architecture Is Perfect for AI: GPU Data Path for a Single Node&lt;/a&gt; (GPUDirect data path)&lt;/li&gt;
&lt;li&gt;Mirantis, &lt;a href="https://www.mirantis.com/blog/gpu-infrastructure-automation-and-strategy/" target="_blank" rel="noopener"&gt;GPU Infrastructure Automation and Strategy&lt;/a&gt; (GPU infra automation and utilization)&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>GPU Utilization Is Breaking: AI Infrastructure Needs a New Definition of Efficiency</title><link>https://jimmysong.io/blog/beyond-gpu-utilization-productive-gpu-hours/</link><pubDate>Wed, 17 Jun 2026 06:22:24 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/beyond-gpu-utilization-productive-gpu-hours/</guid><description>From GPU utilization to productive GPU-hours.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Every GPU should not just be used. It should create value.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="we-have-all-been-chasing-gpu-utilization"&gt;We have all been chasing GPU utilization&lt;/h2&gt;
&lt;p&gt;For the past few years, whether it is Kubernetes GPU scheduling, vGPU, MIG, or &lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;, everyone has really been doing the same thing: pushing one number up.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Utilization&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;It makes sense. GPUs are expensive. An H100 can run anywhere from a few dollars to over ten dollars per GPU-hour, and nobody can afford to let a GPU sit idle. So the entire AI Infra community&amp;rsquo;s narrative for the past few years has boiled down to one line:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Maximize GPU utilization.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I have been involved in the HAMi community for a long time, and a few posts I have written, like &lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes as the GPU Control Plane for AI&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;From GPU to Token: An Eight-Layer Observability Stack for AI Infrastructure&lt;/a&gt;, are really about the same thing: how to slice GPUs finer, share them more thoroughly, and schedule them more sensibly.&lt;/p&gt;
&lt;p&gt;But recently I read Arjun Kaarat&amp;rsquo;s piece in Towards Data Science, &lt;a href="https://towardsdatascience.com/when-gpu-utilization-lies-the-hidden-systems-problem-slowing-modern-ai/" target="_blank" rel="noopener"&gt;&lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI&lt;/em&gt;&lt;/a&gt;, and it made me rethink a question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Is GPU utilization really the metric we should be optimizing for?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There is one line in the article that made me pause for a few seconds the first time I read it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;GPUs can be busy without being productive.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Arjun Kaarat&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It punctures a layer of illusion. The number we have spent so much effort pushing up may have been answering the wrong question from the start.&lt;/p&gt;
&lt;h2 id="gpu-busy-is-not-gpu-productive"&gt;GPU Busy is not GPU Productive&lt;/h2&gt;
&lt;p&gt;Kaarat tells a representative story in the article.&lt;/p&gt;
&lt;p&gt;At 2 AM, an infrastructure team gets paged: inference latency just spiked 60%. They open the monitoring dashboard, and GPU utilization looks perfectly normal:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 79%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 82%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 84%&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Looks healthy. So the usual playbook kicks in: trigger autoscaling, add nodes, add GPUs. The cloud bill climbs, but latency barely improves.&lt;/p&gt;
&lt;p&gt;An hour later, they find the root cause: three nodes had quietly entered a RAID rebuild state, storage throughput was severely dragged down, and the inference tasks around them were starving. The scheduler kept treating these nodes as &amp;ldquo;still healthy enough&amp;rdquo; because the GPU and memory metrics looked fine, but the underlying disk performance had collapsed.&lt;/p&gt;
&lt;p&gt;What strikes me most about this story is that it is not a rare edge case. It is a failure mode that is becoming common.&lt;/p&gt;
&lt;p&gt;Many teams look at their monitoring and see:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 82%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 84%
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU: 79%&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;and think &amp;ldquo;our cluster is busy and healthy.&amp;rdquo; But at the same time:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Latency ↑
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Queue ↑
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Throughput ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Cost ↑&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The real problem is that &lt;strong&gt;the GPU is waiting&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Waiting for what?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The retrieval pipeline to feed embeddings over;&lt;/li&gt;
&lt;li&gt;The SSD to read context out of storage;&lt;/li&gt;
&lt;li&gt;The CPU to prepare the data pipeline;&lt;/li&gt;
&lt;li&gt;The KV cache to have room for new requests;&lt;/li&gt;
&lt;li&gt;Storage I/O to not be squeezed out by background tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The GPU end is full, but the data path feeding it is empty, blocked, or collapsing. From the dashboard the GPU reads 84%, but the actual output may be less than half. Kaarat describes this state precisely in the original piece:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A GPU that appears active may still spend meaningful time waiting for the system around it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is the overlooked gap between &amp;ldquo;busy&amp;rdquo; and &amp;ldquo;productive&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="the-illusion-of-gpu-utilization"&gt;The &amp;ldquo;illusion&amp;rdquo; of GPU utilization&lt;/h2&gt;
&lt;p&gt;The RAID story above is still just a &amp;ldquo;point failure&amp;rdquo;. The more compelling part of Kaarat&amp;rsquo;s article describes a systemic phenomenon: &lt;strong&gt;Fragmentation&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Consider a cluster with three nodes, after running a mixed wave of GenAI workloads:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Node&lt;/th&gt;
&lt;th&gt;GPU compute&lt;/th&gt;
&lt;th&gt;HBM&lt;/th&gt;
&lt;th&gt;Storage bandwidth&lt;/th&gt;
&lt;th&gt;I/O CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;nearly full&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;saturated&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;limited&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;available&lt;/td&gt;
&lt;td&gt;saturated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Residual resources on three nodes after a GenAI wave
&lt;/figcaption&gt;
&lt;p&gt;Now a new inference job arrives, with an ordinary footprint: a little GPU, a little VRAM, decent storage bandwidth, decent I/O capacity.&lt;/p&gt;
&lt;p&gt;In total, the cluster still has plenty of resources. A has GPU and bandwidth, B has VRAM, C has bandwidth and CPU. But no single node can take this job on its own.&lt;/p&gt;
&lt;p&gt;That is fragmentation. I drew it out, roughly like this:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/fragmentation-three-nodes-en.svg" data-img="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/fragmentation-three-nodes-en.svg" alt="Figure 1: The cluster is not short on resources. It is short on resources of the right shape." data-caption="Figure 1: The cluster is not short on resources. It is short on resources of the right shape."
width="702"
height="481"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: The cluster is not short on resources. It is short on resources of the right shape.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The cluster is not empty. It has just been carved into &amp;ldquo;leftovers&amp;rdquo; that can no longer be used productively. Kaarat sums up the phenomenon in one line, which I think is the most memorable sentence in the whole piece:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The cluster is not empty. It is fragmented into leftovers that are difficult to use productively.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Put another way:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The cluster does not lack resources. It lacks resources of the right shape.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This judgment matters a lot for a project like HAMi, which builds a GPU resource control plane. We used to think that &amp;ldquo;slicing the card finely and letting more people share it&amp;rdquo; was the answer to fragmentation. But Kaarat points to a deeper problem: &lt;strong&gt;fragmentation is not just a GPU-layer concern&lt;/strong&gt;. It spans GPU, HBM, storage bandwidth, and I/O CPU. You can slice the GPU as finely as you like, but if the storage dimension is choked, that node is still unavailable for the next genuinely useful task.&lt;/p&gt;
&lt;h2 id="what-hami-solves"&gt;What HAMi solves&lt;/h2&gt;
&lt;p&gt;Let me directly answer a question: what role does HAMi play in this chain?&lt;/p&gt;
&lt;p&gt;HAMi solves a very specific, and very foundational, problem:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Can the GPU be used by more people.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;What it does can be summarized in one line: &lt;strong&gt;reduce fragmentation at the GPU layer&lt;/strong&gt;. The concrete forms include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPU Sharing&lt;/strong&gt;: letting multiple Pods share one card instead of one card per Pod;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;vGPU / HAMi-core&lt;/strong&gt;: doing memory isolation and compute throttling in userspace, slicing one card into MB-level virtual devices;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MIG integration&lt;/strong&gt;: managing NVIDIA MIG hardware partitions in software;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Heterogeneous GPU abstraction&lt;/strong&gt;: abstracting more than a dozen device families, including NVIDIA, Ascend, Cambricon, Hygon, and Vastai, into semantics the scheduler can consume uniformly;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DRA compatibility&lt;/strong&gt;: keeping pace with the evolution of the Kubernetes resource model.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the first layer of efficiency. It answers the question &amp;ldquo;can the GPU be put to use&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;At this layer, HAMi&amp;rsquo;s value is already clear: in real environments with mixed domestic and foreign GPUs, mixed training and inference, and multi-tenant sharing, HAMi lets one card serve more workloads, so the cluster is no longer wasted by a coarse-grained &amp;ldquo;one card per Pod&amp;rdquo; model.&lt;/p&gt;
&lt;p&gt;But note: this is only the first chapter of the efficiency story.&lt;/p&gt;
&lt;h2 id="what-comes-after-hami"&gt;What comes after HAMi&lt;/h2&gt;
&lt;p&gt;If I step back and look at it from a higher vantage point, &amp;ldquo;improving GPU utilization&amp;rdquo; is actually solved across three distinct layers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer one: can the GPU be sliced, shared, and allocated?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is what a project like HAMi solves. It corresponds to the GPU resource control plane.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer two: who gets to use the GPU? Who runs first, who queues, who has priority?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is what schedulers like &lt;a href="https://volcano.sh/en/" target="_blank" rel="noopener"&gt;Volcano&lt;/a&gt;, &lt;a href="https://kueue.sigs.k8s.io/" target="_blank" rel="noopener"&gt;Kueue&lt;/a&gt;, and &lt;a href="https://github.com/NVIDIA/KAI-Scheduler" target="_blank" rel="noopener"&gt;KAI Scheduler&lt;/a&gt; solve. It corresponds to job queuing, fair share, priority, and gang scheduling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer three: can the GPU actually run? Are the data path, storage I/O, and KV cache keeping up?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is exactly where Kaarat&amp;rsquo;s article sounds the alarm. When a node&amp;rsquo;s RAID is rebuilding, its SSD queue is exploding, and its I/O CPU is eaten by background tasks, no matter how much GPU you allocate to it and no matter how elegantly the scheduler queues its tasks, it still cannot produce effective compute.&lt;/p&gt;
&lt;p&gt;Draw these three layers together, and you get what next-generation AI infrastructure should actually look like:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/productive-gpu-hours-stack-en.svg" data-img="https://assets.jimmysong.io/images/blog/beyond-gpu-utilization-productive-gpu-hours/productive-gpu-hours-stack-en.svg" alt="Figure 2: From GPU utilization to productive GPU-hours" data-caption="Figure 2: From GPU utilization to productive GPU-hours"
width="695"
height="552"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: From GPU utilization to productive GPU-hours&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I think this diagram is the single most important one for understanding the whole picture.&lt;/p&gt;
&lt;p&gt;For the past few years, almost all of the industry&amp;rsquo;s attention has been on the bottom two layers: Kubernetes and HAMi. Those two layers have essentially solved &amp;ldquo;can the GPU be put to use&amp;rdquo;. Volcano, Kueue, and KAI are also mature at layer two, solving queuing and priority.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;But layer three, storage / I/O-aware scheduling, is currently almost a blank.&lt;/strong&gt; And that is precisely the layer Kaarat&amp;rsquo;s article keeps emphasizing, the one that is becoming increasingly valuable in modern GenAI systems. Because for workloads like RAG, long context, and multimodal, the bottleneck has long since shifted from &amp;ldquo;is the GPU enough&amp;rdquo; to &amp;ldquo;is the data path feeding the GPU clear&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;To put it bluntly: HAMi slices the GPU finely, and Volcano queues the jobs well, but if a node assigned a task cannot feed its GPU, then all that upstream effort is just pumping blood into an idle endpoint.&lt;/p&gt;
&lt;h2 id="what-should-next-gen-ai-infrastructure-optimize-for"&gt;What should next-gen AI infrastructure optimize for&lt;/h2&gt;
&lt;p&gt;Based on the layering above, I want to make one clear point: &lt;strong&gt;we should upgrade the optimization target from &amp;ldquo;GPU utilization&amp;rdquo; to &amp;ldquo;Productive GPU-Hours&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is not wordplay. It is a shift that translates directly into money.&lt;/p&gt;
&lt;p&gt;Kaarat does the math in the article. A 1000-H100 cluster, at a blended cost of about $3 per GPU-hour, runs around $26 million a year. If fragmentation and I/O stall quietly waste 10% of the effective GPU time, that is roughly $2.6 million a year of wasted spend. Not because the GPUs are missing, but because the system failed to use them efficiently.&lt;/p&gt;
&lt;p&gt;That math can be translated into a simple contrast.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The past target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize GPU Utilization&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Today&amp;rsquo;s target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize Productive GPU-Hours&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;The future target:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Maximize Productive Compute
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Across Heterogeneous AI Clusters&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This evolution maps exactly onto the three-layer structure in the diagram above. That final line, &amp;ldquo;Across Heterogeneous AI Clusters&amp;rdquo;, is the heterogeneous narrative HAMi has been pushing all along: in the future you will not optimize just one kind of GPU. You will maximize effective compute uniformly across completely different cards from NVIDIA, Ascend, Cambricon, and Hygon.&lt;/p&gt;
&lt;p&gt;In other words, HAMi&amp;rsquo;s long-term value should not be boxed into the old &amp;ldquo;improve GPU utilization&amp;rdquo; narrative. Its real direction is: &lt;strong&gt;a resource control plane that lets Productive GPU-Hours be maximized across heterogeneous AI clusters.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="from-utilization-to-productive-gpu-hours"&gt;From utilization to productive GPU-hours&lt;/h2&gt;
&lt;p&gt;If I had to summarize this whole line of thinking in one sentence, I would put it like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;HAMi solves &amp;ldquo;how to let more GPUs be used&amp;rdquo;, and next-generation AI infrastructure has to solve &amp;ldquo;how to let every GPU actually create value&amp;rdquo;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is also the deepest line Kaarat&amp;rsquo;s article left me with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The real question is no longer &amp;ldquo;Are the GPUs busy?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;It is: &amp;ldquo;Are they productively busy?&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It also makes me rethink what the HAMi community&amp;rsquo;s tagline for the next phase should be. We used to say &amp;ldquo;let GPUs be shared by more people&amp;rdquo;. Next, we should probably move toward:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Turn GPU utilization into productive GPU-hours.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Let every GPU not just be used, but genuinely create value.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;GPU utilization was never the endpoint. It was only the first chapter of the efficiency story. Only when we start talking about Productive GPU-Hours, and start caring about storage, I/O, the data path, and the aggregate output of heterogeneous clusters, does AI infrastructure truly enter its second chapter.&lt;/p&gt;
&lt;h2 id="acknowledgments-and-references"&gt;Acknowledgments and references&lt;/h2&gt;
&lt;p&gt;Parts of this post were inspired by Arjun Kaarat&amp;rsquo;s piece in Towards Data Science, &lt;a href="https://towardsdatascience.com/when-gpu-utilization-lies-the-hidden-systems-problem-slowing-modern-ai/" target="_blank" rel="noopener"&gt;&lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI&lt;/em&gt;&lt;/a&gt;, and quote its published paper and article. My thanks to the author. Both diagrams in this post (resource fragmentation, the efficiency layering) were redrawn by the author based on the ideas in the original, not copied from it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Arjun Kaarat. &lt;em&gt;When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI.&lt;/em&gt; Towards Data Science, 2026.&lt;/li&gt;
&lt;li&gt;Kaarat, A., Batthula, V. J. R., &amp;amp; Segall, R. &lt;em&gt;Fitting the Void: Residual-Aware Geometric Packing for GenAI Workloads.&lt;/em&gt; IEEE, 2025.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Further reading:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra"&gt;Kubernetes as the GPU Control Plane: HAMi v2.9 and Next-Gen AI Infra&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;From GPU to Token: An Eight-Layer Observability Stack for AI Infrastructure&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ai-inference-on-kubernetes"&gt;Why AI Inference Belongs on Kubernetes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>When an Agent Becomes a Distributed State Machine: Agentic AI Infrastructure Reliability</title><link>https://jimmysong.io/blog/agentic-ai-infrastructure-reliability/</link><pubDate>Tue, 16 Jun 2026 12:45:16 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/agentic-ai-infrastructure-reliability/</guid><description>A practical AI Infra review of Agentic AI reliability, covering a five-dimension framework, fault tolerance, recovery, observability, and hybrid architecture design.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Statement
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This is my reading and critique of the paper &amp;ldquo;AI Infrastructure Reliability Features and Architecture for Agentic AI&amp;rdquo; (June 2026), shared by Hesham ElBakoury in the &lt;a href="https://www.opencompute.org" target="_blank" rel="noopener"&gt;Open Compute Project (OCP)&lt;/a&gt; community. It blends my personal engineering perspective from the Kubernetes / AI Infra space, and is not a translation of the original paper.
&lt;/div&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;An Agent is not a single inference request but a long-running distributed state machine; therefore, Agent reliability is fundamentally a distributed-systems problem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="why-i-wanted-to-write-about-this-paper"&gt;Why I wanted to write about this paper&lt;/h2&gt;
&lt;p&gt;After years in cloud native, I have a stubborn instinct: &lt;strong&gt;whether a new technology is mature is not judged by how powerful its model is or how flashy the demo looks, but by whether it has a systematic language for reliability.&lt;/strong&gt; Kubernetes won not on the container runtime, but on liveness/readiness probes, the reconciler loop, PDs/PDBs, and the whole SRE vocabulary that made &amp;ldquo;long-running distributed state&amp;rdquo; legible.&lt;/p&gt;
&lt;p&gt;So when I came across the paper Hesham ElBakoury shared in the &lt;a href="https://www.opencompute.org" target="_blank" rel="noopener"&gt;Open Compute Project (OCP)&lt;/a&gt; community, it caught my eye. It proposes no new algorithm or framework; the entire paper does one thing: &lt;strong&gt;systematically translate the traditional SRE reliability vocabulary onto Agentic AI.&lt;/strong&gt; This is exactly what I have been circling around for the past six months in &lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt;, without a canonical reference to align against. So I decided to write a dedicated post: unpack its framework, then give it a critical evaluation from the AI Infra practitioner&amp;rsquo;s point of view.&lt;/p&gt;
&lt;p&gt;One-line positioning: &lt;strong&gt;it reads more like an SRE white paper for Agentic AI than a systems paper.&lt;/strong&gt; Its value is not in novelty but in &amp;ldquo;building consensus&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="one-line-summary"&gt;One-line summary&lt;/h2&gt;
&lt;p&gt;Traditional AI cares about model accuracy, while Agentic AI must care about &amp;ldquo;reliability during long-running operation&amp;rdquo;, so fault tolerance, recovery, monitoring, security, and state management must be elevated to first-class citizens of architectural design.&lt;/p&gt;
&lt;h2 id="the-fundamental-split-between-traditional-ai-and-agentic-ai"&gt;The fundamental split between Traditional AI and Agentic AI&lt;/h2&gt;
&lt;p&gt;The paper argues the fundamental difference between Traditional AI and Agentic AI lies in the &lt;strong&gt;execution model&lt;/strong&gt;. Traditional AI is one-shot request-response, focused on accuracy, latency, and throughput; Agentic AI is a continuous loop, where &lt;strong&gt;a single error is no longer just a wrong answer but a wrong decision that may change every subsequent action of the Agent.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/execution-model-comparison-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/execution-model-comparison-en.svg" alt="Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop" data-caption="Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop"
width="859"
height="290"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Traditional AI’s request-response vs Agentic AI’s perceive→think→act→observe loop&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The comparison table below states this most clearly. I suggest you focus on the last row, &amp;ldquo;failure impact&amp;rdquo;, because that is the thesis of the whole paper:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Traditional AI&lt;/th&gt;
&lt;th&gt;Agentic AI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Execution model&lt;/td&gt;
&lt;td&gt;Request-response&lt;/td&gt;
&lt;td&gt;Continuous loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State management&lt;/td&gt;
&lt;td&gt;Stateless&lt;/td&gt;
&lt;td&gt;Stateful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger&lt;/td&gt;
&lt;td&gt;Human-triggered&lt;/td&gt;
&lt;td&gt;Self-initiated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time span&lt;/td&gt;
&lt;td&gt;Single interaction&lt;/td&gt;
&lt;td&gt;Long-running session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure impact&lt;/td&gt;
&lt;td&gt;Single wrong output&lt;/td&gt;
&lt;td&gt;Cascading behavior changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource usage&lt;/td&gt;
&lt;td&gt;Bursty&lt;/td&gt;
&lt;td&gt;Continuous&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Execution model: Traditional AI vs Agentic AI
&lt;/figcaption&gt;
&lt;p&gt;The point of this table is not the technical details but the shift in mental model: when the impact of failure escalates from &amp;ldquo;one wrong answer&amp;rdquo; to &amp;ldquo;behavior-level cascading errors&amp;rdquo;, the weight of reliability must be reassigned from the very bottom of the architecture. &lt;strong&gt;Engineering fault tolerance for a stateless API and for a self-directing, continuously running state machine are simply not the same problem.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-core-contribution-a-five-dimension-agent-reliability-framework"&gt;The core contribution: a five-dimension Agent reliability framework&lt;/h2&gt;
&lt;p&gt;This is the most memorable part of the paper. The author decomposes Agent reliability into five dimensions, forming a complete evaluation coordinate system. I drew it out: &amp;ldquo;Agent Reliability&amp;rdquo; sits in the center, with the five dimensions fanning out like petals:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/five-dim-framework-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/five-dim-framework-en.svg" alt="Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it" data-caption="Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it"
width="838"
height="525"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The five-dimension Agent reliability framework: the first four describe the Agent’s own behavior, the fifth describes the platform that runs it&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One by one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Functional Reliability&lt;/strong&gt;: can the Agent complete the task? Concerns correctness, accuracy, consistency, completeness. Plainly: &lt;em&gt;can this Agent get the job done?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Temporal Reliability&lt;/strong&gt;: can the Agent stay stable over time? Concerns timeliness, responsiveness, stability, durability. Plainly: &lt;em&gt;can it keep getting the job done?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environmental Reliability&lt;/strong&gt;: is it still effective when the environment changes? Concerns adaptability, robustness, portability. Plainly: &lt;em&gt;can it still work in a different environment?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Social Reliability&lt;/strong&gt;: this is an Agent-unique dimension. Concerns safety, trustworthiness, collaborativity, explainability. Plainly: &lt;em&gt;can it collaborate safely with humans and other Agents?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systemic Reliability&lt;/strong&gt;: the infrastructure layer. Concerns availability, fault tolerance, scalability, security. Plainly: &lt;em&gt;is the platform underneath the Agent solid?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
Why this framework matters
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
The real value of these five dimensions is that they turn &amp;ldquo;Agent reliability&amp;rdquo; from a vague concept into &lt;strong&gt;a decomposable, measurable, separately-owned engineering problem.&lt;/strong&gt; The first four describe behavioral attributes of the Agent itself; the fifth describes the platform that carries it, &lt;strong&gt;and that fifth dimension is exactly the landing point those of us doing AI Infra / Kubernetes should pick up.&lt;/strong&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="treat-the-agent-as-a-stateful-service"&gt;Treat the Agent as a &amp;ldquo;stateful service&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;The paper&amp;rsquo;s second core judgment is one I strongly agree with: &lt;strong&gt;the biggest infrastructure challenge for an Agent is not inference, it is state.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;An Agent simultaneously holds Memory, Goal, Plan, Context, and tool state. Treating it as &lt;code&gt;HTTP API + Model&lt;/code&gt; to operate is the root cause of why most demos today cannot survive production. Its true shape is closer to a composite of database + workflow engine + LLM + distributed system:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/stateful-runtime-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/stateful-runtime-en.svg" alt="Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system" data-caption="Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system"
width="892"
height="271"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Wrong ops mindset (left) vs right ops mindset (right): an Agent is fundamentally a stateful distributed system&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is entirely consistent with where the industry is heading: LangGraph, Temporal, and the OpenAI Responses API are all promoting &amp;ldquo;stateful, long-running logic&amp;rdquo; to a first-class citizen. &lt;strong&gt;When an Agent runs for hours or even days, its state is more critical than the model parameters themselves, and far easier to lose for good on a restart.&lt;/strong&gt; A model can be reloaded, but a three-hour planning context, once lost, is usually lost.&lt;/p&gt;
&lt;h2 id="fault-tolerance-the-golden-trio-of-agent-systems"&gt;Fault tolerance: the &amp;ldquo;golden trio&amp;rdquo; of Agent systems&lt;/h2&gt;
&lt;p&gt;The paper provides a catalog of Agent fault-tolerance patterns:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Redundancy&lt;/td&gt;
&lt;td&gt;Standby Agent takes over at any time&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpointing&lt;/td&gt;
&lt;td&gt;Save state for recovery&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heartbeat&lt;/td&gt;
&lt;td&gt;Liveness detection&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Circuit Breaker&lt;/td&gt;
&lt;td&gt;Isolate abnormal Agents, prevent cascading&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rollback&lt;/td&gt;
&lt;td&gt;Revert a wrong decision&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quorum&lt;/td&gt;
&lt;td&gt;Multiple Agents vote to reach consensus&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-Healing&lt;/td&gt;
&lt;td&gt;Automatically detect and correct errors&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: Agent fault-tolerance pattern catalog: mechanism, effect, and implementation complexity
&lt;/figcaption&gt;
&lt;p&gt;The author values the &lt;strong&gt;Checkpointing + Redundancy + Heartbeat&lt;/strong&gt; trio most, calling it the &amp;ldquo;golden trio&amp;rdquo; of Agent systems. I drew a closed loop to show how they relate; none of the three can be missing, and drop any one link and the recovery chain is broken:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/golden-trio-loop-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/golden-trio-loop-en.svg" alt="Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance" data-caption="Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance"
width="838"
height="429"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: The Agent fault-tolerance golden trio: heartbeat detects the issue → checkpointing saves the scene → redundancy switches the instance&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
The closed-loop logic of the golden trio
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
Heartbeat &lt;strong&gt;detects problems fast&lt;/strong&gt;, checkpointing &lt;strong&gt;saves recoverable state&lt;/strong&gt;, and redundancy &lt;strong&gt;takes over seamlessly on failure.&lt;/strong&gt; Together they form the closed loop of &amp;ldquo;detect the problem → save the scene → switch the instance&amp;rdquo;. This is also why these mechanisms all sound like old friends from distributed systems: because Agent reliability is, by nature, a distributed-systems problem.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="recovery-porting-traditional-dr-thinking-onto-the-agent"&gt;Recovery: porting traditional DR thinking onto the Agent&lt;/h2&gt;
&lt;p&gt;This section essentially ports the RTO/RPO thinking of traditional DR (disaster recovery) onto the Agent scenario. The paper compares several recovery approaches:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Recovery approach&lt;/th&gt;
&lt;th&gt;RTO&lt;/th&gt;
&lt;th&gt;RPO&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cold Start&lt;/td&gt;
&lt;td&gt;minutes to hours&lt;/td&gt;
&lt;td&gt;High (total loss)&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm Start&lt;/td&gt;
&lt;td&gt;seconds to minutes&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hot Start&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;td&gt;Low (minimal loss)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpoint Recovery&lt;/td&gt;
&lt;td&gt;seconds to minutes&lt;/td&gt;
&lt;td&gt;Medium (since last checkpoint)&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State Reconstruction&lt;/td&gt;
&lt;td&gt;minutes to hours&lt;/td&gt;
&lt;td&gt;Low (fully recoverable)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Agent recovery approaches compared: RTO, RPO, and implementation complexity
&lt;/figcaption&gt;
&lt;p&gt;The author&amp;rsquo;s conclusion: &lt;strong&gt;the Agent fits Checkpoint Recovery best.&lt;/strong&gt; The reason was foreshadowed above: an Agent&amp;rsquo;s state is far more important than its model parameters. Checkpoint Recovery strikes the most balanced trade-off among RTO, RPO, and implementation complexity. That is also why checkpointing sits at the center of the &amp;ldquo;golden trio&amp;rdquo; earlier.&lt;/p&gt;
&lt;h2 id="agent-observability--infrastructure-metrics--behavior-analysis"&gt;Agent observability = infrastructure metrics + behavior analysis&lt;/h2&gt;
&lt;p&gt;This section resonates with me the most. The paper argues that traditional monitoring is far from enough: traditional monitoring watches CPU, memory, latency, and QPS, while an Agent must additionally monitor the behavior itself.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agent Health&lt;/strong&gt;: heartbeat, resource utilization, error rate, decision latency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Behavioral&lt;/strong&gt;: action frequency, decision patterns, state transitions, goal progress.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;System&lt;/strong&gt;: Agent count, communication volume, infrastructure health.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The &amp;ldquo;behavior&amp;rdquo; layer is unique to Agents.&lt;/p&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
The easiest trap to fall into
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
An Agent&amp;rsquo;s CPU may be low and its latency normal, but if its &amp;ldquo;action frequency&amp;rdquo; suddenly spikes or its &amp;ldquo;decision pattern&amp;rdquo; deviates from baseline, &lt;strong&gt;that is often the truly dangerous signal.&lt;/strong&gt; In other words: &lt;strong&gt;Agent Observability = Infra Metrics + Behavior Analytics.&lt;/strong&gt; Watching only infrastructure metrics means you will perfectly miss the Agent&amp;rsquo;s &amp;ldquo;behavioral runaway&amp;rdquo;.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;This lines up completely with the thinking I discussed in &lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;GPU to Token Observability&lt;/a&gt;: in the Agent era, behavioral signals must become first-class citizens of observability, rather than staying stuck at the hardware / inference-metric layer.&lt;/p&gt;
&lt;h2 id="the-recommended-hybrid-architecture"&gt;The recommended hybrid architecture&lt;/h2&gt;
&lt;p&gt;The paper compares four architectures:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;th&gt;Scalability&lt;/th&gt;
&lt;th&gt;Fault isolation&lt;/th&gt;
&lt;th&gt;Suitability for Agents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monolithic&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layered&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microservices&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Event-driven&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 4: Four architectures compared for Agent systems
&lt;/figcaption&gt;
&lt;p&gt;The final recommendation: &lt;strong&gt;production-grade Agent systems adopt a hybrid architecture&lt;/strong&gt;, the combination of &amp;ldquo;layered + microservices + event-driven&amp;rdquo;:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/hybrid-architecture-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/hybrid-architecture-en.svg" alt="Figure 5: The hybrid architecture for production-grade Agent systems: layered &amp;#43; microservices &amp;#43; event-driven" data-caption="Figure 5: The hybrid architecture for production-grade Agent systems: layered &amp;#43; microservices &amp;#43; event-driven"
width="818"
height="443"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: The hybrid architecture for production-grade Agent systems: layered + microservices + event-driven&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Layering provides clear separation of concerns, microservices provide independent scaling and fault isolation, and event-driven provides loosely-coupled async collaboration. &lt;strong&gt;No single architecture can satisfy an Agent system&amp;rsquo;s demands for clarity, elasticity, and collaboration at once, so hybrid is the inevitable conclusion.&lt;/strong&gt; For K8s veterans this layered stack should look very familiar: it is the cloud-native layering of governance, ported verbatim onto the Agent.&lt;/p&gt;
&lt;h2 id="the-biggest-value-of-this-paper-from-the-ai-infra-perspective"&gt;The biggest value of this paper, from the AI Infra perspective&lt;/h2&gt;
&lt;p&gt;From the Kubernetes / AI Infra standpoint, what this paper really conveys is a paradigm shift: &lt;strong&gt;Agent infrastructure will move from &amp;ldquo;Serving&amp;rdquo; to &amp;ldquo;Runtime&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/serving-to-runtime-shift-en.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-infrastructure-reliability/serving-to-runtime-shift-en.svg" alt="Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward" data-caption="Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward"
width="1095"
height="121"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Paradigm shift: from model Serving to Agent Runtime, the focus shifts fully rightward&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The past cared about GPU utilization, throughput, and latency; the future cares about state recovery, long-task continuity, Agent isolation, Agent scheduling, Agent observability, and Agent security governance.&lt;/p&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
My read on the trend
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
&lt;strong&gt;The next phase of AI Infra is not better model Serving, but a more reliable Agent Runtime.&lt;/strong&gt; This is exactly the direction I have kept emphasizing in &lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt; and &lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism&lt;/a&gt;: the Agent is moving from &amp;ldquo;a class you import&amp;rdquo; to &amp;ldquo;a workload you must govern&amp;rdquo;.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="my-evaluation"&gt;My evaluation&lt;/h2&gt;
&lt;p&gt;Strengths:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proposes a fairly complete Agent reliability framework (the five-dimension model), turning a vague concept into a measurable engineering problem.&lt;/li&gt;
&lt;li&gt;Systematically ports traditional SRE thinking onto the Agent scenario, with clear mappings for fault tolerance, recovery, and monitoring.&lt;/li&gt;
&lt;li&gt;Offers strong engineering guidance on architecture selection (layered / microservices / event-driven / hybrid).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Weaknesses:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Lacks real production cases: the three cases in the paper (autonomous vehicle fleets, supply-chain optimization, customer-service bots) read more as illustrative scenarios than verifiable engineering practice.&lt;/li&gt;
&lt;li&gt;The math models (series/parallel reliability, Markov chains, MTBF/MTTR) are essentially standard reliability-engineering textbook material, not innovation.&lt;/li&gt;
&lt;li&gt;It never touches real Agent Runtime implementations like Kubernetes, Ray, Temporal, or LangGraph, so it feels light on engineering grounding.&lt;/li&gt;
&lt;li&gt;Its discussion of GPU and inference systems, the core AI Infra resource layer, is shallow, barely staying at the &amp;ldquo;compute/storage/network&amp;rdquo; level of abstraction.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
My score
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
&lt;p&gt;By &lt;strong&gt;academic novelty&lt;/strong&gt;: &lt;strong&gt;6.5 / 10&lt;/strong&gt;.
By &amp;ldquo;&lt;strong&gt;giving AI Infra practitioners a mental framework for Agent reliability&lt;/strong&gt;&amp;rdquo;: &lt;strong&gt;8 / 10&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Its biggest insight is not some technical detail but that opening line: &lt;em&gt;an Agent is not a single inference request but a long-running distributed state machine.&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The real contribution of this paper is that it &lt;strong&gt;re-categorizes&lt;/strong&gt; the reliability problem of Agentic AI. It is no longer a model problem, nor a prompt-engineering problem, but a distributed-systems problem, and specifically a distributed-systems problem that is stateful, long-running, and self-directing.&lt;/p&gt;
&lt;p&gt;For AI Infra practitioners, this means two things. First, the traditional SRE toolbox (checkpointing, redundancy, heartbeat, circuit breaker, RTO/RPO) can be ported over directly, but it must be extended to the &amp;ldquo;behavior layer&amp;rdquo;. Second, the focus of infrastructure will shift from &amp;ldquo;how to serve models faster&amp;rdquo; to &amp;ldquo;how to run Agents more reliably&amp;rdquo;, which is the Agent Runtime.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The AI platforms of the future must not only run fast, they must run stable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/agentic-runtime-realism"&gt;Agentic Runtime Realism: Insights from McKinsey Ark on 2026 Infrastructure Trends&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ark-agentic-runtime-analysis"&gt;Ark Agentic Runtime, Analyzed&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/gpu-to-token-observability"&gt;GPU to Token Observability: An Eight-Layer Observation System from Hardware to Business&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/ai-2026-infra-agentic-runtime"&gt;AI 2026: Infrastructure, Agents, and the Next Cloud-Native Shift&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>From GPU to Token: The 8-Layer Observability Stack for AI Infrastructure</title><link>https://jimmysong.io/blog/gpu-to-token-observability/</link><pubDate>Tue, 09 Jun 2026 03:31:16 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gpu-to-token-observability/</guid><description>From GPU hardware, Kubernetes scheduling, inference engines to token cost — understanding the 8-layer observability architecture for modern AI infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;GPU utilization is not the destination. Token cost is the true North Star metric for AI infrastructure.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For the past few years, one of the hottest topics in AI infrastructure has been GPU scheduling. Whether it&amp;rsquo;s Kubernetes, Volcano, Kueue, or &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;, they all fundamentally solve the same problem — how to make expensive and scarce GPUs more efficiently utilized.&lt;/p&gt;
&lt;p&gt;But as more enterprises begin running production-grade Large Language Model (LLM) services, a new phenomenon has emerged: GPU utilization is high, but users still complain about slow response times; GPU clusters are near full capacity, but business throughput hasn&amp;rsquo;t grown proportionally; VRAM and compute still have headroom, yet TTFT (Time To First Token) continues to degrade. This reveals a fundamental truth — GPU utilization alone is no longer sufficient to describe the real operational state of modern AI systems.&lt;/p&gt;
&lt;p&gt;For traditional cloud-native applications, we focus on CPU, memory, network, and disk. For AI systems, however, we also need to care about: whether GPUs are truly performing useful computation, whether NCCL (NVIDIA Collective Communications Library) communication has become a bottleneck, whether Kubernetes has correctly allocated resources, whether KV Cache is exhausted, whether Token latency meets user experience requirements, and whether the cost per Token is reasonable. In other words, the observability target for AI infrastructure has expanded from the GPU to the entire inference chain.&lt;/p&gt;
&lt;p&gt;Recently, while participating in the development of an industry standard — &amp;ldquo;Technical Capability Requirements for Computing Power Efficiency Enhancement: Heterogeneous Computing Services&amp;rdquo; organized by the China Academy of Information and Communications Technology (CAICT) — a core question came up repeatedly in discussions with experts: &lt;strong&gt;How do GPU and Token relate? How do we measure GPU output through Tokens?&lt;/strong&gt; The standard introduces the concept of &amp;ldquo;Token as a Service,&amp;rdquo; shifting the unit of measurement for computing services from traditional GPU hours to Tokens, yet a mature practical framework for building a complete observability chain from hardware to Token was still missing.&lt;/p&gt;
&lt;p&gt;Around the same time, I came across a technical article on GPU and LLM observability (&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;GPU &amp;amp; LLM Inference Observability — Layer-by-Layer Coverage&lt;/a&gt;) that proposed a constructive approach: decomposing the AI system into eight observability layers from GPU hardware to business cost, which neatly fills the observability gap between GPU and Token. This article reorganizes and interprets the original content, removes product-specific implementation details and vendor promotion, and extends the model by incorporating Kubernetes, HAMi, and modern LLM inference architectures.&lt;/p&gt;
&lt;p&gt;This article aims to answer one question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After we&amp;rsquo;ve solved GPU scheduling, what should we observe next?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="8-layer-observability-architecture-overview"&gt;8-Layer Observability Architecture Overview&lt;/h2&gt;
&lt;p&gt;The diagram below shows the eight observability layers from GPU hardware to business cost, each corresponding to different observability targets and areas of responsibility:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-to-token-observability/gpu-8-layer-observability.svg" data-img="https://assets.jimmysong.io/images/blog/gpu-to-token-observability/gpu-8-layer-observability.svg" alt="Figure 1: 8-Layer AI Infrastructure Observability Architecture" data-caption="Figure 1: 8-Layer AI Infrastructure Observability Architecture"
width="750"
height="940"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: 8-Layer AI Infrastructure Observability Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;Original image source GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Each layer collects data through the OTel Collector and routes it to metrics, tracing, and logging backends, forming a complete observability loop. Let&amp;rsquo;s examine each layer in detail.&lt;/p&gt;
&lt;h2 id="l1-gpu-hardware-layer"&gt;L1 GPU Hardware Layer&lt;/h2&gt;
&lt;p&gt;This layer focuses on the GPU itself. The core question is: &lt;strong&gt;Is the GPU running healthy?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The table below lists the key observability metrics for the GPU hardware layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Key Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Compute&lt;/td&gt;
&lt;td&gt;GPU Utilization, SM Occupancy, Tensor Core Activity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;VRAM Usage, HBM Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interconnect&lt;/td&gt;
&lt;td&gt;NVLink Throughput, PCIe Throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reliability&lt;/td&gt;
&lt;td&gt;ECC Error, XID Error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thermal&lt;/td&gt;
&lt;td&gt;Temperature, Power Draw, Throttle Reason&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: L1 GPU Hardware Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;Many teams only track GPU Utilization, but there&amp;rsquo;s a critical distinction — &lt;strong&gt;GPU Utilization ≠ GPU Efficiency&lt;/strong&gt;. Two GPUs might both show 90% Utilization: one executing Tensor Core operations, the other waiting for memory access. The actual performance can be vastly different. Therefore, SM (Streaming Multiprocessor) Occupancy and Tensor Core Activity are often more valuable than Utilization alone.&lt;/p&gt;
&lt;h2 id="l2-cuda-runtime-and-communication-layer"&gt;L2 CUDA Runtime and Communication Layer&lt;/h2&gt;
&lt;p&gt;For distributed training, GPU computation is usually not the bottleneck — communication is. The core question at this layer is: &lt;strong&gt;Is the GPU computing, or waiting for other GPUs?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The table below lists the key metrics for the CUDA runtime and communication layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NCCL&lt;/td&gt;
&lt;td&gt;AllReduce, AllGather, ReduceScatter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Communication&lt;/td&gt;
&lt;td&gt;Duration, Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernel&lt;/td&gt;
&lt;td&gt;Execution Time, P99 Duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Straggler&lt;/td&gt;
&lt;td&gt;Rank Skew&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: L2 CUDA Runtime and Communication Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;In large-scale training scenarios, a single slow node can drag down the entire training job. By monitoring NCCL communication Duration and Bandwidth, as well as Skew between Ranks, you can quickly pinpoint communication bottlenecks.&lt;/p&gt;
&lt;h2 id="l3-host--os-layer"&gt;L3 Host / OS Layer&lt;/h2&gt;
&lt;p&gt;Many GPU problems are ultimately not GPU problems. When GPU utilization is abnormal, the root cause may lie in the host&amp;rsquo;s CPU, memory, disk, or network.&lt;/p&gt;
&lt;p&gt;The table below lists the key metrics at the host level:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Utilization, IO Wait&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Usage, Swap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk&lt;/td&gt;
&lt;td&gt;Throughput, Latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;td&gt;Retransmit, Bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Process&lt;/td&gt;
&lt;td&gt;RSS, Thread Count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: L3 Host / OS Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;A common misconception: when GPU utilization is low, the first reaction is that the GPU isn&amp;rsquo;t fast enough. In reality, the GPU is often &lt;strong&gt;waiting for data&lt;/strong&gt; — a slow DataLoader, insufficient CPU, network congestion, or inadequate storage performance can all leave the GPU idle.&lt;/p&gt;
&lt;h2 id="l4-kubernetes-and-scheduling-layer"&gt;L4 Kubernetes and Scheduling Layer&lt;/h2&gt;
&lt;p&gt;This layer is the core of cloud-native AI infrastructure. The question to answer is: &lt;strong&gt;Who owns the GPU?&lt;/strong&gt; You need to know which Pod is using the GPU, which Namespace consumes the most resources, which team has the highest cost, and which model occupies the most VRAM.&lt;/p&gt;
&lt;p&gt;The table below lists the key observability dimensions at the scheduling layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Key Attributes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ownership&lt;/td&gt;
&lt;td&gt;Pod, Container&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workload&lt;/td&gt;
&lt;td&gt;Deployment, Job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Organization&lt;/td&gt;
&lt;td&gt;Namespace, Team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling&lt;/td&gt;
&lt;td&gt;GPU Sharing, GPU Partition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Topology&lt;/td&gt;
&lt;td&gt;Node, AZ, Region&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 4: L4 Kubernetes and Scheduling Layer Key Dimensions
&lt;/figcaption&gt;
&lt;p&gt;Projects like HAMi, Volcano, and Kueue primarily operate at this layer. HAMi solves core problems like GPU Sharing, GPU Partitioning, and heterogeneous GPU scheduling, while observability answers the question: &lt;strong&gt;Has scheduling actually improved resource utilization?&lt;/strong&gt; The observability data at this layer is the foundation for resource auditing and cost allocation.&lt;/p&gt;
&lt;h2 id="l5-training-runtime-layer"&gt;L5 Training Runtime Layer&lt;/h2&gt;
&lt;p&gt;The training phase requires monitoring the model&amp;rsquo;s runtime state. The table below lists the key metrics for the training runtime:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Metrics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Efficiency&lt;/td&gt;
&lt;td&gt;MFU, TFLOPS, Step Time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gradient&lt;/td&gt;
&lt;td&gt;Norm, NaN, Inf&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Loss&lt;/td&gt;
&lt;td&gt;Training Loss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;DataLoader Wait&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checkpoint&lt;/td&gt;
&lt;td&gt;Save Time, Restore Time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: L5 Training Runtime Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;The most important metric here is &lt;strong&gt;MFU (Model FLOPs Utilization)&lt;/strong&gt;. Many training jobs show high GPU utilization but only 30% MFU, meaning a significant amount of GPU time is not being converted into actual training progress. MFU is the golden metric for measuring training efficiency — it directly reflects the ratio of hardware compute converted into effective training computation.&lt;/p&gt;
&lt;h2 id="l6-inference-engine-layer"&gt;L6 Inference Engine Layer&lt;/h2&gt;
&lt;p&gt;This is the most critical layer in production environments and the one seeing the fastest growth in industry attention. The table below lists the core metrics for inference engines:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;Time To First Token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ITL&lt;/td&gt;
&lt;td&gt;Inter Token Latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Queue Wait&lt;/td&gt;
&lt;td&gt;Request queuing time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;Token throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch Size&lt;/td&gt;
&lt;td&gt;Batch processing efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV Cache Usage&lt;/td&gt;
&lt;td&gt;KV Cache utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: L6 Inference Engine Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;Many teams still focus solely on GPU Utilization, but the true capacity metric for inference systems is often &lt;strong&gt;KV Cache (Key-Value Cache) Utilization&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="why-kv-cache-matters-more-than-gpu-utilization"&gt;Why KV Cache Matters More Than GPU Utilization&lt;/h3&gt;
&lt;p&gt;In modern LLM inference systems, GPU compute usually has headroom, but KV Cache often runs out first. A typical scenario: GPU utilization is only 60%, KV Cache is at 95%, new requests start queuing, and TTFT spikes rapidly.&lt;/p&gt;
&lt;p&gt;For inference systems, KV Cache is more akin to a database&amp;rsquo;s Buffer Pool — it often determines the capacity ceiling of the entire system. When KV Cache approaches saturation, the system is forced to perform Eviction, leading to context loss, request retries, and ultimately latency spikes and throughput drops.&lt;/p&gt;
&lt;h2 id="l7-genai-api-layer"&gt;L7 GenAI API Layer&lt;/h2&gt;
&lt;p&gt;With the development of OpenTelemetry GenAI Semantic Conventions, the industry is converging on unified AI observability standards. The table below lists the key metrics at the API layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input Tokens&lt;/td&gt;
&lt;td&gt;Input token count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output Tokens&lt;/td&gt;
&lt;td&gt;Output token count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request Duration&lt;/td&gt;
&lt;td&gt;Request latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;First Token latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token Throughput&lt;/td&gt;
&lt;td&gt;Token throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 7: L7 GenAI API Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;This layer marks the shift from infrastructure-centric to application-centric observability. API-layer metrics directly face end users and business systems, serving as the core data source for SLA/SLO quality assessment.&lt;/p&gt;
&lt;h2 id="l8-business-and-cost-layer"&gt;L8 Business and Cost Layer&lt;/h2&gt;
&lt;p&gt;Ultimately, enterprises don&amp;rsquo;t care about GPU utilization — they care about: &lt;strong&gt;Is the GPU creating value?&lt;/strong&gt; The table below lists the core metrics for the business and cost layer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cost per GPU Hour&lt;/td&gt;
&lt;td&gt;GPU hourly cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per Token&lt;/td&gt;
&lt;td&gt;Token cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per Request&lt;/td&gt;
&lt;td&gt;Request cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idle GPU Cost&lt;/td&gt;
&lt;td&gt;Idle resource cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per Watt&lt;/td&gt;
&lt;td&gt;Inference efficiency&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 8: L8 Business and Cost Layer Key Metrics
&lt;/figcaption&gt;
&lt;p&gt;The core competitive metric for future AI infrastructure may not be GPU Utilization, but rather &lt;strong&gt;Cost per Useful Token&lt;/strong&gt; — the comprehensive cost of producing one useful Token. This metric unifies hardware cost, energy consumption, inference efficiency, and business value into a single measurement framework.&lt;/p&gt;
&lt;h2 id="cross-layer-troubleshooting"&gt;Cross-Layer Troubleshooting&lt;/h2&gt;
&lt;p&gt;Real production issues often span multiple layers. The table below lists common cross-layer failure symptoms and their potential causes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Possible Causes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT Increase&lt;/td&gt;
&lt;td&gt;CPU IO Wait, Queue Depth, KV Cache Pressure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput Drop&lt;/td&gt;
&lt;td&gt;NCCL, Batch Size, Network&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency Spike&lt;/td&gt;
&lt;td&gt;KV Cache Eviction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OOM&lt;/td&gt;
&lt;td&gt;Insufficient VRAM, Oversized Batch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training Stall&lt;/td&gt;
&lt;td&gt;NCCL Straggler&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 9: Cross-Layer Troubleshooting Reference
&lt;/figcaption&gt;
&lt;p&gt;Modern AI operations are evolving from single-point monitoring to cross-layer observability. The root cause of a single-layer metric anomaly often lies in another layer. Only by building cross-layer correlation capabilities can teams achieve rapid root-cause identification and precise troubleshooting.&lt;/p&gt;
&lt;h2 id="from-gpu-control-plane-to-ai-observability-plane"&gt;From GPU Control Plane to AI Observability Plane&lt;/h2&gt;
&lt;p&gt;Over the past few years, the industry has focused primarily on projects like Kubernetes Scheduler, Volcano, Kueue, and HAMi, solving the problem of &lt;strong&gt;how to allocate GPUs&lt;/strong&gt;. In the coming years, the industry will start asking &lt;strong&gt;whether GPUs are truly generating value&lt;/strong&gt;, leading to the formation of a new AI infrastructure technology stack:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Hardware
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Kubernetes Control Plane
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Inference Runtime
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Observability Plane
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Optimization Plane&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;GPU scheduling determines how resources are allocated. Observability determines whether resources are being used correctly. Together, they form the next-generation AI Infrastructure Stack. The evolution from GPU Control Plane to AI Observability Plane marks a new era where AI infrastructure transitions from &amp;ldquo;resource management&amp;rdquo; to &amp;ldquo;value management.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;AI infrastructure observability is undergoing a fundamental paradigm shift. We used to look only at GPU utilization; now we need to build a complete observability system across eight layers — from GPU hardware, CUDA runtime, host OS, Kubernetes scheduling, training runtime, inference engine, GenAI API, to business cost.&lt;/p&gt;
&lt;p&gt;Key takeaways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;L1-L3 focus on hardware and systems&lt;/strong&gt;: GPU health, CUDA communication efficiency, and host resource adequacy&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L4 focuses on resource allocation&lt;/strong&gt;: Kubernetes scheduling and GPU sharing, with tools like HAMi solving allocation problems&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L5-L6 focus on model runtime&lt;/strong&gt;: Training efficiency (MFU) and inference capacity (KV Cache) are the core metrics&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;L7-L8 focus on business value&lt;/strong&gt;: From Token throughput to cost per Token, ultimately measuring whether GPUs are creating value&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;GPU scheduling is the starting point for AI infrastructure, but not the end goal. The 8-layer observability stack from GPU to Token is the critical closed loop that ensures AI infrastructure truly delivers business value.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/last9/gpu-telemetry/blob/main/docs/GPU_LLM_OBSERVABILITY.md" target="_blank" rel="noopener"&gt;GPU &amp;amp; LLM Inference Observability — Layer-by-Layer Coverage - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>My Personal AI Stack: Building a Continuously Running Personal AI Infrastructure for ~$100/Month</title><link>https://jimmysong.io/blog/personal-ai-stack/</link><pubDate>Sun, 07 Jun 2026 16:58:18 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/personal-ai-stack/</guid><description>How I built a personal AI infrastructure using ChatGPT, OpenClaw, Obsidian, GitHub, Lark, GLM-5.1, and a Mac mini M4.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;A personal AI infrastructure is not a single tool — it&amp;rsquo;s a system of long-term synergy.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/banner.webp" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/banner.webp" alt="Figure 17: My Personal AI Stack" data-caption="Figure 17: My Personal AI Stack"
width="1774"
height="887"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 17: My Personal AI Stack&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Many people talk about AI Agents, Second Brain, Personal Knowledge Management (PKM), and digital avatars.&lt;/p&gt;
&lt;p&gt;But over the past year, I&amp;rsquo;ve come to realize that what I&amp;rsquo;m actually building is not some AI assistant — it&amp;rsquo;s a continuously running Personal AI Infrastructure.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not a single product, nor a single model. It&amp;rsquo;s a set of systems that work together over the long term.&lt;/p&gt;
&lt;p&gt;This system helps me think, research, write, code, manage knowledge, process emails, maintain my website, and accumulate long-term memory — every single day.&lt;/p&gt;
&lt;p&gt;If I had to summarize it in one sentence:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ChatGPT handles thinking, OpenClaw handles execution, Obsidian handles memory, GitHub handles publishing.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="from-tools-to-infrastructure"&gt;From Tools to Infrastructure&lt;/h2&gt;
&lt;p&gt;Most AI workflow articles follow a similar structure: start with the model, then the plugins, then the editor.&lt;/p&gt;
&lt;p&gt;But I increasingly feel that tools are not the point.&lt;/p&gt;
&lt;p&gt;What matters is how these tools work together.&lt;/p&gt;
&lt;p&gt;My work spans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI infrastructure research&lt;/li&gt;
&lt;li&gt;Open source community operations&lt;/li&gt;
&lt;li&gt;Technical writing&lt;/li&gt;
&lt;li&gt;Developer Relations&lt;/li&gt;
&lt;li&gt;Product and ecosystem building&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every day produces a massive amount of information:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ChatGPT conversations&lt;/li&gt;
&lt;li&gt;Technical research&lt;/li&gt;
&lt;li&gt;GitHub activity&lt;/li&gt;
&lt;li&gt;Community discussions&lt;/li&gt;
&lt;li&gt;Email newsletters&lt;/li&gt;
&lt;li&gt;Hacker News&lt;/li&gt;
&lt;li&gt;Discord&lt;/li&gt;
&lt;li&gt;WeChat and Lark messages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The problem is never a lack of information — it&amp;rsquo;s how to organize it.&lt;/p&gt;
&lt;p&gt;So I gradually built a Personal AI Stack around my own work.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Workflow Perspective
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
This is not a single tool, not a single model. It emphasizes continuous synergy across four layers: thinking, memory, execution, and publishing.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="overall-architecture"&gt;Overall Architecture&lt;/h2&gt;
&lt;p&gt;The architecture diagram below shows how my Personal AI Stack works in layers.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/ddbe89d6ae15e708024e83d0dc1bd038.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/ddbe89d6ae15e708024e83d0dc1bd038.svg" alt="Figure 18: Personal AI Stack Architecture Layers" data-caption="Figure 18: Personal AI Stack Architecture Layers"
width="4610"
height="2138"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 18: Personal AI Stack Architecture Layers&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The core components for each layer are listed below.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Components&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interface Layer&lt;/td&gt;
&lt;td&gt;ChatGPT, Telegram, Discord, Lark, WeChat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning Layer&lt;/td&gt;
&lt;td&gt;ChatGPT, GLM-5.1, Claude Code, Codex&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory Layer&lt;/td&gt;
&lt;td&gt;Obsidian, Markdown, iCloud&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Execution Layer&lt;/td&gt;
&lt;td&gt;OpenClaw, Gmail, Calendar, Lark CLI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Publishing Layer&lt;/td&gt;
&lt;td&gt;GitHub, Hugo, Cloudflare Pages&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 5: Personal AI Stack Layers and Components
&lt;/figcaption&gt;
&lt;h2 id="chatgpt-my-thinking-system"&gt;ChatGPT: My Thinking System&lt;/h2&gt;
&lt;p&gt;Although many workflows revolve around Agents, the tool I use most frequently is actually ChatGPT.&lt;/p&gt;
&lt;p&gt;I primarily use it for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deep Research&lt;/li&gt;
&lt;li&gt;Technical analysis&lt;/li&gt;
&lt;li&gt;Architecture discussions&lt;/li&gt;
&lt;li&gt;Content planning&lt;/li&gt;
&lt;li&gt;Writing assistance&lt;/li&gt;
&lt;li&gt;Career decisions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Rather than calling it an assistant, it&amp;rsquo;s more like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Research Partner&lt;/li&gt;
&lt;li&gt;Technical Advisor&lt;/li&gt;
&lt;li&gt;Thinking Companion&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Years of accumulated conversations have helped it gradually understand my background, projects, and long-term goals.&lt;/p&gt;
&lt;p&gt;Many articles, talks, and technical judgments actually originate from these ongoing conversations.&lt;/p&gt;
&lt;h2 id="openclaw-my-execution-system"&gt;OpenClaw: My Execution System&lt;/h2&gt;
&lt;p&gt;OpenClaw is the OpenClaw Agent I deployed on my Mac mini M4 at home.&lt;/p&gt;
&lt;p&gt;I mainly interact with OpenClaw through Telegram.&lt;/p&gt;
&lt;p&gt;To avoid context mixing, I use separate Telegram groups for different topics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;HAMi&lt;/li&gt;
&lt;li&gt;AI Handbook&lt;/li&gt;
&lt;li&gt;Personal&lt;/li&gt;
&lt;li&gt;Work&lt;/li&gt;
&lt;li&gt;Research&lt;/li&gt;
&lt;li&gt;Blog&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This naturally creates Context Isolation.&lt;/p&gt;
&lt;p&gt;OpenClaw handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Gmail email management&lt;/li&gt;
&lt;li&gt;Apple Calendar scheduling&lt;/li&gt;
&lt;li&gt;Scheduled tasks&lt;/li&gt;
&lt;li&gt;Obsidian operations&lt;/li&gt;
&lt;li&gt;GitHub operations&lt;/li&gt;
&lt;li&gt;Website maintenance&lt;/li&gt;
&lt;li&gt;Lark knowledge base operations&lt;/li&gt;
&lt;li&gt;Automated workflows&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It&amp;rsquo;s more like a Chief of Staff than a chatbot.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
OpenClaw&amp;rsquo;s Positioning
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
OpenClaw&amp;rsquo;s value lies in unifying scattered execution entry points into structured workflows, rather than being a simple chat interface.
&lt;/div&gt;
&lt;/div&gt;
&lt;h3 id="openclaws-workflow"&gt;OpenClaw&amp;rsquo;s Workflow&lt;/h3&gt;
&lt;p&gt;OpenClaw&amp;rsquo;s main path unfolds through the Telegram entry point for multi-platform execution.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/c89afbc470d0c515198eb344d8dc1e4e.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/c89afbc470d0c515198eb344d8dc1e4e.svg" alt="Figure 19: OpenClaw Execution Path" data-caption="Figure 19: OpenClaw Execution Path"
width="3119"
height="1058"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 19: OpenClaw Execution Path&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="lark-company-workflow-entry-point"&gt;Lark: Company Workflow Entry Point&lt;/h2&gt;
&lt;p&gt;I had barely used Lark before.&lt;/p&gt;
&lt;p&gt;After joining my current company, which heavily promotes Lark adoption, all daily workflows, knowledge bases, and collaborative communication happen in Lark.&lt;/p&gt;
&lt;p&gt;At first, I just treated it as an enterprise IM tool.&lt;/p&gt;
&lt;p&gt;But looking at it now, Lark is more of a company-level workflow entry point:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Group chats&lt;/li&gt;
&lt;li&gt;Documents&lt;/li&gt;
&lt;li&gt;Knowledge bases&lt;/li&gt;
&lt;li&gt;Approvals&lt;/li&gt;
&lt;li&gt;Tasks&lt;/li&gt;
&lt;li&gt;Automation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For me, what truly changed the experience was Lark CLI.&lt;/p&gt;
&lt;p&gt;Through Lark CLI, I can more easily manage the company&amp;rsquo;s knowledge bases, documents, and some process-driven information.&lt;/p&gt;
&lt;p&gt;This turns Lark from merely a chat tool into something that OpenClaw can incorporate into its automation system.&lt;/p&gt;
&lt;p&gt;From the perspective of a Personal AI Stack, Lark serves as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The organizational workflow layer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It connects information, knowledge, and tasks in a company context.&lt;/p&gt;
&lt;h2 id="obsidian-working-memory"&gt;Obsidian: Working Memory&lt;/h2&gt;
&lt;p&gt;Obsidian is my most frequently used knowledge tool.&lt;/p&gt;
&lt;p&gt;But I don&amp;rsquo;t consider it my final knowledge base.&lt;/p&gt;
&lt;p&gt;For me:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Obsidian
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Working Memory&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This is where I store:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Daily Notes&lt;/li&gt;
&lt;li&gt;Weekly Reports&lt;/li&gt;
&lt;li&gt;Research Notes&lt;/li&gt;
&lt;li&gt;Inbox&lt;/li&gt;
&lt;li&gt;Drafts&lt;/li&gt;
&lt;li&gt;Fleeting thoughts&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Starting three years ago, I developed a habit of consistently writing weekly reports.&lt;/p&gt;
&lt;p&gt;These reports record:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Work progress&lt;/li&gt;
&lt;li&gt;Learning content&lt;/li&gt;
&lt;li&gt;Community activities&lt;/li&gt;
&lt;li&gt;Project evolution&lt;/li&gt;
&lt;li&gt;Personal reflections&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In a sense, weekly reports constitute my daily memory.&lt;/p&gt;
&lt;p&gt;And Obsidian is the carrier of these memories.&lt;/p&gt;
&lt;h2 id="github-and-jimmysongio-long-term-memory"&gt;GitHub and jimmysong.io: Long-term Memory&lt;/h2&gt;
&lt;p&gt;Many people think of a blog as a content publishing platform.&lt;/p&gt;
&lt;p&gt;But for me, jimmysong.io is closer to a long-term memory system.&lt;/p&gt;
&lt;p&gt;This website has been continuously maintained for nearly ten years.&lt;/p&gt;
&lt;p&gt;What&amp;rsquo;s recorded here is not just technical articles, but more importantly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;My viewpoints&lt;/li&gt;
&lt;li&gt;My judgments&lt;/li&gt;
&lt;li&gt;My experiences&lt;/li&gt;
&lt;li&gt;My growth trajectory&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unlike Obsidian, content that makes it to the website typically goes through:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Research
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Thinking
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Validation
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Writing
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Revision
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Publishing&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Therefore:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Obsidian
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Working Memory
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;jimmysong.io
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;=
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Long-term Memory&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The following flow shows the basic path from Obsidian to website publishing.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/e43d18197c002d3343a02ba47e100ece.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/e43d18197c002d3343a02ba47e100ece.svg" alt="Figure 20: Long-term Memory Publishing Flow" data-caption="Figure 20: Long-term Memory Publishing Flow"
width="2627"
height="172"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 20: Long-term Memory Publishing Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="information-flow"&gt;Information Flow&lt;/h2&gt;
&lt;p&gt;There are clear boundaries between content sources, curation methods, and output channels. This diagram reveals my information flow logic.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/769ffd8bfc324f83203fa12b51796292.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/769ffd8bfc324f83203fa12b51796292.svg" alt="Figure 21: Personal AI Stack Information Flow" data-caption="Figure 21: Personal AI Stack Information Flow"
width="2518"
height="1349"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 21: Personal AI Stack Information Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="github-execution-and-publishing-layer"&gt;GitHub: Execution and Publishing Layer&lt;/h2&gt;
&lt;p&gt;For me, GitHub is no longer just a code repository.&lt;/p&gt;
&lt;p&gt;It also handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Blog content&lt;/li&gt;
&lt;li&gt;Website source code&lt;/li&gt;
&lt;li&gt;Documentation system&lt;/li&gt;
&lt;li&gt;AI Handbook&lt;/li&gt;
&lt;li&gt;AI Native Landscape&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All content eventually makes its way into GitHub.&lt;/p&gt;
&lt;p&gt;GitHub Actions automatically handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Building&lt;/li&gt;
&lt;li&gt;Testing&lt;/li&gt;
&lt;li&gt;Publishing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Finally, everything is served through Cloudflare Pages.&lt;/p&gt;
&lt;p&gt;The diagram below shows the publishing path from Markdown to static site.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/e6cc6be64df6f178466a37b70068a3b5.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/e6cc6be64df6f178466a37b70068a3b5.svg" alt="Figure 22: GitHub Publishing Pipeline" data-caption="Figure 22: GitHub Publishing Pipeline"
width="2352"
height="141"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 22: GitHub Publishing Pipeline&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="my-development-toolchain"&gt;My Development Toolchain&lt;/h2&gt;
&lt;p&gt;I currently use three main AI development tools.&lt;/p&gt;
&lt;h3 id="claude-code"&gt;Claude Code&lt;/h3&gt;
&lt;p&gt;My daily development workhorse.&lt;/p&gt;
&lt;p&gt;Despite the name Claude Code, I primarily use Zhipu&amp;rsquo;s GLM-5.1 model.&lt;/p&gt;
&lt;p&gt;It handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Coding&lt;/li&gt;
&lt;li&gt;Refactoring&lt;/li&gt;
&lt;li&gt;Debugging&lt;/li&gt;
&lt;li&gt;Documentation maintenance&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="codex"&gt;Codex&lt;/h3&gt;
&lt;p&gt;Mainly used for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Project initialization&lt;/li&gt;
&lt;li&gt;Large-scale code generation&lt;/li&gt;
&lt;li&gt;Automated execution of complex tasks&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="openclaw"&gt;OpenClaw&lt;/h3&gt;
&lt;p&gt;Handles automation work beyond development.&lt;/p&gt;
&lt;p&gt;The three form a clear division of labor.&lt;/p&gt;
&lt;p&gt;The diagram below shows how my toolchain collaborates.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/1d9a2ef112b60d5c6d3014e75e449be3.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/1d9a2ef112b60d5c6d3014e75e449be3.svg" alt="Figure 23: Development Toolchain Collaboration" data-caption="Figure 23: Development Toolchain Collaboration"
width="654"
height="1342"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 23: Development Toolchain Collaboration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-not-claude"&gt;Why Not Claude?&lt;/h2&gt;
&lt;p&gt;This is one of the questions I get asked most often.&lt;/p&gt;
&lt;p&gt;Objectively speaking, Claude is indeed very strong in code generation and code understanding.&lt;/p&gt;
&lt;p&gt;I seriously considered using Claude as my primary model.&lt;/p&gt;
&lt;p&gt;But ultimately, I chose not to.&lt;/p&gt;
&lt;p&gt;The reason is not about model capability — it&amp;rsquo;s about overall return on investment.&lt;/p&gt;
&lt;p&gt;First, there&amp;rsquo;s the account issue.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve registered Claude accounts multiple times in the past, and each time I ran into restrictions and bans.&lt;/p&gt;
&lt;p&gt;Second, there&amp;rsquo;s the cost issue.&lt;/p&gt;
&lt;p&gt;Because I pay for all my AI tools out of pocket.&lt;/p&gt;
&lt;p&gt;So I care more about:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Capability / Cost&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Rather than:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Absolute Capability&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For enterprise users, Claude Max might be a very reasonable choice.&lt;/p&gt;
&lt;p&gt;But for individual paying users, the conclusion may be different.&lt;/p&gt;
&lt;p&gt;This article discusses a Personal AI Stack that is entirely self-funded.&lt;/p&gt;
&lt;p&gt;If the company reimburses expenses, or if you have an enterprise budget, many choices would change.&lt;/p&gt;
&lt;h2 id="why-not-self-host-large-models"&gt;Why Not Self-host Large Models?&lt;/h2&gt;
&lt;p&gt;This is another frequently asked question.&lt;/p&gt;
&lt;p&gt;Many people believe:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Buy a GPU + Open-source Model = Free AI&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In reality, this is often not the case.&lt;/p&gt;
&lt;p&gt;If the goal is simply to get a stable, powerful AI assistant, I lean toward:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Subscription &amp;gt; Self-hosting&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The reasons include:&lt;/p&gt;
&lt;h3 id="time-cost"&gt;Time Cost&lt;/h3&gt;
&lt;p&gt;Maintaining a model is work in itself.&lt;/p&gt;
&lt;p&gt;Including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CUDA&lt;/li&gt;
&lt;li&gt;Drivers&lt;/li&gt;
&lt;li&gt;Inference frameworks&lt;/li&gt;
&lt;li&gt;Model upgrades&lt;/li&gt;
&lt;li&gt;Networking issues&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All of these require time.&lt;/p&gt;
&lt;p&gt;And I&amp;rsquo;d rather spend my time on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Writing&lt;/li&gt;
&lt;li&gt;Community&lt;/li&gt;
&lt;li&gt;Product&lt;/li&gt;
&lt;li&gt;Technical research&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="cost-issue"&gt;Cost Issue&lt;/h3&gt;
&lt;p&gt;Currently:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ChatGPT Plus&lt;/li&gt;
&lt;li&gt;GLM Coding Plan Max&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total monthly cost is under 600 CNY.&lt;/p&gt;
&lt;p&gt;While a high-end GPU:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;RTX 5090D&lt;/li&gt;
&lt;li&gt;RTX PRO&lt;/li&gt;
&lt;li&gt;Enterprise GPU&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Typically costs tens of thousands of CNY.&lt;/p&gt;
&lt;p&gt;Plus:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity&lt;/li&gt;
&lt;li&gt;Depreciation&lt;/li&gt;
&lt;li&gt;Maintenance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For my use case, it&amp;rsquo;s simply not worth it.&lt;/p&gt;
&lt;h3 id="model-upgrade-speed"&gt;Model Upgrade Speed&lt;/h3&gt;
&lt;p&gt;Cloud models upgrade every month.&lt;/p&gt;
&lt;p&gt;Local models require manual follow-up.&lt;/p&gt;
&lt;p&gt;For knowledge workers:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Using the latest model is more important than owning a model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Subscription vs. Self-hosting
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
For individual users, subscription services typically offer advantages in time cost, upgrade speed, and stability. They let you focus on results rather than infrastructure maintenance.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="daily-tools"&gt;Daily Tools&lt;/h2&gt;
&lt;p&gt;Beyond the core systems described above, I also rely on some daily tools to complete the workflow.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Atlas Browser&lt;/strong&gt;: My primary browser for web reading, research, and information gathering. Noteworthy content is saved to my knowledge base via Obsidian Clipper.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Warp&lt;/strong&gt;: My most frequently used terminal tool. Its modern interactive experience and AI capabilities make command-line work more efficient.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Typora&lt;/strong&gt;: A Markdown editor I&amp;rsquo;ve used for a long time, ideal for immersive writing and long-form editing. Many blog posts and documents are completed here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CodexBar&lt;/strong&gt;: Used to monitor usage of ChatGPT, Codex, Claude Code, and other tools. For heavy AI users, token consumption has become a resource metric worth tracking.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sogou Input Method&lt;/strong&gt;: My primary voice input tool. Compared to keyboard input, voice better matches my thinking habits, especially when working remotely, writing, and communicating with AI.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These tools are not the core of the system themselves, but they form the foundational experience layer of the entire Personal AI Stack, making information acquisition, content creation, and daily development smoother.&lt;/p&gt;
&lt;h2 id="my-ai-usage-scale"&gt;My AI Usage Scale&lt;/h2&gt;
&lt;p&gt;Current approximate consumption:&lt;/p&gt;
&lt;h3 id="glm-51"&gt;GLM-5.1&lt;/h3&gt;
&lt;p&gt;Weekly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;600 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Monthly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2.4 billion Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="codex-1"&gt;Codex&lt;/h3&gt;
&lt;p&gt;Weekly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;200 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Monthly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;800 million Tokens&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Approximately 3.2 billion Tokens per month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That&amp;rsquo;s roughly 100 million Tokens consumed per day. Through ChatGPT Plus and GLM Coding Plan subscriptions, this is much more cost-effective than paying per token — otherwise, these tokens would cost at least $500 per month.&lt;/p&gt;
&lt;h2 id="how-much-does-this-system-cost-per-month"&gt;How Much Does This System Cost Per Month?&lt;/h2&gt;
&lt;p&gt;The table below compares the fixed monthly costs.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th style="text-align: right"&gt;Monthly Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT Plus&lt;/td&gt;
&lt;td style="text-align: right"&gt;$19.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM Coding Plan Max&lt;/td&gt;
&lt;td style="text-align: right"&gt;¥422.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iCloud+&lt;/td&gt;
&lt;td style="text-align: right"&gt;$0.99&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 6: Monthly Cost Breakdown
&lt;/figcaption&gt;
&lt;p&gt;Approximately:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥573 / month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;OpenClaw runs on my Mac mini M4.&lt;/p&gt;
&lt;p&gt;Hardware includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mac mini M4&lt;/li&gt;
&lt;li&gt;Samsung 990 Pro 1TB&lt;/li&gt;
&lt;li&gt;HAGIBIS Dock&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Total investment:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥4,469&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Amortized over four years:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Approximately ¥100 / month&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="total-cost"&gt;Total Cost&lt;/h3&gt;
&lt;p&gt;The entire Personal AI Stack&amp;rsquo;s fixed cost is approximately:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;¥700 CNY / month (~$100)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This cost pie chart shows the monthly spending structure of the Personal AI Stack.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/personal-ai-stack/b67330204f6917024f61be43bba7004c.svg" data-img="https://assets.jimmysong.io/images/blog/personal-ai-stack/b67330204f6917024f61be43bba7004c.svg" alt="Figure 24: Personal AI Stack Monthly Cost Breakdown" data-caption="Figure 24: Personal AI Stack Monthly Cost Breakdown"
width="618"
height="384"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 24: Personal AI Stack Monthly Cost Breakdown&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="faq"&gt;FAQ&lt;/h2&gt;
&lt;h3 id="is-this-the-most-powerful-setup"&gt;Is This the Most Powerful Setup?&lt;/h3&gt;
&lt;p&gt;No.&lt;/p&gt;
&lt;p&gt;This is not a &amp;ldquo;most powerful AI tool configuration guide.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This is a completely self-funded Personal AI Stack designed for long-term individual use.&lt;/p&gt;
&lt;p&gt;If your company reimburses expenses, or if you have a higher budget, you could choose Claude Max, Cursor, more API services, or even a local GPU workstation.&lt;/p&gt;
&lt;p&gt;But my goal is not to pursue the absolute best — it&amp;rsquo;s to achieve stable, sustainable, and cumulative productivity within a personal budget.&lt;/p&gt;
&lt;h3 id="why-not-sync-all-chatgpt-conversations-to-obsidian"&gt;Why Not Sync All ChatGPT Conversations to Obsidian?&lt;/h3&gt;
&lt;p&gt;Because I don&amp;rsquo;t want to turn Obsidian into a chat log repository.&lt;/p&gt;
&lt;p&gt;What I care about more is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Which content is worth preserving long-term?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Manual saving is itself a curation process.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s more important than automatically syncing everything.&lt;/p&gt;
&lt;h3 id="why-not-use-notion"&gt;Why Not Use Notion?&lt;/h3&gt;
&lt;p&gt;It&amp;rsquo;s not because Notion isn&amp;rsquo;t good.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s because I prefer Markdown First.&lt;/p&gt;
&lt;p&gt;The benefits of Markdown include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Local-first&lt;/li&gt;
&lt;li&gt;Version controllable&lt;/li&gt;
&lt;li&gt;Portable&lt;/li&gt;
&lt;li&gt;AI-friendly&lt;/li&gt;
&lt;li&gt;Suitable for long-term preservation&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="why-use-telegram-only-for-openclaw"&gt;Why Use Telegram Only for OpenClaw?&lt;/h3&gt;
&lt;p&gt;Because Telegram&amp;rsquo;s Bot ecosystem and API are better suited as an Agent entry point.&lt;/p&gt;
&lt;p&gt;WeChat, Lark, and Discord are more for person-to-person communication and community collaboration for me.&lt;/p&gt;
&lt;h3 id="why-emphasize-lark"&gt;Why Emphasize Lark?&lt;/h3&gt;
&lt;p&gt;Because after joining my current company, I truly started using Lark.&lt;/p&gt;
&lt;p&gt;The company&amp;rsquo;s entire workflow revolves around Lark.&lt;/p&gt;
&lt;p&gt;For me, Lark isn&amp;rsquo;t a personal knowledge base — it&amp;rsquo;s a company knowledge and collaboration system.&lt;/p&gt;
&lt;p&gt;Through Lark CLI, it can further become a workflow entry point that OpenClaw can operate.&lt;/p&gt;
&lt;h3 id="why-not-use-a-local-ai-workstation"&gt;Why Not Use a Local AI Workstation?&lt;/h3&gt;
&lt;p&gt;Because my main work is not training models or running inference services.&lt;/p&gt;
&lt;p&gt;My main work is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Thinking&lt;/li&gt;
&lt;li&gt;Research&lt;/li&gt;
&lt;li&gt;Writing&lt;/li&gt;
&lt;li&gt;Open source community&lt;/li&gt;
&lt;li&gt;Software development&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Subscription models + a Mac mini is more than enough.&lt;/p&gt;
&lt;p&gt;A local AI workstation requires significant investment, complex maintenance, and the model upgrade speed may not keep up with cloud services.&lt;/p&gt;
&lt;h3 id="why-not-buy-a-high-end-nvidia-gpu"&gt;Why Not Buy a High-end NVIDIA GPU?&lt;/h3&gt;
&lt;p&gt;If the primary purpose is learning CUDA, GPU scheduling, or AI infrastructure, buying a GPU has value.&lt;/p&gt;
&lt;p&gt;But if the primary purpose is daily productivity, a high-end GPU isn&amp;rsquo;t necessarily cost-effective.&lt;/p&gt;
&lt;p&gt;For me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Subscribing to models is more important than owning a GPU.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="workflow-matters-more-than-models"&gt;Workflow Matters More Than Models&lt;/h2&gt;
&lt;p&gt;My biggest takeaway from the past year is:&lt;/p&gt;
&lt;p&gt;Many people believe the core of AI is the model.&lt;/p&gt;
&lt;p&gt;But my experience suggests the opposite.&lt;/p&gt;
&lt;p&gt;What truly affects productivity is often not the model leaderboard, but workflow design.&lt;/p&gt;
&lt;p&gt;A 95-score model in an excellent workflow is usually more valuable than a 100-score model in a chaotic workflow.&lt;/p&gt;
&lt;p&gt;In a sense, this is very similar to the evolution of the cloud-native world.&lt;/p&gt;
&lt;p&gt;GPUs are important, but scheduling systems are equally important.&lt;/p&gt;
&lt;p&gt;Models are important, but workflows are equally important.&lt;/p&gt;
&lt;p&gt;Agents are important, but long-term memory and execution systems are equally important.&lt;/p&gt;
&lt;p&gt;For me, the ultimate goal of the Personal AI Stack is not to replace people.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s to connect thinking, memory, and execution — freeing up more time for what truly matters.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This article defines my Personal AI Stack as a long-term runnable infrastructure, rather than a single tool or a single model.&lt;/p&gt;
&lt;p&gt;The core conclusion is: in a self-funded scenario, stable workflows, continuous memory systems, and executable automation are more productive than chasing the most powerful model.&lt;/p&gt;
&lt;p&gt;If your goal is to accumulate long-term value, investing time in &amp;ldquo;how things work together&amp;rdquo; often yields better returns than investing time in &amp;ldquo;model rankings.&amp;rdquo;&lt;/p&gt;</content:encoded></item><item><title>Token Is More Than a Billing Unit, It's Becoming the Resource Unit of the AI Era</title><link>https://jimmysong.io/blog/tokenomics-foundation-new-resource-unit/</link><pubDate>Thu, 04 Jun 2026 06:42:40 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/tokenomics-foundation-new-resource-unit/</guid><description>The Linux Foundation&amp;#39;s Tokenomics Foundation signals a shift: tokens are becoming a core resource in the AI era, much like CPUs in the cloud era.</description><content:encoded>
&lt;p&gt;The &lt;a href="https://www.linuxfoundation.org/" target="_blank" rel="noopener"&gt;Linux Foundation&lt;/a&gt; recently announced plans to establish the &lt;a href="https://www.tokeneconomics.com/insights/launch-tokenomics-foundation/" target="_blank" rel="noopener"&gt;Tokenomics Foundation&lt;/a&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/tokenomics-foundation-new-resource-unit/banner.webp" data-img="https://assets.jimmysong.io/images/blog/tokenomics-foundation-new-resource-unit/banner.webp" alt="Figure 1: Tokenomics Foundation" data-caption="Figure 1: Tokenomics Foundation"
width="1774"
height="887"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Tokenomics Foundation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;When I first saw the news, my immediate reaction wasn&amp;rsquo;t:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Yet another foundation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI infrastructure is shifting from &amp;ldquo;managing GPUs&amp;rdquo; to &amp;ldquo;managing Tokens.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That shift may matter more than the foundation itself.&lt;/p&gt;
&lt;p&gt;The official rationale from the Linux Foundation is straightforward:&lt;/p&gt;
&lt;p&gt;As enterprises begin deploying generative AI and agents at scale, the Token has become the new unit of technology spend. The foundation will partner with the &lt;a href="https://www.finops.org/" target="_blank" rel="noopener"&gt;FinOps Foundation&lt;/a&gt; to establish Token cost management, benchmarks, open standards, and best practices.&lt;/p&gt;
&lt;p&gt;If you stop there, most people would reduce it to:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;FinOps for AI&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Or:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI cost management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;But I think there&amp;rsquo;s more to it than that.&lt;/p&gt;
&lt;h2 id="the-cloud-era-managed-resources"&gt;The Cloud Era Managed Resources&lt;/h2&gt;
&lt;p&gt;For the past two decades, the infrastructure industry has been managing resources.&lt;/p&gt;
&lt;p&gt;We discussed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CPU&lt;/li&gt;
&lt;li&gt;Memory&lt;/li&gt;
&lt;li&gt;Storage&lt;/li&gt;
&lt;li&gt;Network&lt;/li&gt;
&lt;li&gt;GPU&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://kubernetes.io/" target="_blank" rel="noopener"&gt;Kubernetes&lt;/a&gt; was no different.&lt;/p&gt;
&lt;p&gt;Whether it was the scheduler, autoscaling, or resource quotas, everything fundamentally answered one question:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How do we allocate resources?&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That was the central question of the cloud era.&lt;/p&gt;
&lt;h2 id="the-ai-era-starts-managing-outcomes"&gt;The AI Era Starts Managing Outcomes&lt;/h2&gt;
&lt;p&gt;With the rise of AI, an interesting shift began.&lt;/p&gt;
&lt;p&gt;Enterprises increasingly don&amp;rsquo;t care about:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How many GPUs were used&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;They care far more about:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How many Tokens were produced&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;For most enterprises:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPU is cost.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Token is output.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;GPU, CPU, and memory are merely means of production.&lt;/li&gt;
&lt;li&gt;Token is the final deliverable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/jrstorment_tokenomics-finops-share-7467935235771367424-XEy5/" target="_blank" rel="noopener"&gt;J.R. Storment&lt;/a&gt; (former Executive Director of the FinOps Foundation) said something on LinkedIn that stuck with me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Tokens are what all the hardware is being built to produce.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;All hardware is ultimately producing Tokens.&lt;/p&gt;
&lt;p&gt;From this perspective:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Inference
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;A new value chain is forming.&lt;/p&gt;
&lt;h2 id="token-could-become-the-new-resource-model"&gt;Token Could Become the New Resource Model&lt;/h2&gt;
&lt;p&gt;Over the past few years, I&amp;rsquo;ve been tracking &lt;a href="https://jimmysong.io/blog/ai-inference-on-kubernetes/"&gt;how Kubernetes is evolving in the AI era&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;More and more signals suggest:&lt;/p&gt;
&lt;p&gt;Kubernetes is evolving from a &lt;strong&gt;Compute Control Plane&lt;/strong&gt; into an &lt;strong&gt;AI Control Plane&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Once AI becomes the dominant workload, the resources we manage may no longer be just:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;cpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;8&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;32Gi&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;gpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;They may gradually become:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;token
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;context
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;latency
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;throughput&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Metrics much closer to business value.&lt;/p&gt;
&lt;p&gt;The most interesting thing about the Tokenomics Foundation isn&amp;rsquo;t whether it will define new standards.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s that it implicitly acknowledges one thing:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token has started to become a resource.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Just like CPU in the cloud era.&lt;/p&gt;
&lt;h2 id="implications-for-ai-infrastructure"&gt;Implications for AI Infrastructure&lt;/h2&gt;
&lt;p&gt;For teams building AI infrastructure, this shift deserves serious thought.&lt;/p&gt;
&lt;p&gt;Today, many projects (including &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;) focus on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPU utilization&lt;/li&gt;
&lt;li&gt;GPU sharing&lt;/li&gt;
&lt;li&gt;GPU scheduling&lt;/li&gt;
&lt;li&gt;GPU virtualization&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are all important.&lt;/p&gt;
&lt;p&gt;But from an enterprise perspective, they ultimately don&amp;rsquo;t buy GPU utilization.&lt;/p&gt;
&lt;p&gt;They buy:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Cost Per Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Or even further:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Value Per Token&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Given the same pool of GPUs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whoever can produce more Tokens,&lt;/li&gt;
&lt;li&gt;Whoever can reduce the cost per million Tokens,&lt;/li&gt;
&lt;li&gt;Whoever can increase the business value of each Token,&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Will capture greater commercial value.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why I believe the significance of the Tokenomics Foundation lies beyond the standards themselves.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s driving the entire industry from:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Resource management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Toward:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Value management&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="my-take"&gt;My Take&lt;/h2&gt;
&lt;p&gt;I don&amp;rsquo;t think the Tokenomics Foundation will become the next Kubernetes.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s more like the &lt;a href="https://www.finops.org/" target="_blank" rel="noopener"&gt;FinOps Foundation&lt;/a&gt;, &lt;a href="https://www.opencost.io/" target="_blank" rel="noopener"&gt;OpenCost&lt;/a&gt;, or &lt;a href="https://opentelemetry.io/" target="_blank" rel="noopener"&gt;OpenTelemetry&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;What it&amp;rsquo;s trying to define isn&amp;rsquo;t software.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s the &lt;strong&gt;metering system for the AI era&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The past decade:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;CPU was the language of infrastructure&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The next decade:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Token may become the language of AI infrastructure&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If this trend holds, then the GPUs, inference frameworks, scheduling systems, and Agent Runtimes we discuss today are ultimately just parts of a larger Token economy.&lt;/p&gt;
&lt;p&gt;And that may be the real signal the Linux Foundation is sending by launching the Tokenomics Foundation.&lt;/p&gt;</content:encoded></item><item><title>AI Native Landscape Launches as a Standalone Site</title><link>https://jimmysong.io/blog/ai-native-landscape-launch/</link><pubDate>Thu, 04 Jun 2026 06:12:10 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-native-landscape-launch/</guid><description>AI Native Landscape has moved to landscape.jimmysong.io with 600+ curated open-source projects, AI skill search support, and a call for community contributions.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;AI Native Landscape now has its own home. 600+ curated open-source projects, a brand-new standalone site, and direct search from your AI coding tools.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="how-it-started"&gt;How It Started&lt;/h2&gt;
&lt;p&gt;Last July, I started collecting and organizing AI open-source projects on my personal website. The AI space was moving incredibly fast, with new projects popping up every day. GitHub is a treasure trove, full of tools and frameworks worth studying and using. I wanted to organize them systematically so others (and myself) could find the truly useful ones.&lt;/p&gt;
&lt;p&gt;At first, everything lived on jimmysong.io. But as the catalog grew, maintaining it became a pain. The main site already had plenty of content, and AI-related pages mixed in meant that updating a single project required rebuilding the entire site. So this year I decided to spin it out as a standalone open-source project: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;ai-native-landscape&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;All pages previously at &lt;code&gt;jimmysong.io/en/ai/*&lt;/code&gt; now redirect to the new site at &lt;a href="https://landscape.jimmysong.io" target="_blank" rel="noopener"&gt;landscape.jimmysong.io&lt;/a&gt;. Existing links and bookmarks still work.&lt;/p&gt;
&lt;h2 id="not-just-a-list-but-a-scored-one"&gt;Not Just a List, But a Scored One&lt;/h2&gt;
&lt;p&gt;There are plenty of AI project directories out there, but most just list a name and a link and call it a day. That&amp;rsquo;s not enough.&lt;/p&gt;
&lt;p&gt;GitHub has thousands upon thousands of projects. You see a name, click through, and find out the last commit was six months ago and nobody&amp;rsquo;s responding to Issues. You just spent time researching something that&amp;rsquo;s essentially abandoned.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why AI Native Landscape does something different: &lt;strong&gt;every project gets a score&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The scoring is based on four dimensions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Activity&lt;/strong&gt;: commit frequency, release cadence, Issue response time&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community&lt;/strong&gt;: contributor count, PR merge speed, discussion engagement&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality&lt;/strong&gt;: Stars, Forks, dependency relationships&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sustainability&lt;/strong&gt;: maintenance history, team stability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These four dimensions combine into an overall health score. At a glance, you can tell whether a project is worth your time.&lt;/p&gt;
&lt;p&gt;And the scores &lt;strong&gt;update daily&lt;/strong&gt;, automatically. Today&amp;rsquo;s score might differ from tomorrow&amp;rsquo;s.&lt;/p&gt;
&lt;p&gt;The project list itself is also &lt;strong&gt;human-curated&lt;/strong&gt;. Projects that have gone inactive don&amp;rsquo;t make the cut. I want every project on this list to be &amp;ldquo;alive,&amp;rdquo; so the time you spend studying them won&amp;rsquo;t be wasted.&lt;/p&gt;
&lt;h2 id="whats-covered"&gt;What&amp;rsquo;s Covered&lt;/h2&gt;
&lt;p&gt;The catalog spans eight major areas of AI:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agent Systems&lt;/td&gt;
&lt;td&gt;Frameworks, orchestration, and workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge &amp;amp; Context&lt;/td&gt;
&lt;td&gt;RAG, vector databases, document processing, knowledge graphs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Models &amp;amp; Modalities&lt;/td&gt;
&lt;td&gt;Foundation models, toolkits, speech and vision generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference &amp;amp; Runtime&lt;/td&gt;
&lt;td&gt;Model serving, inference engines, sandboxes, GPU acceleration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training, Evaluation &amp;amp; Optimization&lt;/td&gt;
&lt;td&gt;Frameworks, fine-tuning, benchmarks, observability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Tooling&lt;/td&gt;
&lt;td&gt;MCP protocols, coding agents, IDE tools, SDKs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Applications &amp;amp; Experience&lt;/td&gt;
&lt;td&gt;Chat interfaces, workflow automation, low-code platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform &amp;amp; Infrastructure&lt;/td&gt;
&lt;td&gt;Cloud-native AI, data platforms, security and operations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Currently &lt;strong&gt;600+&lt;/strong&gt; open-source projects, each with bilingual descriptions in English and Chinese.&lt;/p&gt;
&lt;h2 id="ai-skill-search"&gt;AI Skill Search&lt;/h2&gt;
&lt;p&gt;The landscape supports direct search from AI coding tools. No browser needed. Install with one command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;npx landscape-search&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Works with Claude Code, Cursor, GitHub Copilot, OpenAI Codex, Cline, Aider, and other popular AI coding tools. After installation, just search in natural language:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Find me an MCP-compatible agent framework&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;No API key required. No local data files. Your AI tool fetches the live project catalog and returns ranked results.&lt;/p&gt;
&lt;h2 id="tech-stack"&gt;Tech Stack&lt;/h2&gt;
&lt;p&gt;Built with &lt;a href="https://astro.build" target="_blank" rel="noopener"&gt;Astro&lt;/a&gt;, deployed on &lt;a href="https://pages.cloudflare.com" target="_blank" rel="noopener"&gt;Cloudflare Pages&lt;/a&gt;, with &lt;a href="https://www.typescriptlang.org/" target="_blank" rel="noopener"&gt;TypeScript&lt;/a&gt; for type safety. Project data lives in bilingual Markdown files. The build pipeline validates, indexes, and generates OG images in one pass.&lt;/p&gt;
&lt;h2 id="get-involved"&gt;Get Involved&lt;/h2&gt;
&lt;p&gt;GitHub is everyone&amp;rsquo;s treasure trove, and I hope this project can grow with the community&amp;rsquo;s help.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Star the repo&lt;/strong&gt;: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;github.com/rootsongjc/ai-native-landscape&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Submit a project&lt;/strong&gt;: If you know an AI open-source project worth listing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/rootsongjc/ai-native-landscape/issues/new?template=add-project.md" target="_blank" rel="noopener"&gt;Open an Issue&lt;/a&gt; on GitHub&lt;/li&gt;
&lt;li&gt;Or submit a PR directly, add bilingual Markdown files under &lt;code&gt;data/projects/&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;See the &lt;a href="https://github.com/rootsongjc/ai-native-landscape/blob/main/docs/contributing.md" target="_blank" rel="noopener"&gt;contributing guide&lt;/a&gt; for details.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Live site: &lt;a href="https://landscape.jimmysong.io" target="_blank" rel="noopener"&gt;landscape.jimmysong.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;GitHub repo: &lt;a href="https://github.com/rootsongjc/ai-native-landscape" target="_blank" rel="noopener"&gt;rootsongjc/ai-native-landscape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Skill install: &lt;code&gt;npx landscape-search&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Kubernetes as the GPU Control Plane: HAMi v2.9 and Next-Gen AI Infra</title><link>https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/</link><pubDate>Thu, 14 May 2026 06:34:19 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/</guid><description>Observations on the evolution of AI infrastructure control planes, focusing on HAMi v2.9, GPU scheduling, and Kubernetes resource models.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Recently, I&amp;rsquo;ve been following the progress of domestic GPU scheduling and Kubernetes AI resource models. With the release of &lt;a href="https://project-hami.io/blog/hami-v2-9-0-release" target="_blank" rel="noopener"&gt;HAMi v2.9&lt;/a&gt;, I want to share several observations on how the AI Infra control plane is evolving.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/banner.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/banner.webp" alt="Figure 1: Kubernetes as the GPU Control Plane for AI" data-caption="Figure 1: Kubernetes as the GPU Control Plane for AI"
width="1983"
height="793"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Kubernetes as the GPU Control Plane for AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-discuss-this-topic-now"&gt;Why Discuss This Topic Now&lt;/h2&gt;
&lt;p&gt;When DeepSeek R1 was released in early 2025, most people focused on the fact that it trained a model competitive with OpenAI o1 for just $5.6 million. What struck me, however, was that as inference costs plummeted, GPU utilization issues would quickly come to the forefront.&lt;/p&gt;
&lt;p&gt;As models became more useful and inference demand exploded, &amp;ldquo;one GPU per model&amp;rdquo; rapidly became a luxury. Meanwhile, the NVIDIA H200 export saga accelerated the adoption of domestic compute. First, a sales ban; then, at the end of 2025, a 25% tariff under Trump; and by January 2026, Chinese customs had cleared zero units. Policy now mandates that over 40% of data center chips must be domestically produced by 2026.&lt;/p&gt;
&lt;p&gt;The reality is harsh: not only are GPUs scarce, but you must also learn to use NVIDIA, Ascend, Cambricon, Hygon, and other very different platforms simultaneously.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why I believe HAMi v2.9 is more significant than it appears on the surface.&lt;/p&gt;
&lt;h2 id="gpus-are-no-longer-just-about-the-card"&gt;GPUs Are No Longer Just About the Card&lt;/h2&gt;
&lt;p&gt;Kubernetes has always managed GPUs in a rather crude way:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;resources&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;nvidia.com/gpu&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This was sufficient in 2019, when the main question was simply whether a Pod needed a GPU or not.&lt;/p&gt;
&lt;p&gt;But that&amp;rsquo;s no longer enough. An inference service might only need 4GB of VRAM, multiple small models can share a single card, training jobs care about GPU topology and interconnect bandwidth, and multi-tenancy requires fault domain isolation. Treating GPUs as integer resources is like using an abacus for statistical analysis—not impossible, but the mental model is out of sync with reality.&lt;/p&gt;
&lt;p&gt;The most notable feature in HAMi v2.9 is the HAMi-core mode for the Ascend 910C. Previously, sharing Ascend cards relied on SR-IOV hardware virtualization, which was coarse-grained and inflexible. HAMi-core takes a different approach: it uses &lt;code&gt;LD_PRELOAD&lt;/code&gt; to intercept ACL calls in user space, enabling memory isolation at the MB level and compute throttling by percentage.&lt;/p&gt;
&lt;p&gt;In short: it&amp;rsquo;s managed by software, not hardware slicing.&lt;/p&gt;
&lt;p&gt;This is reminiscent of how SDN abstracted the network control plane from hardware devices—GPU partitioning is shifting from a hardware capability to a cluster control plane capability. Considering that Huawei shipped 810,000 Ascend 910C cards last year—nearly half of all domestic chips—this capability has significant real-world impact.&lt;/p&gt;
&lt;h2 id="dra-kubernetes-finally-has-a-robust-device-resource-model"&gt;DRA: Kubernetes Finally Has a Robust Device Resource Model&lt;/h2&gt;
&lt;p&gt;Kubernetes v1.34 (September 2025) officially promoted DRA (Dynamic Resource Allocation) to GA, and Red Hat OpenShift 4.21 followed suit. This is a big deal.&lt;/p&gt;
&lt;p&gt;The Device Plugin solved &amp;ldquo;how to connect GPUs to K8s,&amp;rdquo; but not &amp;ldquo;how to express complex AI resource requirements.&amp;rdquo; Device Plugins only know how many cards are on a node, not how much VRAM you need, what topology, or what isolation level.&lt;/p&gt;
&lt;p&gt;DRA standardizes device resource declaration, allocation, and management via &lt;code&gt;ResourceClaim&lt;/code&gt; and &lt;code&gt;DeviceClass&lt;/code&gt;. HAMi-DRA takes a pragmatic approach: it doesn&amp;rsquo;t require users to change how they declare resources. Instead, it uses a Mutating Webhook to automatically convert existing Device Plugin-style declarations into the DRA model. Legacy systems don&amp;rsquo;t need to change, but can still leverage new capabilities.&lt;/p&gt;
&lt;p&gt;I liken this to what CSI did for storage: it didn&amp;rsquo;t eliminate vendor differences, but allowed Kubernetes to consume different storage capabilities in a unified way. DRA does the same for AI accelerators—NVIDIA, Ascend, AMD, Vastai cards will never be identical, but the scheduling layer should speak a common language.&lt;/p&gt;
&lt;h2 id="a-complete-control-plane-path"&gt;A Complete Control Plane Path&lt;/h2&gt;
&lt;p&gt;If we look beyond individual features and consider HAMi-core, DRA, CDI, and the scheduler together, they actually correspond to different layers of GPU resource management:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi-core&lt;/strong&gt;: How to partition and isolate devices internally&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DRA&lt;/strong&gt;: How to declare, allocate, and bind resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CDI&lt;/strong&gt;: How to standardize device injection into container runtimes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduler/Webhook&lt;/strong&gt;: How to schedule, admit, and observe&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Connecting these layers, from top to bottom, forms the complete Kubernetes GPU Control Plane:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/kubernetes-gpu-control-plane-en.svg" data-img="https://assets.jimmysong.io/images/blog/kubernetes-gpu-control-plane-hami-v29-ai-infra/kubernetes-gpu-control-plane-en.svg" alt="Figure 2: Kubernetes as the GPU Control Plane for AI Workloads" data-caption="Figure 2: Kubernetes as the GPU Control Plane for AI Workloads"
width="1296"
height="1302"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Kubernetes as the GPU Control Plane for AI Workloads&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is a complete control plane path. In v2.9, Volcano vGPU was upgraded to v0.19 with enhanced CDI support. While this may seem like a minor improvement in device injection, it actually completes a critical link in this chain.&lt;/p&gt;
&lt;h2 id="heterogeneity-is-the-main-battlefield-for-domestic-ai-infra"&gt;Heterogeneity Is the Main Battlefield for Domestic AI Infra&lt;/h2&gt;
&lt;p&gt;The reality for domestic AI clusters: you can&amp;rsquo;t build infrastructure around just one type of GPU.&lt;/p&gt;
&lt;p&gt;Enterprise environments often have NVIDIA, Ascend, Biren, Cambricon, Hygon, Muxi, Kunlunxin, Vastai, and other devices coexisting. Each card has different drivers, runtimes, virtualization capabilities, and monitoring methods. HAMi v2.9 adds support for Vastai, covering more than ten types of heterogeneous compute devices. Mixed training and inference, online and offline workloads, domestic and overseas GPUs, multi-team and multi-tenant resource pools—in these scenarios, unified scheduling is far more important than single-card performance.&lt;/p&gt;
&lt;h2 id="key-judgments"&gt;Key Judgments&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;GPU sharing will shift from a cost-saving measure to a default requirement.&lt;/strong&gt; After the explosion of inference workloads, not every workload deserves exclusive access to an entire card. Exclusive allocation will increasingly become a luxury.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DRA is the way forward, but migration will be gradual.&lt;/strong&gt; The Device Plugin ecosystem is too large to disappear overnight. HAMi-DRA&amp;rsquo;s compatibility layer shows the project team understands this.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous scheduling will become the core challenge for AI Infra.&lt;/strong&gt; Whoever can abstract different vendor devices into a unified scheduling language will control the key position in the control plane.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes will be reshaped by AI workloads.&lt;/strong&gt; From scheduling semantics to resource models, AI requires much greater expressiveness than traditional web services. DRA, CDI, topology-aware scheduling—these are not isolated evolutions, but all point to one thing: Kubernetes is evolving from a container orchestrator to the control plane for AI computing.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The significance of HAMi v2.9 is not just in supporting a particular device or partitioning method, but in making one thing clear: the next generation of AI infrastructure competition is not just about model frameworks or GPU counts, but about the &lt;strong&gt;control plane&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;GPUs are shifting from external devices on nodes to native resources within the Kubernetes control plane. Whoever defines the resource model for the AI era will define the long-term boundaries of AI Infra.&lt;/p&gt;</content:encoded></item><item><title>Kubernetes's Anxiety and Rebirth in the AI Wave</title><link>https://jimmysong.io/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/</link><pubDate>Fri, 03 Apr 2026 05:20:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/</guid><description>At KubeCon EU 2026, I witnessed Kubernetes&amp;#39; anxiety and transformation in the AI era. This article explores the challenges and future opportunities for Kubernetes in the age of AI.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Kubernetes hasn&amp;rsquo;t been replaced by AI, but it&amp;rsquo;s being redefined by it. Anxiety is the prelude to rebirth.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After attending KubeCon EU 2026 in Amsterdam, I&amp;rsquo;ve been pondering a key question: Kubernetes isn&amp;rsquo;t obsolete, but it&amp;rsquo;s no longer &amp;ldquo;enough&amp;rdquo;; it hasn&amp;rsquo;t been replaced by AI, but it&amp;rsquo;s being redefined by AI.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/keep-cloud-native-moving.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/keep-cloud-native-moving.webp" alt="Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date." data-caption="Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date."
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: KubeCon EU 2026 slogan: Keep Cloud Native Moving. This event had over 13,000 registrations, making it the largest KubeCon to date.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This was my third time attending KubeCon in Europe. Over the past few years, you can actually see the community&amp;rsquo;s mindset shift through the event slogans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;2024 Paris: &lt;strong&gt;La vie en Cloud Native&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ Cloud Native has become a &amp;ldquo;way of life,&amp;rdquo; the default state&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;2025 London: &lt;strong&gt;No slogan, just the 10th anniversary&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ Kubernetes reached a milestone, focusing on retrospection rather than moving forward&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;2026 Amsterdam: &lt;strong&gt;Keep Cloud Native Moving&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;→ But the question is: where is it moving?&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The absence of a slogan in 2025 was a signal in itself:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;When an ecosystem starts commemorating the past instead of defining the future, it&amp;rsquo;s already at an inflection point.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This article doesn&amp;rsquo;t recap the talks, but instead distills my observations at KubeCon into insights about Kubernetes&amp;rsquo; anxiety and rebirth in the AI wave.&lt;/p&gt;
&lt;h2 id="the-root-of-anxiety-is-kubernetes-facing-a-crisis"&gt;The Root of Anxiety: Is Kubernetes Facing a &amp;ldquo;Crisis&amp;rdquo;?&lt;/h2&gt;
&lt;p&gt;The biggest change at KubeCon was that &lt;strong&gt;AI has completely replaced traditional cloud native topics&lt;/strong&gt;. The focus shifted from service optimization and microservices management to how to deploy and manage AI workloads on Kubernetes, especially inference tasks and GPU scheduling.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainer-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainer-summit.webp" alt="Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI." data-caption="Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI."
width="4000"
height="2668"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Before KubeCon officially started, the Maintainer Summit was all about AI.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Kubernetes, as the foundational infrastructure, was once the core of the cloud native world. With the explosive growth of AI models, &lt;strong&gt;the question now is whether Kubernetes can still serve as a &amp;ldquo;universal&amp;rdquo; platform for everything&lt;/strong&gt;, which has become a new source of anxiety.&lt;/p&gt;
&lt;p&gt;The AI boom brings real challenges: &lt;strong&gt;Can Kubernetes&amp;rsquo; &amp;ldquo;universality&amp;rdquo; adapt to the complexity of AI workloads?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-focus-brought-by-the-ai-boom"&gt;The Focus Brought by the AI Boom&lt;/h2&gt;
&lt;p&gt;AI&amp;rsquo;s popularity has shifted the cloud native spotlight entirely to artificial intelligence. AI coding, OpenClaw, large language models, and generative models have all drawn widespread attention. AI has become the core computing demand in the real world.&lt;/p&gt;
&lt;p&gt;This surge in demand raises the question: Can Kubernetes continue to serve as the infrastructure platform for complex tasks? Especially with issues like GPU sharing, inference model scheduling, VRAM allocation, and device attribute selection, is the traditional Kubernetes resource model sufficient?&lt;/p&gt;
&lt;p&gt;In the past, Kubernetes handled compute, storage, and networking as foundational infrastructure. But with the rapid development of AI, its &amp;ldquo;universality&amp;rdquo; is being challenged. Particularly for inference tasks, Kubernetes&amp;rsquo; model appears thin.&lt;/p&gt;
&lt;h2 id="comparing-with-openstack-will-kubernetes-repeat-history"&gt;Comparing with OpenStack: Will Kubernetes Repeat History?&lt;/h2&gt;
&lt;p&gt;OpenStack once aimed to be a complete open-source cloud platform, but ultimately failed to sustain growth due to &lt;strong&gt;complexity&lt;/strong&gt; and a lack of &lt;strong&gt;flexibility&lt;/strong&gt; in adapting to new technologies.&lt;/p&gt;
&lt;p&gt;Will Kubernetes follow the same path? I believe Kubernetes has different strengths: as a container and microservices orchestration platform, it&amp;rsquo;s widely adopted and has strong community and vendor support. It doesn&amp;rsquo;t try to replace all cloud provider capabilities but serves as an infrastructure control plane to help users manage resources.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainers-summit-group-photo.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/maintainers-summit-group-photo.webp" alt="Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality." data-caption="Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality."
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Cloud native contributors remain active. The crowd at the KubeCon EU 2026 Maintainer Summit shows the community’s vitality.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;However, as AI workloads become mainstream, Kubernetes must find a new position to avoid being replaced by &amp;ldquo;AI-optimized platforms.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="kubernetes-challenge-the-gpu-resource-management-gap"&gt;Kubernetes&amp;rsquo; Challenge: The GPU Resource Management Gap&lt;/h3&gt;
&lt;p&gt;At KubeCon, NVIDIA announced the donation of the &lt;strong&gt;&lt;a href="https://github.com/kubernetes-sigs/nvidia-dra-driver-gpu" target="_blank" rel="noopener"&gt;GPU DRA&lt;/a&gt; (Dynamic Resource Allocation) driver&lt;/strong&gt; to the CNCF, marking the upstreaming of GPU resource management. GPU sharing and scheduling have become urgent issues for Kubernetes.&lt;/p&gt;
&lt;p&gt;Traditionally, Kubernetes relied on the &lt;strong&gt;Device Plugin&lt;/strong&gt; model to schedule GPUs, only supporting allocation by device count (e.g., &lt;code&gt;nvidia.com/gpu: 1&lt;/code&gt;). But for AI inference tasks, more information is needed for resource scheduling, such as &lt;strong&gt;VRAM size&lt;/strong&gt;, &lt;strong&gt;GPU topology&lt;/strong&gt;, and &lt;strong&gt;sharing strategies&lt;/strong&gt;. NVIDIA DRA makes GPU resource management more flexible and intelligent, gradually easing the &amp;ldquo;GPU resource crunch&amp;rdquo; in AI workloads.&lt;/p&gt;
&lt;p&gt;This shift means Kubernetes is no longer just a &amp;ldquo;container orchestration platform,&amp;rdquo; but is becoming the &lt;strong&gt;infrastructure layer for AI-specific resource scheduling&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Against this backdrop, both the community and industry are exploring finer-grained GPU resource abstraction and scheduling mechanisms. For example, the open-source project &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; is building a GPU resource management layer for AI workloads on top of Kubernetes, supporting GPU sharing, VRAM-level allocation, and heterogeneous device scheduling.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/hami-kubecon-demo.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/hami-kubecon-demo.webp" alt="Figure 4: HAMi demo at KubeCon EU 2026 Keynote" data-caption="Figure 4: HAMi demo at KubeCon EU 2026 Keynote"
width="2048"
height="1365"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi demo at KubeCon EU 2026 Keynote&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;These efforts are not about replacing Kubernetes, but about filling the resource model gaps for the AI era. In the long run, this layer may evolve into a &amp;ldquo;GPU Abstraction Layer&amp;rdquo; similar to CNI/CSI, becoming a key part of AI-native infrastructure.&lt;/p&gt;
&lt;h3 id="the-production-gap-many-ai-pocs-few-in-production"&gt;The Production &amp;ldquo;Gap&amp;rdquo;: Many AI PoCs, Few in Production&lt;/h3&gt;
&lt;p&gt;A common post-event summary was: &lt;strong&gt;Many PoCs, but &amp;ldquo;everyday production deployments&amp;rdquo; are still rare&lt;/strong&gt;. Pulumi summarized it as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;lots of working demos, very few production setups people trust&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This shows that while many AI workload solutions succeed in technical demos, the transition from &lt;strong&gt;experimentation to production&lt;/strong&gt; remains difficult. Whether it&amp;rsquo;s GPU resource sharing or inference request scheduling, &lt;strong&gt;whether Kubernetes as the foundation can support this transformation&lt;/strong&gt; is still an open question.&lt;/p&gt;
&lt;h2 id="the-rise-of-inference-systems-kubernetes-scheduling-boundaries-are-challenged"&gt;The Rise of Inference Systems: Kubernetes&amp;rsquo; Scheduling Boundaries Are Challenged&lt;/h2&gt;
&lt;p&gt;Another major event at this KubeCon was &lt;a href="https://github.com/llm-d/llm-d" target="_blank" rel="noopener"&gt;llm-d&lt;/a&gt; being contributed to the CNCF as a Sandbox project.&lt;/p&gt;
&lt;p&gt;If GPU DRA represents the upstreaming of device resource models, then llm-d represents another critical evolution: &lt;strong&gt;Distributed LLM inference capabilities are moving from proprietary engineering implementations to standardized, community-driven collaboration in cloud native.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is significant not just because it&amp;rsquo;s another open-source project, but because it shows that Kubernetes&amp;rsquo; challenges in the AI era are no longer just about &amp;ldquo;how to schedule GPUs,&amp;rdquo; but also &amp;ldquo;how to host inference systems themselves.&amp;rdquo; As prefill/decode separation, request routing, KV cache management, and throughput optimization move into the infrastructure layer, Kubernetes&amp;rsquo; boundaries are being redefined.&lt;/p&gt;
&lt;p&gt;Traditionally, the Kubernetes scheduler focused on Pod scheduling. But in AI inference scenarios, scheduling is not just about picking a node—it&amp;rsquo;s about &lt;strong&gt;selecting the most suitable inference instance based on request characteristics&lt;/strong&gt;. Factors like model state, request queue depth, and cache hit rate all need to be considered. This process is increasingly managed by inference runtimes, forming new &amp;ldquo;request-level scheduling&amp;rdquo; systems.&lt;/p&gt;
&lt;p&gt;This leads to an &lt;strong&gt;overlap between the Kubernetes scheduler and inference systems&lt;/strong&gt;, forcing Kubernetes to rethink its role: should it keep expanding, or collaborate with inference systems?&lt;/p&gt;
&lt;h2 id="ai-native-infrastructure-the-key-challenge-for-production"&gt;AI-Native Infrastructure: The Key Challenge for Production&lt;/h2&gt;
&lt;p&gt;At the &lt;strong&gt;AI Native Summit&lt;/strong&gt;, the real needs for AI-native infrastructure were especially clear. The focus was no longer &amp;ldquo;can it run on Kubernetes,&amp;rdquo; but how to make AI workloads routine, stable, and production-ready on Kubernetes.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/ai-native-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubernetes-in-ai-wave-anxiety-and-rebirth/ai-native-summit.webp" alt="Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era." data-caption="Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era."
width="2970"
height="1980"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: At the AI Native Summit after KubeCon, Linux Foundation Chairman Jonathan said cloud native is entering the AI-native era.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The core challenge is &lt;strong&gt;delivery&lt;/strong&gt;. Unlike traditional apps, AI model weights are often huge—tens of GB or even TB—making model delivery and data management extremely complex. Traditional container delivery systems (like image layers) struggle with such massive data and complex versioning.&lt;/p&gt;
&lt;p&gt;A key direction for Kubernetes is to &lt;strong&gt;standardize model weight and data delivery&lt;/strong&gt;, using &lt;strong&gt;ImageVolume&lt;/strong&gt; and &lt;strong&gt;OCI artifacts&lt;/strong&gt; to solve AI model delivery and version management on Kubernetes. This not only reduces &amp;ldquo;cold start&amp;rdquo; times but also provides infrastructure support for multi-tenancy and compliance.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Kubernetes won&amp;rsquo;t be replaced by AI, but it&amp;rsquo;s being reshaped as the core of infrastructure. This anxiety is the force driving its evolution—it&amp;rsquo;s moving from a &lt;strong&gt;&amp;ldquo;general-purpose infrastructure platform&amp;rdquo;&lt;/strong&gt; to an &lt;strong&gt;&amp;ldquo;AI-powered multifunctional base&amp;rdquo;&lt;/strong&gt;. Some even call it the AI operating system.&lt;/p&gt;
&lt;p&gt;In the future, Kubernetes&amp;rsquo; core competitiveness will no longer be just container management, but &lt;strong&gt;how effectively it can schedule and manage AI workloads&lt;/strong&gt;, and how it can make AI a routine part of operations. This was my biggest takeaway from the AI Native Summit and KubeCon, and it&amp;rsquo;s what I look forward to in the Kubernetes ecosystem over the next few years.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blogs.nvidia.com/blog/nvidia-at-kubecon-2026/" target="_blank" rel="noopener"&gt;Advancing Open Source AI, NVIDIA Donates Dynamic Resource Allocation Driver for GPUs to Kubernetes Community - blog.nvidia.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.pulumi.com/blog/kubecon-eu-2026-recap/" target="_blank" rel="noopener"&gt;KubeCon EU 2026 Recap: The Year AI Moved Into Production on Kubernetes - pulumi.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Day One in Amsterdam: Kubernetes Is Rethinking AI</title><link>https://jimmysong.io/blog/kubecon-eu-2026-day1-ai-infra/</link><pubDate>Sun, 22 Mar 2026 20:41:19 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kubecon-eu-2026-day1-ai-infra/</guid><description>KubeCon Europe 2026 Day One: How Kubernetes is adapting to the AI infrastructure wave and the evolution of the GPU resource layer.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Today marks my first day at &lt;strong&gt;KubeCon Europe 2026&lt;/strong&gt;. The most striking feeling is: the world is vast, but this community is truly small.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/jimmy-at-kubecon-eu.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/jimmy-at-kubecon-eu.webp" alt="Figure 11: Jimmy on the first day of KubeCon EU 2026" data-caption="Figure 11: Jimmy on the first day of KubeCon EU 2026"
width="2400"
height="2400"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 11: Jimmy on the first day of KubeCon EU 2026&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One strong impression stands out:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The world is big, but this circle is really small.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="old-friends-new-cycle"&gt;&lt;strong&gt;Old Friends, New Cycle&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;At the Maintainer Summit, I met many familiar faces—&lt;/p&gt;
&lt;p&gt;Colleagues from Ant Group, friends from Tetrate, and some people I&amp;rsquo;ve known for nearly a decade. Together, we&amp;rsquo;ve journeyed from the early days of Kubernetes, Service Mesh, and cloud native infrastructure to today.&lt;/p&gt;
&lt;p&gt;In a sense, this generation has fully experienced:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The rise of Kubernetes&lt;/li&gt;
&lt;li&gt;The standardization of Cloud Native&lt;/li&gt;
&lt;li&gt;The microservices and service mesh boom&lt;/li&gt;
&lt;li&gt;And now, the era of AI Infrastructure&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This isn&amp;rsquo;t about &amp;ldquo;new people entering the field,&amp;rdquo; but rather—&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The same group stepping into a new technology cycle.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="what-is-the-maintainer-summit-discussing"&gt;&lt;strong&gt;What Is the Maintainer Summit Discussing?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If you ask:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What is the Kubernetes community most concerned about right now?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Today&amp;rsquo;s answer is very clear:&lt;/p&gt;
&lt;p&gt;👉 &lt;strong&gt;How to run AI workloads better on Kubernetes&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit.webp" alt="Figure 12: The Maintainer Summit’s main topic is AI Infra" data-caption="Figure 12: The Maintainer Summit’s main topic is AI Infra"
width="1440"
height="960"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 12: The Maintainer Summit’s main topic is AI Infra&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Many topics at the Maintainer Summit revolved around:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Scheduling models for LLM / AI workloads&lt;/li&gt;
&lt;li&gt;GPU / accelerator resource management&lt;/li&gt;
&lt;li&gt;Integrating inference systems with Kubernetes&lt;/li&gt;
&lt;li&gt;Redefining the roles of data plane vs. control plane&lt;/li&gt;
&lt;li&gt;How observability tools like OTel monitor AI workloads&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes hasn&amp;rsquo;t been replaced by AI; it&amp;rsquo;s actively &amp;ldquo;absorbing&amp;rdquo; AI.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="key-signal-gpus-are-becoming-the"&gt;&lt;strong&gt;Key Signal: GPUs Are Becoming the &amp;ldquo;Infrastructure Layer&amp;rdquo;&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Today, I had an in-depth discussion with CNCF TOC, Red Hat, and the vLLM community.&lt;/p&gt;
&lt;p&gt;The core question was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How should GPUs be &amp;ldquo;platformized&amp;rdquo;?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Some consensus is already clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPUs are no longer just devices&lt;/li&gt;
&lt;li&gt;They are now &lt;strong&gt;a schedulable, partitionable, and shareable resource layer&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/toc-meeting.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/toc-meeting.webp" alt="Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration" data-caption="Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration"
width="2400"
height="1648"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 13: TOC meeting discussing GPU resource management and LLM Serving integration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;At the Maintainer Summit in Amsterdam, we had deep discussions with CNCF TOC, Red Hat, and the vLLM community about GPU resource management and LLM Serving integration in Kubernetes scenarios, and explored potential collaboration between vLLM and HAMi.&lt;/p&gt;
&lt;p&gt;Behind this is a major paradigm shift:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Past&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Now&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPU = Node resource&lt;/td&gt;
&lt;td&gt;GPU = Infrastructure layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exclusive use&lt;/td&gt;
&lt;td&gt;Multi-tenant sharing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Static binding&lt;/td&gt;
&lt;td&gt;Dynamic scheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Managed within frameworks&lt;/td&gt;
&lt;td&gt;Unified management at the platform layer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This is exactly what we&amp;rsquo;ve been working on in &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="hami-from"&gt;&lt;strong&gt;HAMi: From &amp;ldquo;Project&amp;rdquo; to &amp;ldquo;Reference Pattern&amp;rdquo;&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Another interesting change today:&lt;/p&gt;
&lt;p&gt;HAMi is no longer just a &amp;ldquo;community project&amp;rdquo;—it&amp;rsquo;s becoming:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A reference implementation (reference pattern) for AI Infra&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit-hami.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/kubecon-eu-maintainer-summit-hami.webp" alt="Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit" data-caption="Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit"
width="1440"
height="960"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 14: Li Mengxuan, CTO of Dynamia, sharing HAMi’s design and practice at KubeCon EU 2026 Maintainer Summit&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is reflected in several ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Invited to present at the Maintainer Summit&lt;/li&gt;
&lt;li&gt;Participating in CNCF TOC discussions&lt;/li&gt;
&lt;li&gt;Involved in incubating review demos&lt;/li&gt;
&lt;li&gt;Exploring joint content with the vLLM community (even discussing a joint blog 👀)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Especially in conversations with Red Hat and vLLM, a clear trend emerged:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GPU resource management and LLM serving are becoming coupled&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Upper layer: vLLM / inference frameworks&lt;/li&gt;
&lt;li&gt;Lower layer: GPU scheduling / sharing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A new &amp;ldquo;interface layer&amp;rdquo; is gradually forming.&lt;/p&gt;
&lt;p&gt;This is a direction worth betting on.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/incubating-review.webp" data-img="https://assets.jimmysong.io/images/blog/kubecon-eu-2026-day1-ai-infra/incubating-review.webp" alt="Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo" data-caption="Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo"
width="2400"
height="1489"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 15: At the TAG Workshop, HAMi was discussed as an Incubating demo&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="a-caution-the-ai-infra-startup-boom-hasn"&gt;&lt;strong&gt;A Caution: The AI Infra Startup Boom Hasn&amp;rsquo;t Really Begun&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;At the same time, I have a somewhat &amp;ldquo;counterintuitive&amp;rdquo; observation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;We haven&amp;rsquo;t yet seen a large wave of AI Infra (K8s-focused) startups.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Most companies I saw today:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many are pivoting from CI/CD, Service Mesh, or Gateway&lt;/li&gt;
&lt;li&gt;Many are traditional cloud vendors extending into AI&lt;/li&gt;
&lt;li&gt;Many are working on models, agents, or even lower-level tech&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But those truly focused on:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Making AI workloads run better on Kubernetes&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There are actually not many startups at this layer.&lt;/p&gt;
&lt;p&gt;This could mean two things:&lt;/p&gt;
&lt;h3 id="1-this-layer-isn"&gt;&lt;strong&gt;1) This Layer Isn&amp;rsquo;t Fully Formed Yet&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Currently, most activity is at:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The model layer (LLM / foundation models)&lt;/li&gt;
&lt;li&gt;The application layer (Agent / Copilot)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But not at:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The scheduling layer&lt;/li&gt;
&lt;li&gt;The resource layer&lt;/li&gt;
&lt;li&gt;The runtime layer&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="2-or-the-barrier-to-entry-is-very-high"&gt;&lt;strong&gt;2) Or, the Barrier to Entry Is Very High&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Because at its core, this is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The intersection of Cloud Native × GPU × AI workload&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It&amp;rsquo;s not just &amp;ldquo;wrapping AI,&amp;rdquo; but a fundamental re-architecture at the infrastructure level.&lt;/p&gt;
&lt;h2 id="my-take"&gt;&lt;strong&gt;My Take&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If we break down the AI technology stack:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Agent / Application
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;LLM Serving (vLLM, etc.)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI Runtime / Scheduling
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU Resource Layer
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Hardware&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Most innovation today is concentrated in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The top two layers (Agent / LLM)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But the real long-term moat lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The middle two layers (Runtime + Resource Layer)&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And Kubernetes is very likely to remain:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The default platform for this middle layer&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Today&amp;rsquo;s takeaway:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes is not obsolete; it&amp;rsquo;s being redefined.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And our generation is shifting from:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Cloud Native Builders&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;AI Infrastructure Builders&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;More to come tomorrow.&lt;/p&gt;</content:encoded></item><item><title>HAMi Website Refactor: Why HAMi Docs and Website Underwent a Complete Redesign</title><link>https://jimmysong.io/blog/hami-website-redesign/</link><pubDate>Tue, 17 Mar 2026 08:55:52 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/hami-website-redesign/</guid><description>A systematic upgrade to HAMi’s website and docs, improving community visibility, content structure, search, and usability.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;This redesign is more than a style update—it&amp;rsquo;s a step toward clearer technical communication and better user experience. Try the new HAMi website at &lt;a href="https://project-hami.io" target="_blank" rel="noopener"&gt;https://project-hami.io&lt;/a&gt; and submit issues &lt;a href="https://github.com/Project-HAMi/website/issues" target="_blank" rel="noopener"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Over the past two months, I conducted a thorough refactor of the documentation website (see &lt;a href="https://github.com/Project-HAMi/website/pulls?q=is%3Apr&amp;#43;is%3Aclosed&amp;#43;author%3Arootsongjc" target="_blank" rel="noopener"&gt;GitHub&lt;/a&gt;). Externally, it looks like a &amp;ldquo;visual redesign&amp;rdquo;, but from the perspective of community maintainers and content builders, it&amp;rsquo;s a comprehensive upgrade of information architecture, content system, and frontend experience.&lt;/p&gt;
&lt;p&gt;This article aims to systematically explain three things: why we did this refactor, what exactly changed, and what these changes mean for the HAMi community.&lt;/p&gt;
&lt;h2 id="why-refactor-the-website-and-documentation"&gt;Why Refactor the Website and Documentation&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; is a CNCF-hosted open source project initiated and contributed by &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt;, with growing influence in GPU virtualization, heterogeneous compute scheduling, and AI infrastructure. The community content is expanding, and user types are becoming more diverse: from first-time visitors to engineers and enterprise users seeking deployment docs, architecture diagrams, case studies, and ecosystem information.&lt;/p&gt;
&lt;p&gt;The original site was functional, but as content grew, several issues became apparent:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The homepage lacked information density, making it hard to quickly grasp the project&amp;rsquo;s overall value.&lt;/li&gt;
&lt;li&gt;Connections between docs, blogs, and community info were not smooth; content entry points were scattered.&lt;/li&gt;
&lt;li&gt;Search experience was unstable; external solutions were not ideal in practice.&lt;/li&gt;
&lt;li&gt;Mobile experience had many details needing improvement, especially navigation, card layouts, and footer areas.&lt;/li&gt;
&lt;li&gt;Visual style was inconsistent, making it hard to convey community influence and engineering maturity.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For a fast-evolving open source community, the website is not just a &amp;ldquo;place for docs&amp;rdquo;, but the public interface of the community. It needs to serve as project introduction, knowledge gateway, adoption proof, community connector, and brand expression.&lt;/p&gt;
&lt;p&gt;So the goal of this refactor was clear: not just superficial beautification, but to truly upgrade the website into HAMi&amp;rsquo;s systematic community entry point.&lt;/p&gt;
&lt;h2 id="what-was-done-in-this-refactor"&gt;What Was Done in This Refactor&lt;/h2&gt;
&lt;p&gt;This update was not a single-point change, but a series of systematic improvements.&lt;/p&gt;
&lt;h3 id="homepage-redesign-and-complete-information-architecture-overhaul"&gt;Homepage Redesign and Complete Information Architecture Overhaul&lt;/h3&gt;
&lt;p&gt;The most obvious change is the homepage.&lt;/p&gt;
&lt;p&gt;We redesigned the homepage structure, moving away from simply stacking content blocks, and instead organizing the page around the main narrative: &amp;ldquo;Project Positioning → Core Capabilities → Ecosystem Entry → Content Accumulation → Community Trust&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Specifically, the homepage received several key upgrades:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Rebuilt the Hero section to strengthen first-screen information delivery and action entry.&lt;/li&gt;
&lt;li&gt;Optimized CTA design so users can quickly access docs, blogs, and resources.&lt;/li&gt;
&lt;li&gt;Added and enhanced multiple homepage sections to showcase project value and community reach in a more structured way.&lt;/li&gt;
&lt;li&gt;Adjusted visual hierarchy, background atmosphere, and scroll rhythm, transforming the homepage from a &amp;ldquo;content list&amp;rdquo; into a &amp;ldquo;narrative page&amp;rdquo;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These changes include Hero animations and atmosphere layers, research/story sections, new resource entry sections, refreshed CTAs, unified background design, and ongoing reduction of visual noise. Together, they solve a core problem: enabling visitors to understand what HAMi is and why it&amp;rsquo;s worth exploring further within seconds.&lt;/p&gt;
&lt;h3 id="architecture-diagrams"&gt;Architecture Diagrams&lt;/h3&gt;
&lt;p&gt;Key diagrams were redrawn for clearer technical communication. This helps users grasp HAMi&amp;rsquo;s role in AI infrastructure.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-hero-diagram.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-hero-diagram.webp" alt="Figure 1: HAMi website homepage architecture diagram" data-caption="Figure 1: HAMi website homepage architecture diagram"
width="3160"
height="1714"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: HAMi website homepage architecture diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For HAMi, this change is critical. The community faces not just a single feature, but a set of system-level challenges involving Kubernetes, schedulers, GPU Operators, heterogeneous devices, and enterprise platforms. Improved diagrams make the website a better technical entry point.&lt;/p&gt;
&lt;h3 id="added-case-studies-community-and-ecosystem-sections-to-make-impact-visible"&gt;Added Case Studies, Community, and Ecosystem Sections to Make Impact Visible&lt;/h3&gt;
&lt;p&gt;Another important direction was strengthening the &amp;ldquo;community proof&amp;rdquo; layer.&lt;/p&gt;
&lt;p&gt;Many open source project sites fall into the trap of having complete docs, but users can&amp;rsquo;t tell if the project is truly adopted, if the community is active, or if the ecosystem is expanding. The HAMi website redesign consciously addresses this.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/ecosystem.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/ecosystem.webp" alt="Figure 2: HAMi ecosystem and device support" data-caption="Figure 2: HAMi ecosystem and device support"
width="2200"
height="454"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: HAMi ecosystem and device support&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/adopters.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/adopters.webp" alt="Figure 3: HAMi adopters" data-caption="Figure 3: HAMi adopters"
width="3688"
height="1534"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: HAMi adopters&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/contributors.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/contributors.webp" alt="Figure 4: HAMi contributor organizations" data-caption="Figure 4: HAMi contributor organizations"
width="3662"
height="674"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi contributor organizations&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="blog--reading-experience"&gt;Blog &amp;amp; Reading Experience&lt;/h3&gt;
&lt;p&gt;Blog cards, lists, and metadata were unified for easier reading and sharing. Blogs are now a core communication layer.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-blog.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/hami-blog.webp" alt="Figure 5: HAMi website blog list page" data-caption="Figure 5: HAMi website blog list page"
width="2318"
height="1088"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: HAMi website blog list page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="mobile-optimization"&gt;Mobile Optimization&lt;/h3&gt;
&lt;p&gt;Navigation, card layouts, footer, and search were improved for smoother mobile browsing.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/mobile.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/mobile.webp" alt="Figure 6: HAMi website mobile view" data-caption="Figure 6: HAMi website mobile view"
width="654"
height="1418"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: HAMi website mobile view&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="footer--search"&gt;Footer &amp;amp; Search&lt;/h3&gt;
&lt;p&gt;Footer layout was enhanced for better navigation and credibility. Built-in search replaced unreliable external solutions, improving content accessibility.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/footer.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/footer.webp" alt="Figure 7: HAMi website footer" data-caption="Figure 7: HAMi website footer"
width="2472"
height="718"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: HAMi website footer&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/hami-website-redesign/search.webp" data-img="https://assets.jimmysong.io/images/blog/hami-website-redesign/search.webp" alt="Figure 8: HAMi website built-in search" data-caption="Figure 8: HAMi website built-in search"
width="1330"
height="1214"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 8: HAMi website built-in search&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="what-this-redesign-means-for-the-hami-community"&gt;What This Redesign Means for the HAMi Community&lt;/h2&gt;
&lt;p&gt;From screenshots, it looks like &amp;ldquo;the website looks better&amp;rdquo;. But from a community-building perspective, its significance is deeper.&lt;/p&gt;
&lt;p&gt;First, HAMi&amp;rsquo;s external expression is more systematic.&lt;/p&gt;
&lt;p&gt;The website is no longer just a collection of scattered pages, but is forming a complete narrative chain: users can understand project value from the homepage, capability details from docs, practical paths from blogs, and community impact from ecosystem modules.&lt;/p&gt;
&lt;p&gt;Second, community content assets are reorganized.&lt;/p&gt;
&lt;p&gt;Previously, valuable articles, diagrams, and explanations existed but were hard to find. Now, through homepage sections, navigation, and search refactor, these contents are more effectively connected.&lt;/p&gt;
&lt;p&gt;Third, HAMi&amp;rsquo;s community image is more mature.&lt;/p&gt;
&lt;p&gt;A mature open source project needs not just an active code repository, but clear, stable, and sustainable website expression. Structure, style, and usability are part of the community&amp;rsquo;s engineering capability.&lt;/p&gt;
&lt;p&gt;Fourth, this lays the foundation for expanding case studies, adopters, contributors, and ecosystem content.&lt;/p&gt;
&lt;p&gt;With the framework sorted, adding more case studies, collaboration entry points, or showcasing more adopters and partners will be more natural and easier for users to understand.&lt;/p&gt;
&lt;h2 id="as-a-community-contributor-my-top-three-takeaways-from-this-redesign"&gt;As a Community Contributor, My Top Three Takeaways from This Redesign&lt;/h2&gt;
&lt;p&gt;In summary, I believe this refactor got three things right:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Upgraded the website from a &amp;ldquo;content dump&amp;rdquo; to a &amp;ldquo;community gateway&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Combined visual optimization with information architecture adjustment, not just a skin change.&lt;/li&gt;
&lt;li&gt;Improved basic experiences like search, mobile, navigation, and footer.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These may not be as flashy as launching a new feature, but they directly impact content dissemination, user comprehension, and the project&amp;rsquo;s long-term image.&lt;/p&gt;
&lt;p&gt;For infrastructure projects like HAMi, technical capability is fundamental, but clearly communicating, organizing, and continuously presenting that capability is also a form of infrastructure.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;This HAMi documentation and website refactor is essentially an upgrade to the community&amp;rsquo;s &amp;ldquo;expression layer&amp;rdquo; infrastructure.&lt;/p&gt;
&lt;p&gt;It improves visual and reading experience, reorganizes content, homepage narrative, search paths, mobile access, and community signal display. Homepage redesign, architecture diagram redraw, unified blog style, mobile optimization, enhanced footer, and switching from external to built-in search together constitute a true &amp;ldquo;refactor&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Externally, it helps more people quickly understand HAMi; internally, it provides a stable platform for the community to accumulate case studies, expand the ecosystem, and serve adopters and contributors.&lt;/p&gt;
&lt;p&gt;The website is not an accessory to the open source community, but part of its long-term influence. HAMi&amp;rsquo;s redesign is about taking this seriously.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re interested in Kubernetes GPU virtualization, add me on WeChat &lt;code&gt;jimmysong&lt;/code&gt; or scan the QR code below.&lt;/p&gt;
&lt;div class="cta-group"&gt;
&lt;a href="https://github.com/project-hami/hami" class="btn btn-sm btn-primary"&gt;Check out the HAMi project on GitHub&lt;/a&gt;
&lt;/div&gt;</content:encoded></item><item><title>GTC 2026 Eve: AI is Becoming the New Infrastructure</title><link>https://jimmysong.io/blog/gtc-2026-ai-native-infrastructure/</link><pubDate>Sun, 15 Mar 2026 11:34:06 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gtc-2026-ai-native-infrastructure/</guid><description>On the eve of GTC 2026, rethinking whether AI is becoming the new infrastructure from NVIDIA&amp;#39;s AI Five-Layer Cake, the rise of agent runtime, to AI-native infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;AI is quietly reshaping the infrastructure landscape, and GTC 2026 may become a key node in this transformation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Next week, one of the most important technology conferences in the AI industry, &lt;a href="https://www.nvidia.com/gtc/" target="_blank" rel="noopener"&gt;&lt;strong&gt;NVIDIA GTC 2026&lt;/strong&gt;&lt;/a&gt;, will be held in San Jose, USA.&lt;/p&gt;
&lt;p&gt;For many people, GTC is just a GPU technology conference. But if you follow the development of the AI industry over the past few years, you&amp;rsquo;ll find an interesting phenomenon:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Many important narratives about AI infrastructure are gradually taking shape at GTC.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From CUDA, DGX, to AI Factory, and most recently Jensen Huang&amp;rsquo;s proposed &lt;strong&gt;AI Five-Layer Cake&lt;/strong&gt;, NVIDIA is constantly attempting to redefine the computing infrastructure of the AI era.&lt;/p&gt;
&lt;p&gt;This is why many people call GTC:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI&amp;rsquo;s &amp;ldquo;Woodstock.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-gtc.webp" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-gtc.webp" alt="Figure 1: NVIDIA GTC Conference" data-caption="Figure 1: NVIDIA GTC Conference"
width="2212"
height="1152"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: NVIDIA GTC Conference&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This year&amp;rsquo;s GTC (March 16-19) is expected to cover various levels of the AI stack, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Chips&lt;/li&gt;
&lt;li&gt;AI Data Centers&lt;/li&gt;
&lt;li&gt;AI Agents&lt;/li&gt;
&lt;li&gt;Robotics&lt;/li&gt;
&lt;li&gt;Inference Computing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;According to &lt;a href="https://blogs.nvidia.com/blog/gtc-2026-news/" target="_blank" rel="noopener"&gt;NVIDIA&amp;rsquo;s official blog&lt;/a&gt;, this year&amp;rsquo;s keynote will focus on &lt;strong&gt;the complete AI stack from chips to applications&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;If we put these signals together, we can actually see a larger trend:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is transforming from an &amp;ldquo;applied technology&amp;rdquo; into &amp;ldquo;infrastructure.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-perspective-of-industrial-revolutions"&gt;The Perspective of Industrial Revolutions&lt;/h2&gt;
&lt;p&gt;From a longer time scale, the technological revolutions in human history are essentially infrastructure revolutions.&lt;/p&gt;
&lt;p&gt;We usually divide industrial revolutions into four times.&lt;/p&gt;
&lt;p&gt;In the table below, you can see the infrastructure corresponding to each industrial revolution:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Industrial Revolution&lt;/th&gt;
&lt;th&gt;Infrastructure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Steam Revolution&lt;/td&gt;
&lt;td&gt;Steam Engine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Electrical Revolution&lt;/td&gt;
&lt;td&gt;Power Grid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Digital Revolution&lt;/td&gt;
&lt;td&gt;Computer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internet Era&lt;/td&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Industrial Revolutions and Corresponding Infrastructure
&lt;/figcaption&gt;
&lt;h3 id="first-industrial-revolution-steam"&gt;First Industrial Revolution: Steam&lt;/h3&gt;
&lt;p&gt;The steam engine allowed humans to utilize mechanical power on a large scale for the first time. Production no longer relied on human or animal power, but on machines.&lt;/p&gt;
&lt;h3 id="second-industrial-revolution-electricity"&gt;Second Industrial Revolution: Electricity&lt;/h3&gt;
&lt;p&gt;Electricity changed not only the source of power, but also the organization of production. Assembly lines, large-scale manufacturing, and modern industrial systems are all built on the foundation of the power grid.&lt;/p&gt;
&lt;h3 id="third-industrial-revolution-computers"&gt;Third Industrial Revolution: Computers&lt;/h3&gt;
&lt;p&gt;Computers allowed information to be processed digitally. Software became a production tool.&lt;/p&gt;
&lt;h3 id="fourth-industrial-revolution-internet-and-intelligence"&gt;Fourth Industrial Revolution: Internet and Intelligence&lt;/h3&gt;
&lt;p&gt;The internet connects all computers together. Cloud computing transforms computing resources into infrastructure. And AI gives machines a certain degree of &amp;ldquo;cognitive ability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-true-significance-of-ai"&gt;The True Significance of AI&lt;/h2&gt;
&lt;p&gt;If we observe these industrial revolutions, we discover a pattern:&lt;/p&gt;
&lt;p&gt;Each industrial revolution produces a new &lt;strong&gt;General Purpose Infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;And AI is likely to become the next-generation infrastructure.&lt;/p&gt;
&lt;p&gt;NVIDIA even directly stated in a &lt;a href="https://blogs.nvidia.com/blog/ai-5-layer-cake/" target="_blank" rel="noopener"&gt;recent article&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI is essential infrastructure, like electricity and the internet.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;p&gt;AI is no longer just an applied technology, but a &lt;strong&gt;new factor of production&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="nvidias-five-layer-cake"&gt;NVIDIA&amp;rsquo;s Five-Layer Cake&lt;/h2&gt;
&lt;p&gt;Recently, Jensen Huang proposed a very interesting concept: &lt;strong&gt;AI Five-Layer Cake&lt;/strong&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-five-layer-cake.webp" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-five-layer-cake.webp" alt="Figure 2: AI Five Layer Cake (Image source: &amp;lt;a href=&amp;#34;https://blogs.nvidia.com/blog/ai-5-layer-cake/&amp;#34; target=&amp;#34;_blank&amp;#34; rel=&amp;#34;noopener&amp;#34;&amp;gt;NVIDIA&amp;lt;/a&amp;gt;)" data-caption="Figure 2: AI Five Layer Cake (Image source: &amp;lt;a href=&amp;#34;https://blogs.nvidia.com/blog/ai-5-layer-cake/&amp;#34; target=&amp;#34;_blank&amp;#34; rel=&amp;#34;noopener&amp;#34;&amp;gt;NVIDIA&amp;lt;/a&amp;gt;)"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: AI Five Layer Cake (Image source: &lt;a href="https://blogs.nvidia.com/blog/ai-5-layer-cake/" target="_blank" rel="noopener"&gt;NVIDIA&lt;/a&gt;)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;AI is broken down into five layers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Energy&lt;/li&gt;
&lt;li&gt;Chips&lt;/li&gt;
&lt;li&gt;AI Infrastructure&lt;/li&gt;
&lt;li&gt;Models&lt;/li&gt;
&lt;li&gt;Applications&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This model actually illustrates one thing:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is a complete industrial system.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Jensen Huang even described AI at Davos as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;One of the largest-scale infrastructure constructions in human history.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="signals-gtc-2026-may-release"&gt;Signals GTC 2026 May Release&lt;/h2&gt;
&lt;p&gt;This year&amp;rsquo;s GTC is expected to release several important directions.&lt;/p&gt;
&lt;h3 id="inference-computing"&gt;Inference Computing&lt;/h3&gt;
&lt;p&gt;The focus of AI in the past was training. But the main load of AI in the future is likely to be &lt;strong&gt;Inference&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Analysts expect that by 2030, &lt;strong&gt;75% of computing demand in the AI data center market will come from inference&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="agentic-ai"&gt;Agentic AI&lt;/h3&gt;
&lt;p&gt;The past AI model was:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;User → Model → Answer&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The Agent model is more complex:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;User → Agent → Tools → Model → Action&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The flowchart below shows the main interaction paths in the Agent model:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agentic-ai-interaction-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agentic-ai-interaction-en.svg" alt="Figure 3: Agentic AI Interaction Flow" data-caption="Figure 3: Agentic AI Interaction Flow"
width="936"
height="536"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Agentic AI Interaction Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;AI is no longer just answering questions, but &lt;strong&gt;executing tasks&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="agent-platform"&gt;Agent Platform&lt;/h3&gt;
&lt;p&gt;Recent media reports suggest that NVIDIA may launch a new Agent platform: &lt;strong&gt;NemoClaw&lt;/strong&gt;, aimed at helping enterprises deploy AI Agents.&lt;/p&gt;
&lt;p&gt;If this project is truly released, it means NVIDIA&amp;rsquo;s stack will become the following structure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-agent-platform-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/nvidia-agent-platform-en.svg" alt="Figure 4: NVIDIA Agent Platform Architecture" data-caption="Figure 4: NVIDIA Agent Platform Architecture"
width="416"
height="816"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: NVIDIA Agent Platform Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is actually a complete AI stack.&lt;/p&gt;
&lt;h2 id="agents-change-computing-workloads"&gt;Agents Change Computing Workloads&lt;/h2&gt;
&lt;p&gt;The emergence of Agents brings new computing workload issues.&lt;/p&gt;
&lt;p&gt;Past AI workloads were mainly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training&lt;/li&gt;
&lt;li&gt;Inference&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But Agents bring a third type of workload:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent Workloads&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The figure below shows the diverse workload types related to Agents:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agent-workloads-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/agent-workloads-en.svg" alt="Figure 5: Agent Workloads Structure" data-caption="Figure 5: Agent Workloads Structure"
width="1376"
height="316"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Agent Workloads Structure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The characteristic of this workload is &lt;strong&gt;highly fragmented&lt;/strong&gt;. GPUs are no longer occupied for long periods, but rather face many small requests. This poses new challenges for infrastructure.&lt;/p&gt;
&lt;h2 id="ai-native-infrastructure"&gt;AI-Native Infrastructure&lt;/h2&gt;
&lt;p&gt;For the past few years, I&amp;rsquo;ve been thinking about a question:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is AI-native infrastructure?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It is clearly not just &amp;ldquo;Kubernetes with GPUs.&amp;rdquo; I&amp;rsquo;m more inclined to believe it needs to possess several characteristics.&lt;/p&gt;
&lt;h3 id="gpu-as-a-first-class-resource"&gt;GPU as a First-Class Resource&lt;/h3&gt;
&lt;p&gt;In the cloud computing era, CPU is the core resource. In the AI era, &lt;strong&gt;GPU is the core resource&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="heterogeneous-computing"&gt;Heterogeneous Computing&lt;/h3&gt;
&lt;p&gt;Real-world AI chips are not limited to NVIDIA:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;NVIDIA&lt;/li&gt;
&lt;li&gt;Ascend&lt;/li&gt;
&lt;li&gt;Cambricon&lt;/li&gt;
&lt;li&gt;Metax&lt;/li&gt;
&lt;li&gt;Moore Threads&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Future AI infrastructure must be able to manage &lt;strong&gt;heterogeneous computing&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="gpu-sharing"&gt;GPU Sharing&lt;/h3&gt;
&lt;p&gt;GPU is a very expensive resource. If it cannot be shared, utilization will be very low. This is why GPU virtualization and slicing are becoming increasingly important.&lt;/p&gt;
&lt;h3 id="ai-scheduling"&gt;AI Scheduling&lt;/h3&gt;
&lt;p&gt;AI scheduling includes not only traditional CPU and Memory, but also:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;GPU
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;VRAM
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Topology
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Bandwidth&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="a-possible-ai-tech-stack"&gt;A Possible AI Tech Stack&lt;/h2&gt;
&lt;p&gt;Combining the above trends, the future AI stack may present the following structure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-tech-stack-en.svg" data-img="https://assets.jimmysong.io/images/blog/gtc-2026-ai-native-infrastructure/ai-tech-stack-en.svg" alt="Figure 6: AI Tech Stack Evolution" data-caption="Figure 6: AI Tech Stack Evolution"
width="416"
height="956"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: AI Tech Stack Evolution&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This structure is very close to NVIDIA&amp;rsquo;s Five-Layer Cake.&lt;/p&gt;
&lt;h2 id="my-judgment"&gt;My Judgment&lt;/h2&gt;
&lt;p&gt;Combining signals from GTC, AI Factory, Agents, and AI Five-Layer Cake, we can see a very obvious trend:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI is rewriting computing infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Future competition may not just be &amp;ldquo;who has the best model,&amp;rdquo; but:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Who has the best AI Infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Just like the past few decades:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity determines industrial capability&lt;/li&gt;
&lt;li&gt;Internet determines information capability&lt;/li&gt;
&lt;li&gt;Cloud computing determines software capability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The future may be:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Infrastructure determines intelligence capability.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;If we stretch the time scale a bit longer, we may be in a new historical stage.&lt;/p&gt;
&lt;p&gt;AI is no longer just a technological tool. It is becoming &lt;strong&gt;new infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Just like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Electricity&lt;/li&gt;
&lt;li&gt;Internet&lt;/li&gt;
&lt;li&gt;Cloud computing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And AI-native infrastructure is likely to become one of the most important technology directions for the next decade.&lt;/p&gt;</content:encoded></item><item><title>When GPUs Move Toward Open Scheduling: Structural Shifts in AI Native Infrastructure</title><link>https://jimmysong.io/blog/gpu-open-scheduling-hami-2025/</link><pubDate>Fri, 13 Feb 2026 14:32:46 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gpu-open-scheduling-hami-2025/</guid><description>A CTO/VP view on open GPU scheduling: CDI, Kubernetes DRA, virtualization data planes, ecosystem governance, and lock-in risk.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The future of GPU scheduling isn&amp;rsquo;t about whose implementation is more &amp;ldquo;black-box&amp;rdquo;—it&amp;rsquo;s about who can standardize device resource contracts into something governable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/banner.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/banner.webp" alt="Figure 1: GPU Open Scheduling" data-caption="Figure 1: GPU Open Scheduling"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: GPU Open Scheduling&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Have you ever wondered: why are GPUs so expensive, yet overall utilization often hovers around 10–20%?&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/underutilization.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/underutilization.webp" alt="Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization" data-caption="Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: GPU Utilization Problem: Expensive GPUs with only 10-20% utilization&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This isn&amp;rsquo;t a problem you solve with &amp;ldquo;better scheduling algorithms.&amp;rdquo; It&amp;rsquo;s a &lt;strong&gt;structural problem&lt;/strong&gt; - GPU scheduling is undergoing a shift from &amp;ldquo;proprietary implementation&amp;rdquo; to &amp;ldquo;open scheduling,&amp;rdquo; similar to how networking converged on CNI and storage converged on CSI.&lt;/p&gt;
&lt;p&gt;In the &lt;a href="https://dynamia.ai/blog/hami-2025-recap" target="_blank" rel="noopener"&gt;HAMi 2025 Annual Review&lt;/a&gt;, we noted: &amp;ldquo;HAMi 2025 is no longer just about GPU sharing tools—it&amp;rsquo;s a more structural signal: GPUs are moving toward open scheduling.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;By 2025, the signals of this shift became visible: Kubernetes Dynamic Resource Allocation (DRA) graduated to GA and became enabled by default, NVIDIA GPU Operator started defaulting to &lt;a href="https://github.com/cncf-tags/container-device-interface" target="_blank" rel="noopener"&gt;CDI&lt;/a&gt; (Container Device Interface), and HAMi&amp;rsquo;s production-grade case studies under CNCF are moving &amp;ldquo;GPU sharing&amp;rdquo; from experimental capability to operational excellence.&lt;/p&gt;
&lt;p&gt;This post analyzes this structural shift from an AI Native Infrastructure perspective, and what it means for &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; and the industry.&lt;/p&gt;
&lt;h2 id="why-open-scheduling-matters"&gt;Why &amp;ldquo;Open Scheduling&amp;rdquo; Matters&lt;/h2&gt;
&lt;p&gt;In multi-cloud and hybrid cloud environments, GPU model diversity significantly amplifies operational costs. One large internet company&amp;rsquo;s platform spans H200/H100/A100/V100/4090 GPUs across five clusters. If you can only allocate &amp;ldquo;whole GPUs,&amp;rdquo; resource misalignment becomes inevitable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Open scheduling&amp;rdquo; isn&amp;rsquo;t a slogan—it&amp;rsquo;s a set of engineering contracts being solidified into the mainstream stack.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="standardized-resource-expression"&gt;Standardized Resource Expression&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt; GPUs were extended resources. The scheduler didn&amp;rsquo;t understand if they represented memory, compute, or device types.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dra-evolution.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dra-evolution.webp" alt="Figure 3: Open Scheduling Standardization Evolution" data-caption="Figure 3: Open Scheduling Standardization Evolution"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Open Scheduling Standardization Evolution&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Now:&lt;/strong&gt; Kubernetes DRA provides objects like DeviceClass, ResourceClaim, and ResourceSlice. This lets drivers and cluster administrators define device categories and selection logic (including CEL-based selectors), while Kubernetes handles the full loop: match devices → bind claims → place Pods onto nodes with access to allocated devices.&lt;/p&gt;
&lt;p&gt;Even more importantly, Kubernetes 1.34 stated that core APIs in the &lt;code&gt;resource.k8s.io&lt;/code&gt; group graduated to GA, DRA became stable and enabled by default, and the community committed to avoiding breaking changes going forward. This means the ecosystem can invest with confidence in a stable, standard API.&lt;/p&gt;
&lt;h3 id="standardized-device-injection"&gt;Standardized Device Injection&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt; Device injection relied on vendor-specific hooks and runtime class patterns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Now:&lt;/strong&gt; The Container Device Interface (CDI) abstracts device injection into an open specification. NVIDIA&amp;rsquo;s Container Toolkit explicitly describes CDI as an open specification for container runtimes, and NVIDIA GPU Operator 25.10.0 defaults to enabling CDI on install/upgrade—directly leveraging runtime-native CDI support (containerd, CRI-O, etc.) for GPU injection.&lt;/p&gt;
&lt;p&gt;This means &amp;ldquo;devices into containers&amp;rdquo; is also moving toward replaceable, standardized interfaces.&lt;/p&gt;
&lt;h2 id="hami-from-sharing-tool-to-governable-data-plane"&gt;HAMi: From &amp;ldquo;Sharing Tool&amp;rdquo; to &amp;ldquo;Governable Data Plane&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;On this standardization path, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt;&amp;rsquo;s role needs redefinition: &lt;strong&gt;it&amp;rsquo;s not about replacing Kubernetes—it&amp;rsquo;s about turning GPU virtualization and slicing into a declarative, schedulable, governable data plane.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="data-plane-perspective"&gt;Data Plane Perspective&lt;/h3&gt;
&lt;p&gt;HAMi&amp;rsquo;s core contribution expands the allocatable unit from &amp;ldquo;whole GPU integers&amp;rdquo; to finer-grained shares (memory and compute), forming a complete allocation chain:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Device discovery:&lt;/strong&gt; Identify available GPU devices and models&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling placement:&lt;/strong&gt; Use Scheduler Extender to make native schedulers &amp;ldquo;understand&amp;rdquo; vGPU resource models (Filter/Score/Bind phases)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;In-container enforcement:&lt;/strong&gt; Inject share constraints into container runtime&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metric export:&lt;/strong&gt; Provide observable metrics for utilization, isolation, and more&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This transforms &amp;ldquo;sharing&amp;rdquo; from ad-hoc &amp;ldquo;it runs&amp;rdquo; experimentation into engineering capability that can be declared in YAML, scheduled by policy, and validated by metrics.&lt;/p&gt;
&lt;h3 id="scheduling-mechanism-enhancement-not-replacement"&gt;Scheduling Mechanism: Enhancement, Not Replacement&lt;/h3&gt;
&lt;p&gt;HAMi&amp;rsquo;s scheduling doesn&amp;rsquo;t replace Kubernetes—it uses a &lt;strong&gt;Scheduler Extender&lt;/strong&gt; pattern to let the native scheduler understand vGPU resource models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Filter:&lt;/strong&gt; Filter nodes based on memory, compute, device type, topology, and other constraints&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Score:&lt;/strong&gt; Apply configurable policies like binpack, spread, topology-aware scoring&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bind:&lt;/strong&gt; Complete final device-to-Pod binding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This architecture positions HAMi naturally as an execution layer under higher-level &amp;ldquo;AI control planes&amp;rdquo; (queuing, quotas, priorities)—working alongside Volcano, Kueue, Koordinator, and others.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/hami-scheduler-extender.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/hami-scheduler-extender.webp" alt="Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)" data-caption="Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: HAMi Scheduling Architecture (Filter → Score → Bind)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="production-evidence-from-can-we-share-to-can-we-operate"&gt;Production Evidence: From &amp;ldquo;Can We Share?&amp;rdquo; to &amp;ldquo;Can We Operate?&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.cncf.io/case-studies/?_sft_lf_project=hami" target="_blank" rel="noopener"&gt;CNCF public case studies&lt;/a&gt; provide concrete answers: &lt;strong&gt;in a hybrid, multi-cloud platform built on Kubernetes and HAMi, 10,000+ Pods run concurrently, and GPU utilization improves from 13% to 37% (nearly 3×).&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/case-studies.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/case-studies.webp" alt="Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%&amp;#43; utilization, SF Technology 57% savings" data-caption="Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%&amp;#43; utilization, SF Technology 57% savings"
width="2466"
height="1508"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: CNCF Production Case Studies: Ke Holdings 13%→37%, DaoCloud 80%+ utilization, SF Technology 57% savings&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Here are highlights from several cases:&lt;/p&gt;
&lt;h3 id="case-study-1-ke-holdings-february-5-2026"&gt;Case Study 1: Ke Holdings (February 5, 2026)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Environment:&lt;/strong&gt; 5 clusters spanning public and private clouds&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU models:&lt;/strong&gt; H200/H100/A100/V100/4090 and more&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Architecture:&lt;/strong&gt; Separate &amp;ldquo;GPU clusters&amp;rdquo; for large training tasks (dedicated allocation) vs &amp;ldquo;vGPU clusters&amp;rdquo; with HAMi fine-grained memory slicing for high-density inference&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Concurrent scale:&lt;/strong&gt; 10,000+ Pods&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcome:&lt;/strong&gt; Overall GPU utilization improved from 13% to 37% (nearly 3×)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-2-daocloud-december-2-2025"&gt;Case Study 2: DaoCloud (December 2, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hard constraints:&lt;/strong&gt; Must remain cloud-native, vendor-agnostic, and compatible with CNCF toolchain&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adoption outcomes:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Average GPU utilization: 80%+&lt;/li&gt;
&lt;li&gt;GPU-related operating cost reduction: 20–30%&lt;/li&gt;
&lt;li&gt;Coverage: 10+ data centers, 10,000+ GPUs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit benefit:&lt;/strong&gt; Unified abstraction layer across NVIDIA and domestic GPUs, reducing vendor dependency&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-3-prep-edu-august-20-2025"&gt;Case Study 3: Prep EDU (August 20, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Negative experience:&lt;/strong&gt; Isolation failures in other GPU-sharing approaches caused memory conflicts and instability&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Positive outcome:&lt;/strong&gt; HAMi&amp;rsquo;s vGPU scheduling, GPU type/UUID targeting, and compatibility with NVIDIA GPU Operator and RKE2 became decisive factors for production adoption&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environment:&lt;/strong&gt; Heterogeneous RTX 4070/4090 cluster&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="case-study-4-sf-technology-september-18-2025"&gt;Case Study 4: SF Technology (September 18, 2025)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Project:&lt;/strong&gt; EffectiveGPU (built on HAMi)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use cases:&lt;/strong&gt; Large model inference, test services, speech recognition, domestic AI hardware (Huawei Ascend, Baidu Kunlun, etc.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcomes:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;GPU savings: Large model inference runs 65 services on 28 GPUs (37 saved); test cluster runs 19 services on 6 GPUs (13 saved)&lt;/li&gt;
&lt;li&gt;Overall savings: Up to 57% GPU savings for production and test clusters&lt;/li&gt;
&lt;li&gt;Utilization improvement: Up to 100% GPU utilization improvement with GPU virtualization&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Highlights:&lt;/strong&gt; Cross-node collaborative scheduling, priority-based preemption, memory over-subscription&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These cases demonstrate a consistent pattern: &lt;strong&gt;GPU virtualization becomes economically meaningful only when it participates in a governable contract—where utilization, isolation, and policy can be expressed, measured, and improved over time.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="strategic-implications-for-dynamia"&gt;Strategic Implications for Dynamia&lt;/h2&gt;
&lt;p&gt;From Dynamia&amp;rsquo;s perspective (and as VP of Open Source Ecosystem), the strategic value of HAMi becomes clear:&lt;/p&gt;
&lt;h3 id="two-layer-architecture-open-source-vs-commercial"&gt;Two-Layer Architecture: Open Source vs Commercial&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi (CNCF open source project):&lt;/strong&gt; Responsible for &amp;ldquo;adoption and trust,&amp;rdquo; focused on GPU virtualization and compute efficiency&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamia enterprise products and services:&lt;/strong&gt; Responsible for &amp;ldquo;production and scale,&amp;rdquo; providing commercial distributions and enterprise services built on HAMi&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dynamia-hami-dual-mechanism.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/dynamia-hami-dual-mechanism.webp" alt="Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial" data-caption="Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Dynamia Dual Mechanism: Open Source vs Commercial&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This boundary is the foundation for long-term trust—project and company offerings remain separate, with commercial distributions and services built on the open source project.&lt;/p&gt;
&lt;h3 id="global-narrative-strategy"&gt;Global Narrative Strategy&lt;/h3&gt;
&lt;p&gt;The internal alignment memo recommends a bilingual approach:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First layer:&lt;/strong&gt; Lead globally with &amp;ldquo;GPU virtualization / sharing / utilization&amp;rdquo; (Chinese can directly use &amp;ldquo;GPU virtualization and heterogeneous scheduling,&amp;rdquo; but English first layer should avoid &amp;ldquo;heterogeneous&amp;rdquo; as a headline)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second layer:&lt;/strong&gt; When users discuss mixed GPUs or workload diversity, introduce &amp;ldquo;heterogeneous&amp;rdquo; to confirm capability boundaries—never as the opening hook&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core anchor:&lt;/strong&gt; Maintain &amp;ldquo;HAMi (project and community) ≠ company products&amp;rdquo; as the non-negotiable baseline for long-term positioning&lt;/p&gt;
&lt;h3 id="the-right-commercialization-landing"&gt;The Right Commercialization Landing&lt;/h3&gt;
&lt;p&gt;DaoCloud&amp;rsquo;s case study already set vendor-agnostic and CNCF toolchain compatibility as hard constraints, framing vendor dependency reduction as a business and operational benefit—not just a technical detail. Project-HAMi&amp;rsquo;s official documentation lists &amp;ldquo;avoid vendor lock&amp;rdquo; as a core value proposition.&lt;/p&gt;
&lt;p&gt;In this context, &lt;strong&gt;the right commercialization landing isn&amp;rsquo;t &amp;ldquo;closed-source scheduling&amp;rdquo;—it&amp;rsquo;s productizing capabilities around real enterprise complexity:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Systematic compatibility matrix&lt;/li&gt;
&lt;li&gt;SLO and tail-latency governance&lt;/li&gt;
&lt;li&gt;Metering for billing&lt;/li&gt;
&lt;li&gt;RBAC, quotas, multi-cluster governance&lt;/li&gt;
&lt;li&gt;Upgrade and rollback safety&lt;/li&gt;
&lt;li&gt;Faster path-to-production for DRA/CDI and other standardization efforts&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="forward-view-the-next-23-years"&gt;Forward View: The Next 2–3 Years&lt;/h2&gt;
&lt;p&gt;My strong judgment: &lt;strong&gt;over the next 2–3 years, GPU scheduling competition will shift from &amp;ldquo;whose implementation is more black-box&amp;rdquo; to &amp;ldquo;whose contract is more open.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The reasons are practical:&lt;/p&gt;
&lt;h3 id="hardware-form-factors-and-supply-chains-are-diversifying"&gt;Hardware Form Factors and Supply Chains Are Diversifying&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI&amp;rsquo;s February 12, 2026 &amp;ldquo;GPT‑5.3‑Codex‑Spark&amp;rdquo; release emphasizes ultra-low latency serving, including persistent WebSockets and a dedicated serving tier on Cerebras hardware&lt;/li&gt;
&lt;li&gt;Large-scale GPU-backed financing announcements (for pan-European deployments) illustrate the infrastructure scale and financial engineering surrounding accelerator fleets&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These signals suggest that heterogeneity will grow: mixed accelerators, mixed clouds, mixed workload types.&lt;/p&gt;
&lt;h3 id="low-latency-inference-tiers-will-force-systematic-scheduling"&gt;Low-Latency Inference Tiers Will Force Systematic Scheduling&lt;/h3&gt;
&lt;p&gt;Low-latency inference tiers (beyond just GPUs) will force resource scheduling toward &amp;ldquo;multi-accelerator, multi-layer cache, multi-class node&amp;rdquo; architectural design—scheduling must inherently be heterogeneous.&lt;/p&gt;
&lt;h3 id="open-scheduling-is-risk-management-not-idealism"&gt;Open Scheduling Is Risk Management, Not Idealism&lt;/h3&gt;
&lt;p&gt;In this world, &amp;ldquo;open scheduling&amp;rdquo; isn&amp;rsquo;t idealism—it&amp;rsquo;s risk management. Building schedulable governable &amp;ldquo;control plane + data plane&amp;rdquo; combinations around DRA/CDI and other solidifying open interfaces, ones that are pluggable, multi-tenant governable, and co-evolvable with the ecosystem—this looks like the truly sustainable path for AI Native Infrastructure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The next battleground isn&amp;rsquo;t &amp;ldquo;whose scheduling is smarter&amp;rdquo;—it&amp;rsquo;s &amp;ldquo;who can standardize device resource contracts into something governable.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;When you place HAMi 2025 back in the broader AI Native Infrastructure context, it&amp;rsquo;s no longer just the year of &amp;ldquo;GPU sharing tools&amp;rdquo;—it&amp;rsquo;s a more structural signal: &lt;strong&gt;GPUs are moving toward open scheduling.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/future-vision-open-scheduling.webp" data-img="https://assets.jimmysong.io/images/blog/gpu-open-scheduling-hami-2025/future-vision-open-scheduling.webp" alt="Figure 7: Open Scheduling Future Vision" data-caption="Figure 7: Open Scheduling Future Vision"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: Open Scheduling Future Vision&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The driving forces come from both ends:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Upstream:&lt;/strong&gt; Standards like DRA/CDI continue to solidify&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downstream:&lt;/strong&gt; Scale and diversity (multi-cloud, multi-model, even accelerators beyond GPUs)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For Dynamia, HAMi&amp;rsquo;s significance has transcended &amp;ldquo;GPU sharing tool&amp;rdquo;: it turns GPU virtualization and slicing into declarative, schedulable, measurable data planes—letting queues, quotas, priorities, and multi-tenancy actually close the governance loop.&lt;/p&gt;</content:encoded></item><item><title>AI Learning Resources: 44 Curated Collections from Our Cleanup</title><link>https://jimmysong.io/blog/ultimate-ai-learning-resources/</link><pubDate>Sun, 08 Feb 2026 12:20:05 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ultimate-ai-learning-resources/</guid><description>A curated collection of AI learning resources we removed from the AI Resources list: awesome lists, courses, tutorials, and cookbooks. These educational materials deserve their own spotlight.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;The best way to learn AI is to start building. These resources will guide your journey.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ultimate-ai-learning-resources/banner.webp" data-img="https://assets.jimmysong.io/images/blog/ultimate-ai-learning-resources/banner.webp" alt="Figure 1: AI Learning Resources Collection" data-caption="Figure 1: AI Learning Resources Collection"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: AI Learning Resources Collection&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In my ongoing effort to keep the AI Resources list focused on &lt;strong&gt;production-ready tools and frameworks&lt;/strong&gt;, I&amp;rsquo;ve removed &lt;strong&gt;44 collection-type projects&lt;/strong&gt;—courses, tutorials, awesome lists, and cookbooks.&lt;/p&gt;
&lt;p&gt;These resources aren&amp;rsquo;t gone—they&amp;rsquo;ve been moved here. This post is a &lt;strong&gt;curated collection&lt;/strong&gt; of those educational materials, organized by type and topic. Whether you&amp;rsquo;re a complete beginner or an experienced practitioner, you&amp;rsquo;ll find something valuable here.&lt;/p&gt;
&lt;h2 id="why-remove-collections-from-ai-resources"&gt;Why Remove Collections from AI Resources?&lt;/h2&gt;
&lt;p&gt;My AI Resources list now focuses on &lt;strong&gt;concrete tools and frameworks&lt;/strong&gt;—projects you can directly use in production. Collections, while valuable, serve a different purpose: &lt;strong&gt;education and discovery&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;By separating them, I:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Keep the resources list actionable and focused&lt;/li&gt;
&lt;li&gt;Create a dedicated space for learning materials&lt;/li&gt;
&lt;li&gt;Make it easier to find what you need&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-awesome-lists-14-collections"&gt;📚 Awesome Lists (14 Collections)&lt;/h2&gt;
&lt;p&gt;Awesome lists are community-curated collections of the best resources. They&amp;rsquo;re perfect for discovering new tools and staying updated.&lt;/p&gt;
&lt;h3 id="must-explore-awesome-lists"&gt;Must-Explore Awesome Lists&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/filipecalegario/awesome-generative-ai" target="_blank" rel="noopener"&gt;Awesome Generative AI&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Models, tools, tutorials, and research papers&lt;/li&gt;
&lt;li&gt;Great for: Comprehensive overview of generative AI landscape&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/hannibal046/awesome-llm" target="_blank" rel="noopener"&gt;Awesome LLM&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM resources: papers, tools, datasets, applications&lt;/li&gt;
&lt;li&gt;Great for: Deep dive into large language models&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/arindam200/awesome-ai-apps" target="_blank" rel="noopener"&gt;Awesome AI Apps&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Practical LLM applications, RAG examples, agent implementations&lt;/li&gt;
&lt;li&gt;Great for: Real-world implementation examples&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/hesreallyhim/awesome-claude-code" target="_blank" rel="noopener"&gt;Awesome Claude Code&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Claude Code commands, files, and workflows&lt;/li&gt;
&lt;li&gt;Great for: Maximizing Claude Code productivity&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/punkpeye/awesome-mcp-servers" target="_blank" rel="noopener"&gt;Awesome MCP Servers&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP servers for modular AI backend systems&lt;/li&gt;
&lt;li&gt;Great for: Building with Model Context Protocol&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="specialized-awesome-lists"&gt;Specialized Awesome Lists&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/f/awesome-chatgpt-prompts" target="_blank" rel="noopener"&gt;Awesome ChatGPT Prompts&lt;/a&gt;&lt;/strong&gt; - Prompt examples for various scenarios&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/shubhamsaboo/awesome-llm-apps" target="_blank" rel="noopener"&gt;Awesome LLM Apps&lt;/a&gt;&lt;/strong&gt; - LLM applications with code examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/bradyfu/awesome-multimodal-large-language-models" target="_blank" rel="noopener"&gt;Awesome Multimodal LLM&lt;/a&gt;&lt;/strong&gt; - Multimodal model resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/punkpeye/awesome-mcp-clients" target="_blank" rel="noopener"&gt;Awesome MCP Clients&lt;/a&gt;&lt;/strong&gt; - MCP client tools and SDKs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/composiohq/awesome-claude-skills" target="_blank" rel="noopener"&gt;Awesome Claude Skills&lt;/a&gt;&lt;/strong&gt; - Claude Skills and workflows&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/github/awesome-copilot" target="_blank" rel="noopener"&gt;Awesome GitHub Copilot&lt;/a&gt;&lt;/strong&gt; - Copilot customizations&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/zerolu/awesome-nanobanana-pro" target="_blank" rel="noopener"&gt;Awesome Nano Banana Pro&lt;/a&gt;&lt;/strong&gt; - Image model prompts and examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/alchemyst-ai/awesome-saas" target="_blank" rel="noopener"&gt;Awesome SaaS&lt;/a&gt;&lt;/strong&gt; - AI platform templates&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/voltagent/awesome-claude-code-subagents" target="_blank" rel="noopener"&gt;Awesome Claude Code Subagents&lt;/a&gt;&lt;/strong&gt; - Claude Code subagents&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-courses--tutorials-9-curricula"&gt;🎓 Courses &amp;amp; Tutorials (9 Curricula)&lt;/h2&gt;
&lt;p&gt;Structured learning paths from universities and tech companies.&lt;/p&gt;
&lt;h3 id="microsofts-ai-curriculum"&gt;Microsoft&amp;rsquo;s AI Curriculum&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ai-for-beginners" target="_blank" rel="noopener"&gt;AI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;12 weeks, 24 lessons covering neural networks, deep learning, CV, NLP&lt;/li&gt;
&lt;li&gt;Great for: Complete AI foundation&lt;/li&gt;
&lt;li&gt;Format: Lessons, quizzes, projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ml-for-beginners" target="_blank" rel="noopener"&gt;Machine Learning for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;12-week, 26-lesson curriculum on classic ML&lt;/li&gt;
&lt;li&gt;Great for: ML fundamentals without deep math&lt;/li&gt;
&lt;li&gt;Format: Project-based exercises&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/generative-ai-for-beginners" target="_blank" rel="noopener"&gt;Generative AI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;18 lessons on building GenAI applications&lt;/li&gt;
&lt;li&gt;Great for: Practical GenAI development&lt;/li&gt;
&lt;li&gt;Format: Hands-on projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/ai-agents-for-beginners" target="_blank" rel="noopener"&gt;AI Agents for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;11 lessons on agent systems&lt;/li&gt;
&lt;li&gt;Great for: Understanding autonomous agents&lt;/li&gt;
&lt;li&gt;Format: Project-driven learning&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/edgeai-for-beginners" target="_blank" rel="noopener"&gt;EdgeAI for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Optimization, deployment, and real-world Edge AI&lt;/li&gt;
&lt;li&gt;Great for: On-device AI applications&lt;/li&gt;
&lt;li&gt;Format: Practical tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/microsoft/mcp-for-beginners" target="_blank" rel="noopener"&gt;MCP for Beginners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model Context Protocol curriculum&lt;/li&gt;
&lt;li&gt;Great for: Building with MCP&lt;/li&gt;
&lt;li&gt;Format: Cross-language examples and labs&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="official-platform-courses"&gt;Official Platform Courses&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/course" target="_blank" rel="noopener"&gt;Hugging Face Learn Center&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Free courses on LLMs, deep RL, CV, audio&lt;/li&gt;
&lt;li&gt;Great for: Hands-on Hugging Face ecosystem&lt;/li&gt;
&lt;li&gt;Format: Interactive notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/openai/openai-cookbook" target="_blank" rel="noopener"&gt;OpenAI Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Runnable examples using OpenAI API&lt;/li&gt;
&lt;li&gt;Great for: OpenAI API best practices&lt;/li&gt;
&lt;li&gt;Format: Code examples and guides&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/pytorch/tutorials" target="_blank" rel="noopener"&gt;PyTorch Tutorials&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Basics to advanced deep learning&lt;/li&gt;
&lt;li&gt;Great for: PyTorch mastery&lt;/li&gt;
&lt;li&gt;Format: Comprehensive tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-cookbooks--example-collections-5-collections"&gt;🍳 Cookbooks &amp;amp; Example Collections (5 Collections)&lt;/h2&gt;
&lt;p&gt;Practical code examples and recipes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/anthropics/claude-cookbooks" target="_blank" rel="noopener"&gt;Claude Cookbooks&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Notebooks and examples for building with Claude&lt;/li&gt;
&lt;li&gt;Great for: Anthropic Claude integration&lt;/li&gt;
&lt;li&gt;Format: Jupyter notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/cookbook" target="_blank" rel="noopener"&gt;Hugging Face Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Practical AI cookbook with Jupyter notebooks&lt;/li&gt;
&lt;li&gt;Great for: Open models and tools&lt;/li&gt;
&lt;li&gt;Format: Hands-on examples&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/RationaleInstitute/tinker-cookbook" target="_blank" rel="noopener"&gt;Tinker Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training and fine-tuning examples&lt;/li&gt;
&lt;li&gt;Great for: Fine-tuning workflows&lt;/li&gt;
&lt;li&gt;Format: Platform-specific recipes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/e2b-dev/e2b-cookbook" target="_blank" rel="noopener"&gt;E2B Cookbook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Examples for building LLM apps&lt;/li&gt;
&lt;li&gt;Great for: LLM application development&lt;/li&gt;
&lt;li&gt;Format: Recipes and tutorials&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/jamwithai/arxiv-paper-curator" target="_blank" rel="noopener"&gt;arXiv Paper Curator&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;6-week course on RAG systems&lt;/li&gt;
&lt;li&gt;Great for: Production-ready RAG&lt;/li&gt;
&lt;li&gt;Format: Project-based learning&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-guides--handbooks-5-resources"&gt;📖 Guides &amp;amp; Handbooks (5 Resources)&lt;/h2&gt;
&lt;p&gt;In-depth guides on specific topics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dair-ai/prompt-engineering-guide" target="_blank" rel="noopener"&gt;Prompt Engineering Guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Comprehensive prompt engineering resources&lt;/li&gt;
&lt;li&gt;Great for: Mastering prompt design&lt;/li&gt;
&lt;li&gt;Format: Guides, papers, lectures, notebooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/huggingface/evaluation-guidebook" target="_blank" rel="noopener"&gt;Evaluation Guidebook&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM evaluation best practices from Hugging Face&lt;/li&gt;
&lt;li&gt;Great for: Assessing LLM performance&lt;/li&gt;
&lt;li&gt;Format: Practical guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/davidkimai/context-engineering" target="_blank" rel="noopener"&gt;Context Engineering&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Design and optimize context beyond prompt engineering&lt;/li&gt;
&lt;li&gt;Great for: Advanced context management&lt;/li&gt;
&lt;li&gt;Format: Practical handbook&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/coleam00/context-engineering-intro" target="_blank" rel="noopener"&gt;Context Engineering Intro&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Template and guide for context engineering&lt;/li&gt;
&lt;li&gt;Great for: Providing project context to AI assistants&lt;/li&gt;
&lt;li&gt;Format: Template + guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/IIETER/IIETER" target="_blank" rel="noopener"&gt;Vibe-Coding Workflow&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;5-step prompt template for building MVPs with LLMs&lt;/li&gt;
&lt;li&gt;Great for: Rapid prototyping with AI&lt;/li&gt;
&lt;li&gt;Format: Workflow template&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-template--workflow-collections"&gt;🗂️ Template &amp;amp; Workflow Collections&lt;/h2&gt;
&lt;p&gt;Reusable templates and workflows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/davila7/claude-code-templates" target="_blank" rel="noopener"&gt;Claude Code Templates&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Code templates for various programming scenarios&lt;/li&gt;
&lt;li&gt;Great for: Claude AI development&lt;/li&gt;
&lt;li&gt;Format: Template collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/zie619/n8n-workflows" target="_blank" rel="noopener"&gt;n8n Workflows&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2,000+ professionally organized n8n workflows&lt;/li&gt;
&lt;li&gt;Great for: Workflow automation&lt;/li&gt;
&lt;li&gt;Format: Searchable catalog&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/nusquama/n8nworkflows.xyz" target="_blank" rel="noopener"&gt;N8N Workflows Catalog&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Community-driven reusable workflow templates&lt;/li&gt;
&lt;li&gt;Great for: Workflow import and versioning&lt;/li&gt;
&lt;li&gt;Format: Template catalog&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-research--evaluation"&gt;📊 Research &amp;amp; Evaluation&lt;/h2&gt;
&lt;p&gt;Academic and evaluation resources.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/amberljc/llmsys-paperlist" target="_blank" rel="noopener"&gt;LLMSys PaperList&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Curated list of LLM systems papers&lt;/li&gt;
&lt;li&gt;Great for: Research on training, inference, serving&lt;/li&gt;
&lt;li&gt;Format: Paper collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/cheahjs/free-llm-api-resources" target="_blank" rel="noopener"&gt;Free LLM API Resources&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM providers with free/trial API access&lt;/li&gt;
&lt;li&gt;Great for: Experimentation without cost&lt;/li&gt;
&lt;li&gt;Format: Provider list&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-other-notable-resources"&gt;🎨 Other Notable Resources&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/x1xhlol/system-prompts-and-models-of-ai-tools" target="_blank" rel="noopener"&gt;System Prompts and Models of AI Tools&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Community-curated collection of system prompts and AI tool examples&lt;/li&gt;
&lt;li&gt;Great for: Prompt and agent engineering&lt;/li&gt;
&lt;li&gt;Format: Resource collection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/epfml/ml_course" target="_blank" rel="noopener"&gt;ML Course CS-433&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;EPFL Machine Learning Course&lt;/li&gt;
&lt;li&gt;Great for: Academic ML foundation&lt;/li&gt;
&lt;li&gt;Format: Lectures, labs, projects&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/stas00/ml-engineering" target="_blank" rel="noopener"&gt;Machine Learning Engineering&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ML engineering open-book: compute, storage, networking&lt;/li&gt;
&lt;li&gt;Great for: Production ML systems&lt;/li&gt;
&lt;li&gt;Format: Comprehensive guide&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/neural-maze/realtime-phone-agents-course" target="_blank" rel="noopener"&gt;Realtime Phone Agents Course&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Build low-latency voice agents&lt;/li&gt;
&lt;li&gt;Great for: Voice AI applications&lt;/li&gt;
&lt;li&gt;Format: Hands-on course&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/johnma2006/m3-workshop" target="_blank" rel="noopener"&gt;LLMs from Scratch&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Build a working LLM from first principles&lt;/li&gt;
&lt;li&gt;Great for: Understanding LLM internals&lt;/li&gt;
&lt;li&gt;Format: Repository + book materials&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="-how-to-use-this-collection"&gt;💡 How to Use This Collection&lt;/h2&gt;
&lt;h3 id="for-complete-beginners"&gt;For Complete Beginners&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Start with&lt;/strong&gt;: Microsoft&amp;rsquo;s AI for Beginners&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Practice with&lt;/strong&gt;: PyTorch Tutorials&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explore&lt;/strong&gt;: Awesome AI Apps for inspiration&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-developers"&gt;For Developers&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Build skills&lt;/strong&gt;: OpenAI Cookbook + Claude Cookbooks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find tools&lt;/strong&gt;: Awesome Generative AI + Awesome LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn workflows&lt;/strong&gt;: n8n Workflows Catalog&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-researchers"&gt;For Researchers&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Stay updated&lt;/strong&gt;: Awesome Generative AI + LLMSys PaperList&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deep dive&lt;/strong&gt;: Awesome LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Implement&lt;/strong&gt;: Hugging Face Cookbook&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="for-product-builders"&gt;For Product Builders&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Find examples&lt;/strong&gt;: Awesome AI Apps&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn workflows&lt;/strong&gt;: n8n Workflows Catalog&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Study patterns&lt;/strong&gt;: Awesome LLM Apps&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="-what-was-not-removed"&gt;🔄 What Was NOT Removed&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Agent frameworks and production tools remain in the AI Resources list&lt;/strong&gt;, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AutoGen&lt;/strong&gt; - Microsoft&amp;rsquo;s multi-agent framework&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CrewAI&lt;/strong&gt; - High-performance multi-agent orchestration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LangGraph&lt;/strong&gt; - Stateful multi-agent applications&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flowise&lt;/strong&gt; - Visual agent platform&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Langflow&lt;/strong&gt; - Visual workflow builder&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;And 80+ more agent frameworks&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are &lt;strong&gt;functional tools&lt;/strong&gt; you can use to build applications, not educational collections. They belong in the AI Resources list.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="-summary"&gt;📝 Summary&lt;/h2&gt;
&lt;p&gt;I removed &lt;strong&gt;44 collection-type projects&lt;/strong&gt; from the AI Resources list to keep it focused on production tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;14 Awesome Lists&lt;/strong&gt; - Discover new tools and stay updated&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;9 Courses &amp;amp; Tutorials&lt;/strong&gt; - Structured learning paths&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;5 Cookbooks&lt;/strong&gt; - Practical code examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;5 Guides &amp;amp; Handbooks&lt;/strong&gt; - In-depth resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;4 Template Collections&lt;/strong&gt; - Reusable workflows&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;7 Other Resources&lt;/strong&gt; - Research and evaluation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These resources remain &lt;strong&gt;incredibly valuable&lt;/strong&gt; for learning and discovery. They just serve a different purpose than the production-focused tools in my AI Resources list.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Next Steps&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Bookmark this post for future reference&lt;/li&gt;
&lt;li&gt;Explore the AI Resources list for production tools (agent frameworks, databases, etc.)&lt;/li&gt;
&lt;li&gt;Check out my blog for more AI engineering insights&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Acknowledgments&lt;/strong&gt;: This collection was compiled during my AI Resources cleanup initiative. Special thanks to all the maintainers of these awesome lists, courses, and collections for their invaluable contributions to the AI community.&lt;/p&gt;</content:encoded></item><item><title>Standing on Giants' Shoulders: The Traditional Infrastructure Powering Modern AI</title><link>https://jimmysong.io/blog/giants-beneath-ai-feet/</link><pubDate>Sun, 08 Feb 2026 08:00:00 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/giants-beneath-ai-feet/</guid><description>Before ChatGPT and TensorFlow, there was Hadoop, Kafka, and Kubernetes. This post honors the traditional open source infrastructure that became the foundation of today&amp;#39;s AI revolution.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;If I have seen further, it is by standing on the shoulders of giants.&amp;rdquo; — Isaac Newton&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/giants-beneath-ai-feet/banner.webp" data-img="https://assets.jimmysong.io/images/blog/giants-beneath-ai-feet/banner.webp" alt="Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI" data-caption="Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Standing on Giants’ Shoulders: The Traditional Infrastructure Powering Modern AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the excitement surrounding LLMs, vector databases, and AI agents, it&amp;rsquo;s easy to forget that modern AI didn&amp;rsquo;t emerge from a vacuum. Today&amp;rsquo;s AI revolution stands upon decades of infrastructure work—distributed systems, data pipelines, search engines, and orchestration platforms that were built long before &amp;ldquo;AI Native&amp;rdquo; became a buzzword.&lt;/p&gt;
&lt;p&gt;This post is a tribute to those traditional open source projects that became the invisible foundation of AI infrastructure. They&amp;rsquo;re not &amp;ldquo;AI projects&amp;rdquo; per se, but without them, the AI revolution as we know it wouldn&amp;rsquo;t exist.&lt;/p&gt;
&lt;h2 id="the-evolution-from-big-data-to-ai"&gt;The Evolution: From Big Data to AI&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Era&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;Core Technologies&lt;/th&gt;
&lt;th&gt;AI Connection&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2000s&lt;/td&gt;
&lt;td&gt;Web Search &amp;amp; Indexing&lt;/td&gt;
&lt;td&gt;Lucene, Elasticsearch&lt;/td&gt;
&lt;td&gt;Semantic search foundations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Big Data &amp;amp; Distributed Computing&lt;/td&gt;
&lt;td&gt;Hadoop, Spark, Kafka&lt;/td&gt;
&lt;td&gt;Data processing at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Cloud Native&lt;/td&gt;
&lt;td&gt;Docker, Kubernetes&lt;/td&gt;
&lt;td&gt;Model deployment platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2010s&lt;/td&gt;
&lt;td&gt;Stream Processing&lt;/td&gt;
&lt;td&gt;Flink, Storm, Pulsar&lt;/td&gt;
&lt;td&gt;Real-time ML inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2020s&lt;/td&gt;
&lt;td&gt;AI Native&lt;/td&gt;
&lt;td&gt;Transformers, Vector DBs&lt;/td&gt;
&lt;td&gt;Built on everything above&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Evolution of Data Infrastructure
&lt;/figcaption&gt;
&lt;h2 id="big-data-frameworks-the-data-engines"&gt;Big Data Frameworks: The Data Engines&lt;/h2&gt;
&lt;p&gt;Before we could train models on petabytes of data, we needed ways to store, process, and move that data.&lt;/p&gt;
&lt;h3 id="apache-hadoop-2006"&gt;Apache Hadoop (2006)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/hadoop" target="_blank" rel="noopener"&gt;https://github.com/apache/hadoop&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Hadoop democratized big data by making distributed computing accessible. Its HDFS filesystem and MapReduce paradigm proved that commodity hardware could process web-scale datasets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Modern ML training datasets live in HDFS-compatible storage&lt;/li&gt;
&lt;li&gt;Data lakes built on Hadoop became training data reservoirs&lt;/li&gt;
&lt;li&gt;Proved that distributed computing could scale horizontally&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="apache-kafka-2011"&gt;Apache Kafka (2011)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/kafka" target="_blank" rel="noopener"&gt;https://github.com/apache/kafka&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Kafka redefined data streaming with its log-based architecture. It became the nervous system for real-time data flows in enterprises worldwide.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Real-time feature pipelines for ML models&lt;/li&gt;
&lt;li&gt;Event-driven architectures for AI agent systems&lt;/li&gt;
&lt;li&gt;Streaming inference pipelines&lt;/li&gt;
&lt;li&gt;Model telemetry and monitoring backbones&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="apache-spark-2014"&gt;Apache Spark (2014)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/spark" target="_blank" rel="noopener"&gt;https://github.com/apache/spark&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Spark brought in-memory computing to big data, making iterative algorithms (like ML training) practical at scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MLlib made ML accessible to data engineers&lt;/li&gt;
&lt;li&gt;Distributed data processing for model training&lt;/li&gt;
&lt;li&gt;Spark ML became the de facto standard for big data ML&lt;/li&gt;
&lt;li&gt;Proved that in-memory computing could accelerate ML workloads&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="search-engines-the-retrieval-foundation"&gt;Search Engines: The Retrieval Foundation&lt;/h2&gt;
&lt;p&gt;Before RAG (Retrieval-Augmented Generation) became a buzzword, search engines were solving retrieval at scale.&lt;/p&gt;
&lt;h3 id="elasticsearch-2010"&gt;Elasticsearch (2010)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/elastic/elasticsearch" target="_blank" rel="noopener"&gt;https://github.com/elastic/elasticsearch&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Elasticsearch made full-text search accessible and scalable. Its distributed architecture and RESTful API became the standard for search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;pioneered distributed inverted index structures&lt;/li&gt;
&lt;li&gt;Proved that horizontal scaling was possible for search workloads&lt;/li&gt;
&lt;li&gt;Many &amp;ldquo;AI search&amp;rdquo; systems actually use Elasticsearch under the hood&lt;/li&gt;
&lt;li&gt;Query DSL influenced modern vector database query languages&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="opensearch-2021"&gt;OpenSearch (2021)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/opensearch-project/opensearch" target="_blank" rel="noopener"&gt;https://github.com/opensearch-project/opensearch&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;When AWS forked Elasticsearch, it ensured search infrastructure remained truly open. OpenSearch continues the mission of accessible, scalable search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maintains open source innovation in search&lt;/li&gt;
&lt;li&gt;Vector search capabilities added in 2023&lt;/li&gt;
&lt;li&gt;Demonstrates community fork resilience&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="databases-from-sql-to-vectors"&gt;Databases: From SQL to Vectors&lt;/h2&gt;
&lt;p&gt;The evolution from relational databases to vector databases represents a paradigm shift—but both have AI relevance.&lt;/p&gt;
&lt;h3 id="traditional-databases-that-paved-the-way"&gt;Traditional Databases That Paved the Way&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dgraph&lt;/strong&gt; (2015) - Graph database proving that specialized data structures enable new use cases&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TDengine&lt;/strong&gt; (2019) - Time-series database for IoT ML workloads&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OceanBase&lt;/strong&gt; (2021) - Distributed database showing ACID transactions could scale&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proved that specialized database engines could outperform general-purpose ones&lt;/li&gt;
&lt;li&gt;Database internals (indexing, sharding, replication) are now applied to vector databases&lt;/li&gt;
&lt;li&gt;Multi-model databases (graph + vector + relational) are becoming the norm for AI apps&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="cloud-native-the-runtime-foundation"&gt;Cloud Native: The Runtime Foundation&lt;/h2&gt;
&lt;p&gt;When Docker and Kubernetes emerged, they weren&amp;rsquo;t built for AI—but AI couldn&amp;rsquo;t scale without them.&lt;/p&gt;
&lt;h3 id="docker-2013--kubernetes-2014"&gt;Docker (2013) &amp;amp; Kubernetes (2014)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/kubernetes/kubernetes" target="_blank" rel="noopener"&gt;https://github.com/kubernetes/kubernetes&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes became the operating system for cloud-native applications. Its declarative API and controller pattern made it perfect for AI workloads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model deployment platforms (KServe, Seldon Core) run on K8s&lt;/li&gt;
&lt;li&gt;GPU orchestration (NVIDIA GPU Operator, Volcano, HAMi) extends K8s&lt;/li&gt;
&lt;li&gt;Kubeflow made K8s the standard for ML pipelines&lt;/li&gt;
&lt;li&gt;Microservice patterns enable modular AI agent architectures&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="service-mesh--serverless"&gt;Service Mesh &amp;amp; Serverless&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Istio&lt;/strong&gt; (2016), &lt;strong&gt;Knative&lt;/strong&gt; (2018) - Service mesh and serverless platforms that proved:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Network-level observability applies to AI model calls&lt;/li&gt;
&lt;li&gt;Scale-to-zero is essential for cost-effective inference&lt;/li&gt;
&lt;li&gt;Traffic splitting enables A/B testing of ML models&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Gateway patterns evolved from API gateways + service mesh&lt;/li&gt;
&lt;li&gt;Serverless inference platforms use Knative-style autoscaling&lt;/li&gt;
&lt;li&gt;Observability patterns (tracing, metrics) are now standard for ML systems&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="api-gateways-from-rest-to-llm"&gt;API Gateways: From REST to LLM&lt;/h2&gt;
&lt;p&gt;API gateways weren&amp;rsquo;t designed for AI, but they became the foundation of AI Gateway patterns.&lt;/p&gt;
&lt;h3 id="kong-apisix-kgateway"&gt;Kong, APISIX, KGateway&lt;/h3&gt;
&lt;p&gt;These API gateways solved rate limiting, auth, and routing at scale. When LLMs emerged, the same patterns applied:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Gateway Evolution&lt;/strong&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Traditional API Gateway (2010s)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Rate Limiting → Token Bucket Rate Limiting
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Auth → API Key + Organization Management
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Routing → Model Routing (GPT-4 → Claude → Local Models)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Observability → LLM-specific Telemetry (token usage, cost)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AI Gateway (2024)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Proved that centralized API management scales&lt;/li&gt;
&lt;li&gt;Plugin architectures enable LLM-specific features&lt;/li&gt;
&lt;li&gt;Traffic management patterns apply to prompt routing&lt;/li&gt;
&lt;li&gt;Security patterns (mTLS, JWT) now protect AI endpoints&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="workflow-orchestration-the-pipeline-backbone"&gt;Workflow Orchestration: The Pipeline Backbone&lt;/h2&gt;
&lt;p&gt;Data engineering needs pipelines. ML engineering needs pipelines. AI agents need workflows.&lt;/p&gt;
&lt;h3 id="apache-airflow-2015"&gt;Apache Airflow (2015)&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/apache/airflow" target="_blank" rel="noopener"&gt;https://github.com/apache/airflow&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Airflow made pipeline orchestration accessible with its DAG-based approach. It became the standard for ETL and data engineering.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ML pipeline orchestration (feature engineering, training, evaluation)&lt;/li&gt;
&lt;li&gt;Proved that DAG-based workflow definition works at scale&lt;/li&gt;
&lt;li&gt;Prompt engineering pipelines use Airflow-style orchestration&lt;/li&gt;
&lt;li&gt;Scheduler patterns are now applied to AI agent workflows&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="n8n-prefect-flyte"&gt;n8n, Prefect, Flyte&lt;/h3&gt;
&lt;p&gt;Modern workflow platforms that evolved from Airflow&amp;rsquo;s foundations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;n8n&lt;/strong&gt; (2019) - Visual workflow automation with AI capabilities&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prefect&lt;/strong&gt; (2018) - Python-native workflow orchestration for ML&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flyte&lt;/strong&gt; (2019) - Kubernetes-native workflow orchestration for ML/data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-modal agents need workflow orchestration&lt;/li&gt;
&lt;li&gt;RAG pipelines are essentially ETL pipelines for embeddings&lt;/li&gt;
&lt;li&gt;Prompt chaining is DAG-based orchestration&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="data-formats-the-lakehouse-foundation"&gt;Data Formats: The Lakehouse Foundation&lt;/h2&gt;
&lt;p&gt;Before we could train on massive datasets, we needed formats that supported ACID transactions and schema evolution.&lt;/p&gt;
&lt;h3 id="delta-lake-apache-iceberg-apache-hudi"&gt;Delta Lake, Apache Iceberg, Apache Hudi&lt;/h3&gt;
&lt;p&gt;These table formats brought reliability to data lakes:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why they matter for AI&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training datasets need versioning and reproducibility&lt;/li&gt;
&lt;li&gt;Feature stores use Delta/Iceberg as storage formats&lt;/li&gt;
&lt;li&gt;Proved that &amp;ldquo;big data&amp;rdquo; could have transactional semantics&lt;/li&gt;
&lt;li&gt;Schema evolution handles ML feature drift&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-invisible-thread-why-these-projects-matter"&gt;The Invisible Thread: Why These Projects Matter&lt;/h2&gt;
&lt;p&gt;What do all these projects have in common?&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;They solved scaling first&lt;/strong&gt; - AI training/inference needs horizontal scaling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They proved distributed systems work&lt;/strong&gt; - Modern AI is fundamentally distributed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They created ecosystem patterns&lt;/strong&gt; - Plugin systems, extension points, APIs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They established best practices&lt;/strong&gt; - Observability, security, CI/CD&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;They built developer habits&lt;/strong&gt; - YAML configs, declarative APIs, CLI tools&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="the-ai-native-continuum"&gt;The AI Native Continuum&lt;/h2&gt;
&lt;p&gt;Modern &amp;ldquo;AI Native&amp;rdquo; infrastructure didn&amp;rsquo;t replace these projects—it builds on them:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Traditional Project&lt;/th&gt;
&lt;th&gt;AI Native Evolution&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hadoop HDFS&lt;/td&gt;
&lt;td&gt;Distributed model storage&lt;/td&gt;
&lt;td&gt;HDFS for datasets, S3 for checkpoints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kafka&lt;/td&gt;
&lt;td&gt;Real-time feature pipelines&lt;/td&gt;
&lt;td&gt;Kafka → Feature Store → Model Serving&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spark ML&lt;/td&gt;
&lt;td&gt;Distributed ML training&lt;/td&gt;
&lt;td&gt;MLlib → PyTorch Distributed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Elasticsearch&lt;/td&gt;
&lt;td&gt;Vector search&lt;/td&gt;
&lt;td&gt;ES → Weaviate/Qdrant/Milvus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubernetes&lt;/td&gt;
&lt;td&gt;ML orchestration&lt;/td&gt;
&lt;td&gt;K8s → Kubeflow/KServe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Istio&lt;/td&gt;
&lt;td&gt;AI Gateway service mesh&lt;/td&gt;
&lt;td&gt;Istio → LLM Gateway with mTLS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Airflow&lt;/td&gt;
&lt;td&gt;ML pipeline orchestration&lt;/td&gt;
&lt;td&gt;Airflow → Prefect/Flyte for ML&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: From Traditional to AI Native
&lt;/figcaption&gt;
&lt;h2 id="why-were-removing-them-from-ai-resources-list"&gt;Why We&amp;rsquo;re Removing Them from AI Resources List&lt;/h2&gt;
&lt;p&gt;This post honors these projects, but we&amp;rsquo;re also removing them from our AI Resources list. Here&amp;rsquo;s why:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;They&amp;rsquo;re not &amp;ldquo;AI Projects&amp;rdquo;—they&amp;rsquo;re foundational infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hadoop, Kafka, Spark&lt;/strong&gt; are data engineering tools, not ML frameworks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Elasticsearch&lt;/strong&gt; is search, not semantic search&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kubernetes&lt;/strong&gt; is general-purpose orchestration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API gateways&lt;/strong&gt; serve REST/GraphQL, not just LLMs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;But their absence doesn&amp;rsquo;t diminish their importance.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;By removing them, we acknowledge that:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;AI has its own ecosystem&lt;/strong&gt; - Transformers, vector DBs, LLM ops&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Traditional infra has its own domain&lt;/strong&gt; - Data engineering, cloud native&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The intersection is where innovation happens&lt;/strong&gt; - AI-native data platforms, LLM ops on K8s&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="the-giants-we-stand-on"&gt;The Giants We Stand On&lt;/h2&gt;
&lt;p&gt;The next time you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deploy a model on Kubernetes&lt;/li&gt;
&lt;li&gt;Stream features through Kafka&lt;/li&gt;
&lt;li&gt;Search embeddings with a vector database&lt;/li&gt;
&lt;li&gt;Orchestrate a RAG pipeline with Prefect&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Remember: You&amp;rsquo;re standing on the shoulders of Hadoop, Kafka, Elasticsearch, Kubernetes, and countless others. They built the roads we now drive on.&lt;/p&gt;
&lt;h2 id="the-future-building-new-giants"&gt;The Future: Building New Giants&lt;/h2&gt;
&lt;p&gt;Just as Hadoop and Kafka enabled modern AI, today&amp;rsquo;s AI infrastructure will become tomorrow&amp;rsquo;s foundation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Vector databases&lt;/strong&gt; may become the new standard for all search&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM observability&lt;/strong&gt; may evolve into general distributed tracing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI agent orchestration&lt;/strong&gt; may reinvent workflow automation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU scheduling&lt;/strong&gt; may influence general-purpose resource management&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The cycle continues. The giants of today will be the foundations of tomorrow.&lt;/p&gt;
&lt;h2 id="conclusion-gratitude-and-continuity"&gt;Conclusion: Gratitude and Continuity&lt;/h2&gt;
&lt;p&gt;As we clean up our AI Resources list to focus on AI-native projects, we don&amp;rsquo;t forget where we came from. Traditional big data and cloud native infrastructure made the AI revolution possible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;To the Hadoop committers, Kafka maintainers, Kubernetes contributors, and all who built the foundation: Thank you.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Your work enabled ChatGPT, enabled Transformers, enabled everything we now call &amp;ldquo;AI.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Standing on your shoulders, we see further.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Acknowledgments&lt;/strong&gt;: This post was inspired by the need to refactor our AI Resources list. The 27 projects mentioned here are being removed—not because they&amp;rsquo;re unimportant, but because they deserve their own category: &lt;strong&gt;The Foundation&lt;/strong&gt;.&lt;/p&gt;</content:encoded></item><item><title>My First Month at Dynamia: Why AI Native Infra Is Worth the Investment</title><link>https://jimmysong.io/blog/why-i-join-dynamia-ai-native-infra/</link><pubDate>Fri, 06 Feb 2026 12:56:35 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/why-i-join-dynamia-ai-native-infra/</guid><description>Observations from my first month at Dynamia: From cloud native to AI Native Infra, why this direction is worth investing in, and the key issues and opportunities in compute governance.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Time flies—it&amp;rsquo;s already been a month since I joined Dynamia. In this article, I want to share my observations from this past month: why AI Native Infra is a direction worth investing in, and some considerations for those thinking about their own career or technical direction.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;After nearly five years of remote work, I officially joined &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; last month as VP of Open Source Ecosystem. This decision was not sudden, but a natural extension of my journey from cloud native to AI Native Infra.&lt;/p&gt;
&lt;p&gt;But this article is not just about my personal choice. I want to answer a more universal question: &lt;strong&gt;In the wave of AI infrastructure startups, why is compute governance a direction worth investing in?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For the past decade, I have worked continuously in the infrastructure space: from Kubernetes to Service Mesh, and now to AI Infra. I am increasingly convinced that the core challenge in the AI era is not &amp;ldquo;can the model run,&amp;rdquo; but &amp;ldquo;can compute resources be run efficiently, reliably, and in a controlled manner.&amp;rdquo; This conviction has only grown stronger through my observations and reflections during this first month at Dynamia.&lt;/p&gt;
&lt;p&gt;This article answers three questions: What is AI Native Infra? Why is GPU virtualization a necessity? Why did I choose Dynamia and HAMi?&lt;/p&gt;
&lt;h2 id="what-is-ai-native-infra"&gt;What Is AI Native Infra&lt;/h2&gt;
&lt;p&gt;The core of &lt;a href="https://jimmysong.io/book/ai-native-infra/"&gt;AI Native Infrastructure&lt;/a&gt; is not about adding another platform layer, but about redefining the governance target: expanding from &amp;ldquo;services and containers&amp;rdquo; to &amp;ldquo;model behaviors and compute assets.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I summarize it as three key shifts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Models as execution entities&lt;/strong&gt;: Governance now includes not just processes, but also model behaviors.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compute as a scarce asset&lt;/strong&gt;: GPU, memory, and bandwidth must be scheduled and metered precisely.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Uncertainty as the default&lt;/strong&gt;: Systems must remain observable and recoverable amid fluctuations.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In essence, AI Native Infra is about upgrading compute governance from &amp;ldquo;resource allocation&amp;rdquo; to &amp;ldquo;sustainable business capability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="why-gpu-virtualization-is-essential"&gt;Why GPU Virtualization Is Essential&lt;/h2&gt;
&lt;p&gt;Many teams focus on model inference optimization, but in production, enterprises first encounter the problem of &amp;ldquo;underutilized GPUs.&amp;rdquo; This is where GPU virtualization delivers value.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Structural idleness&lt;/strong&gt;: Small tasks monopolize large GPUs, leaving them idle for long periods.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pseudo-isolation risks&lt;/strong&gt;: Native sharing lacks hard boundaries, so a single task OOM can cause cascading failures.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling failures&lt;/strong&gt;: Some users queue for GPUs while others occupy but do not use them, leading to both shortages and idleness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fragmentation waste&lt;/strong&gt;: There may be enough total GPU, but not enough full cards, making efficient packing impossible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vendor lock-in anxiety&lt;/strong&gt;: Proprietary, tightly coupled solutions make migration costs uncontrollable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In short: GPUs must not only be allocatable, but also splittable, isolatable, schedulable, and governable.&lt;/p&gt;
&lt;h2 id="the-relationship-between-hami-and-dynamia"&gt;The Relationship Between HAMi and Dynamia&lt;/h2&gt;
&lt;p&gt;This is the most frequently asked question. Here is the shortest answer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HAMi&lt;/strong&gt;: A CNCF-hosted open source project and community focused on GPU virtualization and heterogeneous compute scheduling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamia&lt;/strong&gt;: The founding and leading company behind HAMi, providing enterprise-grade products and services based on HAMi.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open source projects are not the same as company products, but the two evolve together. HAMi drives industry adoption and technical trust, while Dynamia brings these capabilities into enterprise production environments at scale. This &amp;ldquo;dual engine&amp;rdquo; approach is what makes Dynamia unique.&lt;/p&gt;
&lt;h2 id="what-hami-provides"&gt;What HAMi Provides&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; (&lt;em&gt;Heterogeneous AI Computing Virtualization Middleware&lt;/em&gt;) delivers three key capabilities on Kubernetes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Virtualization and partitioning&lt;/strong&gt;: Split physical GPUs into logical resources on demand to improve utilization.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scheduling and topology awareness&lt;/strong&gt;: Place workloads optimally based on topology to reduce communication bottlenecks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Isolation and observability&lt;/strong&gt;: Support quotas, policies, and monitoring to reduce production risks.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Currently, HAMi has attracted over 360 contributors from 16 countries, with more than 200 enterprise end users, and its international influence continues to grow.&lt;/p&gt;
&lt;h2 id="market-trends-the-ai-infrastructure-startup-wave"&gt;Market Trends: The AI Infrastructure Startup Wave&lt;/h2&gt;
&lt;p&gt;AI infrastructure is experiencing a new wave of startups. The vLLM team&amp;rsquo;s company raised $150 million, SGLang&amp;rsquo;s commercial spin-off RadixArk is valued at $4 billion, and Databricks acquired MosaicML for $1.3 billion—all pointing to a consensus: &lt;strong&gt;Whoever helps enterprises run large models more efficiently and cost-effectively will hold the keys to next-generation AI infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Against this backdrop, &lt;strong&gt;the positioning of Dynamia and HAMi&lt;/strong&gt; is even clearer. Many teams focus on &amp;ldquo;model performance acceleration&amp;rdquo; and &amp;ldquo;inference optimization&amp;rdquo; (like vLLM, SGLang), while we focus on &lt;strong&gt;&amp;ldquo;resource scheduling and virtualization&amp;rdquo;&lt;/strong&gt;—enabling better orchestration of existing accelerated hardware resources.&lt;/p&gt;
&lt;p&gt;The two are complementary: the former makes individual models run faster and cheaper, while the latter ensures that compute allocation at the cluster level is efficient, fair, and controllable. This is similar to extending Kubernetes&amp;rsquo; CPU/memory scheduling philosophy to GPU and heterogeneous compute management in the AI era.&lt;/p&gt;
&lt;h2 id="why-ai-native-infra-is-worth-the-investment"&gt;Why AI Native Infra Is Worth the Investment&lt;/h2&gt;
&lt;p&gt;My observations this month have convinced me that &lt;strong&gt;compute governance is the most undervalued yet most promising area in AI infrastructure&lt;/strong&gt;. If you are considering a career or technical investment, here is my assessment:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, this is a real and urgent pain point&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Model training and inference optimization attract a lot of attention, but in production, enterprises first encounter the problem of &amp;ldquo;underutilized GPUs&amp;rdquo;—structural idleness, scheduling failures, fragmentation waste, and vendor lock-in anxiety. Without solving these problems, even the fastest models cannot scale in production. GPU virtualization and heterogeneous compute scheduling are the &amp;ldquo;infrastructure below infrastructure&amp;rdquo; for enterprise AI transformation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, this is a clear long-term track&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Frameworks like vLLM and SGLang emerge constantly, making individual models run faster. But who ensures that compute allocation at the cluster level is efficient, fair, and controllable? This is similar to extending Kubernetes&amp;rsquo; success in CPU/memory scheduling to GPU and heterogeneous compute management in the AI era. This is not something that can be finished in a year or two, but a direction for continuous construction over the next five to ten years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Third, this is an open and verifiable path&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Dynamia chose to build on HAMi as an open source foundation, first solving general capabilities, then supporting enterprise adoption. This means the technical direction is transparent and verifiable in the community. You can form your own judgment by participating in open source, observing adoption, and evaluating the ecosystem—rather than relying on the black-box promises of proprietary solutions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fourth, this is a window of opportunity that is opening now&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;AI infrastructure is being redefined. Investing in its construction today will continue to yield value in the coming years. The vLLM team&amp;rsquo;s company raised $150 million, SGLang&amp;rsquo;s commercial spin-off RadixArk is valued at $4 billion, Databricks acquired MosaicML for $1.3 billion—all validating the same trend: &lt;strong&gt;Whoever helps enterprises run large models more efficiently will hold the keys to next-generation AI infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I hope to bring my experience in cloud native and open source communities to the next stage of HAMi and Dynamia: turning GPU resources from a &amp;ldquo;cost center&amp;rdquo; into an &amp;ldquo;operational asset.&amp;rdquo; This is not just my career choice, but my judgment and investment in the direction of next-generation infrastructure.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Join the HAMi Community
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Add me on WeChat (&lt;code&gt;jimmysong&lt;/code&gt;) to join the &lt;a href="https://github.com/project-hami/hami" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; community focused on GPU virtualization and heterogeneous compute scheduling.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;If you are also interested in HAMi, GPU virtualization, AI Native Infra, or Dynamia, feel free to reach out.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;From cloud native to AI Native Infra, my observations this month have only strengthened my conviction: &lt;strong&gt;The true upper limit of AI applications is determined by the infrastructure&amp;rsquo;s ability to govern compute resources.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;HAMi addresses the fundamental issues of GPU virtualization and heterogeneous compute scheduling, while Dynamia is driving these capabilities into large-scale production. If you are also looking for a technical direction worth long-term investment, AI Native Infra—especially compute governance and scheduling—is a track with real pain points, a clear path, an open ecosystem, and an opening window of opportunity.&lt;/p&gt;
&lt;p&gt;Joining Dynamia is not just a career choice, but a commitment to building the next generation of infrastructure. I hope the observations and reflections in this article can provide some reference for you as you evaluate technical directions and career opportunities.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;If you are also interested in HAMi, GPU virtualization, AI Native Infra, or Dynamia, feel free to reach out.&lt;/em&gt;&lt;/p&gt;</content:encoded></item><item><title>The True Inflection Point of ADD: When Spec Becomes the Core Asset of AI-Era Software</title><link>https://jimmysong.io/blog/add-inflection-point-spec-as-core-asset/</link><pubDate>Tue, 20 Jan 2026 07:51:36 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/add-inflection-point-spec-as-core-asset/</guid><description>Exploring how Spec becomes the governable core asset in Agent-Driven Development (ADD) and the trend toward control-plane engineering systems.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The role of Spec is undergoing a fundamental transformation, becoming the governance anchor of engineering systems in the AI era.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="the-essence-of-software-engineering-and-the-cost-structure-shift-brought-by-ai"&gt;The Essence of Software Engineering and the Cost Structure Shift Brought by AI&lt;/h2&gt;
&lt;p&gt;From first principles, software engineering has always been about one thing: &lt;strong&gt;stably, controllably, and reproducibly transforming human intent into executable systems.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Artificial Intelligence (AI) does not change this engineering essence, but it dramatically alters the cost structure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Implementation costs plummet:&lt;/strong&gt; Code, tests, and boilerplate logic are rapidly commoditized.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consistency costs rise sharply:&lt;/strong&gt; Intent drift, hidden conflicts, and cross-module inconsistencies become more frequent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance costs are amplified:&lt;/strong&gt; As agents can act directly, auditability, accountability, and explainability become hard constraints.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, in the era of Agent-Driven Development (ADD), the core issue is not &amp;ldquo;can agents do the work,&amp;rdquo; but how to maintain controllability and intent preservation in engineering systems under highly autonomous agents.&lt;/p&gt;
&lt;h2 id="the-add-era-inflection-point-three-structural-preconditions"&gt;The ADD Era Inflection Point: Three Structural Preconditions&lt;/h2&gt;
&lt;p&gt;Many attribute the &amp;ldquo;explosion&amp;rdquo; of ADD to more mature multi-agent systems, stronger models, or more automated tools. In reality, the true structural inflection point arises only when these three conditions are met:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agents have acquired multi-step execution capabilities&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;With frameworks like LangChain, LangGraph, and CrewAI, agents are no longer just prompt invocations, but long-lived entities capable of planning, decomposition, execution, and rollback.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agents are entering real enterprise delivery pipelines&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Once in enterprise R&amp;amp;D, the question shifts from &amp;ldquo;can it generate&amp;rdquo; to &amp;ldquo;who approved it, is it compliant, can it be rolled back.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traditional engineering tools lack a control plane for the agent era&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Tools like Git, CI, and Issue Trackers were designed for &amp;ldquo;human developer collaboration,&amp;rdquo; not for &amp;ldquo;agent execution.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;When these three factors converge, ADD inevitably shifts from an &amp;ldquo;efficiency tool&amp;rdquo; to a &amp;ldquo;governance system.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-changing-role-of-spec-from-documentation-to-system-constraint"&gt;The Changing Role of Spec: From Documentation to System Constraint&lt;/h2&gt;
&lt;p&gt;In the context of ADD, Spec is undergoing a fundamental shift:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Spec is no longer &amp;ldquo;documentation for humans,&amp;rdquo; but &amp;ldquo;the source of constraints and facts for systems and agents to execute.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Spec now serves at least three roles:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Verifiable expression of intent and boundaries&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Requirements, acceptance criteria, and design principles are no longer just text, but objects that can be checked, aligned, and traced.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stable contracts for organizational collaboration&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When agents participate in delivery, verbal consensus and tacit knowledge quickly fail. Versioned, auditable artifacts become the foundation of collaboration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy surface for agent execution&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Agents can write code, modify configurations, and trigger pipelines. Spec must become the constraint on &amp;ldquo;what can and cannot be done.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;From this perspective, the status of Spec is approaching that of the &lt;strong&gt;Control Plane&lt;/strong&gt; in AI-native infrastructure.&lt;/p&gt;
&lt;h2 id="the-reality-of-multi-agent-workflows-orchestration-and-governance-first"&gt;The Reality of Multi-Agent Workflows: Orchestration and Governance First&lt;/h2&gt;
&lt;p&gt;In recent systems (such as &lt;a href="https://apoxai.com" target="_blank" rel="noopener"&gt;APOX&lt;/a&gt; and other enterprise products), an industry consensus is emerging:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-agent collaboration no longer pursues &amp;ldquo;full automation,&amp;rdquo; but is staged and gated.&lt;/li&gt;
&lt;li&gt;Frameworks like LangGraph are used to build persistent, debuggable agent workflows.&lt;/li&gt;
&lt;li&gt;RAG (e.g., based on Milvus) is used to accumulate historical Specs, decisions, and context as long-term memory.&lt;/li&gt;
&lt;li&gt;The IDE mainly focuses on execution efficiency, not engineering governance.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/add-inflection-point-spec-as-core-asset/apox.webp" data-img="https://assets.jimmysong.io/images/blog/add-inflection-point-spec-as-core-asset/apox.webp" alt="Figure 1: APOX user interface" data-caption="Figure 1: APOX user interface"
width="1400"
height="1045"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: APOX user interface&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;APOX (AI Product Orchestration eXtended) is a multi-agent collaboration workflow platform for enterprise software delivery. Its core goals are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;To connect the entire process from product requirements to executable code with a governable Agentflow and explicit engineering artifact chain.&lt;/li&gt;
&lt;li&gt;To assign dedicated AI agents to each delivery stage (such as PRD, PO, Architecture, Developer, Implementation, Coding, etc.).&lt;/li&gt;
&lt;li&gt;To embed manual approval gates and full audit trails at every step, solving the &amp;ldquo;intent drift and consistency&amp;rdquo; governance problem that traditional AI coding tools cannot address.&lt;/li&gt;
&lt;li&gt;The platform provides a VS Code plugin for real-time sync between local IDE and web artifacts, allowing Specs, code, tasks, and approval statuses to coexist in the repository.&lt;/li&gt;
&lt;li&gt;Supports assigning different base models to different agents according to enterprise needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;APOX is not about simply speeding up code generation, but about elevating &amp;ldquo;Spec&amp;rdquo; from auxiliary documentation to a verifiable, constrainable, and traceable core asset in engineering—building a control plane and workflow governance system suitable for Agent-Driven Development.&lt;/p&gt;
&lt;p&gt;Such systems emphasize:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An explicit artifact chain from PRD → Spec → Task → Implementation.&lt;/li&gt;
&lt;li&gt;Manual confirmation and audit points at every stage.&lt;/li&gt;
&lt;li&gt;Bidirectional sync between Spec, code, repository, and IDE.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is not about &amp;ldquo;smarter AI,&amp;rdquo; but about engineering systems adapting to the agent era.&lt;/p&gt;
&lt;h2 id="the-long-term-value-of-spec-the-core-anchor-of-engineering-assets"&gt;The Long-Term Value of Spec: The Core Anchor of Engineering Assets&lt;/h2&gt;
&lt;p&gt;This is not to devalue code, but to acknowledge reality:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;There will always be long-term differentiation in algorithms and model capabilities.&lt;/li&gt;
&lt;li&gt;General engineering implementation is rapidly homogenizing.&lt;/li&gt;
&lt;li&gt;What is hard to replicate is: how to define problems, constrain systems, and govern change.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the ADD era, the value of Spec is reflected in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Determining what agents can and cannot do.&lt;/li&gt;
&lt;li&gt;Carrying the organization&amp;rsquo;s long-term understanding of the system.&lt;/li&gt;
&lt;li&gt;Serving as the anchor for audit, compliance, and accountability.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Code will be rewritten again and again; Spec is the long-term asset.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="risks-and-challenges-of-add-living-spec-and-governance-constraints"&gt;Risks and Challenges of ADD: Living Spec and Governance Constraints&lt;/h2&gt;
&lt;p&gt;ADD also faces significant risks:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can Spec become a Living Spec&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That is, when key implementation changes occur, can the system detect &amp;ldquo;intent changes&amp;rdquo; and prompt Spec updates, rather than allowing silent drift?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can governance achieve low friction but strong constraints&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If gates are too strict, teams will bypass them; if too loose, the system loses control.&lt;/p&gt;
&lt;p&gt;These two factors determine whether ADD is &amp;ldquo;the next engineering paradigm&amp;rdquo; or &amp;ldquo;just another tool bubble.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-trend-toward-control-planes-in-engineering-systems"&gt;The Trend Toward Control Planes in Engineering Systems&lt;/h2&gt;
&lt;p&gt;From a broader perspective, ADD is the inevitable result of engineering systems becoming &amp;ldquo;control planes&amp;rdquo;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Engineering systems are evolving from &amp;ldquo;human collaboration tools&amp;rdquo; to &amp;ldquo;control systems for agent execution.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In this structure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Agent / IDE is the &lt;strong&gt;execution plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;RAG / Memory is the &lt;strong&gt;state and memory plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Spec is the intent and policy plane&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Gates, audit, and traceability form the governance loop.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This closely aligns with the evolution path of AI-native infrastructure.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The winners of the ADD era will not be the systems with &amp;ldquo;the most agents or the fastest generation,&amp;rdquo; but those that first upgrade Spec from documentation to a governable, auditable, and executable asset. As automation advances, the true scarcity is the long-term control of intent.&lt;/p&gt;</content:encoded></item><item><title>AI Voice Dictation Input Methods Are Becoming the New Shortcut Key for the Programming Era</title><link>https://jimmysong.io/blog/ai-voice-dictation-input-method-comparison/</link><pubDate>Sun, 18 Jan 2026 06:53:08 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-voice-dictation-input-method-comparison/</guid><description>Comparing Miaoyan, Zhipu, and Shandianshuo voice input methods for developers: speed, stability, command capabilities, and cost models.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Voice input methods are not just about being &amp;ldquo;fast&amp;rdquo;—they are becoming a brand new gateway for developers to collaborate with AI.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
Warning
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
On January 12, 2026, due to financial difficulties encountered during operations, the Miaoyan project announced the cessation of operations and the team was disbanded. The application will no longer be updated or maintained, but existing versions can continue to be used on the current device and system, and do not store any audio or transcription content.
&lt;/div&gt;
&lt;/div&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/banner.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/banner.webp" alt="Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience." data-caption="Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience."
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Can voice input become the new shortcut for developers? My in-depth comparison experience.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="ai-voice-input-methods-are-becoming-the-new-shortcut-key-in-the-programming-era"&gt;AI Voice Input Methods Are Becoming the &amp;ldquo;New Shortcut Key&amp;rdquo; in the Programming Era&lt;/h2&gt;
&lt;p&gt;I am increasingly convinced of one thing: &lt;strong&gt;PC-based AI voice input methods are evolving from mere &amp;ldquo;input tools&amp;rdquo; into the foundational interaction layer for the era of programming and AI collaboration.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not just about typing faster—it determines how you deliver your &lt;strong&gt;intent&lt;/strong&gt; to the system, whether you&amp;rsquo;re writing documentation, code, or collaborating with AI in IDEs, terminals, or chat windows.&lt;/p&gt;
&lt;p&gt;Because of this, the differences in voice input method experiences are far more significant than they appear on the surface.&lt;/p&gt;
&lt;h2 id="my-six-evaluation-criteria-for-ai-voice-input-methods"&gt;My Six Evaluation Criteria for AI Voice Input Methods&lt;/h2&gt;
&lt;p&gt;After long-term, high-frequency use, I have developed a set of criteria to assess the real-world performance of AI voice input methods:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Response speed&lt;/strong&gt;: Does text appear quickly enough after pressing the shortcut to keep up with your thoughts?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Continuous input stability&lt;/strong&gt;: Does it remain reliable during extended use, or does it suddenly fail or miss recognition?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mixed Chinese-English and technical terms&lt;/strong&gt;: Can it reliably handle code, paths, abbreviations, and product names?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Developer friendliness&lt;/strong&gt;: Is it truly designed for command line, IDE, and automation scenarios?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interaction restraint&lt;/strong&gt;: Does it avoid introducing distracting features that interfere with input itself?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Subscription and cost structure&lt;/strong&gt;: Is it a standalone paid product, or can it be bundled with existing tool subscriptions?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Based on these criteria, I focused on comparing &lt;strong&gt;Miaoyan&lt;/strong&gt;, &lt;strong&gt;Shandianshuo&lt;/strong&gt;, and &lt;strong&gt;Zhipu AI Voice Input Method&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="miaoyan-currently-the-most-developer-oriented-domestic-product"&gt;Miaoyan: Currently the Most &amp;ldquo;Developer-Oriented&amp;rdquo; Domestic Product&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://miaoyan.cn" target="_blank" rel="noopener"&gt;Miaoyan&lt;/a&gt; was the first domestic AI voice input method I used extensively, and it remains the one I am most willing to use continuously.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/miaoyan.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/miaoyan.webp" alt="Figure 2: Miaoyan is currently my most-used Mac voice input method." data-caption="Figure 2: Miaoyan is currently my most-used Mac voice input method."
width="2272"
height="1624"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Miaoyan is currently my most-used Mac voice input method.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="command-mode-the-key-differentiator-for-developer-productivity"&gt;Command Mode: The Key Differentiator for Developer Productivity&lt;/h3&gt;
&lt;p&gt;It&amp;rsquo;s important to clarify that &lt;strong&gt;Miaoyan&amp;rsquo;s command mode is not about editing text via voice&lt;/strong&gt;. Instead:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You describe your need in natural language, and the system directly generates an &lt;strong&gt;executable command-line command&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is crucial for developers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It&amp;rsquo;s not just about input&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s about turning voice into an automation entry point&lt;/li&gt;
&lt;li&gt;Essentially, it connects voice to the CLI or toolchain&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This design is clearly focused on &lt;strong&gt;engineering efficiency&lt;/strong&gt;, not office document polishing.&lt;/p&gt;
&lt;h3 id="usage-experience-summary"&gt;Usage Experience Summary&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Fast response, nearly instant&lt;/li&gt;
&lt;li&gt;Output is relatively clean, with minimal guessing&lt;/li&gt;
&lt;li&gt;Interaction design is restrained, with no unnecessary concepts&lt;/li&gt;
&lt;li&gt;Developer-friendly mindset&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But there are some practical limitations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It is a &lt;strong&gt;completely standalone product&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Requires a separate subscription&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Still in relatively small-scale use&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a product strategy perspective, it feels more like a &amp;ldquo;pure tool&amp;rdquo; than part of an ecosystem.&lt;/p&gt;
&lt;div class="alert alert-warning-container"&gt;
&lt;div class="alert-warning-title px-2"&gt;
Note
&lt;/div&gt;
&lt;div class="alert-warning px-2"&gt;
On January 12, 2026, due to financial difficulties encountered during operations, the Miaoyan project announced the cessation of operations and the team was disbanded. The application will no longer be updated or maintained, but existing versions can continue to be used on the current device and system, and do not store any audio or transcription content.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="shandianshuo-local-first-approach-developer-experience-depends-on-your-setup"&gt;Shandianshuo: Local-First Approach, Developer Experience Depends on Your Setup&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://shandianshuo.cn" target="_blank" rel="noopener"&gt;Shandianshuo&lt;/a&gt; takes a different approach: it treats voice input as a &amp;ldquo;local-first foundational capability,&amp;rdquo; emphasizing low latency and privacy (at least in its product narrative). The natural advantages of this approach are speed and controllable marginal costs, making it suitable as a &amp;ldquo;system capability&amp;rdquo; that&amp;rsquo;s always available, rather than a cloud service.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/shandianshuo.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/shandianshuo.webp" alt="Figure 3: Shandianshuo settings page" data-caption="Figure 3: Shandianshuo settings page"
width="2556"
height="2080"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Shandianshuo settings page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;However, from a developer&amp;rsquo;s perspective, its upper limit often depends on &amp;ldquo;how you implement enhanced capabilities&amp;rdquo;:&lt;/p&gt;
&lt;p&gt;If you only use it for basic transcription, the experience is more like a high-quality local input tool. But if you want better mixed Chinese-English input, technical term correction, symbol and formatting handling, the common approach is to add optional AI correction/enhancement capabilities, which usually requires extra configuration (such as providing your own API key or subscribing to enhanced features). The key trade-off here is not &amp;ldquo;can it be used,&amp;rdquo; but &amp;ldquo;how much configuration cost are you willing to pay for enhanced capabilities.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;If you want voice input to be a &amp;ldquo;lightweight, stable, non-intrusive&amp;rdquo; foundation, Shandianshuo is worth considering. But if your goal is to make voice input part of your developer workflow (such as command generation or executable actions), it needs to offer stronger productized design at the &amp;ldquo;command layer&amp;rdquo; and in terms of controllability.&lt;/p&gt;
&lt;h2 id="zhipu-ai-voice-input-method-stable-but-with-friction"&gt;Zhipu AI Voice Input Method: Stable but with Friction&lt;/h2&gt;
&lt;p&gt;I also thoroughly tested the &lt;a href="https://autoglm.zhipuai.cn/autotyper/" target="_blank" rel="noopener"&gt;Zhipu AI Voice Input Method&lt;/a&gt;.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/autoglm.webp" data-img="https://assets.jimmysong.io/images/blog/ai-voice-dictation-input-method-comparison/autoglm.webp" alt="Figure 4: Zhipu Voice Input Method settings interface" data-caption="Figure 4: Zhipu Voice Input Method settings interface"
width="2430"
height="1824"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Zhipu Voice Input Method settings interface&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Its strengths include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;More stable for long-term continuous input&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Rarely becomes completely unresponsive&lt;/li&gt;
&lt;li&gt;Good tolerance for longer Chinese input&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But with frequent use, some issues stand out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Idle misrecognition&lt;/strong&gt;: If you press the shortcut but don&amp;rsquo;t speak, it may output random characters, disrupting your input flow&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Occasionally messy output&lt;/strong&gt;: Sometimes adds irrelevant words, making it less controllable than Miaoyan&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Basic recognition errors&lt;/strong&gt;: For example, &amp;ldquo;Zhipu&amp;rdquo; being recognized as &amp;ldquo;Zhipu&amp;rdquo; (with a different character), which is a trust issue for professional users&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Feature-heavy design&lt;/strong&gt;: Various tone and style features increase cognitive load&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="subscription-bundling-zhipus-practical-advantage"&gt;Subscription Bundling: Zhipu&amp;rsquo;s Practical Advantage&lt;/h2&gt;
&lt;p&gt;Although I prefer Miaoyan in terms of experience, &lt;strong&gt;Zhipu has a very practical advantage&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;If you already subscribe to Zhipu&amp;rsquo;s programming package, &lt;strong&gt;the voice input method is included for free&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No need to pay separately for the input method&lt;/li&gt;
&lt;li&gt;Lower psychological and decision-making cost&lt;/li&gt;
&lt;li&gt;More likely to become the &amp;ldquo;default tool&amp;rdquo; that stays&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From a business perspective, this is a very smart strategy.&lt;/p&gt;
&lt;h2 id="main-comparison-table"&gt;Main Comparison Table&lt;/h2&gt;
&lt;p&gt;The following table compares the three products across key dimensions for quick reference.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Miaoyan&lt;/th&gt;
&lt;th&gt;Shandianshuo&lt;/th&gt;
&lt;th&gt;Zhipu AI Voice Input Method&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Response Speed&lt;/td&gt;
&lt;td&gt;Fast, nearly instant&lt;/td&gt;
&lt;td&gt;Usually fast (local-first)&lt;/td&gt;
&lt;td&gt;Slightly slower than Miaoyan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous Stability&lt;/td&gt;
&lt;td&gt;Stable&lt;/td&gt;
&lt;td&gt;Depends on setup and environment&lt;/td&gt;
&lt;td&gt;Very stable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idle Misrecognition&lt;/td&gt;
&lt;td&gt;Rare&lt;/td&gt;
&lt;td&gt;Generally restrained (varies by version)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Obvious: outputs characters even if silent&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output Cleanliness/Control&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;More like an &amp;ldquo;input tool&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Occasionally messy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Differentiator&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Natural language → executable command&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Local-first / optional enhancements&lt;/td&gt;
&lt;td&gt;Ecosystem-attached capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subscription &amp;amp; Cost&lt;/td&gt;
&lt;td&gt;Standalone, separate purchase&lt;/td&gt;
&lt;td&gt;Basic usable; enhancements often require setup/subscription&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Bundled free with programming package&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My Current Preference&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Best experience&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;More like a &amp;ldquo;foundation approach&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Easy to keep but not clean enough&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Core Comparison of Miaoyan, Shandianshuo, and Zhipu AI Voice Input Methods
&lt;/figcaption&gt;
&lt;h2 id="user-loyalty-to-ai-voice-input-methods"&gt;User Loyalty to AI Voice Input Methods&lt;/h2&gt;
&lt;p&gt;The switching cost for voice input methods is actually low: just a shortcut key and a habit of output.&lt;/p&gt;
&lt;p&gt;What really determines whether users stick around is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether the output is controllable&lt;/li&gt;
&lt;li&gt;Whether it keeps causing annoying minor issues&lt;/li&gt;
&lt;li&gt;Whether it integrates into your existing workflow and payment structure&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For me personally:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The best and smoothest experience is still Miaoyan&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The one most likely to stick around is probably Zhipu&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shandianshuo is more of a &amp;ldquo;foundation approach&amp;rdquo; and worth watching for how its enhancements evolve&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These points are not contradictory.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Miaoyan is more mature in &lt;strong&gt;engineering orientation, command capabilities, and input control&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Zhipu has practical advantages in &lt;strong&gt;stability and subscription bundling&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Shandianshuo takes a &lt;strong&gt;local-first + optional enhancement&lt;/strong&gt; approach, with the key being how it balances &amp;ldquo;basic capability&amp;rdquo; and &amp;ldquo;enhancement cost&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Who truly becomes the &amp;ldquo;default gateway&amp;rdquo; depends on reducing distractions, fixing frequent minor issues, and treating voice input as true &amp;ldquo;infrastructure&amp;rdquo; rather than an add-on feature&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The competition among AI voice input methods is no longer about recognition accuracy, but about who can own the shortcut key you press every day.&lt;/strong&gt;&lt;/p&gt;</content:encoded></item><item><title>From Spatial Data to AI Open Source: Technical Standards, Data Sovereignty, and the Global Divide</title><link>https://jimmysong.io/blog/spatial-data-ai-open-source-standards-sovereignty/</link><pubDate>Sun, 11 Jan 2026 03:29:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/spatial-data-ai-open-source-standards-sovereignty/</guid><description>How technical standards and data sovereignty shape AI open source paths and infrastructure competition in the global AI era.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The divide in technical standards and data sovereignty determines the global competitive landscape of infrastructure open source in the AI era.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In this article, I will use the differences in air quality data presentation in Apple Maps and Weather as a starting point to explore how technical standards and data sovereignty influence the open source paths of AI in different countries. I will further analyze why, in the AI era, infrastructure-level open source has become the key battleground for ecosystem dominance.&lt;/p&gt;
&lt;h2 id="authors-note"&gt;Author&amp;rsquo;s Note&lt;/h2&gt;
&lt;p&gt;This article originates from a very everyday observation: Why is air quality data in China shown as &amp;ldquo;points&amp;rdquo; in Apple Maps and Weather, while in other countries it is often displayed as &amp;ldquo;areas&amp;rdquo;?&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/aqi-map.webp" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/aqi-map.webp" alt="Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries" data-caption="Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries"
width="1650"
height="1864"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Air quality map in Apple Weather, showing point-based data in China and area-based data in other countries&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;At first glance, it seems like a product experience difference. But when I reconsidered this issue in the context of engineering, standards, and system design, I realized it actually points to a much bigger question: how different countries understand the relationship between technology, standards, openness, and sovereignty.&lt;/p&gt;
&lt;p&gt;As an engineer who has long worked in cloud native, AI infrastructure, and open source ecosystems, I gradually realized that this difference is not limited to air quality or map data. In the AI era, it is further amplified, directly affecting how we open source models, build infrastructure, and whether we can participate in the formulation of global rules.&lt;/p&gt;
&lt;p&gt;Writing this article is not about judging right or wrong, but about using a concrete example to explain a structural difference and discuss the long-term impact and real opportunities this difference may bring in the AI era.&lt;/p&gt;
&lt;p&gt;What is especially important: at the level of AI infrastructure and infra-level open source, the competition has just begun. China is not without opportunities, but the choice of path will become more critical than ever.&lt;/p&gt;
&lt;h2 id="differences-in-air-quality-data-presentation-a-microcosm-of-technical-standards-and-sovereignty"&gt;Differences in Air Quality Data Presentation: A Microcosm of Technical Standards and Sovereignty&lt;/h2&gt;
&lt;p&gt;The following image illustrates the divide between spatial data, AI open source, and technical standards. By comparing how air quality data is presented in Apple Maps and Weather in different countries, you can intuitively feel the differences in technical standards and sovereignty strategies behind the scenes.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/banner.webp" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/banner.webp" alt="Figure 2: The divide between spatial data, AI open source, and technical standards" data-caption="Figure 2: The divide between spatial data, AI open source, and technical standards"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: The divide between spatial data, AI open source, and technical standards&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;If you regularly use global products such as maps, weather, traffic, or various data services, you may notice a recurring phenomenon that is rarely discussed seriously: the way data is presented in China often differs significantly from global mainstream standards.&lt;/p&gt;
&lt;p&gt;A very intuitive example comes from the air quality display in Apple Maps or Weather. In China, air quality is usually shown as discrete points; in the US, Europe, Japan, and other countries, it is often rendered as continuous coverage areas.&lt;/p&gt;
&lt;p&gt;At first glance, this seems like a product experience difference, and may even lead people to mistakenly believe that &amp;ldquo;China&amp;rsquo;s data is incomplete.&amp;rdquo; But if you treat it as an engineering or system design issue, you will find: this is not a matter of data capability, but a different choice in technical standards, data sovereignty, and openness strategies.&lt;/p&gt;
&lt;p&gt;And this choice is not limited to air quality.&lt;/p&gt;
&lt;h2 id="air-quality-is-just-a-slice-greater-differences-in-spatial-public-data"&gt;Air Quality Is Just a Slice: Greater Differences in Spatial Public Data&lt;/h2&gt;
&lt;p&gt;Air quality is just a highly visible and relatively low-risk example. Similar differences have long existed in broader spatial and public data domains.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maps and coordinate systems&lt;/li&gt;
&lt;li&gt;Surveying and high-precision spatial data&lt;/li&gt;
&lt;li&gt;Real-time traffic and population movement&lt;/li&gt;
&lt;li&gt;Remote sensing, environmental, and urban operation data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In global mainstream systems, such data is usually regarded as public information infrastructure. It is standardized, gridded, API-ified, allows interpolation, modeling, and redistribution, and is widely used in research, business, and product innovation.&lt;/p&gt;
&lt;p&gt;In China, this data often takes another form: hierarchical, discrete, strictly defined, and with centralized interpretation authority.&lt;/p&gt;
&lt;p&gt;This is not a technical preference in a single field, but a systemic logic of technology and governance.&lt;/p&gt;
&lt;h2 id="three-global-paths"&gt;Three Global Paths&lt;/h2&gt;
&lt;p&gt;Placing China in a global context, we can see that there are roughly three different paths worldwide regarding &amp;ldquo;how public data and technical standards are opened.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Engineering-Open Type: Standards and Ecosystem First&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Represented by the US and some European countries, the core features of this system are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Public data prioritized as infrastructure&lt;/li&gt;
&lt;li&gt;Standards and interfaces come first&lt;/li&gt;
&lt;li&gt;Encourages engineering autonomy and ecosystem evolution&lt;/li&gt;
&lt;li&gt;Tolerates model inference and uncertainty&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This path directly shaped the global landscape of foundational software and infrastructure-level open source. Linux, Kubernetes, and the cloud native system are essentially products of openness at the rules layer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Governance-Sovereignty Type: Control and Auditability First&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Represented by China, this path emphasizes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sensitivity of spatial and public data&lt;/li&gt;
&lt;li&gt;Data as part of governance capability&lt;/li&gt;
&lt;li&gt;Standards, definitions, and release methods are highly bound&lt;/li&gt;
&lt;li&gt;Emphasizes traceability, accountability, and controllability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this system, &amp;ldquo;point data&amp;rdquo; is not a sign of technological backwardness, but a governable technical form. When a technical system is designed as a governance system, its primary goal is not reusability, but controllability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compromise-Coordinated Type: Cautious Openness, Engineering Internationalization&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Some countries try to find a balance between the two, maintaining caution in spatial data while being highly internationalized in engineering and industry. This shows that the difference is not about being advanced or backward, but about different objective functions.&lt;/p&gt;
&lt;p&gt;The following diagram compares the core characteristics, typical cases, and advantages/challenges of these three paths from a global perspective. The &amp;ldquo;Engineering-Open Type&amp;rdquo; on the left shapes the global infrastructure software landscape through standards and ecosystems; the &amp;ldquo;Governance-Sovereignty Type&amp;rdquo; in the middle emphasizes data sovereignty and security controllability but has limitations in influence at the rules layer; the &amp;ldquo;Compromise-Coordinated Type&amp;rdquo; on the right attempts to find a balance between security and openness. The divide between these three paths directly affects the infrastructure competition landscape of various countries in the AI era.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/global-three-paths-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/global-three-paths-en.svg" alt="Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards" data-caption="Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards"
width="2663"
height="1862"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Global Perspective: Three Paths for Public Data and Technical Standards&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-essence-of-point-vs-area-in-air-quality"&gt;The Essence of &amp;ldquo;Point&amp;rdquo; vs. &amp;ldquo;Area&amp;rdquo; in Air Quality&lt;/h2&gt;
&lt;p&gt;Among all spatial public data, air quality is an ideal observation window:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Does not directly involve military or core economic security&lt;/li&gt;
&lt;li&gt;Highly visible, updated daily, and perceptible to everyone&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;China does not lack air quality data; on the contrary, the density of monitoring stations is among the highest in the world. The real difference lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether interpolation is allowed&lt;/li&gt;
&lt;li&gt;Whether model inference is allowed&lt;/li&gt;
&lt;li&gt;Whether platforms are allowed to reinterpret the data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;ldquo;Point&amp;rdquo; means authenticity and traceability; &amp;ldquo;area&amp;rdquo; means models, inference, and redistribution of interpretive authority. This is precisely the watershed between technical standards and data sovereignty.&lt;/p&gt;
&lt;p&gt;The following diagram compares two different technical paths. The left side, &amp;ldquo;Governance-Sovereignty Type,&amp;rdquo; emphasizes data traceability and controllability, using discrete point-based data presentation. The right side, &amp;ldquo;Engineering-Open Type,&amp;rdquo; allows model interpolation and inference, providing more user-friendly experience through continuous area-based coverage. The essence of this difference lies not in the level of technical capability, but in the different choices made between data sovereignty, governance capability, and open ecosystems.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/data-sovereignty-comparison-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/data-sovereignty-comparison-en.svg" alt="Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation" data-caption="Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation"
width="2263"
height="1562"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Technical Standards and Sovereignty Divide in Spatial Data Presentation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-amplification-effect-in-the-ai-era"&gt;The Amplification Effect in the AI Era&lt;/h2&gt;
&lt;p&gt;With the above logic in mind, many phenomena in the AI era become less confusing.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Why are Chinese AI companies more willing to open source large language model (LLM) weights, while American companies have clearly shifted toward closed source in recent years?&lt;/li&gt;
&lt;li&gt;Why is foundational software and infrastructure-level open source still mainly led by the US?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key is not &amp;ldquo;whether to open source,&amp;rdquo; but &amp;ldquo;which layer is open sourced.&amp;rdquo;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model weights are static, declarable assets&lt;/li&gt;
&lt;li&gt;Infrastructure, runtimes, protocols, and standards are dynamic, evolving system rules&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open sourcing weights is essentially openness at the asset layer; infrastructure-level open source means relinquishing control over operating rules and interpretive authority.&lt;/p&gt;
&lt;p&gt;The following diagram compares two different layers of AI open source. The left side shows &amp;ldquo;Model Weight Layer Open Source,&amp;rdquo; which is a typical feature of Chinese path—opening static digital assets with low cost and controllable risk, but not involving rule-making. The right side shows &amp;ldquo;Infrastructure Layer Open Source,&amp;rdquo; which is a core strategy of US path—by open sourcing development tools, protocol standards, runtimes, and compute scheduling and other infrastructure, defining how AI is used, thereby mastering ecosystem rules and interpretive authority. Key insight: Open sourcing model weights does not equal mastering AI ecosystem, and the real competitive focus is shifting to the infrastructure layer of &amp;ldquo;how AI runs.&amp;rdquo;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/ai-opensource-layers-en.svg" data-img="https://assets.jimmysong.io/images/blog/spatial-data-ai-open-source-standards-sovereignty/ai-opensource-layers-en.svg" alt="Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure" data-caption="Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure"
width="2363"
height="1862"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Two Layers of AI Era Open Source: Model Weights vs Infrastructure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-us-approach-focusing-on-rules-and-runtime-layers"&gt;The US Approach: Focusing on Rules and Runtime Layers&lt;/h2&gt;
&lt;p&gt;In the past year or two, US-led AI open source and ecosystem initiatives have shown a highly consistent direction: not rushing to open source the strongest models, but focusing on defining &amp;ldquo;how AI is used.&amp;rdquo;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Linux Foundation established &lt;a href="https://aaif.io" target="_blank" rel="noopener"&gt;AAIF&lt;/a&gt; (Agentic AI Foundation), focusing on AI infrastructure, standards, and toolchain collaboration&lt;/li&gt;
&lt;li&gt;Protocols like MCP (Model Context Protocol) aim to define common interaction methods between agents and tools/systems&lt;/li&gt;
&lt;li&gt;Major tech companies are generally focusing on APIs, platforms, runtimes, and ecosystem binding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The commonality of these actions: competing in model capability, but controlling the usage rules.&lt;/p&gt;
&lt;h2 id="chinas-shift-from-model-oriented-to-infrastructure-oriented"&gt;China&amp;rsquo;s Shift: From Model-Oriented to Infrastructure-Oriented&lt;/h2&gt;
&lt;p&gt;It is important to emphasize that this difference does not mean China is unaware of the issue.&lt;/p&gt;
&lt;p&gt;Whether in policy discussions or within industry and research institutions, the risk of &amp;ldquo;only open sourcing models without controlling infrastructure and standard dominance&amp;rdquo; has been repeatedly discussed.&lt;/p&gt;
&lt;p&gt;The real challenge lies in how to achieve a directional shift within the existing governance logic and risk framework. This shift has already appeared in some concrete practices.&lt;/p&gt;
&lt;h2 id="exploration-and-practice-at-the-infrastructure-layer"&gt;Exploration and Practice at the Infrastructure Layer&lt;/h2&gt;
&lt;p&gt;In the AI era, infrastructure often starts with the most engineering-driven problems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HAMi Project&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Projects like &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; do not focus on model capability, but on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Abstraction, allocation, and isolation of GPU resources&lt;/li&gt;
&lt;li&gt;How multi-tenant AI workloads are run&lt;/li&gt;
&lt;li&gt;How computing power transitions from hardware assets to governable system resources&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The significance of such projects is not about being &amp;ldquo;SOTA,&amp;rdquo; but about entering the domain of &amp;ldquo;how AI runs.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Runtime Reconstruction from a System Software Review&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Exploration at the research institution level is also noteworthy. The &lt;a href="https://www.flagos.io" target="_blank" rel="noopener"&gt;FlagOS&lt;/a&gt; initiative by the Beijing Academy of Artificial Intelligence is a clear signal: AI is being redefined as a system software issue, not just a model or algorithm problem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-Term Tech Stack Investment by Industry Players&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the industry, Huawei&amp;rsquo;s strategy reflects a similar direction: not simply open sourcing models, but attempting to build a complete, controllable AI tech stack, from computing power to frameworks, platforms, and ecosystems. This is a slower, heavier, but more infrastructure-competitive path.&lt;/p&gt;
&lt;h2 id="realistic-assessment-the-starting-point-of-ai-infrastructure-competition"&gt;Realistic Assessment: The Starting Point of AI Infrastructure Competition&lt;/h2&gt;
&lt;p&gt;Taking a longer view, we find an easily overlooked fact:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;At the level of AI infrastructure and infra-level open source, there is no settled pattern between China and the US.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The US advantage lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mature engineering culture&lt;/li&gt;
&lt;li&gt;Standard organizations and foundation mechanisms&lt;/li&gt;
&lt;li&gt;High proficiency in openness at the rules layer&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;China&amp;rsquo;s variables include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Huge AI application scenarios&lt;/li&gt;
&lt;li&gt;Extreme demand for computing power and system efficiency&lt;/li&gt;
&lt;li&gt;Ongoing directional adjustments&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The real uncertainty is not &amp;ldquo;whether we can catch up,&amp;rdquo; but whether it is possible to gradually open up space for engineering autonomy and standard co-construction while maintaining governance bottom lines.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The &amp;ldquo;points&amp;rdquo; and &amp;ldquo;areas&amp;rdquo; of air quality, model weights and the world of operations—behind these appearances lies not a simple technical route dispute, but how a country finds its own balance between openness, standards, and sovereignty.&lt;/p&gt;
&lt;p&gt;In the AI era, this issue will not disappear, but will become more concrete and more engineering-driven. And this is precisely where there are still opportunities for China&amp;rsquo;s AI infrastructure open source.&lt;/p&gt;</content:encoded></item><item><title>Joining Dynamia: Embarking on a New Journey in AI Native Infrastructure</title><link>https://jimmysong.io/blog/joining-dynamia/</link><pubDate>Wed, 07 Jan 2026 07:49:21 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/joining-dynamia/</guid><description>Joining Dynamia as Open Source Ecosystem VP to drive AI-native infrastructure ecosystem development, transforming compute from hardware consumption to core asset.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Compute governance is the critical bottleneck for AI scaling. From hardware consumption to core asset, this long-undervalued path needs to be redefined.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/joining-dynamia/banner.webp" data-img="https://assets.jimmysong.io/images/blog/joining-dynamia/banner.webp" alt="Figure 1: Dynamia.ai" data-caption="Figure 1: Dynamia.ai"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Dynamia.ai&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="a-new-beginning"&gt;A New Beginning&lt;/h2&gt;
&lt;p&gt;I have officially joined &lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia&lt;/a&gt; as &lt;strong&gt;Open Source Ecosystem VP&lt;/strong&gt;, responsible for the long-term development of the company in open source, technical narrative, and &lt;strong&gt;AI Native Infrastructure&lt;/strong&gt; ecosystem directions.&lt;/p&gt;
&lt;h2 id="why-i-chose-dynamia"&gt;Why I Chose Dynamia&lt;/h2&gt;
&lt;p&gt;I chose to join Dynamia not because it&amp;rsquo;s a company trying to &amp;ldquo;solve all AI problems,&amp;rdquo; but precisely the opposite—it&amp;rsquo;s because Dynamia &lt;strong&gt;focuses intensely on one unavoidable, yet long-undervalued core issue in AI Native Infrastructure&lt;/strong&gt;: compute, especially &lt;strong&gt;Graphics Processing Units&lt;/strong&gt; (GPU), are evolving from &amp;ldquo;technical resources&amp;rdquo; into infrastructure elements that require refined governance and economic management.&lt;/p&gt;
&lt;p&gt;Through years of practice in cloud native, distributed systems, and AI infrastructure (AI Infra), I&amp;rsquo;ve formed a clear judgment: as Large Language Models (LLM) and &lt;strong&gt;AI Agents&lt;/strong&gt; enter the stage of large-scale deployment, the real bottleneck limiting system scalability and sustainability is no longer just model capability itself, but how compute is measured, allocated, isolated, and scheduled, and how a governable, accountable, and optimizable operational mechanism is formed at the system level. From this perspective, the core challenge of AI infrastructure is essentially evolving into a &amp;ldquo;resource governance and Token economy&amp;rdquo; problem.&lt;/p&gt;
&lt;h2 id="about-dynamia-and-hami"&gt;About Dynamia and HAMi&lt;/h2&gt;
&lt;p&gt;Dynamia is an AI-native infrastructure technology company rooted in open source DNA, driving efficiency leaps in heterogeneous compute through technological innovation. Its leading open source project, &lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi&lt;/a&gt; (Heterogeneous AI Computing Virtualization Middleware), is a &lt;strong&gt;Cloud Native Computing Foundation&lt;/strong&gt; (CNCF) sandbox project providing GPU, NPU and other heterogeneous device virtualization, sharing, isolation, and topology-aware scheduling capabilities, widely adopted by 50+ enterprises and institutions.&lt;/p&gt;
&lt;h2 id="dynamias-technical-approach"&gt;Dynamia&amp;rsquo;s Technical Approach&lt;/h2&gt;
&lt;p&gt;In this context, Dynamia&amp;rsquo;s technical approach—starting from &lt;strong&gt;the GPU layer, which is the most expensive, scarcest, and least unified abstraction layer in AI systems&lt;/strong&gt;, treating compute as a foundational resource that can be measured, partitioned, scheduled, governed, and even &amp;ldquo;tokenized&amp;rdquo; for refined accounting and optimization—aligns highly with my long-term judgment on AI-native infrastructure.&lt;/p&gt;
&lt;p&gt;This path doesn&amp;rsquo;t use &amp;ldquo;model capabilities&amp;rdquo; or &amp;ldquo;application innovation&amp;rdquo; as selling points in the short term, nor is it easily packaged into simple stories. However, with rising compute costs, heterogeneous accelerators becoming the norm, and AI systems moving toward multi-tenant and large-scale operations, these infrastructure-level capabilities are gradually becoming prerequisites for the establishment and expansion of AI systems.&lt;/p&gt;
&lt;h2 id="future-focus"&gt;Future Focus&lt;/h2&gt;
&lt;p&gt;As Dynamia&amp;rsquo;s Open Source Ecosystem VP, I will focus on &lt;strong&gt;technical narrative of AI-native infrastructure, open source ecosystem building, and global developer collaboration&lt;/strong&gt;, promoting compute from &amp;ldquo;hardware resource being consumed&amp;rdquo; to &lt;strong&gt;governable, measurable, and optimizable AI infrastructure core asset&lt;/strong&gt;, laying the foundation for the scaling and sustainable evolution of AI systems in the next stage.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Joining Dynamia is an important milestone in my career and a concrete action demonstrating my long-term optimism about AI-native infrastructure. Compute governance is not a short-term trend that yields quick results, but an infrastructure proposition that cannot be bypassed for AI large-scale deployment. I look forward to exploring, building, and landing solutions on this long-undervalued path with global developers.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dynamia.ai" target="_blank" rel="noopener"&gt;Dynamia Official Website&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Project-HAMi/HAMi" target="_blank" rel="noopener"&gt;HAMi - Heterogeneous AI Computing Virtualization Middleware (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Running Parallel AI Agents on My Mac: Hands-On with Verdent's Standalone App</title><link>https://jimmysong.io/blog/verdent-standalone-app-parallel-agents/</link><pubDate>Sun, 04 Jan 2026 02:25:48 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/verdent-standalone-app-parallel-agents/</guid><description>A hands-on experience with Verdent&amp;#39;s standalone Mac app, exploring how parallel AI agents, isolated workspaces, and task-oriented workflows change real-world development.</description><content:encoded>
&lt;p&gt;I&amp;rsquo;ve been spending more time recently experimenting with vibe coding tools on real projects, not demos. One of those projects is my own website, where I constantly tweak content structure, navigation, and layout.&lt;/p&gt;
&lt;p&gt;During this process, I started using &lt;a href="https://verdent.ai" target="_blank" rel="noopener"&gt;Verdent&amp;rsquo;s standalone Mac app&lt;/a&gt; more seriously. What stood out was not any single feature, but how different the experience felt compared to traditional AI coding tools.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-standalone-app-ui.webp" data-img="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-standalone-app-ui.webp" alt="Figure 1: Verdent Standalone App UI" data-caption="Figure 1: Verdent Standalone App UI"
width="3836"
height="2240"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Verdent Standalone App UI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Verdent doesn&amp;rsquo;t behave like an assistant waiting for instructions. It behaves more like an environment where work happens in parallel.&lt;/p&gt;
&lt;h2 id="a-different-starting-point-tasks-not-chats"&gt;A Different Starting Point: Tasks, Not Chats&lt;/h2&gt;
&lt;p&gt;Most AI coding tools begin with a conversation. Verdent begins with &lt;strong&gt;tasks&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When I opened my website repository in the Verdent app, I didn&amp;rsquo;t start with a long prompt. I created multiple tasks directly: one to rethink navigation and SEO structure, another to explore homepage layout improvements, and a third to review existing content organization.&lt;/p&gt;
&lt;p&gt;Each task immediately spun up its own agent and workspace. From the beginning, the app encouraged me to think in parallel, the same way I normally would when sketching ideas on paper or jumping between files.&lt;/p&gt;
&lt;p&gt;This framing alone changes how you work.&lt;/p&gt;
&lt;h2 id="built-for-multitasking-without-losing-context"&gt;Built for Multitasking, Without Losing Context&lt;/h2&gt;
&lt;p&gt;Switching contexts is unavoidable in real development work. What usually breaks is continuity.&lt;/p&gt;
&lt;p&gt;Verdent handles this well. Each task preserves its full context independently. I could stop one task mid-way, switch to another, and come back later without re-explaining the problem or reloading files.&lt;/p&gt;
&lt;p&gt;For example, while one agent was analyzing my site&amp;rsquo;s navigation structure, another was exploring layout options. I moved between them freely. Nothing was lost. Each agent remembered exactly what it was doing.&lt;/p&gt;
&lt;p&gt;This feels closer to how developers think than how chat-based tools operate.&lt;/p&gt;
&lt;h2 id="safe-parallel-coding-with-workspaces"&gt;Safe Parallel Coding with Workspaces&lt;/h2&gt;
&lt;p&gt;Parallel work only becomes truly safe when code changes are isolated. When parallelism moves from discussion to actual code modification, risk management becomes essential.&lt;/p&gt;
&lt;p&gt;Verdent solves this with &lt;strong&gt;Workspaces&lt;/strong&gt;. Each workspace is an isolated, independent code environment with its own change history, commit log, and branches. This isn&amp;rsquo;t just about separation—it&amp;rsquo;s about making concurrent code changes manageable.&lt;/p&gt;
&lt;p&gt;What this means in practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple tasks can write code simultaneously&lt;/li&gt;
&lt;li&gt;Changes remain isolated from each other&lt;/li&gt;
&lt;li&gt;If conflicts arise, they&amp;rsquo;re visible and cleanly resolvable&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I intentionally let different agents operate on overlapping parts of my project: one modifying Markdown content and links, another adjusting CSS and layout logic. Both ran in parallel. No conflicts emerged. Later, I reviewed the diffs from each workspace and merged only what made sense.&lt;/p&gt;
&lt;p&gt;This kind of isolation removes significant anxiety from AI-assisted coding. You stop worrying about breaking things and start experimenting more freely, knowing that each change exists in its own contained environment.&lt;/p&gt;
&lt;h2 id="parallel-agent-execution-feels-like-delegation"&gt;Parallel Agent Execution Feels Like Delegation&lt;/h2&gt;
&lt;p&gt;Parallelism doesn’t mean that all agents complete the same phase of work at the same time—instead, by isolating and overlapping phases, what was once a strictly sequential process is compressed into a more efficient, collaborative mode.&lt;/p&gt;
&lt;p&gt;In Verdent, each agent runs in its own workspace, essentially an automatically managed branch or worktree. In practice, I often create multiple tasks with different responsibilities for the same requirement, such as planning, implementation, and review. But this doesn’t mean they all complete the same phase simultaneously.&lt;/p&gt;
&lt;p&gt;These tasks are triggered as needed, each running for a period and producing clear artifacts as boundaries for collaboration. The planning task generates planning documents or constraint specifications; the implementation task advances code changes based on those documents and produces diffs; the review task, according to the established planning goals and audit criteria, performs staged reviews of the generated changes. By overlapping phases around artifacts, the originally strict sequential process is compressed into a workflow that more closely resembles team collaboration.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The value of splitting into multiple tasks is not parallel execution, but parallel cognition and clear collaboration boundaries.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;While it’s technically possible to put multiple roles into a single task, this causes planning, implementation, and review to share the same context, which weakens role isolation and the auditability of results.&lt;/p&gt;
&lt;h2 id="configurability-and-design-trade-offs"&gt;Configurability and Design Trade-offs&lt;/h2&gt;
&lt;p&gt;Beyond the workflow model itself, Verdent exposes a surprisingly rich set of configurable capabilities.&lt;/p&gt;
&lt;p&gt;It allows users to customize MCP settings, define subagents with configurable prompts, and create reusable commands via slash (&lt;code&gt;/&lt;/code&gt;) shortcuts. Personal rules can be written to influence agent behavior and response style, and command-level permissions can be configured to enforce basic security boundaries. Verdent also supports multiple mainstream foundation models, including GPT, Claude, Gemini, and K2. For users who prefer a lightweight coding experience without a full IDE, Verdent offers DiffLens as an alternative review-oriented interface. Both &lt;a href="https://www.verdent.ai/pricing" target="_blank" rel="noopener"&gt;subscription-based and credit-based pricing models&lt;/a&gt; are supported.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-settings.webp" data-img="https://assets.jimmysong.io/images/blog/verdent-standalone-app-parallel-agents/verdent-settings.webp" alt="Figure 2: Verdent Settings" data-caption="Figure 2: Verdent Settings"
width="2780"
height="1648"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Verdent Settings&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;That said, Verdent makes a clear set of trade-offs. It is not built around tab-based code completion, nor does it offer a plugin system. If it did, it would start to resemble a traditional IDE - which does not seem to be its goal. Verdent is not designed for direct, fine-grained code manipulation; most changes are mediated through conversational tasks and agent-driven edits. This makes the experience clean and focused, but it also means that for large, highly complex codebases, Verdent may function better as a complementary orchestration layer rather than a full-time development environment.&lt;/p&gt;
&lt;h2 id="where-verdent-fits-today"&gt;Where Verdent Fits Today&lt;/h2&gt;
&lt;p&gt;There are many AI-assisted coding tools emerging right now. Some focus on smarter editors, others on faster generation.&lt;/p&gt;
&lt;p&gt;Verdent feels different because it focuses on &lt;strong&gt;orchestration&lt;/strong&gt;, not just assistance.&lt;/p&gt;
&lt;p&gt;It doesn&amp;rsquo;t try to replace your editor. It sits one level above, coordinating planning, execution, and review across multiple agents.&lt;/p&gt;
&lt;p&gt;That makes it particularly suitable for exploratory work, refactoring, and early-stage design - exactly the kind of work I was doing on my website.&lt;/p&gt;
&lt;h2 id="final-thoughts"&gt;Final Thoughts&lt;/h2&gt;
&lt;p&gt;Using Verdent&amp;rsquo;s standalone app didn&amp;rsquo;t just speed things up. It changed how I structured work.&lt;/p&gt;
&lt;p&gt;Instead of doing everything sequentially, I started thinking in parallel again - and letting the system support that way of thinking.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://verdent.ai" target="_blank" rel="noopener"&gt;Verdent&lt;/a&gt; feels less like an AI feature and more like an environment that assumes AI is already part of how development happens.&lt;/p&gt;
&lt;p&gt;For developers experimenting with AI-native workflows, that shift is worth paying attention to.&lt;/p&gt;</content:encoded></item><item><title>2025 Annual Review: The Transformation Journey from Cloud Native to AI Native</title><link>https://jimmysong.io/blog/2025-annual-review/</link><pubDate>Wed, 31 Dec 2025 10:02:01 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/2025-annual-review/</guid><description>A look back at the major changes in 2025: shifting from Cloud Native to AI Native Infrastructure, AI tool ecosystem, and major website improvements.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The waves of technology keep evolving; only by actively embracing change can we continue to create value. In 2025, I chose to move from Cloud Native to AI Native—this year marked a key turning point for personal growth and system reinvention.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;2025 was a turning point for me. This year, I not only changed my technical direction but also the way I approach problems. Moving from Cloud Native infrastructure to AI Native Infrastructure was not just a migration of content, but an upgrade in mindset.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/2025-annual-review/banner.webp" data-img="https://assets.jimmysong.io/images/blog/2025-annual-review/banner.webp" alt="Figure 1: Farewell 2025!" data-caption="Figure 1: Farewell 2025!"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Farewell 2025!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This year, I conducted a large-scale refactoring of the website and systematically organized the content. Beyond the technical improvements, I want to share my thoughts and changes throughout the year.&lt;/p&gt;
&lt;h2 id="a-bold-shift-embracing-the-ai-native-era"&gt;A Bold Shift: Embracing the AI Native Era&lt;/h2&gt;
&lt;p&gt;At the beginning of 2025, I made an important decision: to reposition myself from a Cloud Native Evangelist to an AI Infrastructure Architect. This was not just a change in title, but a strategic transformation after careful consideration.&lt;/p&gt;
&lt;p&gt;As I witnessed the surge of AI technologies and the rise of Agent-based applications reshaping software, I realized that clinging to the boundaries of Cloud Native might mean missing an era. So, I systematically adjusted the website’s content structure, shifting the focus toward AI Native Infrastructure.&lt;/p&gt;
&lt;p&gt;This transformation was not about abandoning the past, but extending forward from the foundation of Cloud Native. Classic content like Kubernetes and Istio remains and is continuously updated, but new topics such as AI Agent and the AI Native Landscape have been added, forming a more complete knowledge map.&lt;/p&gt;
&lt;h2 id="content-creation-from-technical-details-to-ecosystem-perspective"&gt;Content Creation: From Technical Details to Ecosystem Perspective&lt;/h2&gt;
&lt;h3 id="ai-agent-building-systematic-knowledge"&gt;AI Agent: Building Systematic Knowledge&lt;/h3&gt;
&lt;p&gt;Agents represent a major evolution in software for the AI era. When I tried to understand Agent design principles, I found fragmented information everywhere but lacked a systematic knowledge base.&lt;/p&gt;
&lt;p&gt;So I created content that analyzes the Agent context lifecycle and control loop mechanisms, summarizing several proven architectural patterns. To make complex knowledge easier to digest, I organized it into logical sections so readers can learn step by step.&lt;/p&gt;
&lt;h3 id="ai-tool-ecosystem-mapping-the-open-source-landscape"&gt;AI Tool Ecosystem: Mapping the Open Source Landscape&lt;/h3&gt;
&lt;p&gt;AI tools and frameworks are emerging rapidly, with new projects appearing daily. To help readers quickly grasp the ecosystem, I built a comprehensive AI OSS database.&lt;/p&gt;
&lt;p&gt;This database covers everything from Agent frameworks to development tools and deployment services. I not only included active projects but also established an archive mechanism, preserving detailed information on over 150 historical projects. More importantly, I developed a scoring system to objectively evaluate projects across dimensions like quality and sustainability, helping readers decide which tools are worth investing time in.&lt;/p&gt;
&lt;h3 id="blogging-capturing-technology-trends-faster"&gt;Blogging: Capturing Technology Trends Faster&lt;/h3&gt;
&lt;p&gt;In 2025, I wrote over 120 blog posts. Compared to previous years, these articles focused more on observing and reflecting on technology trends, rather than just technical tutorials.&lt;/p&gt;
&lt;p&gt;I started paying attention to deeper questions: How will AI infrastructure evolve? What does Beijing’s open source initiative mean for the AI industry? What ripple effects might a tech acquisition trigger? These articles allowed me and my readers to not only see &amp;ldquo;what&amp;rdquo; technology is, but also &amp;ldquo;why&amp;rdquo; and &amp;ldquo;what’s next.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="user-experience-making-knowledge-easier-to-discover-and-consume"&gt;User Experience: Making Knowledge Easier to Discover and Consume&lt;/h2&gt;
&lt;p&gt;No matter how good the content is, if it can’t be easily found and read, its value is greatly diminished. In 2025, I invested significant effort into website functionality, with one goal: to provide readers with a smoother reading experience.&lt;/p&gt;
&lt;h3 id="comprehensive-search-upgrade"&gt;Comprehensive Search Upgrade&lt;/h3&gt;
&lt;p&gt;As the volume of content grew, the original search function could no longer meet demand. I redesigned the search system to support fuzzy search and result scoring, and optimized index loading performance. More importantly, the new search interface is more user-friendly, supporting keyboard navigation and category filtering so users can find what they want faster.&lt;/p&gt;
&lt;h3 id="multi-device-experience-optimization"&gt;Multi-Device Experience Optimization&lt;/h3&gt;
&lt;p&gt;Mobile reading experience has improved significantly. I refactored the mobile navigation and table of contents, making reading on phones much smoother. Dark mode is now more refined, fixing several display issues and ensuring images and diagrams look good on dark backgrounds.&lt;/p&gt;
&lt;h3 id="efficiency-revolution-in-content-distribution"&gt;Efficiency Revolution in Content Distribution&lt;/h3&gt;
&lt;p&gt;A major change was optimizing the WeChat Official Account publishing workflow. Previously, publishing website content to WeChat required manual handling of many details; now, it’s almost one-click export. This workflow automatically processes images, metadata, styles, and all details, reducing a half-hour task to just a few minutes.&lt;/p&gt;
&lt;p&gt;Additionally, I added a glossary feature for technical term highlighting and tooltips; improved SEO and social sharing metadata; and cleaned up outdated content. These seemingly minor improvements quietly enhance the user experience.&lt;/p&gt;
&lt;h2 id="content-evolution-more-dimensional-knowledge-expression"&gt;Content Evolution: More Dimensional Knowledge Expression&lt;/h2&gt;
&lt;p&gt;Looking back at content creation in 2025, I found clear changes in several dimensions.&lt;/p&gt;
&lt;h3 id="from-tutorials-to-observations"&gt;From Tutorials to Observations&lt;/h3&gt;
&lt;p&gt;Early content leaned toward technical tutorials and practical guides, showing &amp;ldquo;how to do.&amp;rdquo; This year, I focused more on &amp;ldquo;why&amp;rdquo; and &amp;ldquo;what are the trends.&amp;rdquo; I wrote more technology trend analyses, ecosystem maps, and in-depth case studies. These may not directly teach you how to use an API, but they help you understand the direction of technological evolution.&lt;/p&gt;
&lt;h3 id="from-chinese-to-bilingual"&gt;From Chinese to Bilingual&lt;/h3&gt;
&lt;p&gt;AI is a global wave and cannot be limited to the Chinese-speaking world. In 2025, I wrote bilingual documentation for almost all new AI tools, and important blog posts also have English versions. This increased the workload, but allowed the content to reach a broader audience.&lt;/p&gt;
&lt;h3 id="from-text-to-multimedia"&gt;From Text to Multimedia&lt;/h3&gt;
&lt;p&gt;Text is efficient, but not all knowledge is best expressed in words. This year, I used many architecture and schematic diagrams to explain complex concepts, adding 59 new charts. These visual elements lower the barrier to understanding, making abstract concepts more intuitive. I also optimized image display in dark mode to ensure consistent visual experience.&lt;/p&gt;
&lt;h2 id="development-approach-embracing-ai-assisted-programming"&gt;Development Approach: Embracing AI-Assisted Programming&lt;/h2&gt;
&lt;p&gt;2025 was not only a year of shifting content themes toward AI, but also a year of deep practice in AI-assisted programming.&lt;/p&gt;
&lt;p&gt;I developed a VS Code plugin and created many prompts to automate repetitive tasks. I experimented with various AI programming tools and settled on a toolchain that suits me. I even migrated the website to Cloudflare Pages and used its edge computing services to develop a chatbot. These practices greatly improved development efficiency, giving me more time to focus on thinking and creating rather than mechanical coding.&lt;/p&gt;
&lt;p&gt;This made me realize: AI will not replace developers, but developers who use AI well will replace those who do not. I also shared more insights to help others master AI-assisted programming.&lt;/p&gt;
&lt;h2 id="looking-ahead-to-2026-keep-moving-forward"&gt;Looking Ahead to 2026: Keep Moving Forward&lt;/h2&gt;
&lt;p&gt;Looking back at 2025, the site underwent a profound transformation—from a Cloud Native tech blog to an AI infrastructure knowledge base. But this is just the beginning, not the end.&lt;/p&gt;
&lt;p&gt;Looking forward to 2026, I plan to continue deepening in several areas:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Enhancing the knowledge system&lt;/strong&gt;: Continue to supplement GPU infrastructure and AI Agent content, especially practical cases and performance tuning knowledge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tracking ecosystem evolution&lt;/strong&gt;: AI tools and frameworks iterate rapidly; I need to keep up with this fast-changing ecosystem and update content in a timely manner.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deepening engineering practice&lt;/strong&gt;: Share more practical AI engineering experience to help readers turn theory into practice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exploring knowledge connections&lt;/strong&gt;: Consider building a knowledge graph to connect different content sections, providing smarter navigation and recommendations.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;2025 was a year of change and growth. From Cloud Native to AI Native, from technical practice to ecosystem observation, both the content and functionality of the site have made qualitative leaps.&lt;/p&gt;
&lt;p&gt;What makes me happiest is that this transformation allowed me and my readers to stand at the forefront of the technology wave. We are not just learning new technologies, but thinking about how technology changes the world and the way we write software.&lt;/p&gt;
&lt;p&gt;The waves of technology keep evolving; only by actively embracing change can we continue to create value. Thank you to every reader for your companionship and support. I look forward to sharing more insights and practices in 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Further Reading&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Native Landscape&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/"&gt;2025 Blog Posts&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>The Butterfly Effect After Manus Was Acquired by Meta</title><link>https://jimmysong.io/blog/manus-meta-acquisition-butterfly-effect/</link><pubDate>Tue, 30 Dec 2025 03:30:51 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/manus-meta-acquisition-butterfly-effect/</guid><description>Manus&amp;#39;s acquisition by Meta sparked polarized opinions. This article explores the butterfly effect in AI applications and key lessons for entrepreneurs on growth strategies.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The success or failure of AI applications often lies not in the technology itself, but in the ability to scale delivery and create a closed loop.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/manus-meta-acquisition-butterfly-effect/banner.webp" data-img="https://assets.jimmysong.io/images/blog/manus-meta-acquisition-butterfly-effect/banner.webp" alt="Figure 1: The Butterfly Effect After Manus Was Acquired by Meta" data-caption="Figure 1: The Butterfly Effect After Manus Was Acquired by Meta"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: The Butterfly Effect After Manus Was Acquired by Meta&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="when-those-who-discuss-it-are-not-those-who-pay-for-it"&gt;When &amp;ldquo;Those Who Discuss It&amp;rdquo; Are Not &amp;ldquo;Those Who Pay for It&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;On December 30, 2025, a piece of news went viral: Manus was acquired by Meta for billions of dollars (&lt;a href="https://manus.im/blog/manus-joins-meta-for-next-era-of-innovation" target="_blank" rel="noopener"&gt;Manus Joins Meta for Next Era of Innovation&lt;/a&gt;). This startup, founded in China and under pressure from tech giants since its inception, completed a whirlwind journey in less than a year—from explosive growth, relocating to Singapore, to being acquired by a global giant.&lt;/p&gt;
&lt;p&gt;According to Manus&amp;rsquo;s official statement, its products and subscriptions will continue to be available via the app and website, and the company will remain operational in Singapore. The team will join Meta to provide general Agent capabilities for Meta&amp;rsquo;s consumer and enterprise products (including Meta AI).&lt;/p&gt;
&lt;p&gt;Rather than focusing on &amp;ldquo;who won,&amp;rdquo; I&amp;rsquo;m more interested in the chain reaction this event triggered: it activated completely opposite judgment systems among different groups, and this split is reshaping the growth paths and strategies for AI applications and startups.&lt;/p&gt;
&lt;h2 id="two-public-opinion-arenas-blessings-and-doubts-coexist"&gt;Two Public Opinion Arenas: Blessings and Doubts Coexist&lt;/h2&gt;
&lt;p&gt;After Manus was acquired, the mainstream sentiment in social circles was one of congratulations and excitement. Many saw it as a stellar example of a Chinese team going global—achieving remarkable results in the most competitive field in a very short time.&lt;/p&gt;
&lt;p&gt;Meanwhile, the comment sections of public accounts became &amp;ldquo;venting valves for counter-narratives,&amp;rdquo; with skepticism centering on three main points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether the technology has real barriers (e.g., &amp;ldquo;there are countless similar products,&amp;rdquo; &amp;ldquo;it&amp;rsquo;s not hard for big companies to build their own&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;Valuation and bubble concerns (e.g., &amp;ldquo;another case of the AI bubble&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;Distrust in the buyer&amp;rsquo;s judgment (e.g., &amp;ldquo;giants making desperate bets,&amp;rdquo; &amp;ldquo;history repeating itself&amp;rdquo;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This divergence isn&amp;rsquo;t about who understands AI better, but about different evaluation frameworks: social circles focus on &amp;ldquo;trajectory and outcome,&amp;rdquo; while comment sections focus on &amp;ldquo;legitimacy and worthiness.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="where-does-the-100m-arr-come-from-the-target-users-arent-in-our-social-circles"&gt;Where Does the $100M ARR Come From: The Target Users Aren&amp;rsquo;t in Our Social Circles&lt;/h2&gt;
&lt;p&gt;Many people are impressed by Manus&amp;rsquo;s marketing buzz and controversies, which can lead to skepticism. But if it achieved a &amp;ldquo;strict $100M ARR&amp;rdquo; in 10 months, one fact is clear: &lt;strong&gt;its revenue doesn&amp;rsquo;t depend on broad consensus, but comes from a highly concentrated group of global users with strong willingness to pay.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Manus&amp;rsquo;s core user profile is closer to &amp;ldquo;individuals as production units,&amp;rdquo; including freelancers, indie developers, independent researchers, and key deliverers in small and medium businesses. They don&amp;rsquo;t care about debates over &amp;ldquo;wrapping&amp;rdquo; or not; they care about &amp;ldquo;can I deliver end-to-end tasks,&amp;rdquo; and &amp;ldquo;can this help me hire one less person, work fewer late nights, or avoid juggling ten tools.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This leads to a counterintuitive phenomenon: &lt;strong&gt;those who discuss the most may not pay, while those who pay steadily are often silent.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For these users, tools are not identity badges—they are profit levers.&lt;/p&gt;
&lt;h2 id="three-lessons-for-entrepreneurs-the-growth-paradigm-in-the-ai-application-era-has-changed"&gt;Three Lessons for Entrepreneurs: The Growth Paradigm in the AI Application Era Has Changed&lt;/h2&gt;
&lt;p&gt;Based on the above, the Manus case offers three lessons for entrepreneurs:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Growth No Longer Equals Positive Reviews&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;AI applications can commercialize first and build consensus later. Public opinion can remain divided for a long time, but cash flow doesn&amp;rsquo;t wait for unified recognition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Heavy Marketing&amp;rdquo; Is Becoming a Capability, Not a Stigma&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;As foundational models and capabilities spread rapidly, differentiation is quickly erased. Being seen, understood, and paid for is itself part of the moat. Not all marketing deserves respect, but &amp;ldquo;distribution and mindshare&amp;rdquo; have become unavoidable battlegrounds for AI applications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization Is No Longer a Bonus, but May Be a Survival Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From payment willingness, compliance boundaries, talent density to valuation systems, market structure means many teams &amp;ldquo;can only complete the loop overseas.&amp;rdquo; It&amp;rsquo;s not romantic, but it&amp;rsquo;s reality.&lt;/p&gt;
&lt;h2 id="a-personal-reflection"&gt;A Personal Reflection&lt;/h2&gt;
&lt;p&gt;As someone long engaged in cloud native and AI infrastructure, I&amp;rsquo;m used to evaluating products by their &amp;ldquo;technical barriers.&amp;rdquo; But cases like Manus remind me: at the AI application layer, barriers may not first appear in models or code, but often in &lt;strong&gt;organizational speed, productization capability, delivery loop, and distribution efficiency&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When a system can reliably turn &amp;ldquo;capability&amp;rdquo; into &amp;ldquo;results,&amp;rdquo; it has built a commercial moat—even if its tech stack doesn&amp;rsquo;t meet outsiders&amp;rsquo; ideals of &amp;ldquo;purity.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The biggest butterfly effect of Manus being acquired by Meta may not be the deal itself, but making more entrepreneurs realize: &lt;strong&gt;in the AI era, the winning move is shifting from &amp;ldquo;what model you use&amp;rdquo; to &amp;ldquo;whether you can deliver results at scale.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The acquisition of Manus by Meta is not just a convergence of capital and technology, but also a microcosm of the changing growth paradigm in the AI application era. For entrepreneurs, understanding and mastering &amp;ldquo;user structure,&amp;rdquo; &amp;ldquo;distribution capability,&amp;rdquo; and &amp;ldquo;global closed loops&amp;rdquo; will be key to future competition.&lt;/p&gt;</content:encoded></item><item><title>AI Infra Open Source in China: Analysis of Beijing and Shanghai's Plans</title><link>https://jimmysong.io/blog/beijing-open-source-plan-ai-infra-analysis/</link><pubDate>Thu, 25 Dec 2025 10:01:13 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/beijing-open-source-plan-ai-infra-analysis/</guid><description>Beijing and Shanghai&amp;#39;s open source plans reveal opportunities and challenges for China&amp;#39;s AI infrastructure, balancing technology and governance.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Institutionalized open source marks a new starting point for China&amp;rsquo;s AI Infra, but true breakthroughs and risks lie in the engineering and governance details.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="perspective-on-beijing-and-shanghais-open-source-plans"&gt;Perspective on Beijing and Shanghai&amp;rsquo;s Open Source Plans&lt;/h2&gt;
&lt;p&gt;Using the simultaneous release of open source ecosystem plans by Beijing and Shanghai as a lens, and drawing on China&amp;rsquo;s past foundation practices and international open source governance experience, this article explores the real opportunities, structural constraints, and potential risks as AI Infrastructure (AI Infra, Artificial Intelligence Infrastructure) enters a new phase of institutionalized open source.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/beijing-open-source-plan-ai-infra-analysis/banner.webp" data-img="https://assets.jimmysong.io/images/blog/beijing-open-source-plan-ai-infra-analysis/banner.webp" alt="Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans" data-caption="Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans"
width="1536"
height="1024"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Beijing and Shanghai successively launch open source ecosystem construction plans&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-compare-beijing-and-shanghai-together"&gt;Why Compare Beijing and Shanghai Together&lt;/h2&gt;
&lt;p&gt;It is rare for me to write an article solely because of a local policy document. However, during Christmas, both Beijing and Shanghai&amp;rsquo;s Bureaus of Economy and Information Technology released their respective open source ecosystem construction plans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/9YEL1HORWatsol3nRT596w" target="_blank" rel="noopener"&gt;Building an Open Source Innovation Highland! Beijing Releases Open Source Ecosystem Construction Implementation Plan&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/QZl66fUllKiePwQ7euhiGQ" target="_blank" rel="noopener"&gt;Shanghai&amp;rsquo;s Implementation Plan for Strengthening the Open Source System | Infographic&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This time, the fact that both cities released their plans on the same day sends a signal worth serious attention: China is attempting to advance open source in a more systematic and institutionalized way, especially regarding open source capabilities related to AI Infra.&lt;/p&gt;
&lt;p&gt;If you only look at Beijing&amp;rsquo;s plan, it is easy to interpret it as a local industrial policy upgrade. But when you consider both Beijing and Shanghai&amp;rsquo;s plans together, it looks more like a clearly defined &amp;ldquo;dual-center structure.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The question is no longer whether to develop open source, but:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the AI era, what institutional forms, engineering paths, and governance models will open source take?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="open-source-as-industrial-infrastructure-engineering"&gt;Open Source as &amp;ldquo;Industrial Infrastructure Engineering&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Both Beijing and Shanghai&amp;rsquo;s plans reflect a highly consistent judgment:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Open source is no longer seen as a spontaneous community activity, but as an industrial infrastructure capability that requires systematic construction.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is especially evident in the field of AI Infra.&lt;/p&gt;
&lt;p&gt;Issues such as computing power scheduling, model evaluation, toolchains, data elements, license compliance, and supply chain security—previously hidden in &amp;ldquo;engineering details&amp;rdquo;—are now systematically incorporated into policy language for the first time. This at least shows that decision-makers have realized:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI competition is not only about model parameter scale&lt;/li&gt;
&lt;li&gt;It is even more about toolchains, infrastructure, evaluation systems, and engineering capabilities&lt;/li&gt;
&lt;li&gt;These capabilities are naturally more suitable for building public foundations through open source&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this respect, Beijing and Shanghai are highly aligned.&lt;/p&gt;
&lt;h2 id="two-open-source-paths-infra-vs-platform"&gt;Two Open Source Paths: Infra vs. Platform&lt;/h2&gt;
&lt;p&gt;When we zoom in, the differences between the two plans become clear.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Beijing: &amp;ldquo;Foundation-Oriented&amp;rdquo; Open Source Path for AI Infra&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Beijing&amp;rsquo;s plan focuses on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Heterogeneous computing power scheduling&lt;/li&gt;
&lt;li&gt;Model evaluation toolchains&lt;/li&gt;
&lt;li&gt;Data elements and data governance&lt;/li&gt;
&lt;li&gt;RISC-V software-hardware collaboration&lt;/li&gt;
&lt;li&gt;SBOM, license compatibility, open source compliance&lt;/li&gt;
&lt;li&gt;Supply chain security and industrial resilience&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a typical perspective of &amp;ldquo;treating AI as an infrastructure problem.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;It is less concerned with the number of projects or community size, and more with:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether reusable engineering capabilities can be formed&lt;/li&gt;
&lt;li&gt;Whether these can be trusted by industry and government over the long term&lt;/li&gt;
&lt;li&gt;Whether they can stand up to scrutiny in terms of security, compliance, and governance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To some extent, Beijing is answering the question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;How can open source become a &amp;ldquo;governable, auditable, and scalable public capability&amp;rdquo;?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Shanghai: &amp;ldquo;Scale and Internationalization&amp;rdquo; Path for AI Platform&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In contrast, Shanghai&amp;rsquo;s plan has a different focus:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Building an international open source community for artificial intelligence&lt;/li&gt;
&lt;li&gt;Covering the entire platform chain from development, training, testing, hosting, to operation&lt;/li&gt;
&lt;li&gt;Overseas sites, multilingual support, international activities&lt;/li&gt;
&lt;li&gt;Resource linkage through computing vouchers and model vouchers&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Open source platform first release / global simultaneous release&amp;rdquo; dual-release mechanism&lt;/li&gt;
&lt;li&gt;Clear targets for community, enterprise, and developer scale&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Shanghai cares more about:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How open source can achieve scale effects&lt;/li&gt;
&lt;li&gt;How it can support the growth of commercial enterprises&lt;/li&gt;
&lt;li&gt;How it can be seen and adopted globally&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a path of &amp;ldquo;treating open source as a global digital product and platform capability.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="together-a-complete-but-tension-filled-structure"&gt;Together: A Complete but Tension-Filled Structure&lt;/h2&gt;
&lt;p&gt;When viewed together, Beijing and Shanghai&amp;rsquo;s plans form a more complete picture:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Beijing is responsible for &amp;ldquo;making open source solid,&amp;rdquo; while Shanghai is responsible for &amp;ldquo;taking open source global.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Structurally, this is a clear division of labor:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Beijing focuses on institutions, governance, and foundational capabilities&lt;/li&gt;
&lt;li&gt;Shanghai focuses on community, commercialization, and international communication&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These two paths are not in conflict; in theory, they are even complementary. The real question is whether they can form positive feedback in practice, rather than operating in silos.&lt;/p&gt;
&lt;h2 id="cautious-attitude-toward-institutionalized-platformized-open-source"&gt;Cautious Attitude Toward &amp;ldquo;Institutionalized, Platformized Open Source&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Precisely because both plans are so &amp;ldquo;systematic,&amp;rdquo; I am even more cautious.&lt;/p&gt;
&lt;p&gt;The reason is simple: this is not China&amp;rsquo;s first attempt to promote open source through foundations, associations, or platforms.&lt;/p&gt;
&lt;p&gt;Over the past decade, we have seen similar paths repeatedly, and recurring structural problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The difficulty of establishing neutrality and multi-party trust is extremely high&lt;/li&gt;
&lt;li&gt;There is a huge gap between showcase metrics (quantity, activities, certifications) and ecosystem strength&lt;/li&gt;
&lt;li&gt;Commercialization and long-term maintenance mechanisms are hard to sustain&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These problems will not disappear just because the plans are more comprehensive.&lt;/p&gt;
&lt;h2 id="four-risks-to-watch-under-the-dual-plans"&gt;Four Risks to Watch Under the Dual Plans&lt;/h2&gt;
&lt;p&gt;If we are to &amp;ldquo;listen to their words and watch their actions,&amp;rdquo; I would focus on the following four risks:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will Metrics Hijack Engineering Reality&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When &amp;ldquo;internationally influential projects,&amp;rdquo; &amp;ldquo;star projects,&amp;rdquo; and &amp;ldquo;first-release projects&amp;rdquo; become hard metrics, will this induce packaging, migration, and short-term hype, rather than truly solving engineering problems?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will It Slide Toward Platform Centralism&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The long-term pattern of AI Infra is closer to a model that prioritizes protocols, standards, and interoperability. If it eventually evolves into &amp;ldquo;a few platforms concentrating resources and discourse power,&amp;rdquo; it may be efficient in the short term but will suppress external participation and international collaboration in the long run.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is Internationalization Underestimated as an &amp;ldquo;Operational Issue&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;True international collaboration is never just about language, sites, or events; it also involves governance structures, compliance boundaries, and supply chain trust.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Will Application Demonstrations Become One-Off Projects&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If &amp;ldquo;first plans&amp;rdquo; and &amp;ldquo;computing vouchers&amp;rdquo; are just procurement tactics without continuous iteration and community feedback mechanisms, the long-term benefit to the ecosystem will be very limited.&lt;/p&gt;
&lt;h2 id="what-are-the-hard-results-of-ai-infra-open-source-after-three-years"&gt;What Are the &amp;ldquo;Hard Results&amp;rdquo; of AI Infra Open Source After Three Years&lt;/h2&gt;
&lt;p&gt;If we review the success of this round of institutionalized open source after three years, I would look for three types of results:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether de facto standards and interoperable ecosystems have emerged, including scheduling interfaces, evaluation benchmarks, Agent tool invocation protocols, and observability semantics.&lt;/li&gt;
&lt;li&gt;Whether compliance and supply chain security have become public capabilities—SBOM, license compatibility, vulnerability monitoring—truly productized and service-oriented.&lt;/li&gt;
&lt;li&gt;Whether a sustainable maintenance business mechanism has been established, allowing core maintainers to stay long-term, rather than relying on passion and subsidies.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;If I were to use a North Star metric to measure the success of these plans, it would be the emergence of several outstanding open source commercial companies rooted in China and serving the world.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The open source ecosystem plans of Beijing and Shanghai mark a new phase of institutionalization and engineering for AI Infra open source in China. Over the next three years, the real achievements will not be about meeting targets, but about forming sustainable engineering capabilities, de facto standards, and maintenance mechanisms. Only through continuous participation and practice can open source become the public foundation of AI infrastructure.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jxj.beijing.gov.cn/zwgk/2024zcwj/202512/t20251224_4360437.html" target="_blank" rel="noopener"&gt;Beijing Open Source Ecosystem Construction Implementation Plan (2026–2028) - jxj.beijing.gov.cn&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mp.weixin.qq.com/s/QZl66fUllKiePwQ7euhiGQ" target="_blank" rel="noopener"&gt;Shanghai&amp;rsquo;s Implementation Plan for Strengthening the Open Source System | Infographic - mp.weixin.qq.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>From 2025 Onwards, Software Engineering Shifts from Code-Centric to Runtime and Cost-Centric</title><link>https://jimmysong.io/blog/software-engineering-shift-runtime-cost-2025/</link><pubDate>Wed, 24 Dec 2025 14:59:11 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/software-engineering-shift-runtime-cost-2025/</guid><description>In 2025, software engineering shifts from code-centric to runtime and cost governance. AI and Agents move complexity to runtime, compute, and budget layers, reshaping engineering value.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;In 2025, the core of software engineering is no longer just about code itself, but about runtime controllability and cost governance. This shift is fundamentally reshaping the industry&amp;rsquo;s underlying logic.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Looking back at 2025, I became increasingly aware that this year was not about &amp;ldquo;code becoming unimportant,&amp;rdquo; but rather that &lt;strong&gt;the value coordinates of engineering have shifted as a whole&lt;/strong&gt;. For more than a decade, software engineering has focused on code quality, architectural evolution, and delivery efficiency. But starting in 2025, the key to system success is shifting—&lt;strong&gt;towards whether the runtime is controllable and whether costs are governable&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This is not just a slogan, but a conclusion repeatedly validated by my real-world experiences throughout the year.&lt;/p&gt;
&lt;h2 id="my-2025-from-platform-engineering-to-runtime-challenges"&gt;My 2025: From &amp;ldquo;Platform Engineering&amp;rdquo; to &amp;ldquo;Runtime Challenges&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;In my annual review, I noted a clear change: I spent less time on &amp;ldquo;how to write a good system,&amp;rdquo; and more time on &amp;ldquo;how to keep the system running stably, reliably, and affordably.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This shift in focus is a natural extension of a decade of cloud native evolution.&lt;/p&gt;
&lt;p&gt;The following timeline diagram illustrates how my focus has changed over recent years:
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/focus-shift-timeline-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/focus-shift-timeline-en.svg" alt="Figure 1: My Focus Shift Timeline" data-caption="Figure 1: My Focus Shift Timeline"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: My Focus Shift Timeline&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;My focus shifted from cloud native platform engineering to LLM application engineering, then to AI infrastructure, and finally to Agentic Runtime with governance and cost control.&lt;/p&gt;
&lt;p&gt;When AI workloads truly enter business scenarios, the core challenges engineers face also change:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Are inference, training, and evaluation competing for the same compute pool?&lt;/li&gt;
&lt;li&gt;Is GPU utilization consistently below expectations?&lt;/li&gt;
&lt;li&gt;Does cost scale linearly and uncontrollably with concurrency?&lt;/li&gt;
&lt;li&gt;Does the system have failure isolation and replay capabilities?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These issues go far beyond the code level.&lt;/p&gt;
&lt;h2 id="industry-consensus-ai-is-shifting-the-focus-of-engineering"&gt;Industry Consensus: AI Is Shifting the Focus of Engineering&lt;/h2&gt;
&lt;p&gt;By 2025, an industry consensus is emerging: AI is rewriting software engineering. But the real change is not happening in the IDE or code completion speed—it is reflected in &lt;strong&gt;the migration of engineering complexity&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Previously, complexity was concentrated in code and interfaces, and problems were solved through abstraction, refactoring, and testing.&lt;/p&gt;
&lt;p&gt;Now, complexity has shifted to the runtime, resource, and cost layers, and must be addressed through scheduling, isolation, observability, and governance.&lt;/p&gt;
&lt;p&gt;This is why the same AI tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Serve as &amp;ldquo;accelerators&amp;rdquo; for junior engineers&lt;/li&gt;
&lt;li&gt;But act as &amp;ldquo;magnifiers&amp;rdquo; for senior engineers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;AI tools amplify whether you truly understand how systems run in production.&lt;/p&gt;
&lt;h2 id="why-cost-becomes-a-first-principle"&gt;Why &amp;ldquo;Cost&amp;rdquo; Becomes a First Principle&lt;/h2&gt;
&lt;p&gt;In traditional cloud native systems, low CPU utilization is often just an efficiency issue; but in AI systems, &lt;strong&gt;low GPU utilization is often a cash flow problem&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In 2025, I repeatedly encountered scenarios like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Resources &amp;ldquo;seem insufficient,&amp;rdquo; but utilization is not actually high&lt;/li&gt;
&lt;li&gt;Scaling up to solve queuing issues ends up increasing unit costs&lt;/li&gt;
&lt;li&gt;The system lacks clear budget and quota boundaries, so throttling becomes the only way to stop the bleeding&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The root cause of these phenomena is not model selection, but &lt;strong&gt;the lack of a runtime and cost control plane tailored for AI workloads&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The following flowchart visually illustrates the cyclical relationship between GPU resources and cost pressures:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/gpu-cost-cycle-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/gpu-cost-cycle-en.svg" alt="Figure 2: GPU Resource and Cost Cycle in AI Systems" data-caption="Figure 2: GPU Resource and Cost Cycle in AI Systems"
width="2263"
height="320"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: GPU Resource and Cost Cycle in AI Systems&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In AI systems, limited GPU supply leads to queuing and waiting, which causes throughput to drop. Attempts to solve this through blind scaling only increase unit costs and create budget pressure, ultimately forcing the adoption of finer scheduling and governance strategies.&lt;/p&gt;
&lt;p&gt;Engineering problems ultimately manifest as cost issues.&lt;/p&gt;
&lt;h2 id="the-rise-of-agents-the-real-challenge-is-at-runtime"&gt;The Rise of Agents: The Real Challenge Is at Runtime&lt;/h2&gt;
&lt;p&gt;In 2025, Agent (Intelligent Agent, Agent, Intelligent Agent) became a hot topic; by 2026, it will enter the &amp;ldquo;can it actually run&amp;rdquo; stage.&lt;/p&gt;
&lt;p&gt;The challenge for Agents has never been about &amp;ldquo;how smart they are,&amp;rdquo; but rather:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether there are clear permission and data boundaries&lt;/li&gt;
&lt;li&gt;Whether they run in an isolated execution environment&lt;/li&gt;
&lt;li&gt;Whether they can be observed, evaluated, and replayed&lt;/li&gt;
&lt;li&gt;Whether they are subject to explicit cost and budget constraints&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These capabilities form the outline of &lt;strong&gt;Agentic Runtime (Agentic Runtime, Intelligent Agent Runtime)&lt;/strong&gt; that I have been trying to clarify throughout the year.&lt;/p&gt;
&lt;p&gt;The following flowchart shows the core capability layers of Agentic Runtime:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/agentic-runtime-layers-en.svg" data-img="https://assets.jimmysong.io/images/blog/software-engineering-shift-runtime-cost-2025/agentic-runtime-layers-en.svg" alt="Figure 3: Agentic Runtime Capability Layers" data-caption="Figure 3: Agentic Runtime Capability Layers"
width="463"
height="983"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Agentic Runtime Capability Layers&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Agentic Runtime builds from the foundation of Agents and workflows, connecting through orchestration and tool protocols, with the runtime managing state, memory, and evaluation. It provides secure execution environments (Sandbox and Policy), and ultimately implements a resource and cost control plane that unifies GPU, quota, and billing management.&lt;/p&gt;
&lt;p&gt;Without a runtime, an Agent is just a demo; without cost constraints, an Agent is just a risk amplifier.&lt;/p&gt;
&lt;h2 id="outlook-for-2026-the-foundation-of-engineering-matters-again"&gt;Outlook for 2026: The &amp;ldquo;Foundation&amp;rdquo; of Engineering Matters Again&lt;/h2&gt;
&lt;p&gt;Looking ahead to 2026, I remain cautiously optimistic.&lt;/p&gt;
&lt;p&gt;I do not believe the future belongs to &amp;ldquo;those who write the best prompts,&amp;rdquo; but more likely to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Those who understand runtime boundaries&lt;/li&gt;
&lt;li&gt;Those who can govern compute as a constrained resource&lt;/li&gt;
&lt;li&gt;Those who design AI systems as long-running systems&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From 2025 onwards, software engineering is no longer code-centric, but &lt;strong&gt;runtime and cost-centric&lt;/strong&gt;. This is not a regression, but a return: a return to being responsible for the whole system and for real-world constraints.&lt;/p&gt;
&lt;p&gt;For me personally, this is both a year-end summary and the direction I will continue to invest in for the coming years.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;In 2025, the focus of software engineering has shifted from code itself to runtime and cost governance. The rise of AI and Agents has not diminished the value of engineering, but has pushed complexity to a higher level. In the future, understanding runtime, managing compute and cost will become the new core competencies for engineers. I hope this year-end review provides some inspiration and reflection for fellow professionals.&lt;/p&gt;</content:encoded></item><item><title>From Cloud Native to AI Native: Why Kubernetes Is the Foundation for Next-Gen AI Agents</title><link>https://jimmysong.io/blog/ai-native-from-cloud-native/</link><pubDate>Wed, 24 Dec 2025 12:25:52 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-native-from-cloud-native/</guid><description>Explores why AI Agents need Kubernetes infrastructure and how Agent orchestration, MCP services, and AI gateways enable production-ready AI architectures.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;As a long-time practitioner in the cloud native field, I am increasingly convinced of one thing: &lt;strong&gt;AI Agents are not just a change in application form, but a migration of infrastructure paradigms.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;As artificial intelligence evolves from demos and copilots to systems that truly take on tasks and responsibilities, &lt;strong&gt;AI Agents&lt;/strong&gt; are becoming the new execution units in enterprise IT architectures. They not only &amp;ldquo;think,&amp;rdquo; but also &lt;strong&gt;act&lt;/strong&gt;: they can invoke tools, access systems, and collaborate to achieve goals.&lt;/p&gt;
&lt;p&gt;This raises an important question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What kind of infrastructure should such systems run on?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In my view, Kubernetes remains a solid choice for large-scale scenarios—but only if we &lt;strong&gt;reimagine Kubernetes in an AI-native way&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="cloud-native-challenges-for-production-grade-ai-agents"&gt;Cloud Native Challenges for Production-Grade AI Agents&lt;/h2&gt;
&lt;p&gt;In real production environments, AI Agents expose infrastructure needs that are fundamentally different from traditional microservices. Agents are not &amp;ldquo;just another HTTP service&amp;rdquo;; they have three distinct characteristics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Behavior is non-deterministic&lt;/strong&gt; (driven by model inference)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Execution paths are dynamic&lt;/strong&gt; (tool invocation cannot be fully enumerated in advance)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decisions must be auditable, constrained, and reviewable&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If we simply apply existing cloud native infrastructure, we quickly hit bottlenecks.&lt;/p&gt;
&lt;p&gt;The following table summarizes the main challenges and risks AI Agents face in cloud native environments:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Challenge Category&lt;/th&gt;
&lt;th&gt;Real Needs of Agents&lt;/th&gt;
&lt;th&gt;What Happens If Missing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Policy &amp;amp; Security&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dynamic control of tool and data access based on context, identity, and task&lt;/td&gt;
&lt;td&gt;Agents have &amp;ldquo;superuser&amp;rdquo; privileges, risks are uncontrollable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not just &amp;ldquo;did it succeed,&amp;rdquo; but also &lt;strong&gt;why was this decision made&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hard to debug, hard to review, hard to hold accountable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance &amp;amp; Consistency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Platform-level guardrails enforce organizational policies&lt;/td&gt;
&lt;td&gt;Each Agent could become a &amp;ldquo;shadow AI&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Challenges and Risks for AI Agents in Cloud Native Environments
&lt;/figcaption&gt;
&lt;p&gt;All these issues point to one conclusion:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI Agents must be treated as first-class citizens in Kubernetes, not just ordinary workloads.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="core-architecture-making-agents-native-kubernetes-objects"&gt;Core Architecture: Making Agents Native Kubernetes Objects&lt;/h2&gt;
&lt;p&gt;Looking back at the evolution of cloud native technologies, we&amp;rsquo;ve gone through similar stages:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Physical machines → Virtual machines&lt;/li&gt;
&lt;li&gt;Virtual machines → Containers&lt;/li&gt;
&lt;li&gt;Containers → Microservices&lt;/li&gt;
&lt;li&gt;Microservices → Declarative, governable platforms&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;AI Agents are simply the next step.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A production-ready AI Agent architecture requires at least three layers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Agent Orchestration Layer&lt;/strong&gt;: Declaratively define Agents&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool Service-ization Layer (MCP Services)&lt;/strong&gt;: Turn capabilities into governable services&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Native Data Plane / Gateway&lt;/strong&gt;: Unify policy, security, and protocols&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="agent-orchestration-layer-declarative-agent-management"&gt;Agent Orchestration Layer: Declarative Agent Management&lt;/h2&gt;
&lt;p&gt;Agents should no longer be &amp;ldquo;runtime objects&amp;rdquo; inside an SDK—they should be managed like Pods or Deployments.&lt;/p&gt;
&lt;p&gt;Key concepts:&lt;/p&gt;
&lt;h3 id="agents-as-kubernetes-resources"&gt;Agents as Kubernetes Resources&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Agents are defined using &lt;strong&gt;CRD (CustomResourceDefinition)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Lifecycle managed via &lt;code&gt;kubectl&lt;/code&gt; or GitOps&lt;/li&gt;
&lt;li&gt;Agent &lt;strong&gt;models, tools, and policies&lt;/strong&gt; are all explicitly declared&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A typical Agent definition includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agent logic&lt;/strong&gt; (inference loop)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model configuration&lt;/strong&gt; (specifying which large language model to use)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Callable toolset&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;This closely mirrors how we once decomposed &amp;ldquo;applications&amp;rdquo; into Deployments, Services, and ConfigMaps.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="tool-service-ization-layer-mcp-services-are-essential"&gt;Tool Service-ization Layer: MCP Services Are Essential&lt;/h2&gt;
&lt;p&gt;In Agent architectures, &lt;strong&gt;tools&lt;/strong&gt; are where real &amp;ldquo;actions&amp;rdquo; happen.&lt;/p&gt;
&lt;p&gt;Early MCP tools were often:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Local processes&lt;/li&gt;
&lt;li&gt;Tightly coupled to a single Agent&lt;/li&gt;
&lt;li&gt;Lacking versioning, permissions, and auditing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is unsustainable in enterprise environments.&lt;/p&gt;
&lt;h3 id="the-essence-of-mcp-service-ization"&gt;The Essence of MCP Service-ization&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Tools → &lt;strong&gt;Remote services&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Services → &lt;strong&gt;Kubernetes native workloads&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Capabilities → &lt;strong&gt;Reusable, governable, auditable&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This step is fundamentally similar to how we once turned scripts into microservices.&lt;/p&gt;
&lt;h2 id="ai-native-gateway-the-control-plane-entry-for-the-agent-world"&gt;AI Native Gateway: The &amp;ldquo;Control Plane Entry&amp;rdquo; for the Agent World&lt;/h2&gt;
&lt;p&gt;As the number of Agents grows and tools/models diversify, &lt;strong&gt;connectivity itself becomes a system risk&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Traditional API Gateways do not understand scenarios like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP&lt;/li&gt;
&lt;li&gt;Agent-to-Agent (A2A) communication&lt;/li&gt;
&lt;li&gt;Model invocation context&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thus, we need an &lt;strong&gt;AI native gateway&lt;/strong&gt; dedicated to mediation and governance.&lt;/p&gt;
&lt;p&gt;It must understand at least three types of traffic:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A2T&lt;/strong&gt;: Agent → Tool&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A2L&lt;/strong&gt;: Agent → LLM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A2A&lt;/strong&gt;: Agent ↔ Agent&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And enforce, across these paths:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Identity and authorization&lt;/li&gt;
&lt;li&gt;Policy and guardrails&lt;/li&gt;
&lt;li&gt;Auditing and rate limiting&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="architecture-overview"&gt;Architecture Overview&lt;/h2&gt;
&lt;p&gt;The diagram below illustrates the core layers and traffic paths of an AI-native system on Kubernetes:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-native-from-cloud-native/5be5cb784d4b228006abdf024bb99d6f.svg" data-img="https://assets.jimmysong.io/images/blog/ai-native-from-cloud-native/5be5cb784d4b228006abdf024bb99d6f.svg" alt="Figure 3: AI Native Architecture Layers and Traffic Paths" data-caption="Figure 3: AI Native Architecture Layers and Traffic Paths"
width="1311"
height="1642"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: AI Native Architecture Layers and Traffic Paths&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;AI Agents do not negate cloud native; on the contrary:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI Agents are the natural extension of cloud native in the era of intelligence.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Declarative → Agent definitions&lt;/li&gt;
&lt;li&gt;Service → MCP Services&lt;/li&gt;
&lt;li&gt;Service Mesh → AI Native Gateway&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If Kubernetes is the &amp;ldquo;automated factory,&amp;rdquo; then AI Agents are the &lt;strong&gt;intelligent workers who actually get things done&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;And the AI native gateway is the &lt;strong&gt;security and governance system tailored for these intelligent workers&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This is not an optional architecture—it is &lt;strong&gt;the only path for AI to reach production&lt;/strong&gt;.&lt;/p&gt;</content:encoded></item><item><title>AI Open Source Landscape: A One-Stop Guide to AI Project Navigation and Scoring System</title><link>https://jimmysong.io/blog/ai-oss-landscape-intro/</link><pubDate>Tue, 23 Dec 2025 08:34:05 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-oss-landscape-intro/</guid><description>Comprehensive introduction to the AI Open Source Landscape&amp;#39;s positioning, interface, scoring model, and data mechanisms to help developers efficiently discover quality AI projects.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The AI Open Source Landscape is not just a project directory, but an innovative attempt to bring transparency and quantifiability to the AI open source ecosystem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note: This article is intended for general readers and focuses on platform features and usage scenarios. If you want to see the technical details and formulas behind the scoring, please refer to:&lt;/strong&gt; AI Project Scoring and Inclusion Criteria.&lt;/p&gt;
&lt;h2 id="project-background-and-positioning"&gt;Project Background and Positioning&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape aims to provide developers, researchers, and enterprise users with a one-stop navigation and evaluation platform for AI open source projects. With the rapid development of large language models (LLM, Large Language Model), multimodal models (Multimodal Model), and other AI technologies, the open source community has seen a surge of innovative projects. However, information is scattered and quality varies, making it difficult for users to filter and make decisions.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/ai-oss-landscape.webp" data-img="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/ai-oss-landscape.webp" alt="Figure 1: AI Open Source Landscape" data-caption="Figure 1: AI Open Source Landscape"
width="3653"
height="2494"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: AI Open Source Landscape&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The AI Open Source Landscape systematically collects mainstream AI open source projects. As of the time of writing, it has included 851 open source projects. This landscape combines a multi-dimensional scoring system to help users efficiently discover, compare, and select the most suitable AI tools and frameworks for their needs. The platform not only focuses on models themselves, but also covers datasets, inference engines, evaluation tools, application frameworks, and the entire ecosystem chain, striving to promote transparency, quantifiability, and sustainable development in the AI open source ecosystem.&lt;/p&gt;
&lt;h2 id="main-interface-and-feature-highlights"&gt;Main Interface and Feature Highlights&lt;/h2&gt;
&lt;p&gt;The platform homepage presents project distribution in both landscape and list views, supporting category filtering, keyword search, and tag navigation to help users quickly locate target projects.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/project-details.webp" data-img="https://assets.jimmysong.io/images/blog/ai-oss-landscape-intro/project-details.webp" alt="Figure 2: Open Source Project Detail Page" data-caption="Figure 2: Open Source Project Detail Page"
width="2780"
height="2915"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Open Source Project Detail Page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For general readers, the main experience points include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Card view: One-sentence overview, star rating, and overall score for quick browsing and comparison.&lt;/li&gt;
&lt;li&gt;Health card: Displays overall health and key dimensions (activity, community, influence, sustainability) on the project page or sidebar, with the latest update marked for easy assessment of maintenance status.&lt;/li&gt;
&lt;li&gt;Detail page: Provides more background information, project links, and application scenarios to help you evaluate suitability for your needs.&lt;/li&gt;
&lt;li&gt;Smart badges: Visually display labels such as &amp;ldquo;Active&amp;rdquo;, &amp;ldquo;New Project&amp;rdquo;, &amp;ldquo;Popular&amp;rdquo;, &amp;ldquo;Archived&amp;rdquo; on cards, helping you quickly capture key project features.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you are interested in the specific rules for badge determination or scoring, detailed explanations are available on the Scoring Rules Page.&lt;/p&gt;
&lt;h2 id="scoring-and-ranking-mechanism"&gt;Scoring and Ranking Mechanism&lt;/h2&gt;
&lt;p&gt;The platform uses multi-dimensional scores to reflect the overall health and popularity of projects. The main dimensions include: &lt;strong&gt;Activity&lt;/strong&gt;, &lt;strong&gt;Community&lt;/strong&gt;, &lt;strong&gt;Quality&lt;/strong&gt;, &lt;strong&gt;Sustainability&lt;/strong&gt;, and the comprehensive &lt;strong&gt;Health&lt;/strong&gt; score. These scores help you quickly judge whether a project is suitable for production or experimentation.&lt;/p&gt;
&lt;h2 id="data-sources-and-update-mechanism"&gt;Data Sources and Update Mechanism&lt;/h2&gt;
&lt;p&gt;The platform&amp;rsquo;s data mainly comes from GitHub, project lists, official documentation, and community recommendations. We regularly and automatically synchronize and update metrics to ensure that the &amp;ldquo;last updated&amp;rdquo; and scores displayed on the interface reflect the current maintenance status of projects. Projects that have not been updated for a long time or are determined to be &amp;ldquo;inactive&amp;rdquo; are moved to the Archived Page. Archived projects remain searchable and retain historical scores, but will not appear in the default view of active rankings, making it easier for readers to focus on projects that are still maintained and active.&lt;/p&gt;
&lt;p&gt;For general readers, the key points are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The page displays key metrics and &amp;ldquo;last updated&amp;rdquo; time, helping you quickly judge whether a project is still maintained.&lt;/li&gt;
&lt;li&gt;The AI Open Source Landscape continuously iterates on the scoring model to improve fairness and differentiation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="how-to-contribute-and-correct-data"&gt;How to Contribute and Correct Data&lt;/h2&gt;
&lt;p&gt;If you want a project to be included or its data updated, you can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/rootsongjc/rootsongjc.github.io/issues/new?template=ai-resource.md" target="_blank" rel="noopener"&gt;Submit an AI Open Source Project Inclusion Request&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Keep the project&amp;rsquo;s README, License, documentation, and other information complete in the repository to facilitate our data collection and assessment.&lt;/li&gt;
&lt;li&gt;For faster synchronization or if you encounter data issues, contact the maintainers via project issues or raise a request in the site discussion area.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="typical-use-cases-or-user-feedback"&gt;Typical Use Cases or User Feedback&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape has been widely used in various scenarios such as AI developer selection, enterprise technology research, and academic studies. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Developers can quickly filter models or tools that meet their needs through the platform, saving significant research time.&lt;/li&gt;
&lt;li&gt;Enterprise technical teams use the ranking lists for competitor analysis and technology planning.&lt;/li&gt;
&lt;li&gt;Educational and research institutions refer to the landscape to understand trends in the AI open source ecosystem, supporting course design and topic selection.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Some users have commented that the platform is &amp;ldquo;comprehensive, well-structured, and fair in scoring,&amp;rdquo; greatly improving the efficiency of AI project selection and learning. Community suggestions continue to drive ongoing improvements in platform features and content.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The AI Open Source Landscape systematically and quantitatively organizes the AI open source ecosystem: the backend worker is responsible for reliable data collection and scoring calculations (supporting backfill and migration), while frontend components handle fast rendering and visualization (including smart badges, health cards, and metric explanations).&lt;/p&gt;
&lt;p&gt;If you want to learn more about the scoring details or participate in improvements:&lt;/p&gt;
&lt;p&gt;The community is welcome to join in evaluation, backfilling historical data, and refining scoring rules, working together to make the AI open source ecosystem more transparent and sustainable.&lt;/p&gt;</content:encoded></item><item><title>AI 2026: Infrastructure, Agents, and the Next Cloud-Native Shift</title><link>https://jimmysong.io/blog/ai-2026-infra-agentic-runtime/</link><pubDate>Fri, 19 Dec 2025 03:54:31 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-2026-infra-agentic-runtime/</guid><description>2026 AI&amp;#39;s turning point: not models, but infrastructure, agentic runtimes, GPU efficiency, and new organizational forms.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The real turning point for AI in 2026 is not autonomy, but the maturity of infrastructure - where agentic runtimes, GPU efficiency, and organizational design will decide who wins.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="introduction-2026-is-not-an-ai-moment-it-is-an-infrastructure-moment"&gt;Introduction: 2026 Is Not an AI Moment, It Is an Infrastructure Moment&lt;/h2&gt;
&lt;p&gt;Over the past fifteen years, every major shift in software has followed a familiar arc. Microservices were adopted not out of love for distributed systems, but because monoliths reached organizational limits. Kubernetes succeeded not because containers were novel, but because infrastructure finally matched how teams operated. Cloud native was never about YAML—it was about &lt;strong&gt;operability at scale&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;AI now stands at a similar inflection point.&lt;/p&gt;
&lt;p&gt;The central question for 2026 is not whether models will become more autonomous. That debate overlooks the core issue. Instead, the real question is whether AI can become &lt;strong&gt;operable, governable, and economically sustainable&lt;/strong&gt; within real systems.&lt;/p&gt;
&lt;p&gt;Most organizations today are limited not by intelligence, but by infrastructure: inefficient GPU utilization, escalating inference costs, fragile agent demos, and a tendency to treat AI as a feature rather than a runtime. The next phase of AI will be shaped not by model breakthroughs, but by the maturity of AI infrastructure and its ability to absorb responsibility.&lt;/p&gt;
&lt;h2 id="from-automation-to-capability-multiplication--a-familiar-cloud-native-pattern"&gt;From Automation to Capability Multiplication — A Familiar Cloud-Native Pattern&lt;/h2&gt;
&lt;p&gt;Reflecting on early cloud adoption, the dominant narrative was cost reduction: fewer servers, lower CapEx, elastic scaling. Yet, the true payoff emerged later, when teams realized cloud enabled &lt;strong&gt;entirely new operating models&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;AI is repeating this pattern.&lt;/p&gt;
&lt;p&gt;The following diagram illustrates the shift from automation to capability multiplication.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/from-automation-to-capability-multilication.svg" data-img="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/from-automation-to-capability-multilication.svg" alt="Figure 1: From Automation to Capability Multiplication" data-caption="Figure 1: From Automation to Capability Multiplication"
width="1642"
height="1214"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: From Automation to Capability Multiplication&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The first wave of AI focused on labor replacement. The second wave reframes AI as &lt;strong&gt;capability multiplication&lt;/strong&gt;: the same team, observing more signals, covering broader areas, and acting sooner.&lt;/p&gt;
&lt;p&gt;This mirrors the evolution of monitoring, tracing, and SRE practices. Rather than reducing engineers, these systems enabled continuous observation instead of occasional sampling.&lt;/p&gt;
&lt;p&gt;Preemptive AI systems—monitoring every interaction, log, and signal—are only viable if the underlying infrastructure can support them. This exposes a critical constraint: &lt;strong&gt;AI capability scales faster than AI infrastructure&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Without efficient scheduling, isolation, and utilization, multiplying capability simply multiplies cost.&lt;/p&gt;
&lt;h2 id="agents-are-becoming-distributed-systems-whether-we-admit-it-or-not"&gt;Agents Are Becoming Distributed Systems, Whether We Admit It or Not&lt;/h2&gt;
&lt;p&gt;The industry often discusses agents as products. In reality, agents are evolving into &lt;strong&gt;distributed systems&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The diagram below highlights this architectural shift.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/agents-are-becoming-distributed-systems.svg" data-img="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/agents-are-becoming-distributed-systems.svg" alt="Figure 2: Agents Are Becoming Distributed Systems" data-caption="Figure 2: Agents Are Becoming Distributed Systems"
width="1102"
height="1382"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Agents Are Becoming Distributed Systems&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Single-agent designs resemble early monoliths: impressive demos, fragile behavior, and opaque failure modes. As tasks grow in complexity, systems must decompose work into planning, execution, verification, and review—making coordination inevitable.&lt;/p&gt;
&lt;p&gt;This is not merely a philosophical change, but an architectural one.&lt;/p&gt;
&lt;p&gt;Multi-agent systems introduce challenges familiar from the microservices era:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Coordination and orchestration&lt;/li&gt;
&lt;li&gt;Resource contention&lt;/li&gt;
&lt;li&gt;Fault isolation&lt;/li&gt;
&lt;li&gt;Observability and rollback&lt;/li&gt;
&lt;li&gt;Deterministic artifacts between stages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Labeling this as &amp;ldquo;multi-agent collaboration&amp;rdquo; can be misleading. What is actually occurring is &lt;strong&gt;workload decomposition and control-plane emergence&lt;/strong&gt;. Agents are transitioning from tools to workloads competing for limited resources.&lt;/p&gt;
&lt;p&gt;Recognizing this clarifies why agent progress is inseparable from infrastructure maturity.&lt;/p&gt;
&lt;h2 id="ai-infra-is-the-missing-layer-between-models-and-organizations"&gt;AI Infra Is the Missing Layer Between Models and Organizations&lt;/h2&gt;
&lt;p&gt;Cloud native taught us that abstractions only scale when a control plane exists.&lt;/p&gt;
&lt;p&gt;Currently, AI lacks a mature control plane.&lt;/p&gt;
&lt;p&gt;The following image demonstrates the gap between models and organizations.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/ai-infra-is-the-missing-layer-between-models-and-organizations.svg" data-img="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/ai-infra-is-the-missing-layer-between-models-and-organizations.svg" alt="Figure 3: AI Infra Is the Missing Layer Between Models and Organizations" data-caption="Figure 3: AI Infra Is the Missing Layer Between Models and Organizations"
width="1102"
height="1260"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: AI Infra Is the Missing Layer Between Models and Organizations&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Models are powerful, but the surrounding infrastructure—scheduling, isolation, quota enforcement, cost attribution, observability—remains primitive, especially at the GPU layer.&lt;/p&gt;
&lt;p&gt;GPUs are expensive, scarce, and often underutilized. In many environments, utilization remains below 30–40%, while inference costs continue to rise. Training pipelines monopolize resources, inference workloads spike unpredictably, and organizations must choose between waste and throttling innovation.&lt;/p&gt;
&lt;p&gt;This is not a model problem. It is fundamentally an &lt;strong&gt;AI infrastructure problem&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The next phase of AI will depend on treating GPUs as we learned to treat CPUs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Fine-grained allocation&lt;/li&gt;
&lt;li&gt;Fair sharing&lt;/li&gt;
&lt;li&gt;Preemption and prioritization&lt;/li&gt;
&lt;li&gt;Clear ownership and accounting&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Until GPU utilization becomes a primary design goal, AI systems will remain economically fragile.&lt;/p&gt;
&lt;h2 id="domain-expertise-matters-because-infrastructure-finally-exposes-it"&gt;Domain Expertise Matters Because Infrastructure Finally Exposes It&lt;/h2&gt;
&lt;p&gt;As models plateau in general reasoning, differentiation shifts elsewhere.&lt;/p&gt;
&lt;p&gt;The diagram below illustrates how infrastructure exposes domain expertise.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/domain-expertise-matters-because-infrastructure-finally-exposes-it.svg" data-img="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/domain-expertise-matters-because-infrastructure-finally-exposes-it.svg" alt="Figure 4: Domain Expertise Matters Because Infrastructure Finally Exposes It" data-caption="Figure 4: Domain Expertise Matters Because Infrastructure Finally Exposes It"
width="1482"
height="1302"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Domain Expertise Matters Because Infrastructure Finally Exposes It&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In cloud-native systems, competitive advantage eventually moved from frameworks to &lt;strong&gt;operational excellence&lt;/strong&gt;: superior runbooks, incident response, and cost control. AI is following a similar trajectory.&lt;/p&gt;
&lt;p&gt;High-value AI systems must operate within dense, rule-heavy domains such as finance, healthcare, manufacturing, and infrastructure operations. What matters is not abstract intelligence, but the ability to encode domain constraints, exceptions, and failure patterns.&lt;/p&gt;
&lt;p&gt;Here, domain experts become central—not as prompt engineers, but as &lt;strong&gt;system shapers&lt;/strong&gt;. Their decisions define agent permissions, human intervention points, and error containment strategies.&lt;/p&gt;
&lt;p&gt;Infrastructure determines whether this expertise can be safely operationalized.&lt;/p&gt;
&lt;h2 id="simulation-is-becoming-the-new-staging-environment-for-ai"&gt;Simulation Is Becoming the New Staging Environment for AI&lt;/h2&gt;
&lt;p&gt;One of the most important lessons from cloud-native operations: distributed systems are not tested in production.&lt;/p&gt;
&lt;p&gt;AI systems that act, plan, and modify state are no exception.&lt;/p&gt;
&lt;p&gt;The following image shows simulation as the new staging environment.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/simulation-is-becoming-the-new-staging-environment-for-ai.svg" data-img="https://assets.jimmysong.io/images/blog/ai-2026-infra-agentic-runtime/simulation-is-becoming-the-new-staging-environment-for-ai.svg" alt="Figure 5: Simulation Is Becoming the New Staging Environment for AI" data-caption="Figure 5: Simulation Is Becoming the New Staging Environment for AI"
width="1062"
height="1482"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Simulation Is Becoming the New Staging Environment for AI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Training and validating agents directly in live environments is unsustainable. The future lies in &lt;strong&gt;simulation-first AI development&lt;/strong&gt;—sandboxed environments that mirror real systems, workloads, and constraints.&lt;/p&gt;
&lt;p&gt;This approach is analogous to staging clusters, chaos engineering, and load testing, but elevated for decision-making systems. Evaluation shifts from static benchmarks to behavioral metrics: intervention rates, rollback frequency, and cost impact.&lt;/p&gt;
&lt;p&gt;Organizations that build these environments will advance faster and safer. Those that do not may remain limited by conservative deployments and restricted autonomy.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Technological revolutions succeed not on novelty alone, but when infrastructure, tooling, and organizational models align.&lt;/p&gt;
&lt;p&gt;AI is nearing that pivotal moment.&lt;/p&gt;
&lt;p&gt;The leaders in 2026 will be those who:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Treat AI as a runtime, not just a feature&lt;/li&gt;
&lt;li&gt;Optimize for resource efficiency, especially GPUs&lt;/li&gt;
&lt;li&gt;Recognize agents as distributed systems&lt;/li&gt;
&lt;li&gt;Redesign organizations around continuous learning systems&lt;/li&gt;
&lt;li&gt;Invest in infrastructure ahead of autonomy&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;AI is no longer just a model problem. It is an infrastructure challenge—and the next phase will be decided not in labs, but in production systems.&lt;/p&gt;</content:encoded></item><item><title>What I Saw at COSCon'25: The Real State of Open Source in China</title><link>https://jimmysong.io/blog/coscon-2025-china-open-source-observation/</link><pubDate>Thu, 18 Dec 2025 06:14:51 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/coscon-2025-china-open-source-observation/</guid><description>From an engineering and organizer&amp;#39;s perspective, real changes at COSCon&amp;#39;25: AI as the default backdrop, discussions returning to engineering issues, and Chinese open source entering a long-term phase.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Attending COSCon'25 in Beijing, I observed firsthand how open source in China is shifting: AI is now the default context, discussions are grounded in real engineering, and the community is embracing long-term thinking. These are not just trends—they are the new reality.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In early December this year, I attended COSCon'25, the China Open Source Annual Conference, in Beijing. Although I have worked in open source for many years, this was my first time participating in an event organized by the Open Source Society—and I joined as a sub-forum producer. Previously, I thought such conferences were too high-level or disconnected from reality, but after actually taking part, I found there was much to gain.&lt;/p&gt;
&lt;p&gt;A quick note: &lt;strong&gt;this article is not an official conference summary or review&lt;/strong&gt;. The organizers have already published detailed information about the event&amp;rsquo;s scale, attendee numbers, and forum sessions. If you&amp;rsquo;re interested in those details, please refer to the official article:
&lt;a href="https://mp.weixin.qq.com/s/1Q5xBUEmSN9MXon03P00lA" target="_blank" rel="noopener"&gt;COSCon'25: The 10th China Open Source Annual Conference Successfully Concludes in Beijing—A Comprehensive Recap!&lt;/a&gt;&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/coscon-2025-china-open-source-observation/banner.webp" data-img="https://assets.jimmysong.io/images/blog/coscon-2025-china-open-source-observation/banner.webp" alt="Figure 1: 10th COSCon Venue" data-caption="Figure 1: 10th COSCon Venue"
width="1080"
height="716"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: 10th COSCon Venue&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What I want to share is this: &lt;strong&gt;Standing on site, on the engineering front lines, and as an organizer rather than an audience member, I saw real changes happening in Chinese open source.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="this-coscon-no-more-trying-to-prove-open-source-matters"&gt;This COSCon: No More Trying to &amp;ldquo;Prove Open Source Matters&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;One clear impression:
&lt;strong&gt;Almost no one spent time arguing &amp;ldquo;why do open source&amp;rdquo; anymore.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In earlier years, common narratives included:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Is open source safe?&lt;/li&gt;
&lt;li&gt;Can open source be commercialized?&lt;/li&gt;
&lt;li&gt;Can China create its own open source projects?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But at COSCon'25, these questions were basically assumed as &amp;ldquo;background conditions.&amp;rdquo; The focus shifted to &lt;strong&gt;those already doing open source, and what comes next&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This doesn&amp;rsquo;t mean the issues have disappeared, but it does mean:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In China&amp;rsquo;s engineering circles, open source is no longer a &amp;ldquo;philosophical choice&amp;rdquo;—it&amp;rsquo;s a practical way of working.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="ai-as-background-noise-not-the-main-character"&gt;AI as Background Noise, Not the Main Character&lt;/h2&gt;
&lt;p&gt;The theme of this year&amp;rsquo;s conference was Open Source × Open Intelligence, but interestingly, &lt;strong&gt;AI did not take center stage&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Instead, it was more like background noise—
Almost every topic touched on AI, but no one was giving talks solely &amp;ldquo;about AI.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;You would see it repeatedly in areas like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cloud native scheduling, focusing on GPU / NPU / heterogeneous resources&lt;/li&gt;
&lt;li&gt;Storage and data, focusing on data paths for training and inference&lt;/li&gt;
&lt;li&gt;Serverless, focusing on LLM cold starts and elasticity&lt;/li&gt;
&lt;li&gt;Observability, focusing on what to do when system complexity gets out of hand&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;AI was not treated as a &amp;ldquo;hot trend,&amp;rdquo; but as &lt;strong&gt;a new workload reality&lt;/strong&gt;.
This is a significant change, though not one easily captured in press releases.&lt;/p&gt;
&lt;h2 id="real-impressions-as-a-cloud-native-sub-forum-producer"&gt;Real Impressions as a Cloud Native Sub-forum Producer&lt;/h2&gt;
&lt;p&gt;I helped organize the cloud native open source sub-forum at this year&amp;rsquo;s conference. This role gave me a perspective very different from that of a typical attendee.&lt;/p&gt;
&lt;h3 id="first-topics-clearly-converged-on-engineering-problems"&gt;First, Topics Clearly Converged on &amp;ldquo;Engineering Problems&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;There were almost no talks about Kubernetes concepts;
Very few about &amp;ldquo;architectural philosophies.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Instead, the focus was on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What pitfalls did you encounter at what scale?&lt;/li&gt;
&lt;li&gt;Why did you choose this solution over another?&lt;/li&gt;
&lt;li&gt;Which problems remain unsolved?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Many presentations weren&amp;rsquo;t &amp;ldquo;pleasant to hear,&amp;rdquo; but they were very real.&lt;/p&gt;
&lt;h3 id="second-the-boundary-between-academia-and-industry-is-thinning"&gt;Second, The Boundary Between Academia and Industry Is Thinning&lt;/h3&gt;
&lt;p&gt;This was especially evident this year.&lt;/p&gt;
&lt;p&gt;Some talks from universities and research institutes were no longer just &amp;ldquo;from a paper&amp;rsquo;s perspective,&amp;rdquo; but directly addressed core issues in industrial systems, such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cold start of serverless LLMs&lt;/li&gt;
&lt;li&gt;The real value of RDMA in inference paths&lt;/li&gt;
&lt;li&gt;Whether prefill/decode separation is truly feasible in engineering&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These topics may not be immediately applicable, but &lt;strong&gt;they are now colliding head-on with engineering problems&lt;/strong&gt;, rather than talking past each other.&lt;/p&gt;
&lt;h3 id="third-open-source-is-no-longer-just-about-code"&gt;Third, Open Source Is No Longer Just About Code&lt;/h3&gt;
&lt;p&gt;In many discussions, &amp;ldquo;governance,&amp;rdquo; &amp;ldquo;maintenance cost,&amp;rdquo; and &amp;ldquo;community collaboration&amp;rdquo; came up frequently.&lt;/p&gt;
&lt;p&gt;This is a signal:
When a project is truly being used, &lt;strong&gt;code is no longer the hardest part&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="main-forum-more-questions-not-answers"&gt;Main Forum: More Questions, Not Answers&lt;/h2&gt;
&lt;p&gt;If I had to sum up the main forum in one sentence:
&lt;strong&gt;It kept raising questions, but wasn&amp;rsquo;t in a hurry to provide answers.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Has the boundary of open source changed in the AI era?&lt;/li&gt;
&lt;li&gt;Should models, data, and chips become part of the open source core?&lt;/li&gt;
&lt;li&gt;Are developers&amp;rsquo; roles being redefined?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There are no standard answers to these questions, but the fact that they are being raised repeatedly shows they have become common concerns, not just the thoughts of a few.&lt;/p&gt;
&lt;h2 id="exhibition-area-and-sub-forums-closer-to-the-real-ecosystem"&gt;Exhibition Area and Sub-forums: Closer to the Real Ecosystem&lt;/h2&gt;
&lt;p&gt;Compared to the main forum, I personally paid more attention to the sub-forums and exhibition area.&lt;/p&gt;
&lt;p&gt;There, you would see:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many projects no longer emphasize &amp;ldquo;who they want to replace&amp;rdquo;&lt;/li&gt;
&lt;li&gt;More discussions about &amp;ldquo;who they can work with&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Several communities are seriously discussing long-term maintenance, not just releasing versions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This may not be glamorous, but it&amp;rsquo;s important.&lt;/p&gt;
&lt;h2 id="a-personal-judgment"&gt;A Personal Judgment&lt;/h2&gt;
&lt;p&gt;If I had to make a judgment about COSCon'25, I would say:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Chinese open source is shifting from &amp;ldquo;can we do it&amp;rdquo; to &amp;ldquo;can we sustain it for the long term.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is a more difficult, but also more realistic, stage.&lt;/p&gt;
&lt;p&gt;This COSCon did not try to create a grand narrative. Instead, it felt like a &amp;ldquo;status exposure&amp;rdquo; at a particular stage:
There are more questions, participants are more diverse, but the discussions are also closer to the real world.&lt;/p&gt;
&lt;p&gt;Open source doesn&amp;rsquo;t depend on a single conference to move forward, but being on site helps you see more clearly:
&lt;strong&gt;Where exactly are we standing right now?&lt;/strong&gt;&lt;/p&gt;</content:encoded></item><item><title>Decoding Goose: Why It Joined AAIF and What This Means for Agentic Runtime</title><link>https://jimmysong.io/blog/goose-aaif-agentic-runtime/</link><pubDate>Fri, 12 Dec 2025 08:16:48 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/goose-aaif-agentic-runtime/</guid><description>An analysis of Block&amp;#39;s Goose project, why it became one of the first Agentic AI Foundation (AAIF) projects, and what this means for Agentic Runtime and the evolution of AI-Native infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Goose is not a project that excites you at first glance in this wave of Agent innovation, but its entry into AAIF signals a deeper shift in how we think about Agentic Runtime and AI-Native infrastructure.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;At first glance, &lt;a href="https://github.com/block/goose" target="_blank" rel="noopener"&gt;Goose&lt;/a&gt; is not a project that immediately excites people.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/goose-aaif-agentic-runtime/goose.webp" data-img="https://assets.jimmysong.io/images/blog/goose-aaif-agentic-runtime/goose.webp" alt="Figure 1: Goose App UI" data-caption="Figure 1: Goose App UI"
width="2622"
height="2360"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Goose App UI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It doesn&amp;rsquo;t have flashy demos, nor does it showcase overwhelming multimodal capabilities, and it certainly doesn&amp;rsquo;t look like an AI product aimed at consumers. Yet, this seemingly &amp;ldquo;plain&amp;rdquo; project became one of the first donations to the Agentic AI Foundation (AAIF), standing alongside Anthropic&amp;rsquo;s MCP and OpenAI&amp;rsquo;s AGENTS.md.&lt;/p&gt;
&lt;p&gt;This fact alone is worth a closer look.&lt;/p&gt;
&lt;p&gt;This article does not aim to prove how powerful Goose is, but rather to answer three more practical questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What overlooked but long-term critical problems does Goose actually solve?&lt;/li&gt;
&lt;li&gt;Why was it Goose, and not another Agent framework, that entered AAIF?&lt;/li&gt;
&lt;li&gt;What does this mean for Agentic Runtime and AI-Native infrastructure, which I care about?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="gooses-true-positioning-its-not-an-ide-or-a-chatbot"&gt;Goose&amp;rsquo;s True Positioning: It&amp;rsquo;s Not an IDE or a Chatbot&lt;/h2&gt;
&lt;p&gt;If you only look at its surface features, Goose is easily mistaken for one of two things:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &amp;ldquo;multi-model AI desktop client&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Or an &amp;ldquo;intelligent programming assistant that can run commands&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But inside Block, it was never designed as a &amp;ldquo;tool&amp;rdquo; from the start.&lt;/p&gt;
&lt;p&gt;Goose&amp;rsquo;s origin is closely tied to Block&amp;rsquo;s engineering environment.&lt;/p&gt;
&lt;p&gt;Block (formerly Square) is a classic engineering-driven company: complex systems, high automation needs, many internal tools, and very high execution costs in real production environments. In its recent AI transformation, Block did not focus on &amp;ldquo;which model to choose&amp;rdquo; or &amp;ldquo;which AI tool to introduce,&amp;rdquo; but directly targeted the engineering execution layer itself.&lt;/p&gt;
&lt;p&gt;Goose was born in this context.&lt;/p&gt;
&lt;p&gt;Its goal is not to &amp;ldquo;help people code faster,&amp;rdquo; but to enable models to &lt;strong&gt;stably and controllably take action&lt;/strong&gt;: run tests, modify code, drive UIs, call internal systems, and operate reliably in real engineering environments.&lt;/p&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Goose is more like an executable Agent Runtime than a conversation-centric product.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="blocks-ai-transformation-started-with-organization-not-tools"&gt;Block&amp;rsquo;s AI Transformation Started with Organization, Not Tools&lt;/h2&gt;
&lt;p&gt;To understand Goose, you can&amp;rsquo;t ignore a key organizational shift at Block.&lt;/p&gt;
&lt;p&gt;In an interview with Block&amp;rsquo;s CTO, one signal was very clear: the starting point for AI transformation was not buying tools or stacking models, but the organizational structure itself.&lt;/p&gt;
&lt;p&gt;Block shifted from a business-line GM model to a more functionally oriented structure, making engineering and design the company&amp;rsquo;s core scheduling units again. This is essentially a proactive response to Conway&amp;rsquo;s Law.&lt;/p&gt;
&lt;p&gt;If the organizational structure doesn&amp;rsquo;t allow technical capabilities to be orchestrated centrally, Agents will ultimately remain &amp;ldquo;personal assistants&amp;rdquo; or &amp;ldquo;engineering toys.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;From this perspective, Goose is not just a tool, but a &lt;strong&gt;cultural signal&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Every employee can use AI to build and execute real system behaviors.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This also explains a fact many overlook:
Goose was not packaged as SaaS, nor was it rushed to commercialization, but was open-sourced and rapidly standardized.&lt;/p&gt;
&lt;p&gt;Because its role inside Block is closer to an &amp;ldquo;operating system for execution models&amp;rdquo; than a product that can be sold separately.&lt;/p&gt;
&lt;h2 id="why-did-goose-enter-aaif-not-because-its-technically-strongest"&gt;Why Did Goose Enter AAIF? Not Because It&amp;rsquo;s &amp;ldquo;Technically Strongest&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;This is what confuses outsiders the most.&lt;/p&gt;
&lt;p&gt;If you only look at flashy features, model support, or community popularity, Goose doesn&amp;rsquo;t stand out. But AAIF&amp;rsquo;s choice was not about &amp;ldquo;maximum capability,&amp;rdquo; but about &lt;strong&gt;whether the position is right&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Looking at the first batch of AAIF projects, a clear chain emerges:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP (Anthropic): Defines how models safely and standardly call tools&lt;/li&gt;
&lt;li&gt;AGENTS.md (OpenAI): Defines behavioral conventions for Agents in code repositories&lt;/li&gt;
&lt;li&gt;Goose (Block): A real, runnable, open-source Agent execution framework&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Goose&amp;rsquo;s role is not to set new protocols, but to serve as the &lt;strong&gt;practical carrier and reference implementation&lt;/strong&gt; for these protocols.&lt;/p&gt;
&lt;p&gt;It proves one thing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MCP is not just a paper standard&lt;/li&gt;
&lt;li&gt;Agents are not just research concepts&lt;/li&gt;
&lt;li&gt;In real enterprise environments, they can actually run&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From this angle, Goose&amp;rsquo;s &amp;ldquo;ordinariness&amp;rdquo; is actually an advantage.&lt;/p&gt;
&lt;p&gt;It is not tied to Block&amp;rsquo;s business moat, nor does it have irreplaceable private APIs; it can be forked, replaced, audited—&amp;ldquo;boring&amp;rdquo; enough, and neutral enough.&lt;/p&gt;
&lt;p&gt;And that is the most important trait of public infrastructure.&lt;/p&gt;
&lt;h2 id="gooses-value-lies-not-in-today-but-in-23-years"&gt;Goose&amp;rsquo;s Value Lies Not in Today, But in 2–3 Years&lt;/h2&gt;
&lt;p&gt;From a longer-term perspective, Goose&amp;rsquo;s value becomes clearer.&lt;/p&gt;
&lt;p&gt;What we&amp;rsquo;re experiencing now is much like the early days of containers:
Most Agent projects today are demos, IDE plugins, or workflow wrappers, but what&amp;rsquo;s really missing is a &lt;strong&gt;sustainable, schedulable, observable execution layer&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Goose is already moving in this direction.&lt;/p&gt;
&lt;p&gt;Block&amp;rsquo;s metrics for Goose&amp;rsquo;s success are straightforward:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How many human hours are saved each week&lt;/li&gt;
&lt;li&gt;How much non-technical teams reduce their dependence on engineering teams&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Behind this is a judgment I&amp;rsquo;m increasingly convinced of:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What enterprises truly need is not &amp;ldquo;smarter models,&amp;rdquo; but &amp;ldquo;cheaper execution.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The long-term value of Agents is not in generation quality, but in execution substitution rate.&lt;/p&gt;
&lt;h2 id="aaif-is-an-attempt-at-infrastructure-level-consensus"&gt;AAIF Is an Attempt at Infrastructure-Level Consensus&lt;/h2&gt;
&lt;p&gt;Just as CNCF did for cloud native, AAIF is not guaranteed to succeed.&lt;/p&gt;
&lt;p&gt;But it at least marks a shift:
Agents are no longer just application-layer innovations, but are beginning to enter the stage of infrastructure-layer collaboration.&lt;/p&gt;
&lt;p&gt;As a reference implementation, Goose is likely to remain in this ecosystem for a long time—even if it is replaced, rewritten, or evolved in the future.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;If you see Goose as a &amp;ldquo;product,&amp;rdquo; it is indeed not dazzling.&lt;/p&gt;
&lt;p&gt;But if you place it in the long-term evolution path of Agentic AI, its significance becomes clear:&lt;/p&gt;
&lt;p&gt;It is not the end, but a necessary intermediate state.&lt;/p&gt;
&lt;p&gt;For me, the emergence of Goose further confirms one thing:&lt;/p&gt;
&lt;p&gt;Agentic Runtime is not a conceptual problem, but an engineering and organizational one.&lt;/p&gt;
&lt;p&gt;And that is one of the most worthwhile directions to invest energy in over the next few years.&lt;/p&gt;</content:encoded></item><item><title>ARK: Multi-Agent Systems Are Finally Entering the Engineer's World</title><link>https://jimmysong.io/blog/ark-agentic-runtime-for-kubernetes/</link><pubDate>Thu, 11 Dec 2025 13:19:42 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ark-agentic-runtime-for-kubernetes/</guid><description>How ARK uses cloud-native architecture and declarative runtime to drive engineering adoption of multi-agent systems and shape the Agentic Runtime ecosystem.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The deep integration of cloud native and AI, with the ARK platform, provides a new paradigm for engineering multi-agent systems.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;AI Agents are moving from the &amp;ldquo;single agent demo&amp;rdquo; stage to &amp;ldquo;large-scale operation.&amp;rdquo; The real challenge does not lie in the model itself, but in engineering issues at runtime: model management, tool invocation, state maintenance, elastic scaling, team collaboration, observability, deployment, and upgrades. These are problems that traditional agent libraries struggle to solve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ARK (Agentic Runtime for Kubernetes)&lt;/strong&gt; provides a fully operational, observable, governable, and continuously deliverable multi-agent operating system. It is not a Python library, but a complete runtime platform.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/ark-dashboard-homepage.webp" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/ark-dashboard-homepage.webp" alt="Figure 15: ARK Dashboard" data-caption="Figure 15: ARK Dashboard"
width="3176"
height="1822"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 15: ARK Dashboard&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Note: In this article, ARK refers to McKinsey&amp;rsquo;s open-source &lt;a href="https://github.com/mckinsey/ark-agent-runtime-for-kubernetes" target="_blank" rel="noopener"&gt;ARK Agent Runtime for Kubernetes&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This article, from an engineer&amp;rsquo;s perspective, will reorganize ARK&amp;rsquo;s core capabilities and answer the following questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What engineering challenges does ARK actually solve?&lt;/li&gt;
&lt;li&gt;Why is it worth special attention in the cloud native field?&lt;/li&gt;
&lt;li&gt;How is it fundamentally different from frameworks like LangChain and CrewAI?&lt;/li&gt;
&lt;li&gt;What insights does it offer for the Agentic Runtime ecosystem?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="ark-architecture-treating-agents-as-kubernetes-native-workloads"&gt;ARK Architecture: Treating Agents as Kubernetes-Native Workloads&lt;/h2&gt;
&lt;p&gt;The core idea of ARK is: &lt;strong&gt;An agent is not a script, but a schedulable, governable, and observable Kubernetes workload.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The following architecture diagram illustrates ARK&amp;rsquo;s underlying structure.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/168007ae485fa14769e5483aa20805d3.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/168007ae485fa14769e5483aa20805d3.svg" alt="Figure 16: ARK Overall Architecture" data-caption="Figure 16: ARK Overall Architecture"
width="2060"
height="1146"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 16: ARK Overall Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This diagram highlights ARK&amp;rsquo;s key design points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CRDs declare requirements&lt;/strong&gt; (Agent, Model, Team, Tool, Memory, etc.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Controller translates declarations into actual Pods/Services&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The API provides a unified communication entry point and team orchestration&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory supports long-term state management for agents&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MCP Server enables external systems to become tools&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dashboard provides visual management and observability&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;ARK adopts the typical cloud-native Operator pattern and applies it to multi-agent systems.&lt;/p&gt;
&lt;h2 id="crd-arks-abstraction-layer"&gt;CRD: ARK&amp;rsquo;s &amp;ldquo;Abstraction Layer&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Unlike traditional agent frameworks where &amp;ldquo;code is logic,&amp;rdquo; ARK uses CRDs (Custom Resource Definitions) to abstract the components of agent applications.&lt;/p&gt;
&lt;p&gt;The main CRD types in ARK include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model&lt;/li&gt;
&lt;li&gt;Agent&lt;/li&gt;
&lt;li&gt;Team&lt;/li&gt;
&lt;li&gt;Tool&lt;/li&gt;
&lt;li&gt;Memory&lt;/li&gt;
&lt;li&gt;Evaluation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These CRDs correspond to all the key components of an agent system.&lt;/p&gt;
&lt;p&gt;The following diagram shows the structure of the CRDs:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/b464d2b85b6d664b51fa48a5aed2fbd0.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/b464d2b85b6d664b51fa48a5aed2fbd0.svg" alt="Figure 17: CRD Structure (Simplified)" data-caption="Figure 17: CRD Structure (Simplified)"
width="795"
height="829"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 17: CRD Structure (Simplified)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Through CRDs, ARK achieves the following engineering features:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;All resources are GitOps-ready&lt;/strong&gt;, supporting declarative management&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Changes are auditable, reversible, and continuously deliverable&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The evolution of models, tools, and agents does not require business code changes&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the key gene of ARK&amp;rsquo;s engineering-oriented system.&lt;/p&gt;
&lt;h2 id="agent-execution-flow-from-query-to-tool-invocation"&gt;Agent Execution Flow: From Query to Tool Invocation&lt;/h2&gt;
&lt;p&gt;The following image shows how to view query details in the ARK Dashboard.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/ark-dashboard-queries.webp" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/ark-dashboard-queries.webp" alt="Figure 18: Viewing Query Details in ARK Dashboard" data-caption="Figure 18: Viewing Query Details in ARK Dashboard"
width="3176"
height="1822"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 18: Viewing Query Details in ARK Dashboard&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In ARK, the complete execution flow for an agent receiving a query is as follows:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/67a8b1142ee63f7cacd4d907cd198ce4.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/67a8b1142ee63f7cacd4d907cd198ce4.svg" alt="Figure 19: Agent Execution Flow" data-caption="Figure 19: Agent Execution Flow"
width="1146"
height="591"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 19: Agent Execution Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This flow has the following characteristics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Memory modules are naturally involved in the execution flow, without code specialization&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Large language model (LLM, Large Language Model) and tool invocation are governed by the runtime&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agents can reside in Pods long-term, not just as one-off processes&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This makes ARK more like an &amp;ldquo;agent microservice platform.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Below is an example of a request and response:&lt;/p&gt;
&lt;h2 id="the-true-value-of-multi-agent-team-orchestration"&gt;The True Value of Multi-Agent: Team Orchestration&lt;/h2&gt;
&lt;p&gt;ARK&amp;rsquo;s Team CRD allows multiple agents to be woven into a higher-level &amp;ldquo;system,&amp;rdquo; enabling multi-agent collaboration.&lt;/p&gt;
&lt;p&gt;The following diagram shows the collaboration model of a multi-agent team:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/0fb6990e479cd7b5c0ff3c8e8626693b.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/0fb6990e479cd7b5c0ff3c8e8626693b.svg" alt="Figure 20: Multi-Agent Team Collaboration" data-caption="Figure 20: Multi-Agent Team Collaboration"
width="786"
height="499"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 20: Multi-Agent Team Collaboration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;The engineering value of Team is reflected in:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Making &amp;ldquo;expert collaboration&amp;rdquo; declarative and configurable&lt;/li&gt;
&lt;li&gt;Flexible strategies (such as polling, role assignment, routing, etc.)&lt;/li&gt;
&lt;li&gt;A2A Gateway handles message passing&lt;/li&gt;
&lt;li&gt;The Team itself is observable (every round of collaboration is logged)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For enterprises, this means the &amp;ldquo;agent organizational structure&amp;rdquo; can be standardized, replayed, and tuned.&lt;/p&gt;
&lt;h2 id="fundamental-differences-between-ark-and-other-frameworks"&gt;Fundamental Differences Between ARK and Other Frameworks&lt;/h2&gt;
&lt;p&gt;Many engineers, upon first seeing ARK, may wonder:&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Is it just LangChain or CrewAI wrapped in Kubernetes?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;In fact, there are fundamental differences. The following diagram compares the structural differences between ARK and mainstream agent frameworks:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/8da9b272d1930a2356a6401b6615d134.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-for-kubernetes/8da9b272d1930a2356a6401b6615d134.svg" alt="Figure 21: ARK vs LangChain / AutoGPT / CrewAI" data-caption="Figure 21: ARK vs LangChain / AutoGPT / CrewAI"
width="3454"
height="345"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 21: ARK vs LangChain / AutoGPT / CrewAI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The table below further summarizes the key differences:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Traditional Agent Libraries&lt;/th&gt;
&lt;th&gt;ARK&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core Pattern&lt;/td&gt;
&lt;td&gt;Write Python code&lt;/td&gt;
&lt;td&gt;Write CRDs (declarative)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Local/Container&lt;/td&gt;
&lt;td&gt;Kubernetes-native scheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State&lt;/td&gt;
&lt;td&gt;Managed inside code&lt;/td&gt;
&lt;td&gt;Memory CR + Service&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tools&lt;/td&gt;
&lt;td&gt;Integrated at code level&lt;/td&gt;
&lt;td&gt;Tool CR + MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-Agent&lt;/td&gt;
&lt;td&gt;Dialog managed in code&lt;/td&gt;
&lt;td&gt;Team CR + A2A protocol&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;Almost none&lt;/td&gt;
&lt;td&gt;OTel / Langfuse / Dashboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use Cases&lt;/td&gt;
&lt;td&gt;Demo / Prototype / Single Agent&lt;/td&gt;
&lt;td&gt;Enterprise production / Multi-Agent Systems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: ARK vs Traditional Agent Libraries
&lt;/figcaption&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;LangChain is a &amp;ldquo;library for building agents,&amp;rdquo; while ARK is a &amp;ldquo;platform for running agents.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The two are not in conflict and are, in fact, highly complementary.&lt;/p&gt;
&lt;h2 id="the-engineering-value-of-ark"&gt;The Engineering Value of ARK&lt;/h2&gt;
&lt;p&gt;To summarize ARK&amp;rsquo;s engineering value in simple terms:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Turns agents into &lt;strong&gt;governable workloads&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Unifies models, tools, and memory as &lt;strong&gt;reusable resources&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Makes multi-agent collaboration &lt;strong&gt;structured, observable, and tunable&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Brings agent upgrades and iteration into &lt;strong&gt;CI/CD + GitOps&lt;/strong&gt; mode&lt;/li&gt;
&lt;li&gt;Enables enterprises to &lt;strong&gt;manage agents like microservices&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a clear evolution path:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent → Service → Platform → Runtime → Operating System&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;ARK is currently positioned at the fourth stage: Runtime.&lt;/p&gt;
&lt;h2 id="insights-for-agentic-runtime"&gt;Insights for Agentic Runtime&lt;/h2&gt;
&lt;p&gt;ARK provides three direct insights for building Agentic Runtimes:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unified Scheduling System&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The agent runtime must run on a unified scheduling system (Kubernetes, MicroVM, Wasmtime, etc.)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Declarative Capability Boundaries&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Must use declarative abstractions to split capability boundaries, including:
&lt;ul&gt;
&lt;li&gt;Model Layer&lt;/li&gt;
&lt;li&gt;Tool Layer&lt;/li&gt;
&lt;li&gt;Memory Layer&lt;/li&gt;
&lt;li&gt;Workflow Layer&lt;/li&gt;
&lt;li&gt;Team Layer&lt;/li&gt;
&lt;li&gt;State Layer&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Observability is essential; otherwise, multi-agent systems cannot be engineered
&lt;ul&gt;
&lt;li&gt;Langfuse&lt;/li&gt;
&lt;li&gt;OTel&lt;/li&gt;
&lt;li&gt;Logs / Events&lt;/li&gt;
&lt;li&gt;Structured JSON&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;ARK demonstrates a direction:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multi-agent systems are an engineering problem, not a prompt engineering problem.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;If you only need to build a simple agent, frameworks like LangChain, CrewAI, and AutoGPT are sufficient.&lt;/p&gt;
&lt;p&gt;But if you want to operate a system composed of dozens or hundreds of agents that need to collaborate, run long-term, and support continuous delivery and governance, runtimes like ARK are the inevitable trend.&lt;/p&gt;
&lt;p&gt;It provides Agentic AI with:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A cloud-native runtime model&lt;/li&gt;
&lt;li&gt;Observable execution paths&lt;/li&gt;
&lt;li&gt;Governable abstraction layers&lt;/li&gt;
&lt;li&gt;Extensible, componentized architecture&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, ARK deserves to be regarded as an early model for engineering multi-agent systems.&lt;/p&gt;</content:encoded></item><item><title>Can Open Source Suddenly Disappear? An AI Chat Dev Tool Went 404 Overnight</title><link>https://jimmysong.io/blog/ai-project-lunary-404/</link><pubDate>Thu, 11 Dec 2025 05:20:12 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ai-project-lunary-404/</guid><description>Lunary, an open-source project in the AI DevTool space, suddenly deleted its GitHub repo, exposing the instability of commercial open source projects.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Open source&amp;rdquo; in the AI era is no longer a trustworthy promise. Commercial projects can withdraw their code at any time, and developers must be wary of the gap between appearances and reality.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="the-disappearance-of-lunarys-repository-a-real-case-of-open-source-vanishing"&gt;The Disappearance of Lunary&amp;rsquo;s Repository: A Real Case of Open Source &amp;ldquo;Vanishing&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;While updating the AI open source project library on my website, I encountered a situation that left me stunned for the first time:
An &amp;ldquo;open-source AI tool&amp;rdquo; that still promotes itself, with an active website and commercial services, suddenly vanished from GitHub—its repository went straight to 404.&lt;/p&gt;
&lt;p&gt;The project is called Lunary.&lt;/p&gt;
&lt;p&gt;Original repository address:
&lt;a href="https://github.com/lunary-ai/lunary" target="_blank" rel="noopener"&gt;https://github.com/lunary-ai/lunary&lt;/a&gt;
It now returns a 404 Not Found.&lt;/p&gt;
&lt;p&gt;Notably, the official site lunary.ai remains online, but the core promise of an &amp;ldquo;open-source codebase&amp;rdquo; has disappeared.&lt;/p&gt;
&lt;h2 id="lunarys-positioning-and-features"&gt;Lunary&amp;rsquo;s Positioning and Features&lt;/h2&gt;
&lt;p&gt;Here is an overview of Lunary&amp;rsquo;s main features and positioning to help understand its role in the AI tool ecosystem.&lt;/p&gt;
&lt;p&gt;Lunary claims to be an Observability and Evaluations platform for large language model (LLM, Large Language Model) applications, focusing on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM conversation and feedback logs&lt;/li&gt;
&lt;li&gt;Cost, latency, and metrics analysis&lt;/li&gt;
&lt;li&gt;Prompt version management&lt;/li&gt;
&lt;li&gt;Distributed tracing&lt;/li&gt;
&lt;li&gt;Evaluations&lt;/li&gt;
&lt;li&gt;Supports both self-hosted and managed modes&lt;/li&gt;
&lt;li&gt;Provides JS / Python SDKs, integrates with LangChain&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Its overall positioning is clear:
&amp;ldquo;Development and debugging tools for AI applications.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;In fact, products like this have emerged rapidly over the past year, forming a new AI DevTool track.&lt;/p&gt;
&lt;h2 id="the-reality-and-risks-behind-the-open-source-label"&gt;The Reality and Risks Behind the &amp;ldquo;Open Source&amp;rdquo; Label&lt;/h2&gt;
&lt;p&gt;The core issue is not the tool itself, but its claim to be &amp;ldquo;open source.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Lunary has consistently emphasized:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Lunary is an open-source platform for developers.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This statement is great for attracting users, as open source implies transparency, trustworthiness, self-hosting, and community participation.&lt;/p&gt;
&lt;p&gt;But now the repository is gone, with only the website continuing its promotion—raising many questions.&lt;/p&gt;
&lt;p&gt;Lunary is not a niche hobby project, but a commercial company-led initiative. If an individual suddenly deletes a repo, it&amp;rsquo;s not surprising, but for a company operating publicly, this move is extremely rare.&lt;/p&gt;
&lt;p&gt;This is the first time I&amp;rsquo;ve truly seen a reality in the AI DevTools space: &amp;ldquo;Open source&amp;rdquo; is being used as a branding term, not a commitment.&lt;/p&gt;
&lt;h2 id="possible-industry-reasons-for-repo-deletion"&gt;Possible Industry Reasons for Repo Deletion&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s analyze some common industry reasons for deleting a repository to help developers understand the motivations behind such actions.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Increased commercial pressure&lt;/strong&gt;: These tools often struggle with sustainable business models, prompting teams to shift to closed-source SaaS.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pivoting&lt;/strong&gt;: The company finds the original direction unprofitable and prepares to change course.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Team changes&lt;/strong&gt;: Acquisition, key member departures, or funding issues can all lead to repo shutdowns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compliance or legal risks&lt;/strong&gt;: Observability products involve user data, which may require public code to be taken down.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Regardless of the reason, the impact on users is the same: it is no longer an &amp;ldquo;open-source product.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-pseudo-open-source-phenomenon-in-ai-tools"&gt;The &amp;ldquo;Pseudo Open Source&amp;rdquo; Phenomenon in AI Tools&lt;/h2&gt;
&lt;p&gt;The most noteworthy aspect is not Lunary itself, but the rapid spread of this phenomenon in the AI tool space.&lt;/p&gt;
&lt;p&gt;Many projects use &amp;ldquo;open source&amp;rdquo; as a user acquisition strategy but lack open governance and long-term commitment.&lt;/p&gt;
&lt;p&gt;High substitutability, homogeneity, and commercial pressure mean these DevTools have low survival rates.&lt;/p&gt;
&lt;p&gt;When commercial teams lead open source, a single decision can make the repository disappear instantly.&lt;/p&gt;
&lt;p&gt;In the cloud native era, we&amp;rsquo;ve already seen a wave of &amp;ldquo;pseudo open source.&amp;rdquo; In the AI era, this trend is accelerating.&lt;/p&gt;
&lt;h2 id="three-practical-lessons-for-developers"&gt;Three Practical Lessons for Developers&lt;/h2&gt;
&lt;p&gt;Based on this case, here are three practical lessons for developers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The &amp;ldquo;open source label&amp;rdquo; does not guarantee trustworthiness&lt;/strong&gt;: Open source projects led by commercial companies without community or foundation backing can be withdrawn at any time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI DevTools are far less stable than infrastructure&lt;/strong&gt;: These tools are not essential, highly replaceable, and have short lifecycles.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool usability should take precedence over &amp;ldquo;open source status&amp;rdquo;&lt;/strong&gt;: Because it may stop being open source at any moment.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="my-first-experience-maintaining-an-ai-project-list-and-facing-repo-deletion"&gt;My First Experience Maintaining an AI Project List and Facing Repo Deletion&lt;/h2&gt;
&lt;p&gt;After collecting hundreds of projects over the past two years, this is the first time I&amp;rsquo;ve encountered a &amp;ldquo;commercial open source project disappearing, official repo 404&amp;rdquo; case.&lt;/p&gt;
&lt;p&gt;To me, this is an industry signal: the AI open source world is entering a period of drift, and commercial projects&amp;rsquo; open source commitments are increasingly unstable.&lt;/p&gt;
&lt;p&gt;It also reminds everyone making technical choices: in the AI era, open source is no longer a label you can automatically trust.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The disappearance of the Lunary repository is not an isolated incident, but a reflection of the &amp;ldquo;pseudo open source&amp;rdquo; phenomenon in the AI tool space. Developers should be cautious about the actual commitments behind the &amp;ldquo;open source&amp;rdquo; label, paying attention to project governance and sustainability. In the future, the boundary between open source and commercial will become even more blurred, and rational judgment and risk awareness will be essential for technical decision-making.&lt;/p&gt;
&lt;p&gt;Lunary&amp;rsquo;s sudden disappearance highlights the instability of open source projects in the AI DevTools space. For developers, technical choices should focus more on project usability and community governance, rather than relying solely on the &amp;ldquo;open source&amp;rdquo; label. As the industry evolves, similar incidents may become more frequent. Only rational judgment and risk awareness can help you stand firm in the fast-changing tech landscape.&lt;/p&gt;</content:encoded></item><item><title>CNCF in the AI Native Era? The Agentic AI Foundation Is Officially Established</title><link>https://jimmysong.io/blog/agentic-ai-foundation-cncf-era/</link><pubDate>Wed, 10 Dec 2025 03:25:38 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/agentic-ai-foundation-cncf-era/</guid><description>An analysis of the background, strategic urgency, differences and division of labor between Agentic AI Foundation (AAIF) and CNCF/CNAI, and its significance for the AI Native era.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The standardization and open collaboration of the agent ecosystem is no longer a luxury, but the critical watershed for whether AI Native can be engineered and implemented.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;The establishment of &lt;a href="https://aaif.io/" target="_blank" rel="noopener"&gt;AAIF (Agentic AI Foundation)&lt;/a&gt; is the result of leading vendors staking out the &amp;ldquo;agent protocol layer&amp;rdquo; in advance.&lt;/li&gt;
&lt;li&gt;The real challenge is not technical, but how organizations transition from &amp;ldquo;human execution + AI assistance&amp;rdquo; to &amp;ldquo;agent execution + human supervision&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Successful agent adoption requires a phased adoption path, not just a bunch of protocols and demos.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cncf.io/" target="_blank" rel="noopener"&gt;CNCF&lt;/a&gt; and AAIF are complementary: CNCF manages &amp;ldquo;what infrastructure agents run on&amp;rdquo;, AAIF manages &amp;ldquo;how agents collaborate&amp;rdquo;. This matches the system I am building in &lt;a href="https://arksphere.dev/" target="_blank" rel="noopener"&gt;ArkSphere&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="cloud-native-problems-are-solved-ai-native-problems-are-just-beginning"&gt;Cloud Native Problems Are Solved, AI Native Problems Are Just Beginning&lt;/h2&gt;
&lt;p&gt;Over the past decade, Cloud Native technologies like Kubernetes, Service Mesh, and microservices have standardized &amp;ldquo;how applications run in the cloud&amp;rdquo;.
But AI Native faces a completely different challenge:
&lt;strong&gt;It&amp;rsquo;s not about &amp;ldquo;how to deploy a service&amp;rdquo;, but &amp;ldquo;how many behaviors in the system can be handed over to agents to execute themselves&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;CNCF&amp;rsquo;s Cloud Native AI (CNAI) addresses infrastructure-level issues:
&amp;ldquo;How can model training/inference/RAG run at scale and securely on Kubernetes?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;But what AI Native truly lacks is another layer:
&lt;strong&gt;How do agents collaborate, access tools, get governed, and audited?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is exactly the gap AAIF aims to fill.&lt;/p&gt;
&lt;h2 id="aaifs-three-weapons-protocol--runtime--development-standard"&gt;AAIF&amp;rsquo;s Three Weapons: Protocol + Runtime + Development Standard&lt;/h2&gt;
&lt;p&gt;AAIF hosts three core technologies contributed by its founding members:&lt;/p&gt;
&lt;h3 id="-anthropics-model-context-protocol-mcp"&gt;① Anthropic&amp;rsquo;s Model Context Protocol (MCP)&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://github.com/modelcontextprotocol" target="_blank" rel="noopener"&gt;https://github.com/modelcontextprotocol&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A &amp;ldquo;system call interface for agents&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Unified definition for how agents access databases, APIs, files, and external tools.&lt;/li&gt;
&lt;li&gt;Designed to be more like an AI version of gRPC + OAuth.&lt;/li&gt;
&lt;li&gt;Already integrated by Claude, Cursor, ChatGPT, VS Code, Microsoft Copilot, Gemini, and others.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It may not be the flashiest technology, but it could become the plumbing for the entire Agentic ecosystem.&lt;/p&gt;
&lt;h3 id="-blocks-goose-framework"&gt;② Block&amp;rsquo;s Goose Framework&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://github.com/block/goose" target="_blank" rel="noopener"&gt;https://github.com/block/goose&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Reference runtime for MCP:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Local-first, composable agent workflow engine.&lt;/li&gt;
&lt;li&gt;Enables enterprises to pilot agents in small scopes without betting on a specific vendor.&lt;/li&gt;
&lt;li&gt;Serves as an engineering template for protocol implementation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="-openais-agentsmd"&gt;③ OpenAI&amp;rsquo;s AGENTS.md&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://agents.md/" target="_blank" rel="noopener"&gt;https://agents.md&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A simple but effective standard:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Place an AGENTS.md file in the project repository.&lt;/li&gt;
&lt;li&gt;Clearly document build steps, testing, constraints, and context rules.&lt;/li&gt;
&lt;li&gt;Any agent that understands AGENTS.md can operate the codebase using the same instructions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This makes agent behavior more predictable and auditable.&lt;/p&gt;
&lt;h2 id="why-is-aaif-in-such-a-hurry-this-is-a-race-for-standards"&gt;Why Is AAIF in Such a Hurry? This Is a Race for Standards&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s compare with history:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Kubernetes&amp;rsquo; predecessor Borg ran internally at Google for over a decade; K8s was open sourced and donated to CNCF two years later.&lt;/li&gt;
&lt;li&gt;PyTorch joined the Linux Foundation six years after its release.&lt;/li&gt;
&lt;li&gt;MCP was donated to AAIF just &lt;strong&gt;over one year&lt;/strong&gt; after its launch.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;AAIF is not about &amp;ldquo;mature technology entering a foundation&amp;rdquo;, but &lt;strong&gt;staking out the key position early&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The reasons are practical:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Prevent agent ecosystem fragmentation&lt;/strong&gt;
Today, there are many competing &amp;ldquo;tool invocation protocols&amp;rdquo;, which could become incompatible silos in three years.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Protocol layer is easier to reach global consensus than model layer&lt;/strong&gt;
Model competition is inevitable, but protocols can be standardized, open sourced, and avoid vendor lock-in.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A necessary move in global tech competition&lt;/strong&gt;
Putting the foundational standards for Agentic AI into the Linux Foundation is both a gesture of cooperation and a strategic move.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="aaif-vs-cncf-not-competition-but-two-pieces-of-the-puzzle"&gt;AAIF vs CNCF: Not Competition, But Two Pieces of the Puzzle&lt;/h2&gt;
&lt;p&gt;CNCF&amp;rsquo;s role:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;What infrastructure do agent workloads run on?&amp;rdquo;
Kubernetes, Service Mesh, observability, AI Gateway, RAG Infra—all at this layer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;AAIF&amp;rsquo;s role:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;How do agents collaborate, invoke tools, and get governed?&amp;rdquo;
Protocols, runtimes, and behavioral standards—all at this layer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Analogy:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;th&gt;Responsibilities&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AAIF&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Semantic and collaboration layer of Agentic Runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CNCF/CNAI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Resource and execution layer of AI Native Infra&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: AAIF vs CNCF Comparison
&lt;/figcaption&gt;
&lt;p&gt;This matches the upper semantic and lower infrastructure layers in my &lt;a href="https://arksphere.dev/" target="_blank" rel="noopener"&gt;ArkSphere&lt;/a&gt; architecture diagram.&lt;/p&gt;
&lt;p&gt;In the long run, the two sides will be tightly coupled:
CNCF&amp;rsquo;s KServe, KAgent, and AI Gateway will natively support MCP / AGENTS.md,
AAIF&amp;rsquo;s Runtime will run on Cloud Native infrastructure by default.&lt;/p&gt;
&lt;h2 id="the-real-challenge-not-protocols-but-organizations-and-people"&gt;The Real Challenge: Not Protocols, But Organizations and People&lt;/h2&gt;
&lt;p&gt;Most enterprises will get stuck on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How much responsibility can agents actually take?&lt;/li&gt;
&lt;li&gt;Who is accountable when things go wrong?&lt;/li&gt;
&lt;li&gt;How are audit, SLOs, and compliance defined?&lt;/li&gt;
&lt;li&gt;How is multi-agent collaboration visualized?&lt;/li&gt;
&lt;li&gt;How are tool invocation permissions controlled?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words, &lt;strong&gt;agent adoption is not a &amp;ldquo;technical migration&amp;rdquo;, but an &amp;ldquo;organizational migration&amp;rdquo;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If AAIF cannot provide:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Phased adoption methodologies&lt;/li&gt;
&lt;li&gt;Typical organizational migration paths&lt;/li&gt;
&lt;li&gt;Engineering best practices&lt;/li&gt;
&lt;li&gt;Failure cases and anti-patterns&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It will be difficult for AAIF to achieve the industry impact that CNCF did.&lt;/p&gt;
&lt;h2 id="summary-aaif-is-the-moment-when-boundaries-are-drawn"&gt;Summary: AAIF Is the Moment When Boundaries Are Drawn&lt;/h2&gt;
&lt;p&gt;For me, the establishment of AAIF feels like:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;The battlefield boundaries of the agent world have finally been drawn. Now it&amp;rsquo;s up to the engineering community to make it work.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;CNCF solved &amp;ldquo;how to run Cloud Native&amp;rdquo;,
AAIF is now trying to solve &amp;ldquo;how agents collaborate&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;In the next five years, whoever can truly connect these two worlds
will stand at the gateway to the next generation of infrastructure.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s why I started a dedicated &amp;ldquo;Agentic Runtime + AI Native Infra&amp;rdquo; research track in &lt;a href="https://arksphere.dev" target="_blank" rel="noopener"&gt;ArkSphere&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="the-three-body-architecture-of-the-ai-native-era"&gt;The &amp;lsquo;Three-Body&amp;rsquo; Architecture of the AI Native Era&lt;/h2&gt;
&lt;p&gt;Finally, a personal note—my thoughts on ArkSphere.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-ai-foundation-cncf-era/a2b0ea6c87b10fd78607da5d75c4cd1a.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-ai-foundation-cncf-era/a2b0ea6c87b10fd78607da5d75c4cd1a.svg" alt="Figure 3: AAIF × CNCF: Three-Layer Architecture of Agentic AI in the AI Native Era" data-caption="Figure 3: AAIF × CNCF: Three-Layer Architecture of Agentic AI in the AI Native Era"
width="2677"
height="739"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: AAIF × CNCF: Three-Layer Architecture of Agentic AI in the AI Native Era&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This diagram shows the three-layer structure of the AI Native era:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;CNCF (bottom layer): Provides the Cloud Native foundation required for agent operation, including Kubernetes, Service Mesh, GPU scheduling, and security systems.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AAIF (middle layer): Defines the runtime semantics and standards for agents, including the MCP protocol, Goose reference runtime, and AGENTS.md behavioral standard.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ArkSphere (bridging layer): Aligns the &amp;ldquo;Agentic Runtime semantic layer&amp;rdquo; with the &amp;ldquo;AI Native Infra infrastructure layer&amp;rdquo;, forming an engineerable agent architecture standard.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;p&gt;Infra is responsible for &amp;ldquo;running&amp;rdquo;, Runtime for &amp;ldquo;how to act&amp;rdquo;, and ArkSphere for &amp;ldquo;how to assemble a system&amp;rdquo;.&lt;/p&gt;</content:encoded></item><item><title>Bun Acquired by Anthropic: A Structural Signal for AI-Native Runtimes</title><link>https://jimmysong.io/blog/bun-anthropic-runtime-shift/</link><pubDate>Wed, 03 Dec 2025 05:21:28 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/bun-anthropic-runtime-shift/</guid><description>Bun&amp;#39;s acquisition by Anthropic marks the first time a general-purpose language runtime is integrated into a large model engineering system, revealing a structural trend for AI-native runtimes.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The shifting ownership of runtimes is reshaping the underlying logic of AI programming and infrastructure.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After the &lt;a href="https://bun.com/blog/bun-joins-anthropic" target="_blank" rel="noopener"&gt;announcement of Bun&amp;rsquo;s acquisition by Anthropic&lt;/a&gt;, my focus was not on the deal itself, but on the structural signal it revealed: general-purpose language runtimes are now being drawn into the path dependencies of AI programming systems. This is not just &amp;ldquo;a JS project finding a home,&amp;rdquo; but &amp;ldquo;the first time a language runtime has been actively integrated into the unified engineering system of a leading large model company.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This event deserves a deeper analysis.&lt;/p&gt;
&lt;h2 id="buns-engineering-features-and-current-status"&gt;Bun&amp;rsquo;s Engineering Features and Current Status&lt;/h2&gt;
&lt;p&gt;Before examining &lt;a href="https://bun.com" target="_blank" rel="noopener"&gt;Bun&lt;/a&gt;&amp;rsquo;s industry significance, let&amp;rsquo;s outline its runtime characteristics. The following list summarizes Bun&amp;rsquo;s main engineering capabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;High-performance JavaScript/TypeScript runtime&lt;/li&gt;
&lt;li&gt;Built-in bundler, test framework, and package manager&lt;/li&gt;
&lt;li&gt;Single-file executable&lt;/li&gt;
&lt;li&gt;Extremely fast cold start&lt;/li&gt;
&lt;li&gt;Node compatibility without Node&amp;rsquo;s legacy dependencies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These capabilities have formed measurable performance barriers.&lt;/p&gt;
&lt;p&gt;However, it should be noted that Bun currently lacks the core attributes of an AI Runtime, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Permission model&lt;/li&gt;
&lt;li&gt;Tool isolation&lt;/li&gt;
&lt;li&gt;Capability declaration protocol&lt;/li&gt;
&lt;li&gt;Execution semantics understandable by models&lt;/li&gt;
&lt;li&gt;Sandbox execution environment&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, Bun&amp;rsquo;s &amp;ldquo;AI Native&amp;rdquo; properties have not yet been established, but Anthropic&amp;rsquo;s acquisition provides an opportunity for it to evolve in this direction.&lt;/p&gt;
&lt;h2 id="the-significance-of-a-leading-model-company-acquiring-a-general-purpose-runtime"&gt;The Significance of a Leading Model Company Acquiring a General-Purpose Runtime&lt;/h2&gt;
&lt;p&gt;Historically, it is not uncommon for model companies to acquire editors, plugins, or IDEs, but in known public cases, mainstream large model vendors have never directly acquired a mature general-purpose language runtime. Bun × Anthropic is the first clear event pulling the runtime into the AI programming system landscape. This move sends two engineering-level signals:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The speed of AI code generation continues to increase, amplifying the need for deterministic execution environments. The generate→execute→validate→destroy cycle intensifies the problem of environment non-repeatability.&lt;/li&gt;
&lt;li&gt;Models require a &amp;ldquo;controllable execution substrate&amp;rdquo; rather than a traditional operating system. Agents are not suited to run tools in an uncontrollable, unpredictable OS layer.&lt;/li&gt;
&lt;li&gt;The runtime needs to be embedded into the model&amp;rsquo;s internal engineering pipeline. Future IDEs, agents, and auto-repair pipelines may directly invoke the runtime&amp;rsquo;s API.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is not a short-term business integration, but a manifestation of the trend toward compressed engineering pipelines.&lt;/p&gt;
&lt;h2 id="runtime-requirements-differentiation-in-the-ai-coding-era"&gt;Runtime Requirements Differentiation in the AI Coding Era&lt;/h2&gt;
&lt;p&gt;Based on observations of agentic runtimes over the past year, runtime requirements in the AI coding era are diverging. The following list summarizes the main engineering abstractions trending in this space:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Determinism: AI-generated code is not reviewed line by line; execution results must be consistent across machines and over time.&lt;/li&gt;
&lt;li&gt;Minimal distribution unit: Users no longer install language environments and numerous dependencies. Verifiable, replicable, and portable single execution units are becoming the norm.&lt;/li&gt;
&lt;li&gt;Tool isolation: Models cannot directly access all OS capabilities; the context and permissions visible to tools must be strictly defined.&lt;/li&gt;
&lt;li&gt;Short-lived execution: Agent invocation patterns resemble &amp;ldquo;batch jobs&amp;rdquo; rather than long-running services.&lt;/li&gt;
&lt;li&gt;Capability declaration: The runtime must expose &amp;ldquo;what I can do,&amp;rdquo; rather than the entire OS interface.&lt;/li&gt;
&lt;li&gt;Embeddable self-testing pipeline: After generating code, models need to immediately execute tests, collect errors, and iterate. The runtime must provide observability and diagnostic primitives.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These requirements are not unique to Bun, nor did Bun originate them, but Bun&amp;rsquo;s &amp;ldquo;monolithic and controllable&amp;rdquo; runtime structure is more conducive to evolving in this direction.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/bun-anthropic-runtime-shift/527ff200956d6b73178a0e3521f42fc2.svg" data-img="https://assets.jimmysong.io/images/blog/bun-anthropic-runtime-shift/527ff200956d6b73178a0e3521f42fc2.svg" alt="Figure 3: Minimal execution loop of an AI-native runtime" data-caption="Figure 3: Minimal execution loop of an AI-native runtime"
width="599"
height="1105"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Minimal execution loop of an AI-native runtime&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="buns-potential-role-within-anthropics-system"&gt;Bun&amp;rsquo;s Potential Role Within Anthropic&amp;rsquo;s System&lt;/h2&gt;
&lt;p&gt;If Bun is seen merely as a Node.js replacement, the acquisition is of limited significance. But if it is viewed as the execution foundation for future AI coding systems, the logic becomes clearer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Code is generated by models&lt;/li&gt;
&lt;li&gt;Building is handled by the runtime&amp;rsquo;s built-in toolchain&lt;/li&gt;
&lt;li&gt;Testing, validation, and repair are performed by models repeatedly invoking the runtime&lt;/li&gt;
&lt;li&gt;All execution behaviors are defined by the runtime&amp;rsquo;s semantics&lt;/li&gt;
&lt;li&gt;The runtime forms Anthropic&amp;rsquo;s internal &amp;ldquo;minimal stable layer&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This model is similar to the relationship between Chrome and V8: the execution engine and upper-layer system co-evolve over time, with performance and semantics advancing in sync.&lt;/p&gt;
&lt;p&gt;Whether Bun can fulfill this role depends on Anthropic&amp;rsquo;s architectural choices, but the event itself has opened up possibilities in this direction.&lt;/p&gt;
&lt;h2 id="industry-trends-and-future-evolution"&gt;Industry Trends and Future Evolution&lt;/h2&gt;
&lt;p&gt;Combining facts, signals, and engineering trends, the following directions can be anticipated:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The &amp;ldquo;Agent Runtime&amp;rdquo; category will gradually become more defined&lt;/li&gt;
&lt;li&gt;The boundaries between bundler, runtime, and test runner will continue to blur&lt;/li&gt;
&lt;li&gt;Cloud vendors will launch controllable runtimes with capability declarations&lt;/li&gt;
&lt;li&gt;Permission models and secure sandboxes will move down to the language runtime layer&lt;/li&gt;
&lt;li&gt;Runtimes will become part of the model toolchain, rather than an external environment&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These trends will not all materialize in the short term, but they represent the inevitable path of engineering evolution.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The combination of Bun × Anthropic is not about &amp;ldquo;an open-source project being absorbed,&amp;rdquo; but about a language runtime being actively integrated into the engineering pipeline of a large model system for the first time. Competition at the model layer will continue, but what truly reshapes software is the structural transformation of AI-native runtimes. This is a foundational change worth long-term attention.&lt;/p&gt;</content:encoded></item><item><title>Agentic Runtime Realism: Insights from McKinsey Ark on 2026 Infrastructure Trends</title><link>https://jimmysong.io/blog/agentic-runtime-realism/</link><pubDate>Tue, 02 Dec 2025 12:07:45 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/agentic-runtime-realism/</guid><description>Analyzing Ark from architecture, semantics, community activity, and engineering paradigms to reveal its impact on 2026 AI Infra trends and the ArkSphere community.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Statement
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
ArkSphere has no affiliation or association with McKinsey Ark.
&lt;/div&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;The value of Agentic Runtime lies not in unified interfaces, but in semantic governance and the transformation of engineering paradigms. Ark is just a reflection of the trend; the future belongs to governable Agentic Workloads.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Recently, the &lt;a href="https://jimmysong.io/community/"&gt;ArkSphere community&lt;/a&gt; has been focusing on McKinsey&amp;rsquo;s open-source &lt;a href="https://github.com/mckinsey/agents-at-scale-ark" target="_blank" rel="noopener"&gt;Ark&lt;/a&gt; (Agentic Runtime for Kubernetes). Although the project is still in technical preview, its architecture and semantic model have already become key indicators for the direction of AI Infra in 2026.&lt;/p&gt;
&lt;p&gt;This article analyzes the engineering paradigm and semantic model of Ark, highlighting its industry implications. It avoids repeating the reasons for the failure of unified model APIs and generic infrastructure logic, instead focusing on the unique perspective of the ArkSphere community.&lt;/p&gt;
&lt;h2 id="arks-semantic-model-and-engineering-paradigm"&gt;Ark&amp;rsquo;s Semantic Model and Engineering Paradigm&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s greatest value is in making Agents first-class citizens in Kubernetes, achieving closed-loop tasks through CRD (Custom Resource Definition) and controllers (Reconcilers). This semantic abstraction not only enhances governance capabilities but also aligns closely with the Agentic Runtime strategies of major cloud providers.&lt;/p&gt;
&lt;p&gt;Ark&amp;rsquo;s main resources include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Agent (inference entity)&lt;/li&gt;
&lt;li&gt;Model (model selection and configuration)&lt;/li&gt;
&lt;li&gt;Tools (capability plugins/MCP, Model Capability Plugin)&lt;/li&gt;
&lt;li&gt;Team (multi-agent collaboration)&lt;/li&gt;
&lt;li&gt;Query (task lifecycle)&lt;/li&gt;
&lt;li&gt;Evaluation (assessment)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The diagram below illustrates the semantic relationships in Agentic Runtime:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-runtime-realism/2d76bfcb312694080bd94942b084f210.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-runtime-realism/2d76bfcb312694080bd94942b084f210.svg" alt="Figure 5: Agentic Runtime Semantic Relationships" data-caption="Figure 5: Agentic Runtime Semantic Relationships"
width="1450"
height="589"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Agentic Runtime Semantic Relationships&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="architecture-and-community-activity"&gt;Architecture and Community Activity&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s architecture adopts a standard control plane system, emphasizing unified runtime semantics. The community is highly active, engineer-driven, and the codebase is well-structured, though production readiness is still being improved.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/agentic-runtime-realism/f481241843db17b6e6172e8093a1daa6.svg" data-img="https://assets.jimmysong.io/images/blog/agentic-runtime-realism/f481241843db17b6e6172e8093a1daa6.svg" alt="Figure 6: Ark Architecture and Control Plane Flow" data-caption="Figure 6: Ark Architecture and Control Plane Flow"
width="4130"
height="565"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Ark Architecture and Control Plane Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="arkspheres-boundaries-and-inspirations"&gt;ArkSphere&amp;rsquo;s Boundaries and Inspirations&lt;/h2&gt;
&lt;p&gt;The emergence of Ark has clarified the boundaries of ArkSphere. ArkSphere does not aim for unified model interfaces, multi-cloud abstraction, a collection of miscellaneous tools, or a comprehensive framework layer. Instead, it focuses on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The semantic system of Agentic Runtime (tasks, states, tool invocation, collaboration graphs, etc.)&lt;/li&gt;
&lt;li&gt;Enterprise-grade runtime governance models (permissions, auditing, isolation, multi-tenancy, compliance, cost tracking)&lt;/li&gt;
&lt;li&gt;Integration capabilities for domestic ecosystem tools&lt;/li&gt;
&lt;li&gt;Engineering paradigms from a runtime perspective&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;ArkSphere is an ecosystem and engineering system at the runtime level, not a &amp;ldquo;model abstraction layer&amp;rdquo; or an &amp;ldquo;agent development framework.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="key-changes-in-2026"&gt;Key Changes in 2026&lt;/h2&gt;
&lt;p&gt;2026 will usher in the era of Agentic Runtime, where Agents are no longer just classes but workloads that require governance rather than mere importation. Ark is just one example of this trend, and the direction is clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Semantic models and governability become highlights&lt;/li&gt;
&lt;li&gt;Closed-loop tasks are the core value&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s realism teaches us that the future belongs to runtime, semantics, governability, and workload-level Agents. The industry will no longer pursue unified APIs or framework implementations, but will focus on governable runtime semantics and engineering paradigms.&lt;/p&gt;</content:encoded></item><item><title>In-Depth Analysis of Ark: Kubernetes for the AI Era or a New Engineering Paradigm Shift?</title><link>https://jimmysong.io/blog/ark-agentic-runtime-analysis/</link><pubDate>Tue, 02 Dec 2025 10:54:34 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ark-agentic-runtime-analysis/</guid><description>Analysis of McKinsey&amp;#39;s Ark project: architecture, CRDs, control plane, design paradigms, production readiness, and implications for ArkSphere and AI infrastructure.</description><content:encoded>
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Statement
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
ArkSphere has no affiliation or association with McKinsey Ark.
&lt;/div&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;The greatest value of Ark lies in reshaping engineering paradigms, not just its features. It points the way for AI Infra and leaves vast space for community ecosystems.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Recently, many members in our &lt;a href="https://arksphere.dev/" target="_blank" rel="noopener"&gt;ArkSphere community&lt;/a&gt; have started exploring McKinsey&amp;rsquo;s open-source &lt;a href="https://github.com/mckinsey/agents-at-scale-ark" target="_blank" rel="noopener"&gt;Ark (Agentic Runtime for Kubernetes)&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Some see it as radical, some think it&amp;rsquo;s just a consulting firm&amp;rsquo;s experiment, and others quote a realistic maxim:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What we need now is &amp;ldquo;agentic runtime realism,&amp;rdquo; not &amp;ldquo;unified model romanticism.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I strongly agree with this sentiment.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve spent some time analyzing Ark&amp;rsquo;s source code, architecture, and design philosophy, combined with our community discussions. My conclusion is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ark&amp;rsquo;s significance is not in its features, but in its paradigm.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It&amp;rsquo;s not the answer, but it points toward the future of AI Infra.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Below is my interpretation of Ark, focusing on engineering, architecture, trends, and its inspiration for ArkSphere.&lt;/p&gt;
&lt;h2 id="what-exactly-is-ark"&gt;What Exactly Is Ark?&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s core positioning is: &lt;strong&gt;A runtime that treats Agents as Kubernetes Workloads.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not a framework, not an SDK, not an AutoGen-style multi-agent tool, but a complete system including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Control plane (Controller)&lt;/li&gt;
&lt;li&gt;Custom resource models (CRD, Custom Resource Definition)&lt;/li&gt;
&lt;li&gt;API service&lt;/li&gt;
&lt;li&gt;Dashboard&lt;/li&gt;
&lt;li&gt;CLI&lt;/li&gt;
&lt;li&gt;Python SDK&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Essentially, Ark is the &lt;strong&gt;control plane for Agents&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Ark defines seven core CRDs in Kubernetes. The following flowchart shows the relationships among these resources:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/df35874e6886350db30fdf036a118099.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/df35874e6886350db30fdf036a118099.svg" alt="Figure 7: Ark CRD Resource Relationships" data-caption="Figure 7: Ark CRD Resource Relationships"
width="816"
height="557"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 7: Ark CRD Resource Relationships&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Through this set of CRDs, Ark makes Agent systems resource-oriented and declarative, enabling capabilities such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Lifecycle management&lt;/li&gt;
&lt;li&gt;Multi-tenant isolation&lt;/li&gt;
&lt;li&gt;RBAC (Role-Based Access Control)&lt;/li&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;li&gt;Upgradability&lt;/li&gt;
&lt;li&gt;Extensibility (tools, models, MCP)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words, Ark is not about &amp;ldquo;how to write Agents,&amp;rdquo; but &amp;ldquo;how to operate Agents in enterprise-grade systems.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="three-layer-architecture-mixed-languages-and-components-but-a-complete-system"&gt;Three-Layer Architecture: Mixed Languages and Components, but a Complete System&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s overall architecture is divided into three layers, each with different tech stacks and responsibilities. The following flowchart illustrates the relationships among components in each layer:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/f6ccd732f54e5aa2a0ca3f6283103eb3.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/f6ccd732f54e5aa2a0ca3f6283103eb3.svg" alt="Figure 8: Ark Three-Layer Architecture Components" data-caption="Figure 8: Ark Three-Layer Architecture Components"
width="1532"
height="806"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 8: Ark Three-Layer Architecture Components&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is not a &amp;ldquo;wrapper project,&amp;rdquo; but a fully operational AI Runtime system, with a level of engineering far beyond most agent frameworks on the market.&lt;/p&gt;
&lt;h2 id="is-it-the-kubernetes-of-the-ai-era"&gt;Is It the Kubernetes of the AI Era?&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s revisit Kubernetes&amp;rsquo; core value:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes was never about &amp;ldquo;unifying cloud APIs&amp;rdquo;; it unified the &amp;ldquo;application runtime model.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Cloud provider APIs aren&amp;rsquo;t unified, nor are networking or storage. What is unified: Pod, Deployment, Service—these application models.&lt;/p&gt;
&lt;p&gt;Kubernetes succeeded because:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It provides a stable application abstraction on top of diversity.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Ark&amp;rsquo;s goal is not to unify all large language models (LLMs), MCPs, or tool formats, but rather:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent resource model (CRD) + control plane (Reconciler) + lifecycle.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From this perspective, Ark offers a prototype of a &amp;ldquo;declarative application model&amp;rdquo; for the AI era.&lt;/p&gt;
&lt;p&gt;Whether it will become &amp;ldquo;Kubernetes for AI&amp;rdquo; is still too early to say, but it has already planted a seed.&lt;/p&gt;
&lt;h2 id="comparison-with-other-frameworks-not-on-the-same-level"&gt;Comparison with Other Frameworks: Not on the Same Level&lt;/h2&gt;
&lt;p&gt;Current mainstream agent frameworks like LangChain, CrewAI, AutoGen, MetaGPT, etc., address problems fundamentally different from Ark.&lt;/p&gt;
&lt;p&gt;The table below compares the positioning and limitations of each framework:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;What Problem Does It Solve&lt;/th&gt;
&lt;th&gt;Core Limitation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LangChain&lt;/td&gt;
&lt;td&gt;Agent/Tool composition&lt;/td&gt;
&lt;td&gt;Doesn&amp;rsquo;t address deployment or governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutoGen&lt;/td&gt;
&lt;td&gt;Multi-agent conversations&lt;/td&gt;
&lt;td&gt;Lacks control plane and lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CrewAI&lt;/td&gt;
&lt;td&gt;Workflow-style multi-agent&lt;/td&gt;
&lt;td&gt;Missing scheduling, RBAC, resource model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MetaGPT&lt;/td&gt;
&lt;td&gt;Agent SOP&lt;/td&gt;
&lt;td&gt;Just execution logic, not a platform&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenDevin&lt;/td&gt;
&lt;td&gt;AI IDE/Dev Assistant&lt;/td&gt;
&lt;td&gt;Not an Agent Runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ark&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Agent control plane + resource system&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Functionality not yet mature&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Mainstream Agent Frameworks vs. Ark
&lt;/figcaption&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Other tools focus on &amp;ldquo;how to write Agents.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Ark focuses on &amp;ldquo;how Agents run, schedule, govern, observe, and extend.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That&amp;rsquo;s an architectural difference.&lt;/p&gt;
&lt;h2 id="execution-flow-agents-scheduled-like-pods"&gt;Execution Flow: Agents Scheduled Like Pods&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s execution flow closely resembles the Kubernetes controller model. The following sequence diagram shows the core process:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/1ce387835a38f3380734332ea9e769f7.svg" data-img="https://assets.jimmysong.io/images/blog/ark-agentic-runtime-analysis/1ce387835a38f3380734332ea9e769f7.svg" alt="Figure 9: Ark Agent Execution Flow" data-caption="Figure 9: Ark Agent Execution Flow"
width="961"
height="553"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 9: Ark Agent Execution Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;You can see Ark&amp;rsquo;s process logic is transparent, with a clear engineering path, bringing agent systems into a &amp;ldquo;controllable&amp;rdquo; state.&lt;/p&gt;
&lt;h2 id="production-readiness-right-direction-still-a-tech-preview"&gt;Production Readiness: Right Direction, Still a Tech Preview&lt;/h2&gt;
&lt;p&gt;According to official notes and code maturity, Ark currently offers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Runnable&lt;/li&gt;
&lt;li&gt;Learnable&lt;/li&gt;
&lt;li&gt;Extensible&lt;/li&gt;
&lt;li&gt;But not recommended for large-scale production use yet&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Main reasons include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CRD structures may change&lt;/li&gt;
&lt;li&gt;APIs are not yet stable&lt;/li&gt;
&lt;li&gt;MCP ecosystem is still forming&lt;/li&gt;
&lt;li&gt;Memory service is still basic&lt;/li&gt;
&lt;li&gt;Multi-agent team execution strategies are primitive&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But the engineering system is already taking shape, which is crucial.&lt;/p&gt;
&lt;h2 id="community-activity-small-but-elite-strong-mckinsey-drive"&gt;Community Activity: Small but Elite, Strong McKinsey Drive&lt;/h2&gt;
&lt;p&gt;From GitHub data:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Stars: 222&lt;/li&gt;
&lt;li&gt;Forks: 50&lt;/li&gt;
&lt;li&gt;Contributors: 48&lt;/li&gt;
&lt;li&gt;Commit frequency is steady&lt;/li&gt;
&lt;li&gt;The vast majority of contributions come from within McKinsey&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;Note: Data as of December 2, 2025.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;High stability, but limited openness.&lt;/p&gt;
&lt;p&gt;This is also ArkSphere&amp;rsquo;s opportunity:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The paradigm is right, but the ecosystem needs community-driven growth.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="trends-for-2026-from-framework-era-to-runtime-era"&gt;Trends for 2026: From Framework Era to Runtime Era&lt;/h2&gt;
&lt;p&gt;After deep analysis, I&amp;rsquo;m increasingly convinced:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2023–2024: Large model API call era&lt;/li&gt;
&lt;li&gt;2024–2025: Agent framework era&lt;/li&gt;
&lt;li&gt;2025–2027: Agent Runtime / Control Plane era (Ark&amp;rsquo;s direction)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;While everyone is writing Python scripts for agents, the real value lies in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multi-agent task scheduling&lt;/li&gt;
&lt;li&gt;Tool registration and governance&lt;/li&gt;
&lt;li&gt;Session/Memory lifecycle&lt;/li&gt;
&lt;li&gt;Result reproducibility&lt;/li&gt;
&lt;li&gt;RBAC, auditing, tenant isolation&lt;/li&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;li&gt;Enterprise internal personalized agent systems&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Ark is providing a practical path forward.&lt;/p&gt;
&lt;h2 id="inspiration-for-arksphere"&gt;Inspiration for ArkSphere&lt;/h2&gt;
&lt;p&gt;Ark&amp;rsquo;s inspiration for ArkSphere is both critical and direct:&lt;/p&gt;
&lt;h3 id="arksphere-should-focus-on-paradigm-building-not-feature-stacking"&gt;ArkSphere Should Focus on &amp;ldquo;Paradigm Building,&amp;rdquo; Not &amp;ldquo;Feature Stacking&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;Ark offers a prototype for future Agentic Runtime:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Resource model&lt;/li&gt;
&lt;li&gt;Control plane&lt;/li&gt;
&lt;li&gt;Tool registration&lt;/li&gt;
&lt;li&gt;Multi-agent collaboration&lt;/li&gt;
&lt;li&gt;Evaluation and governance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;ArkSphere&amp;rsquo;s role should be:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Aggregate paradigms, produce standards, incubate ecosystems, not rewrite Ark itself.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is the &amp;ldquo;CNCF (Cloud Native Computing Foundation) for the AI-native era.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="huge-potential-for-localization-in-china"&gt;Huge Potential for Localization in China&lt;/h3&gt;
&lt;p&gt;Localization opportunities include but are not limited to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Integration with domestic large language models (e.g., Qwen, DeepSeek, Zhipu)&lt;/li&gt;
&lt;li&gt;Enterprise privatization scenarios&lt;/li&gt;
&lt;li&gt;Local tool/MCP discovery ecosystem&lt;/li&gt;
&lt;li&gt;Multi-cluster/edge inference&lt;/li&gt;
&lt;li&gt;Enterprise-grade RBAC, auditing, data isolation&lt;/li&gt;
&lt;li&gt;AgentSpec enhancements for industrial scenarios&lt;/li&gt;
&lt;li&gt;Enhanced versions of Runtime/Controller&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Ark solves the &amp;ldquo;model,&amp;rdquo; while ArkSphere can solve the &amp;ldquo;ecosystem.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="what-we-need-is-not-kubernetes-for-the-llm-era-but-an-industry-grade-cognition-system-for-ai-runtime"&gt;What We Need Is Not &amp;ldquo;Kubernetes for the LLM Era,&amp;rdquo; But an &amp;ldquo;Industry-Grade Cognition System for AI Runtime&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;The biggest takeaway from dissecting Ark:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The future of AI-native is not a pile of tools, but an engineering system.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;ArkSphere can be the initiator of this system.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Ark is not a &amp;ldquo;universal runtime,&amp;rdquo; nor is it the &amp;ldquo;ultimate Kubernetes for the AI era.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;But it has done one crucial thing right:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It abstracts all the pain points people faced when writing Python agent scripts into Kubernetes resources and controllers.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It represents engineering, not just a demo.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not mature yet, but it&amp;rsquo;s heading in the right direction.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not the end, but it gives us a clear roadmap.&lt;/p&gt;
&lt;p&gt;For the ArkSphere community I&amp;rsquo;m running, Ark provides a clear inspiration:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The future belongs to Runtime, to Control Plane, to governable agent systems.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;And the ones who can truly scale this system are not McKinsey, but the community.&lt;/strong&gt;&lt;/p&gt;</content:encoded></item><item><title>From Using AI to Relying on AI: Why the Era of AI Engineering Has Yet to Begin</title><link>https://jimmysong.io/blog/from-using-ai-to-building-ai-systems/</link><pubDate>Sat, 29 Nov 2025 12:40:54 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/from-using-ai-to-building-ai-systems/</guid><description>AI&amp;#39;s real turning point is moving from using AI tools to building AI systems. Why the era of AI engineering hasn&amp;#39;t begun, and the developer opportunity in the next three years.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The real inflection point for AI engineering is not &amp;ldquo;how many people use it,&amp;rdquo; but &amp;ldquo;how many people cannot do without it.&amp;rdquo; Only when not using AI leads to direct loss of opportunity and efficiency, can we say the era of AI engineering has truly arrived.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="starting-point-predictions-for-ai-in-2026"&gt;Starting Point: Predictions for AI in 2026&lt;/h2&gt;
&lt;p&gt;Recently, I came across two &lt;a href="https://thenewstack.io/amazon-cto-werner-vogels-predictions-for-2026/" target="_blank" rel="noopener"&gt;predictions for 2026 from Amazon CTO Werner Vogels&lt;/a&gt; that struck me the most:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Renaissance Developer&lt;/strong&gt;: Developers must span code, product, business, and social impact.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Personalized Learning&lt;/strong&gt;: AI will reshape education, focusing on differentiated paths rather than a unified curriculum.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both point to the same trend: AI is not just a tool, but is redefining how people grow and how they are defined.&lt;/p&gt;
&lt;p&gt;There is a gap between prediction and reality, and it is worth exploring.&lt;/p&gt;
&lt;h2 id="correction-will-ai-really-be-saturated-by-2026"&gt;Correction: Will AI Really Be &amp;ldquo;Saturated&amp;rdquo; by 2026?&lt;/h2&gt;
&lt;p&gt;My initial prediction was that AI usage would reach saturation by 2026. Reality has shown me this is too optimistic.&lt;/p&gt;
&lt;p&gt;By the end of 2025, even among internet professionals, most people&amp;rsquo;s use of AI remains at the &amp;ldquo;heard of it&amp;rdquo; or &amp;ldquo;tried it a few times&amp;rdquo; stage. It is still far from being a daily workflow necessity.&lt;/p&gt;
&lt;p&gt;More importantly, this judgment is &lt;strong&gt;conditional&lt;/strong&gt;: infrastructure supply, regulation, and compute costs must not reverse in the next 3–6 years. If any variable breaks down (costs double, models go offline, policy shifts), the adoption curve will be disrupted.&lt;/p&gt;
&lt;h2 id="the-truth-about-the-inflection-point-from-using-to-relying-on"&gt;The Truth About the Inflection Point: From &amp;ldquo;Using&amp;rdquo; to &amp;ldquo;Relying On&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;Relying on&amp;rdquo; is a vague term. A more precise definition requires measurable indicators.&lt;/p&gt;
&lt;p&gt;Here is a diagram that visualizes the metrics for being truly dependent on AI:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This diagram visualizes the quantitative metrics for being truly dependent on AI, comparing target thresholds with current status:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/from-using-ai-to-building-ai-systems/ai-dependency-metrics.svg" data-img="https://assets.jimmysong.io/images/blog/from-using-ai-to-building-ai-systems/ai-dependency-metrics.svg" alt="Figure 1: Quantitative Definition of AI Dependency" data-caption="Figure 1: Quantitative Definition of AI Dependency"
width="1776"
height="616"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Quantitative Definition of AI Dependency&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Most industries have not reached the &amp;ldquo;cannot operate without&amp;rdquo; stage, unlike the internet, mobile, or payment inflection points. Most metrics are still far below the threshold, which is why the most likely outcome for 2026 is: &lt;strong&gt;more people will use AI, but those who truly rely on it will remain a minority&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="using--building-the-five-level-capability-ladder"&gt;Using ≠ Building: The Five-Level Capability Ladder&lt;/h2&gt;
&lt;p&gt;This difference is not binary, but a clear progression.&lt;/p&gt;
&lt;p&gt;The following table shows the five-level model of AI capability maturity.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;Scarcity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Tool User&lt;/td&gt;
&lt;td&gt;ChatGPT/Claude, Coding, Copywriting, Accelerator, Optional&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Integrator&lt;/td&gt;
&lt;td&gt;LLM API + Vector DB, AI layered on existing systems, Usable, not critical&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Settler&lt;/td&gt;
&lt;td&gt;Restructuring data flow, business decisions, AI becomes critical path&lt;/td&gt;
&lt;td&gt;Rising&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Engineering Abstraction&lt;/td&gt;
&lt;td&gt;Extracting frameworks, runtimes, providing infra for ecosystem&lt;/td&gt;
&lt;td&gt;Extremely High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Autonomous System&lt;/td&gt;
&lt;td&gt;Self-feedback, self-optimizing, redefining human-AI relationship&lt;/td&gt;
&lt;td&gt;Future&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Five-Level Model of AI Capability Maturity
&lt;/figcaption&gt;
&lt;p&gt;Currently, the biggest gap is at &lt;strong&gt;Level 3 and Level 4&lt;/strong&gt;. Most people are stuck at Level 1 or 2, with very few reaching Level 4. This means &lt;strong&gt;high-value scarcity will not disappear, but will continue to rise&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="why-the-era-of-ai-engineering-has-not-arrived-three-dimensional-delaying-factors"&gt;Why the Era of AI Engineering Has Not Arrived: Three-Dimensional Delaying Factors&lt;/h2&gt;
&lt;p&gt;It is not technology alone that is holding things back, but constraints in three dimensions.&lt;/p&gt;
&lt;p&gt;The following diagram illustrates the three main constraints delaying AI engineering maturity:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This diagram illustrates the three main constraints (technical, institutional, and organizational) that are delaying AI engineering maturity, along with their delay metrics:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/from-using-ai-to-building-ai-systems/ai-engineering-constraints.svg" data-img="https://assets.jimmysong.io/images/blog/from-using-ai-to-building-ai-systems/ai-engineering-constraints.svg" alt="Figure 2: Three-Dimensional Constraints on AI Engineering Maturity" data-caption="Figure 2: Three-Dimensional Constraints on AI Engineering Maturity"
width="1783"
height="803"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: Three-Dimensional Constraints on AI Engineering Maturity&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The key observation: &lt;strong&gt;If any one dimension is stuck, the entire ecosystem&amp;rsquo;s maturity will be delayed&lt;/strong&gt;. Currently, none of the three dimensions have fully mature solutions.&lt;/p&gt;
&lt;h2 id="the-realistic-window-three-paths-for-capability-advancement"&gt;The Realistic Window: Three Paths for Capability Advancement&lt;/h2&gt;
&lt;p&gt;The next three years will not be &amp;ldquo;winner takes all,&amp;rdquo; but rather a period where multiple capability levels appreciate simultaneously.&lt;/p&gt;
&lt;p&gt;Below is a table comparing the value and bottlenecks of different capability advancement paths:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability Path&lt;/th&gt;
&lt;th&gt;Short-Term Value&lt;/th&gt;
&lt;th&gt;Long-Term Outlook&lt;/th&gt;
&lt;th&gt;Bottleneck&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Level 1→2 (Tool→Integration)&lt;/td&gt;
&lt;td&gt;⭐⭐ Rapid Depreciation&lt;/td&gt;
&lt;td&gt;⭐ Saturation&lt;/td&gt;
&lt;td&gt;Low barrier, fierce competition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Level 2→3 (Integration→Settlement)&lt;/td&gt;
&lt;td&gt;⭐⭐⭐⭐ Scarce&lt;/td&gt;
&lt;td&gt;⭐⭐⭐⭐ Continual Appreciation&lt;/td&gt;
&lt;td&gt;Requires industry depth, long-term iteration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Level 3→4 (Settlement→Abstraction)&lt;/td&gt;
&lt;td&gt;⭐⭐⭐⭐⭐ Extremely Scarce&lt;/td&gt;
&lt;td&gt;⭐⭐⭐⭐⭐ Defines Ecosystem&lt;/td&gt;
&lt;td&gt;Large cognitive leap, needs community influence&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: AI Capability Advancement Paths and Value Comparison
&lt;/figcaption&gt;
&lt;p&gt;&lt;strong&gt;Key conclusion&lt;/strong&gt;: While the number of &amp;ldquo;AI users&amp;rdquo; is rapidly increasing (depressing Level 1 value), due to the three-dimensional delaying factors, scarcity at Level 3 and 4 will only rise.&lt;/p&gt;
&lt;h2 id="what-im-doing-on-arkspheredev"&gt;What I&amp;rsquo;m Doing on arksphere.dev&lt;/h2&gt;
&lt;p&gt;Based on the above judgment, I focus on exploring the architectural evolution of AI Native Infrastructure. The goal is not to catalog model usage, but to study the foundational capability stack supporting scalable intelligent systems: scheduling, storage, inference, Agent Runtime, autonomous control, observability, and reliability.&lt;/p&gt;
&lt;p&gt;The content is no longer a collection of courses or tips, but a continuous record of evolution around Infra → Runtime → System Abstraction. &lt;a href="https://arksphere.dev" target="_blank" rel="noopener"&gt;arksphere.dev&lt;/a&gt; is the site for this experiment and settlement.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The inflection point for the era of AI engineering is not &amp;ldquo;how many people use it,&amp;rdquo; but &amp;ldquo;how many people cannot do without it.&amp;rdquo; The latter requires five measurable indicators to reach their thresholds, and we are still far from that.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Using ≠ Building&amp;rdquo; is not a binary, but a five-level progression. &lt;strong&gt;Scarcity at Level 3 and 4 will rise as the number of Level 1 users increases&lt;/strong&gt;—this is the biggest opportunity window in the next three years.&lt;/p&gt;
&lt;p&gt;But the width of this window depends largely on how technology, institutions, and organizations evolve together. I hope more people working on AI engineering will not only focus on technical innovation, but also invest equal thought into institutional development, talent growth, and risk governance—these &amp;ldquo;invisible engineering&amp;rdquo; challenges.&lt;/p&gt;</content:encoded></item><item><title>Antigravity VS Code Setup Guide: Build a Practical AI IDE Workflow</title><link>https://jimmysong.io/blog/antigravity-vscode-style-ide/</link><pubDate>Thu, 20 Nov 2025 03:55:30 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/antigravity-vscode-style-ide/</guid><description>A practical Antigravity setup guide for developers who want a VS Code-style AI IDE, including marketplace switch, AMP and CodeX installation, and workflow tuning.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The biggest pain point when switching IDEs is user habits. By installing a series of plugins and tweaking configurations, you can make Antigravity feel much more like VS Code—preserving familiar workflows while adding Open Agent Manager capabilities.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you searched for a practical Antigravity VS Code setup, this walkthrough is optimized for that exact use case. The goal is not to replicate VS Code pixel by pixel, but to restore a familiar extension marketplace, keep your daily coding ergonomics, and still use Antigravity&amp;rsquo;s stronger agent-style execution. I focus on the concrete setup steps that materially change productivity: marketplace migration, AMP and CodeX installation, editor behavior alignment, and the trade-offs versus GitHub Copilot in real daily work.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/antigravity-ui.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/antigravity-ui.webp" alt="Figure 1: Antigravity IDE UI" data-caption="Figure 1: Antigravity IDE UI"
width="5120"
height="2880"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Antigravity IDE UI&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="continue-reading"&gt;Continue Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/qoder-alibaba-ai-ide-personal-review/"&gt;Qoder AI IDE review and hands-on comparison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/open-source-ai-agent-workflow-comparison/"&gt;Open-source AI Agent and workflow platform comparison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/blog/vibe-coding-free-tools/"&gt;Free Vibe Coding tools I actually use&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Oh My OpenCode in AI Native Landscape&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Below are the configurations and steps I actually use. Feel free to follow along.&lt;/p&gt;
&lt;h2 id="first-impressions-of-antigravity"&gt;First Impressions of Antigravity&lt;/h2&gt;
&lt;p&gt;A few subjective observations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The interface is split between agent management and editor views, somewhat like AgentHQ + VS Code.&lt;/li&gt;
&lt;li&gt;Agents modify code very quickly, with a much higher completion rate than typical &amp;ldquo;chat-based&amp;rdquo; assistants.&lt;/li&gt;
&lt;li&gt;The editor and context windows are large, ideal for long diffs and logs.&lt;/li&gt;
&lt;li&gt;By default, it uses OpenVSX / OpenVSCode Gallery, so the extension ecosystem isn&amp;rsquo;t identical to my VS Code setup.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All subsequent steps focus on one goal: keep Antigravity&amp;rsquo;s agent features while maintaining my VS Code workflow.&lt;/p&gt;
&lt;h2 id="switching-the-extension-marketplace-to-vs-code-official"&gt;Switching the Extension Marketplace to VS Code Official&lt;/h2&gt;
&lt;p&gt;Antigravity is essentially a VS Code fork, so you can directly change the Marketplace configuration.&lt;/p&gt;
&lt;p&gt;In Antigravity:&lt;/p&gt;
&lt;p&gt;Go to &lt;strong&gt;Settings&lt;/strong&gt; -&amp;gt; &lt;strong&gt;Antigravity Settings&lt;/strong&gt; -&amp;gt; &lt;strong&gt;Editor&lt;/strong&gt;, and update the following URLs to point to VS Code:&lt;/p&gt;
&lt;p&gt;Marketplace Item URL:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;https://marketplace.visualstudio.com/items&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Marketplace Gallery URL:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;https://marketplace.visualstudio.com/_apis/public/gallery&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/vscode-marketplace.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/vscode-marketplace.webp" alt="Figure 2: VSCode Marketplace Configuration" data-caption="Figure 2: VSCode Marketplace Configuration"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: VSCode Marketplace Configuration&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Restart Antigravity.&lt;/p&gt;
&lt;p&gt;After this change, searching and installing extensions works just like the official VS Code Marketplace. Installing AMP, GitHub Theme, VS Code Icon, etc., all follow this process.&lt;/p&gt;
&lt;h2 id="installing-the-amp-extension"&gt;Installing the AMP Extension&lt;/h2&gt;
&lt;p&gt;AMP isn&amp;rsquo;t officially supported on Antigravity yet, but you can install it directly via the VS Code Marketplace.&lt;/p&gt;
&lt;p&gt;Steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Open the Extensions panel (the same icon as in VS Code).&lt;/li&gt;
&lt;li&gt;Search for the AMP extension and install it as usual.&lt;/li&gt;
&lt;li&gt;Log in using your AMP API Key.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Currently, Antigravity doesn&amp;rsquo;t support one-click account login like VS Code; you have to use the API key.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
Summary
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
Once installed, AMP works almost identically in Antigravity as in VS Code—completion and refactoring features are available. The only difference is manual login configuration.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;I recommend AMP because it offers a free mode. In my experience, it&amp;rsquo;s great for writing documentation, running scripts, and as a daily command-line tool. It&amp;rsquo;s fast, and especially useful for optimizing prompts.&lt;/p&gt;
&lt;h2 id="importing-the-codex-extension"&gt;Importing the CodeX Extension&lt;/h2&gt;
&lt;p&gt;CodeX doesn&amp;rsquo;t provide a direct VSIX download link on the web. My approach is to export it from VS Code and then import it into Antigravity.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/codex-extension.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/codex-extension.webp" alt="Figure 3: Exporting Codex Extension in VS Code" data-caption="Figure 3: Exporting Codex Extension in VS Code"
width="3016"
height="2264"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Exporting Codex Extension in VS Code&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Install the CodeX extension in VS Code (if you haven&amp;rsquo;t already).&lt;/li&gt;
&lt;li&gt;In VS Code&amp;rsquo;s extension manager, find CodeX and export it as a &lt;code&gt;.vsix&lt;/code&gt; file.&lt;/li&gt;
&lt;li&gt;Switch to Antigravity, open the Extensions panel, and select &amp;ldquo;Install from VSIX&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Choose the exported &lt;code&gt;codex-x.x.x.vsix&lt;/code&gt; file to complete installation.&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="alert alert-tip-container"&gt;
&lt;div class="alert-tip-title px-2"&gt;
Tip
&lt;/div&gt;
&lt;div class="alert-tip px-2"&gt;
Since my local VS Code is already logged into CodeX, importing it into Antigravity automatically reuses the login state—I didn&amp;rsquo;t need to log in again.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="optimizing-editor-settings"&gt;Optimizing Editor Settings&lt;/h2&gt;
&lt;p&gt;Beyond the marketplace and plugins, a few tweaks make the experience even closer to VS Code:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Theme&lt;/strong&gt;: Choose the same color scheme as VS Code to minimize visual switching. I use GitHub Theme and vscode-icons.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Editor Settings&lt;/strong&gt;: In &amp;ldquo;Open Editor Settings&amp;rdquo;, set indentation, formatting, line width, etc., to match your VS Code preferences. I define these in the workspace&amp;rsquo;s &lt;code&gt;settings.json&lt;/code&gt;, so no migration is needed.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;After these changes, the editing area is essentially &amp;ldquo;VS Code with an agent console&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="remaining-issues"&gt;Remaining Issues&lt;/h2&gt;
&lt;p&gt;To fully migrate from VS Code/GitHub Copilot to Antigravity, I think there are still several key challenges:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Limited Customization&lt;/strong&gt;: Antigravity can&amp;rsquo;t support custom prompts and agents like Copilot Chat. Currently, only &amp;ldquo;rules&amp;rdquo; configuration is available, which limits workflow flexibility.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model Ecosystem Needs Improvement&lt;/strong&gt;: Antigravity hasn&amp;rsquo;t natively integrated the latest models from major vendors (OpenAI, Anthropic, Microsoft, xAI, etc.), whereas GitHub Copilot excels here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost Considerations&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;Future pricing may start at $20/month.&lt;/li&gt;
&lt;li&gt;No free models are supported, unlike GitHub Copilot (even Copilot Pro users have free model options).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stability Issues&lt;/strong&gt;: Agents often encounter &amp;ldquo;Agent terminated due to error&amp;rdquo; during operation, requiring manual retries or new sessions. This affects workflow smoothness, though I expect improvements in the future.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="github-copilot-vs-antigravity"&gt;GitHub Copilot VS. Antigravity&lt;/h2&gt;
&lt;p&gt;Although Antigravity excels in several areas, there is still significant room for improvement compared to the combination of GitHub Copilot and VS Code.&lt;/p&gt;
&lt;p&gt;The large language models (LLMs, Large Language Models) I frequently use are all supported in VS Code:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/models.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/models.webp" alt="Figure 4: Copilot-supported LLMs (partial)" data-caption="Figure 4: Copilot-supported LLMs (partial)"
width="1252"
height="1240"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 4: Copilot-supported LLMs (partial)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;My long-accumulated custom prompts:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/prompts.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/prompts.webp" alt="Figure 5: Copilot Chat enables quick access to custom prompts" data-caption="Figure 5: Copilot Chat enables quick access to custom prompts"
width="1252"
height="852"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: Copilot Chat enables quick access to custom prompts&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;My collection of agents:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/agents.webp" data-img="https://assets.jimmysong.io/images/blog/antigravity-vscode-style-ide/agents.webp" alt="Figure 6: Copilot Chat allows selection of custom agents" data-caption="Figure 6: Copilot Chat allows selection of custom agents"
width="1258"
height="2594"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: Copilot Chat allows selection of custom agents&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Here are some personal experiences using VS Code and Copilot that, for now, are hard to replace with other IDEs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Ask/Edit/Agent/Plan workflow perfectly fits my working habits.&lt;/li&gt;
&lt;li&gt;Support for custom prompts and agents is essential. Many of my prompts and agents have been refined over time and are deeply integrated into my daily workflow—it&amp;rsquo;s hard to find alternatives elsewhere.&lt;/li&gt;
&lt;li&gt;New models are integrated at lightning speed. Whenever a new model is released, GitHub Copilot is among the first to support it.&lt;/li&gt;
&lt;li&gt;The integration with VS Code is seamless—no extra configuration required, making it extremely convenient.&lt;/li&gt;
&lt;li&gt;Frequent updates: just a few days ago, a bug I reported to VS Code was fixed the same night.&lt;/li&gt;
&lt;li&gt;Copilot Chat&amp;rsquo;s keyboard shortcuts make it easy to quickly access various features.&lt;/li&gt;
&lt;li&gt;GitHub has granted me a free Pro account. Although the monthly premium quota is only 300 calls, combining Copilot with other plugins like AMP, Codex, Droid, and Qwen enables a highly efficient workflow. Even if I upgrade to a paid account in the future, the $10/month fee is very cost-effective compared to similar products.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practical-experience"&gt;Practical Experience&lt;/h2&gt;
&lt;p&gt;A few subjective tips from my actual usage—take them as reference:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Don&amp;rsquo;t treat Antigravity as &amp;ldquo;VS Code + chat box&amp;rdquo;. Use its agent features for complete tasks: let the agent propose a plan, then execute changes.&lt;/li&gt;
&lt;li&gt;For major changes, always create a new Git branch and restrict agent actions to that branch. Handle all diffs via standard Pull Request (PR) workflows.&lt;/li&gt;
&lt;li&gt;Ask agents to produce &amp;ldquo;artifacts&amp;rdquo; (plans, proposals, test descriptions), not just final code. This makes it easier to review and track changes.&lt;/li&gt;
&lt;li&gt;Plugins you&amp;rsquo;re already comfortable with in VS Code (like AMP, CodeX) can be migrated directly, reducing cognitive load and letting you focus on new agent workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;My current experience: Antigravity delivers powerful agent capabilities and multi-view consoles. By following these steps to align the interface and plugin ecosystem with VS Code, you can smoothly transition your daily development workflow.&lt;/p&gt;</content:encoded></item><item><title>Cloudflare November 18 Global Outage: The Dangers of Implicit Assumptions in Modern Infrastructure</title><link>https://jimmysong.io/blog/cloudflare-2025-11-18-outage-analysis/</link><pubDate>Wed, 19 Nov 2025 18:56:34 +0800</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/cloudflare-2025-11-18-outage-analysis/</guid><description>An analysis of the Cloudflare global outage on November 18, 2025, exploring implicit assumptions, automated configuration pipelines, and systemic risks in modern infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The greatest risks to modern internet infrastructure often aren&amp;rsquo;t in the code itself, but in those implicit assumptions and automated configuration pipelines that go undefined. Cloudflare&amp;rsquo;s outage is a wake-up call every Infra/AI engineer must heed.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Yesterday (November 18), Cloudflare experienced its largest global outage since 2019. As this site is hosted on Cloudflare, it was also affected—one of the rare times in eight years that the site was inaccessible due to an outage (the last time was a GitHub Pages failure, which happened the year Microsoft acquired GitHub).&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/cloudflare-2025-11-18-outage-analysis/jimmysongio-down.webp" data-img="https://assets.jimmysong.io/images/blog/cloudflare-2025-11-18-outage-analysis/jimmysongio-down.webp" alt="Figure 1: jimmysong.io was down for 27 minutes due to the Cloudflare outage" data-caption="Figure 1: jimmysong.io was down for 27 minutes due to the Cloudflare outage"
width="2694"
height="1424"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: jimmysong.io was down for 27 minutes due to the Cloudflare outage&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This incident was not caused by an attack or a traditional software bug, but by a seemingly &amp;ldquo;safe&amp;rdquo; permissions update that triggered the weakest link in modern infrastructure: &lt;strong&gt;implicit assumptions (Implicit Assumption) and automated configuration pipelines (Automated Configuration Pipeline)&lt;/strong&gt;. Cloudflare has published a blog post &lt;a href="https://blog.cloudflare.com/18-november-2025-outage/" target="_blank" rel="noopener"&gt;Cloudflare outage on November 18, 2025&lt;/a&gt; explaining the cause.&lt;/p&gt;
&lt;p&gt;Here is the chain reaction process of the outage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A permissions adjustment led to metadata changes;&lt;/li&gt;
&lt;li&gt;The metadata change doubled the lines in the feature file;&lt;/li&gt;
&lt;li&gt;The doubled lines triggered the proxy module&amp;rsquo;s memory limit;&lt;/li&gt;
&lt;li&gt;The memory limit caused the core proxy to panic;&lt;/li&gt;
&lt;li&gt;The proxy panic led to a cascade failure in downstream systems.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This kind of chain reaction is the most typical—and dangerous—systemic failure mode at today&amp;rsquo;s internet scale.&lt;/p&gt;
&lt;h2 id="root-cause-implicit-assumptions-are-not-contracts"&gt;Root Cause: Implicit Assumptions Are Not Contracts&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s first look at the core hidden risk in this incident. The Bot Management feature file is automatically generated every five minutes, relying on a default premise:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The system.columns query result contains only the default database.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This assumption was not documented or validated in configuration—it existed only in the engineer&amp;rsquo;s mental model.&lt;/p&gt;
&lt;p&gt;After a ClickHouse permissions update, the underlying r0 tables were exposed, instantly doubling the query results. The file size exceeded the &lt;a href="https://blog.cloudflare.com/20-percent-internet-upgrade/" target="_blank" rel="noopener"&gt;FL2&lt;/a&gt; preset of 200 features in memory, ultimately causing a panic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Once an implicit assumption is broken, the system lacks a buffer and is highly prone to cascading failures.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="configuration-pipelines-are-riskier-than-code-pipelines"&gt;Configuration Pipelines Are Riskier Than Code Pipelines&lt;/h2&gt;
&lt;p&gt;This incident was not caused by code changes, but by data-plane changes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SQL query behavior changed;&lt;/li&gt;
&lt;li&gt;Feature files were automatically generated;&lt;/li&gt;
&lt;li&gt;The files were broadcast across the network.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A typical phenomenon in modern infrastructure: &lt;strong&gt;data, schema, and metadata are far more likely to destabilize systems than code.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Cloudflare&amp;rsquo;s feature file is a &amp;ldquo;supply chain input,&amp;rdquo; not a regular configuration. Anything entering the automated broadcast path is equivalent to a system-level command.&lt;/p&gt;
&lt;h2 id="language-safety-cant-eliminate-boundary-layer-complexity"&gt;Language Safety Can&amp;rsquo;t Eliminate Boundary Layer Complexity&lt;/h2&gt;
&lt;p&gt;A former Cloudflare engineer summarized it well:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Rust can prevent a class of errors, but the complexity of boundary layers, data contracts, and configuration pipelines does not disappear.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The FL2 panic stemmed from a single &lt;code&gt;unwrap()&lt;/code&gt;. This isn&amp;rsquo;t a language issue, but a lack of system contracts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No upper-bound validation for feature count;&lt;/li&gt;
&lt;li&gt;File schema lacked version constraints;&lt;/li&gt;
&lt;li&gt;Feature generation logic depended on implicit behavior;&lt;/li&gt;
&lt;li&gt;Core proxy error mode was panic, not graceful degradation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Most incidents in modern distributed systems (Distributed System) come from &amp;ldquo;bad input,&amp;rdquo; not &amp;ldquo;bad memory.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="core-proxies-need-controllable-failure-paths"&gt;Core Proxies Need Controllable Failure Paths&lt;/h2&gt;
&lt;p&gt;FL/FL2 are Cloudflare&amp;rsquo;s core proxies; all requests must pass through them. Such components should not fail with a panic, but have the following capabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Ignore abnormal features;&lt;/li&gt;
&lt;li&gt;Truncate over-limit fields;&lt;/li&gt;
&lt;li&gt;Roll back to previous versions;&lt;/li&gt;
&lt;li&gt;Fail-open or fail-close;&lt;/li&gt;
&lt;li&gt;Skip the Bot module and continue processing traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As long as the proxy &amp;ldquo;stays alive,&amp;rdquo; the entire network won&amp;rsquo;t be completely paralyzed.&lt;/p&gt;
&lt;h2 id="data-changes-are-more-uncontrollable-than-code-changes"&gt;Data Changes Are More Uncontrollable Than Code Changes&lt;/h2&gt;
&lt;p&gt;The essence of this incident:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Subtle permission changes;&lt;/li&gt;
&lt;li&gt;ClickHouse default behavior changed;&lt;/li&gt;
&lt;li&gt;Query results propagated to distributed systems;&lt;/li&gt;
&lt;li&gt;Automated publishing amplified the error;&lt;/li&gt;
&lt;li&gt;Edge proxies crashed due to uncontrolled input.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Future AI Infra (AI Infrastructure) will be even more complex: models, tokenizers, adapters, RAG indexes, and KV snapshots all require frequent updates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;In future AI infrastructure, data-plane risks will far exceed those of the code-plane.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="recovery-process-shows-engineering-maturity"&gt;Recovery Process Shows Engineering Maturity&lt;/h2&gt;
&lt;p&gt;During the incident, Cloudflare took several measures:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Stopped generating erroneous feature files;&lt;/li&gt;
&lt;li&gt;Force-distributed the previous version of the file;&lt;/li&gt;
&lt;li&gt;Rolled back Bot module configuration;&lt;/li&gt;
&lt;li&gt;Ran Workers KV and Access outside the core proxy;&lt;/li&gt;
&lt;li&gt;Restored traffic in stages.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Restoring hundreds of PoPs worldwide simultaneously demonstrates a high level of engineering maturity.&lt;/p&gt;
&lt;h2 id="lessons-for-infraaicloud-native-engineers"&gt;Lessons for Infra/AI/Cloud Native Engineers&lt;/h2&gt;
&lt;p&gt;The Cloudflare event highlights four common risks in large-scale systems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Implicit assumptions fail;&lt;/li&gt;
&lt;li&gt;Configuration supply chain contamination;&lt;/li&gt;
&lt;li&gt;Automated publishing amplifies errors;&lt;/li&gt;
&lt;li&gt;Core proxies lack graceful degradation paths.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For AI Infra practitioners, these risks are even more relevant:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model weight updates without schema validation;&lt;/li&gt;
&lt;li&gt;Adapter merges may be contaminated;&lt;/li&gt;
&lt;li&gt;RAG index incremental builds are unstable;&lt;/li&gt;
&lt;li&gt;Inference graph configuration may be broken by bad data;&lt;/li&gt;
&lt;li&gt;Automatically rolled-out models may propagate errors network-wide.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;AI engineering is replaying Cloudflare&amp;rsquo;s infrastructure dilemmas—just at greater speed and scale.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary-of-former-cloudflare-engineers-views"&gt;Summary of Former Cloudflare Engineer&amp;rsquo;s Views&lt;/h2&gt;
&lt;p&gt;His insights pinpoint the hardest problems in distributed systems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The issue isn&amp;rsquo;t code, but missing contracts;&lt;/li&gt;
&lt;li&gt;Not the language, but undefined input boundaries;&lt;/li&gt;
&lt;li&gt;Not modules, but lack of validation in the configuration supply chain;&lt;/li&gt;
&lt;li&gt;Not bugs, but absence of fail-safe mechanisms.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This incident proves: &lt;strong&gt;The real fragility in modern infrastructure lies in &amp;ldquo;behavioral boundaries,&amp;rdquo; not &amp;ldquo;memory boundaries.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The Cloudflare November 18 outage was not a coincidence, but an inevitable result of modern internet infrastructure evolving to large-scale, highly automated stages.&lt;/p&gt;
&lt;p&gt;Key takeaways from this event:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;System assumptions must be made explicit;&lt;/li&gt;
&lt;li&gt;Configuration pipelines must be validated;&lt;/li&gt;
&lt;li&gt;Automated publishing needs &amp;ldquo;dead-end&amp;rdquo; mechanisms;&lt;/li&gt;
&lt;li&gt;Core proxies must be designed with controllable failure paths;&lt;/li&gt;
&lt;li&gt;Data-plane contracts must be stricter than code-plane contracts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the AI-native Infra era, these requirements will only become more stringent.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blog.cloudflare.com/18-november-2025-outage/" target="_blank" rel="noopener"&gt;Cloudflare outage on November 18, 2025 - blog.cloudflare.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.cloudflare.com/20-percent-internet-upgrade/" target="_blank" rel="noopener"&gt;20% of the Internet upgraded - blog.cloudflare.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>The Second Half of Cloud Native: The Era of AI Native Platform Engineering Has Arrived</title><link>https://jimmysong.io/blog/cloud-native-second-half-ai-native-platform-engineering/</link><pubDate>Mon, 17 Nov 2025 11:07:40 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/cloud-native-second-half-ai-native-platform-engineering/</guid><description>A decade of cloud native evolution, a look ahead to AI-Native Platform engineering, technical layers, and key changes. KubeCon NA 2025 signals a new era.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The second half of cloud native isn&amp;rsquo;t about being replaced by AI, but being rewritten by it. The future of platform engineering will revolve around models and agents, reshaping the tech stack and developer experience.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Since I first encountered Docker and Kubernetes in 2015, I&amp;rsquo;ve followed the cloud native journey: from writing Deployments in YAML, to exploring Service Mesh and observability, and in recent years, focusing on AI Infra and AI-Native Platforms. Looking back from 2025, the years 2015–2025 can be seen as the &amp;ldquo;first half&amp;rdquo; of cloud native. Marked by &lt;a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/" target="_blank" rel="noopener"&gt;KubeCon / CloudNativeCon NA 2025&lt;/a&gt;, the industry is collectively entering the &amp;ldquo;second half&amp;rdquo;: the era of AI-Native Platform engineering.&lt;/p&gt;
&lt;p&gt;This article reviews the past decade of cloud native, and, combined with KubeCon NA 2025, outlines key turning points and the technical coordinates for the next ten years.&lt;/p&gt;
&lt;h2 id="20152025-the-first-half-of-cloud-native"&gt;2015–2025: The &amp;ldquo;First Half&amp;rdquo; of Cloud Native&lt;/h2&gt;
&lt;p&gt;Over the past decade, cloud native technology themes have evolved through three main stages. The following flowchart illustrates the progression.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The flowchart below illustrates the progression of cloud native technology themes over the past decade:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/cloud-native-second-half-ai-native-platform-engineering/cloud-native-decade-evolution.svg" data-img="https://assets.jimmysong.io/images/blog/cloud-native-second-half-ai-native-platform-engineering/cloud-native-decade-evolution.svg" alt="Figure 1: Cloud Native Decade Technology Evolution Flow" data-caption="Figure 1: Cloud Native Decade Technology Evolution Flow"
width="2543"
height="223"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 1: Cloud Native Decade Technology Evolution Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The first stage focused on containerization and orchestration standardization.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Docker realized the engineering dream of &amp;ldquo;build once, run anywhere&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Kubernetes won the orchestration wars and became the de facto standard&lt;/li&gt;
&lt;li&gt;CNCF was founded, with Prometheus, Envoy, and other projects joining&lt;/li&gt;
&lt;li&gt;Enterprises focused on migrating applications to Kubernetes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Typical tasks during this phase involved moving Java services from VMs to containers and K8s, emphasizing understanding of Deployment, Service, and Ingress.&lt;/p&gt;
&lt;p&gt;The second stage, 2018–2020, saw complexity shift from &amp;ldquo;deployment&amp;rdquo; to &amp;ldquo;communication&amp;rdquo; and &amp;ldquo;operations&amp;rdquo;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Service Mesh (Istio / Linkerd / Consul) addressed east-west traffic management&lt;/li&gt;
&lt;li&gt;The observability trio (Logs / Metrics / Traces) became default configurations&lt;/li&gt;
&lt;li&gt;Multi-cluster and multi-region practices matured&lt;/li&gt;
&lt;li&gt;Enterprises focused on managing large microservice systems&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;During this period, I spent significant time researching Istio, service mesh, and traffic management, and authored Kubernetes and Istio books. The focus shifted to system stability, observability, and reliability.&lt;/p&gt;
&lt;p&gt;The third stage, 2021–2025, is defined by Platform Engineering and GitOps.&lt;/p&gt;
&lt;p&gt;As microservices and tools proliferated, platform complexity began to overwhelm developers, making Platform Engineering a key industry term.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GitOps (Argo CD / Flux) drove declarative delivery processes&lt;/li&gt;
&lt;li&gt;Internal Developer Platforms (IDP) became priorities for large enterprises&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Platform as a product&amp;rdquo; philosophy spread&lt;/li&gt;
&lt;li&gt;FinOps, cost management, and compliance auditing became platform concerns&lt;/li&gt;
&lt;li&gt;DevOps evolved from &amp;ldquo;tool practice&amp;rdquo; to &amp;ldquo;organizational + platform capability&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;My takeaway: simply giving developers a pile of tools isn&amp;rsquo;t enough. End-to-end delivery paths and stable abstraction layers are needed so developers can focus on business, not tool integration.&lt;/p&gt;
&lt;p&gt;The table below summarizes the main features of each &amp;ldquo;first half&amp;rdquo; stage.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Core Challenge&lt;/th&gt;
&lt;th&gt;Key Tech Stack&lt;/th&gt;
&lt;th&gt;Typical Issues&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2015–2017 Orchestration&lt;/td&gt;
&lt;td&gt;Migrating from VM to containers&lt;/td&gt;
&lt;td&gt;Docker, Kubernetes, CNI&lt;/td&gt;
&lt;td&gt;Reliable deployment, rolling upgrades&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2018–2020 Mesh&lt;/td&gt;
&lt;td&gt;Microservice scale, complex communication &amp;amp; observability&lt;/td&gt;
&lt;td&gt;Istio/Linkerd, Prometheus, Jaeger&lt;/td&gt;
&lt;td&gt;Troubleshooting, fragmented observability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2021–2025 Platform&lt;/td&gt;
&lt;td&gt;Tool sprawl, declining developer experience&lt;/td&gt;
&lt;td&gt;GitOps, IDP, FinOps, Policy-as-Code&lt;/td&gt;
&lt;td&gt;Developer fatigue, platform team overload&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 1: Cloud Native First Half Stage Features
&lt;/figcaption&gt;
&lt;h2 id="kubecon-na-2025-signals-of-cloud-natives-second-half"&gt;KubeCon NA 2025: Signals of Cloud Native&amp;rsquo;s &amp;ldquo;Second Half&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;The main theme of KubeCon 2025 is no longer &amp;ldquo;how to use Kubernetes well,&amp;rdquo; but how to reconstruct Kubernetes and the cloud native ecosystem into AI-Native Platforms for the AI era.&lt;/p&gt;
&lt;p&gt;Key signals from KubeCon NA 2025 include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;CNCF released the &lt;a href="https://github.com/cncf/k8s-ai-conformance" target="_blank" rel="noopener"&gt;Certified Kubernetes AI Conformance Program&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Dynamic Resource Allocation (DRA) entered mainstream discussions&lt;/li&gt;
&lt;li&gt;Model Runtime / Agent Runtime projects became conference hotspots&lt;/li&gt;
&lt;li&gt;Vendors focused on AI SRE, AI-assisted development, AI security, and supply chain governance&lt;/li&gt;
&lt;li&gt;Speakers like Alex Zenla openly stated that Kubernetes&amp;rsquo; underlying structure needs rethinking&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Together, these mark a clear dividing line: cloud native has officially entered its &amp;ldquo;second half.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="first-half-vs-second-half-shifting-the-cloud-native-narrative"&gt;First Half vs Second Half: Shifting the Cloud Native Narrative&lt;/h2&gt;
&lt;p&gt;If 2015–2025 is the &amp;ldquo;first half,&amp;rdquo; then 2025–2035 is likely the &amp;ldquo;second half.&amp;rdquo; The table below compares their core differences.&lt;/p&gt;
&lt;p&gt;It highlights changes in platform objects, goals, abstraction layers, and more.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;First Half (2015–2025)&lt;/th&gt;
&lt;th&gt;Second Half (2025–2035, AI Native)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core Objects&lt;/td&gt;
&lt;td&gt;Containers, Pods, Microservices&lt;/td&gt;
&lt;td&gt;Models, inference tasks, Agents, data pipelines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform Goals&lt;/td&gt;
&lt;td&gt;Stable application delivery&lt;/td&gt;
&lt;td&gt;Efficient, continuous AI workload &amp;amp; agent orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Abstraction Layers&lt;/td&gt;
&lt;td&gt;Deployment / Service / Ingress / Job&lt;/td&gt;
&lt;td&gt;Model / Endpoint / Graph / Policy / Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource Scheduling&lt;/td&gt;
&lt;td&gt;CPU / Memory / Node&lt;/td&gt;
&lt;td&gt;GPU / TPU / ASIC / KV Cache / Bandwidth / Power&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering Focus&lt;/td&gt;
&lt;td&gt;DevOps / GitOps / Platform Engineering 1.0&lt;/td&gt;
&lt;td&gt;AI Native Platform Engineering / AI SRE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security &amp;amp; Compliance&lt;/td&gt;
&lt;td&gt;Image security, CVE, supply chain SBOM&lt;/td&gt;
&lt;td&gt;Model security, data security, AI supply chain &amp;amp; &amp;ldquo;hallucination dependencies&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime Forms&lt;/td&gt;
&lt;td&gt;Container + VM + Serverless&lt;/td&gt;
&lt;td&gt;Container + WASM + Nix + Agent Runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 2: Core Differences: First vs Second Half of Cloud Native
&lt;/figcaption&gt;
&lt;p&gt;From a developer&amp;rsquo;s perspective, the most direct change is: future platforms will no longer treat &amp;ldquo;services&amp;rdquo; as first-class citizens, but will center on &amp;ldquo;models + agents.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="example-technical-layers-of-an-ai-native-platform"&gt;Example: Technical Layers of an AI Native Platform&lt;/h2&gt;
&lt;p&gt;To clarify the structure of an AI-Native Platform, the following layered diagram shows the relationships between technical levels.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The layered diagram below shows the relationships between different technical levels in an AI-Native Platform:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/cloud-native-second-half-ai-native-platform-engineering/ai-native-platform-layering.svg" data-img="https://assets.jimmysong.io/images/blog/cloud-native-second-half-ai-native-platform-engineering/ai-native-platform-layering.svg" alt="Figure 2: AI Native Platform Layering Diagram" data-caption="Figure 2: AI Native Platform Layering Diagram"
width="2063"
height="1643"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 2: AI Native Platform Layering Diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Historically, cloud native focused on L0 + L2 (Kubernetes + platform engineering), but in the AI Native era, L1 (Model Runtime, Agent Runtime, heterogeneous resource scheduling) becomes the new battleground.&lt;/p&gt;
&lt;h2 id="key-change-1-from-container-centric-to-model-centric"&gt;Key Change 1: From &amp;ldquo;Container-Centric&amp;rdquo; to &amp;ldquo;Model-Centric&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;In the first half, cloud native&amp;rsquo;s main object was the application process, with containers as packaging. The second half requires handling:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Model version management and canary releases&lt;/li&gt;
&lt;li&gt;Balancing inference performance, latency, and cost&lt;/li&gt;
&lt;li&gt;Multi-model composition, routing, A/B testing&lt;/li&gt;
&lt;li&gt;Relationships between models, data, features, and vector indexes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At KubeCon NA 2025, CNCF&amp;rsquo;s AI Conformance Program aims to standardize model workloads, managing them like Deployments. Platform engineering will gain new abstractions—not just &amp;ldquo;deploying services,&amp;rdquo; but &amp;ldquo;deploying model capabilities.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="key-change-2-dra-and-the-golden-window-for-heterogeneous-resource-scheduling"&gt;Key Change 2: DRA and the Golden Window for Heterogeneous Resource Scheduling&lt;/h2&gt;
&lt;p&gt;Previously, writing a Deployment meant focusing on CPU and memory. Now, GPU inference, training, and Agent Runtime scenarios demand more than static quotas.&lt;/p&gt;
&lt;p&gt;Dynamic Resource Allocation (DRA) brings:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Pluggable resource types (GPU/TPU/FPGA/ASIC)&lt;/li&gt;
&lt;li&gt;Topology-aware, NUMA, and memory fragmentation scheduling&lt;/li&gt;
&lt;li&gt;Binding inference requests to compute allocation for fine-grained QoS&lt;/li&gt;
&lt;li&gt;Cost optimization and power control in scheduling decisions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the most significant &amp;ldquo;resource perspective&amp;rdquo; upgrade since Kubernetes&amp;rsquo; inception. The scheduler is no longer just a cluster component, but the AI platform&amp;rsquo;s policy engine.&lt;/p&gt;
&lt;h2 id="key-change-3-agent-runtime-as-the-new-generation-of-runtime"&gt;Key Change 3: Agent Runtime as the New Generation of Runtime&lt;/h2&gt;
&lt;p&gt;KubeCon showcased several representative projects:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://edera.dev" target="_blank" rel="noopener"&gt;Edera&lt;/a&gt;: Minimal, verifiable runtime redesign&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/flox/flox" target="_blank" rel="noopener"&gt;Flox&lt;/a&gt;: Nix-based &amp;ldquo;uncontained&amp;rdquo; runtime environment&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/golemcloud/golem" target="_blank" rel="noopener"&gt;Golem&lt;/a&gt;: WASM-based large-scale agent orchestration&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The consensus: AI agents aren&amp;rsquo;t suited to traditional container runtime models. Agents have these traits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Strong statefulness: context, memory, sessions&lt;/li&gt;
&lt;li&gt;High concurrency but fine granularity: massive lightweight tasks&lt;/li&gt;
&lt;li&gt;Extremely sensitive to latency and cold starts&lt;/li&gt;
&lt;li&gt;Need to resume after failure&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Next-gen runtimes focus on reliably executing, managing state, and auditing &amp;ldquo;hundreds of thousands of agents,&amp;rdquo; not just &amp;ldquo;spinning up more Pods.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="key-change-4-ai-sre-and-ai-security"&gt;Key Change 4: AI SRE and AI Security&lt;/h2&gt;
&lt;p&gt;At KubeCon NA 2025, security and operations topics were amplified by AI:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Software supply chain attacks and CVEs continue to rise&lt;/li&gt;
&lt;li&gt;LLM-assisted coding introduces &amp;ldquo;hallucination dependencies&amp;rdquo; and &amp;ldquo;vibecoded vulnerabilities&amp;rdquo;&lt;/li&gt;
&lt;li&gt;AI-driven artifact scanning, dependency auditing, and license analysis&lt;/li&gt;
&lt;li&gt;&amp;ldquo;AI SRE&amp;rdquo; is now a formal product category&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Traditional cloud native already emphasized security and SRE, but now must address model weights, datasets, vector stores, and agent workflows. AI-Native Platform engineering must answer:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Are code and dependencies secure?&lt;/li&gt;
&lt;li&gt;Are models and data trustworthy?&lt;/li&gt;
&lt;li&gt;Are agent behaviors controllable?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This will drive deep integration of Policy-as-Code, MCP, graph permission systems, and AI.&lt;/p&gt;
&lt;h2 id="key-change-5-open-source-participation-becomes-a-baseline"&gt;Key Change 5: Open Source Participation Becomes a Baseline&lt;/h2&gt;
&lt;p&gt;In interviews, platform engineering leaders noted:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Hiring increasingly values upstream contributions to Kubernetes and related projects&lt;/li&gt;
&lt;li&gt;Open source involvement shortens ramp-up time&lt;/li&gt;
&lt;li&gt;New AI Native projects (Model Runtime, Agent Runtime, Scheduler) are also open source&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For career growth, contributing to AI Native open source projects will become a basic requirement for platform engineering and AI Infra roles—not just a resume bonus.&lt;/p&gt;
&lt;h2 id="the-contours-of-cloud-natives-second-half"&gt;The Contours of Cloud Native&amp;rsquo;s &amp;ldquo;Second Half&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;The table below summarizes the technical focus and essential differences of the &amp;ldquo;second half.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;It highlights the key coordinates of AI-Native Platform engineering.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Direction&lt;/th&gt;
&lt;th&gt;Technical Focus&lt;/th&gt;
&lt;th&gt;Essential Difference from First Half&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AI Native Platform&lt;/td&gt;
&lt;td&gt;Models/Agents as first-class citizens, unified abstraction &amp;amp; governance&lt;/td&gt;
&lt;td&gt;Objects shift from services to models &amp;amp; inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource Scheduling&lt;/td&gt;
&lt;td&gt;DRA, heterogeneous compute, topology awareness, power &amp;amp; cost&lt;/td&gt;
&lt;td&gt;From static quotas to dynamic, policy-driven&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;Container + WASM + Nix + Agent Runtime&lt;/td&gt;
&lt;td&gt;From &amp;ldquo;process containerization&amp;rdquo; to &amp;ldquo;execution graph containerization&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform Engineering&lt;/td&gt;
&lt;td&gt;IDP + AI SRE + Security + Cost + Compliance&lt;/td&gt;
&lt;td&gt;From toolset to &amp;ldquo;autonomous platform&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security &amp;amp; Supply Chain&lt;/td&gt;
&lt;td&gt;LLM dependencies, model weights, datasets, vector store governance&lt;/td&gt;
&lt;td&gt;Protection expands from images to &amp;ldquo;all AI engineering assets&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Source &amp;amp; Ecosystem&lt;/td&gt;
&lt;td&gt;AI Infra / Model Runtime / Agent Runtime upstream collaboration&lt;/td&gt;
&lt;td&gt;Not just &amp;ldquo;using open source,&amp;rdquo; but &amp;ldquo;building the future in open source&amp;rdquo;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Cloud Native Second Half Technical Coordinates
&lt;/figcaption&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Over the past decade, cloud native evolved from container orchestration to platform engineering 1.0. With KubeCon NA 2025 as a milestone, the industry systematically brings AI into cloud native technology and organizational stacks:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Kubernetes is no longer just &amp;ldquo;infrastructure for microservices,&amp;rdquo; but &amp;ldquo;runtime for AI workloads&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Platform Engineering is no longer just &amp;ldquo;tool integration,&amp;rdquo; but &amp;ldquo;autonomous platforms for models and agents&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Security, SRE, runtime, scheduling, and networking will all be reimagined under AI&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For me, the past ten years were about &amp;ldquo;making applications more stable in the cloud native world.&amp;rdquo; The next ten will focus on &amp;ldquo;making AI better, safer, and more controllable in the cloud native world.&amp;rdquo; This is, in my view, the opening whistle for cloud native&amp;rsquo;s &amp;ldquo;second half.&amp;rdquo;&lt;/p&gt;</content:encoded></item><item><title>NotebookLM: My Most Recommended AI Tool for Learning and Knowledge Organization</title><link>https://jimmysong.io/blog/notebooklm-learning-and-knowledge-organization/</link><pubDate>Mon, 17 Nov 2025 08:44:45 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/notebooklm-learning-and-knowledge-organization/</guid><description>Based on months of deep usage, this article analyzes how NotebookLM helps me learn new technologies, read complex documents, generate teaching outlines, and shares future improvement expectations.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;NotebookLM is the most tailored AI tool I&amp;rsquo;ve used for knowledge workers. It truly helps me structure massive information and dramatically boosts my learning and content creation efficiency.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;As a lifelong learner who reads technical specs and researches open-source projects, I&amp;rsquo;ve always sought a tool that can &amp;ldquo;shortcut&amp;rdquo; my way through mountains of material, reduce mechanical reading, and help me quickly build a global understanding. &lt;a href="https://notebooklm.google.com" target="_blank" rel="noopener"&gt;NotebookLM&lt;/a&gt; has been the smoothest and most reliable experience for me over the past year.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not a traditional &amp;ldquo;chat-style AI tool&amp;rdquo;—it&amp;rsquo;s more like an &lt;strong&gt;AI-native learning and content organization system&lt;/strong&gt; that ingests your materials, organizes them, and presents them in various structured formats. The more I use it, the more I realize its help in learning new technologies, understanding unfamiliar fields, organizing large project documents, and building teaching materials—things that general large language models (LLM, Large Language Model) simply can&amp;rsquo;t match.&lt;/p&gt;
&lt;h2 id="the-core-value-notebooklm-brings-me"&gt;The Core Value NotebookLM Brings Me&lt;/h2&gt;
&lt;p&gt;NotebookLM has significantly improved my workflow, especially in learning new technologies, organizing documents, and content creation.&lt;/p&gt;
&lt;h2 id="quickly-understanding-new-technologies-feed-in-complex-materials-get-a-learnable-version"&gt;Quickly Understanding New Technologies: Feed in Complex Materials, Get a &amp;ldquo;Learnable Version&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;My most frequent and indispensable scenario is &lt;strong&gt;learning a completely unfamiliar technology or development framework&lt;/strong&gt;. Faced with dozens or even hundreds of pages of documentation, my typical approach is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Add official docs, README files, design documents, and architecture diagrams into a single Notebook&lt;/li&gt;
&lt;li&gt;Let NotebookLM generate:
&lt;ul&gt;
&lt;li&gt;Study guides&lt;/li&gt;
&lt;li&gt;Briefings&lt;/li&gt;
&lt;li&gt;Key knowledge points&lt;/li&gt;
&lt;li&gt;FAQs&lt;/li&gt;
&lt;li&gt;Quizzes&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Ultimately, I get a clearly structured &amp;ldquo;learning entry point&amp;rdquo; instead of a flood of raw materials.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following flowchart illustrates how NotebookLM compresses complex documents into a learnable structure:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/notebooklm-learning-and-knowledge-organization/042f3817d5b5c24e7bd54b9638272151.svg" data-img="https://assets.jimmysong.io/images/blog/notebooklm-learning-and-knowledge-organization/042f3817d5b5c24e7bd54b9638272151.svg" alt="Figure 5: NotebookLM Document Structuring Flow" data-caption="Figure 5: NotebookLM Document Structuring Flow"
width="551"
height="833"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 5: NotebookLM Document Structuring Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the end, what I gain is an &amp;ldquo;organized knowledge system&amp;rdquo; rather than a pile of PDFs waiting to be consumed.&lt;/p&gt;
&lt;h2 id="generating-mindmaps-instantly-turning-large-documents-into-structured-knowledge-graphs"&gt;Generating MindMaps: Instantly Turning Large Documents into Structured Knowledge Graphs&lt;/h2&gt;
&lt;p&gt;I rely heavily on MindMaps to build the &amp;ldquo;skeleton of knowledge.&amp;rdquo; NotebookLM&amp;rsquo;s MindMap feature stands out for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Automatically identifying relationships between topics&lt;/li&gt;
&lt;li&gt;Interactive node expansion and collapse&lt;/li&gt;
&lt;li&gt;Integrating multiple source documents&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Although it currently only exports PNG, the logical structure itself is already an excellent &amp;ldquo;knowledge compression.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The table below compares the auto-generation and visualization capabilities of different tools:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Auto-Generation&lt;/th&gt;
&lt;th&gt;Multi-Doc Integration&lt;/th&gt;
&lt;th&gt;Visualization Quality&lt;/th&gt;
&lt;th&gt;Export Formats&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NotebookLM&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;PNG only (SVG not yet supported)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common LLM Tools&lt;/td&gt;
&lt;td&gt;Weak&lt;/td&gt;
&lt;td&gt;Weak&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Depends on tool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MindMap Software (Manual)&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Fully supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 7: Comparison of MindMap Capabilities in Mainstream Tools
&lt;/figcaption&gt;
&lt;p&gt;NotebookLM&amp;rsquo;s greatest advantage is &lt;strong&gt;automation&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="generating-teaching-outlines-training-scripts-and-book-structures-truly-saving-me-time"&gt;Generating Teaching Outlines, Training Scripts, and Book Structures: Truly Saving Me Time&lt;/h2&gt;
&lt;p&gt;NotebookLM is more than just &amp;ldquo;summarization&amp;rdquo;—it can generate &lt;strong&gt;formal teaching structures&lt;/strong&gt; based on my prompts. By feeding in project docs, API references, architecture designs, case studies, videos, and blogs, and prompting it to generate:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Teaching outlines&lt;/li&gt;
&lt;li&gt;Project training manuals&lt;/li&gt;
&lt;li&gt;Course structures&lt;/li&gt;
&lt;li&gt;Book chapter frameworks&lt;/li&gt;
&lt;li&gt;Slide text&lt;/li&gt;
&lt;li&gt;Training case descriptions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For anyone who needs to create content, conduct training, or give presentations, this feature is a huge time-saver.&lt;/p&gt;
&lt;p&gt;Below is a typical prompt I actually use:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Based on the provided content excerpts, write a detailed training manual that systematically explains the core principles covered. The manual should use a professional and instructional tone, breaking down complex concepts into actionable steps and lessons. Ensure all content is strictly based on the source material and covers every aspect mentioned.
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;The training manual should include:
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;1. Training objectives and expected outcomes
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;2. Training content and structure
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;3. Training methods and tools
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;4. Training evaluation and feedback
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;5. Training summary and follow-up actions
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;6. Training cases and examples
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;7. Training resources and references&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The results are often surprisingly good.&lt;/p&gt;
&lt;h2 id="multi-format-input-capability-the-most-stable-ive-seen"&gt;Multi-Format Input Capability: The Most Stable I&amp;rsquo;ve Seen&lt;/h2&gt;
&lt;p&gt;NotebookLM supports direct ingestion of various material types, with extremely stable parsing. The table below summarizes my actual experience:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Input Type&lt;/th&gt;
&lt;th&gt;My Actual Experience&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PDF&lt;/td&gt;
&lt;td&gt;Most stable, clear structure parsing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Docs&lt;/td&gt;
&lt;td&gt;Syncs instantly, very smooth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Word / PPT&lt;/td&gt;
&lt;td&gt;Recognized normally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YouTube Video&lt;/td&gt;
&lt;td&gt;Auto-summary + key content extraction, very useful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Website URL&lt;/td&gt;
&lt;td&gt;Depends on site structure, high success rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plain Text&lt;/td&gt;
&lt;td&gt;No issues&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images&lt;/td&gt;
&lt;td&gt;Partial success, sufficient for screenshots&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 8: NotebookLM Multi-Format Input Experience
&lt;/figcaption&gt;
&lt;p&gt;By contrast, other tools often have format parsing issues, garbled text, missing content, or skipped paragraphs. NotebookLM is especially stable in &amp;ldquo;multi-format ingestion.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="my-most-common-notebooklm-workflow"&gt;My Most Common NotebookLM Workflow&lt;/h2&gt;
&lt;p&gt;The following flowchart shows my daily workflow with NotebookLM:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/notebooklm-learning-and-knowledge-organization/95790bef2620a5625da7e72caea7bb00.svg" data-img="https://assets.jimmysong.io/images/blog/notebooklm-learning-and-knowledge-organization/95790bef2620a5625da7e72caea7bb00.svg" alt="Figure 6: NotebookLM Daily Workflow" data-caption="Figure 6: NotebookLM Daily Workflow"
width="1566"
height="532"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 6: NotebookLM Daily Workflow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Essentially: let AI help me grasp the big picture → then dive deeper → then output content.&lt;/p&gt;
&lt;h2 id="my-suggestions-and-minor-regrets"&gt;My Suggestions and Minor Regrets&lt;/h2&gt;
&lt;p&gt;NotebookLM is already excellent, but I still have some strong expectations for future improvements:&lt;/p&gt;
&lt;h3 id="mindmap-export-formats-should-support-svg-or-text-based-markmap"&gt;MindMap Export Formats Should Support SVG or Text-Based (Markmap)&lt;/h3&gt;
&lt;p&gt;Currently, only PNG is supported, which gets blurry when enlarged. The table below lists my expectations for future features:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Expected Feature&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SVG Export&lt;/td&gt;
&lt;td&gt;For writing books, making slides, scalable without loss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Markmap Output&lt;/td&gt;
&lt;td&gt;Most friendly for Markdown writers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw JSON&lt;/td&gt;
&lt;td&gt;Allows custom rendering&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 9: Expected MindMap Export Formats
&lt;/figcaption&gt;
&lt;p&gt;I&amp;rsquo;m especially looking forward to NotebookLM supporting &lt;a href="https://markmap.js.org" target="_blank" rel="noopener"&gt;Markmap format&lt;/a&gt; export, which would be extremely friendly for users who write blogs and docs in Markdown.&lt;/p&gt;
&lt;p&gt;Recently, Google also launched &lt;a href="https://codewiki.google" target="_blank" rel="noopener"&gt;CodeWiki&lt;/a&gt;, similar to &lt;a href="https://deepwiki.com" target="_blank" rel="noopener"&gt;DeepWiki&lt;/a&gt;, which auto-generates image-rich Wikis for GitHub projects, but currently does not support Mermaid or Markmap.&lt;/p&gt;
&lt;h3 id="conversation-history-should-support-long-term-saving"&gt;Conversation History Should Support Long-Term Saving&lt;/h3&gt;
&lt;p&gt;Currently:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Chats are not persistently saved&lt;/li&gt;
&lt;li&gt;Only manually &amp;ldquo;add to notes&amp;rdquo; preserves results&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This causes some knowledge context to be lost. I hope to see a &amp;ldquo;Notebook conversation history&amp;rdquo; feature in the future.&lt;/p&gt;
&lt;h3 id="slide-generation-should-support-templates-for-content-creators"&gt;Slide Generation Should Support Templates for Content Creators&lt;/h3&gt;
&lt;p&gt;Currently, Video Overview offers various visual styles, but cannot:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Upload custom PPT templates&lt;/li&gt;
&lt;li&gt;Apply enterprise/personal branding templates&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If PPT template support is added, NotebookLM could become the &amp;ldquo;video generation hub&amp;rdquo; for content creators.&lt;/p&gt;
&lt;h3 id="deep-research-should-launch-soon-and-be-fully-open"&gt;Deep Research Should Launch Soon and Be Fully Open&lt;/h3&gt;
&lt;p&gt;I&amp;rsquo;m especially looking forward to this feature, as it could upgrade NotebookLM from a &amp;ldquo;knowledge organization tool&amp;rdquo; to a &amp;ldquo;research-grade tool.&amp;rdquo; I hope it will:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Reliably crawl more public web pages&lt;/li&gt;
&lt;li&gt;Ensure citation quality&lt;/li&gt;
&lt;li&gt;Integrate with existing Notebook materials&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a major upgrade I personally care about.&lt;/p&gt;
&lt;h3 id="mobile-experience-should-be-enhanced-beyond-content-playback"&gt;Mobile Experience Should Be Enhanced Beyond Content Playback&lt;/h3&gt;
&lt;p&gt;Currently, the mobile experience is minimal, only allowing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Listening to audio&lt;/li&gt;
&lt;li&gt;Viewing Notebook Guide summaries&lt;/li&gt;
&lt;li&gt;Simple Q&amp;amp;A&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I hope mobile will soon support:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Editing Notebooks&lt;/li&gt;
&lt;li&gt;Deep conversations&lt;/li&gt;
&lt;li&gt;MindMap interaction&lt;/li&gt;
&lt;li&gt;Content output (generating docs, outlines, etc.)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;NotebookLM is truly one of the AI tools I use every single day because it achieves a critical goal:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Organizing information, structuring knowledge, so I don&amp;rsquo;t have to start from scratch with massive documents.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Whether it&amp;rsquo;s:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Learning new technologies&lt;/li&gt;
&lt;li&gt;Reading long documents&lt;/li&gt;
&lt;li&gt;Creating courses&lt;/li&gt;
&lt;li&gt;Conducting training&lt;/li&gt;
&lt;li&gt;Writing books&lt;/li&gt;
&lt;li&gt;Drafting speeches&lt;/li&gt;
&lt;li&gt;Summarizing content&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It saves me a huge amount of time upfront, letting me focus on &amp;ldquo;understanding&amp;rdquo; and &amp;ldquo;creating.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ll continue to use NotebookLM as one of my essential tools and keep an eye on its progress in Deep Research, template systems, and mobile features.&lt;/p&gt;
&lt;p&gt;This is a tool truly designed for &amp;ldquo;knowledge workers&amp;rdquo; and deserves to be known by more people.&lt;/p&gt;</content:encoded></item><item><title>Helm v4: Paradigm Convergence and Plugin System Rebuild</title><link>https://jimmysong.io/blog/helm-4-delivery-and-plugin-rebuild/</link><pubDate>Fri, 14 Nov 2025 11:18:30 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/helm-4-delivery-and-plugin-rebuild/</guid><description>An analysis of Helm 4&amp;#39;s core changes, including Server-Side Apply, WASM plugin system, kstatus status model, reproducible builds, and content hash caching, with a timeline review of Helm&amp;#39;s history.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The release of Helm 4 is not just a technical upgrade, but a deep convergence of cloud-native delivery paradigms. The rebuilt plugin system and supply chain governance capabilities make Helm once again a driving force in the Kubernetes ecosystem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Since its first release in 2016, Helm has been one of the most important application distribution tools in the Kubernetes ecosystem. &lt;a href="https://github.com/helm/helm/releases/tag/v4.0.0" target="_blank" rel="noopener"&gt;Helm v4&lt;/a&gt; is not a &amp;ldquo;minor enhancement,&amp;rdquo; but a comprehensive update around &lt;strong&gt;delivery methods, extension mechanisms, and supply chain approaches&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This article reconstructs Helm&amp;rsquo;s historical context and focuses on why Helm 4 represents a paradigm-converging release.&lt;/p&gt;
&lt;h2 id="helm-from-tiller-to-declarative-delivery"&gt;Helm: From Tiller to Declarative Delivery&lt;/h2&gt;
&lt;p&gt;Below is a textual timeline showing key milestones from Helm v2 to v4, helping you understand its technical evolution:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;2016: Helm v2 released, using the Tiller architecture.&lt;/li&gt;
&lt;li&gt;2017: Chart Hub expands, major projects begin providing official Charts.&lt;/li&gt;
&lt;li&gt;2018: Security model controversies intensify, Tiller&amp;rsquo;s permission issues become apparent.&lt;/li&gt;
&lt;li&gt;2019: Helm v3 released, Tiller removed, OCI support introduced.&lt;/li&gt;
&lt;li&gt;2021: GitOps becomes widespread, Server-Side Apply (SSA) becomes the mainstream delivery semantic.&lt;/li&gt;
&lt;li&gt;2023: kstatus widely adopted for controller status assessment and health calculation.&lt;/li&gt;
&lt;li&gt;2025: Helm v4 released, bringing SSA, WASM plugins, reproducible builds, and content hash caching.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each major Helm release closely follows Kubernetes paradigms, driving progress in declarative delivery and ecosystem tooling.&lt;/p&gt;
&lt;h2 id="fundamental-changes-in-helm-v4"&gt;Fundamental Changes in Helm v4&lt;/h2&gt;
&lt;p&gt;This section analyzes the core technical upgrades and paradigm shifts in Helm v4.&lt;/p&gt;
&lt;h3 id="delivery-paradigm-update-default-server-side-apply-ssa-server-side-apply"&gt;Delivery Paradigm Update: Default Server-Side Apply (SSA, Server-Side Apply)&lt;/h3&gt;
&lt;p&gt;In Helm v3 and earlier, Helm used a &amp;ldquo;three-way merge&amp;rdquo; model for resource delivery. Helm v4 fully switches to &lt;strong&gt;Server-Side Apply (SSA, Server-Side Apply)&lt;/strong&gt;, meaning the API Server determines field ownership.&lt;/p&gt;
&lt;p&gt;This shift brings several direct results:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Full semantic alignment with &lt;code&gt;kubectl apply&lt;/code&gt; and GitOps controllers (such as Argo, Flux)&lt;/li&gt;
&lt;li&gt;When multiple controllers manage the same object, silent overrides are avoided and conflicts are explainable&lt;/li&gt;
&lt;li&gt;Helm&amp;rsquo;s behavior now follows Kubernetes&amp;rsquo; officially recommended declarative delivery paradigm&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following flowchart compares the delivery semantics of Helm v3 and v4.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/helm-4-delivery-and-plugin-rebuild/f34683c90a9f13678e2cde12ab355e2f.svg" data-img="https://assets.jimmysong.io/images/blog/helm-4-delivery-and-plugin-rebuild/f34683c90a9f13678e2cde12ab355e2f.svg" alt="Figure 3: Helm v3/v4 Delivery Semantics Comparison" data-caption="Figure 3: Helm v3/v4 Delivery Semantics Comparison"
width="2400"
height="377"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Helm v3/v4 Delivery Semantics Comparison&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Helm is now aligned with the delivery semantics of modern Kubernetes versions, improving predictability and safety in resource management.&lt;/p&gt;
&lt;h3 id="kstatus-driven-wait-behavior-and-readiness-annotations"&gt;kstatus-Driven Wait Behavior and Readiness Annotations&lt;/h3&gt;
&lt;p&gt;In Helm 3, &lt;code&gt;--wait&lt;/code&gt; could only make fuzzy status judgments on limited resources, lacking extensibility and explainability.&lt;/p&gt;
&lt;p&gt;Helm 4 introduces &lt;strong&gt;kstatus (Kubernetes Status)&lt;/strong&gt; as the basis for health status parsing, and supports two key annotations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;helm.sh/readiness-success&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;helm.sh/readiness-failure&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Chart authors can precisely define conditions for installation success or failure. Helm&amp;rsquo;s waiting model now offers &amp;ldquo;explainability + extensibility,&amp;rdquo; upgrading from a &amp;ldquo;templating tool&amp;rdquo; to a true &amp;ldquo;deployment orchestrator.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="extension-system-rebuild-wasm-plugin-system"&gt;Extension System Rebuild: WASM Plugin System&lt;/h3&gt;
&lt;p&gt;Helm 4 thoroughly reconstructs the plugin model, mainly including:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Typed and Structured Plugins&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Arbitrary scripts are no longer allowed; plugins must follow typed and structured standards&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;WebAssembly Plugin Runtime (Extism)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;More secure (sandbox isolation)&lt;/li&gt;
&lt;li&gt;Cross-language support&lt;/li&gt;
&lt;li&gt;Easy unified management in CI/CD and enterprise platforms&lt;/li&gt;
&lt;li&gt;Predictable and testable&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Post-renderer Integrated into Plugin System&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Moves beyond the &amp;ldquo;external executable black box&amp;rdquo; era&lt;/li&gt;
&lt;li&gt;Helm becomes a programmable platform, not just a template renderer&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="engineering-capabilities-upgrade-reproducible-builds-content-hash-caching-chart-api-v3"&gt;Engineering Capabilities Upgrade: Reproducible Builds, Content Hash Caching, chart API v3&lt;/h3&gt;
&lt;p&gt;Helm v4 brings the following engineering improvements:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Chart packaging is reproducible (supports signing, SBOM, SLSA, etc. for supply chain governance)&lt;/li&gt;
&lt;li&gt;Local cache uses content hashes, avoiding version-based conflicts&lt;/li&gt;
&lt;li&gt;chart API v3 (experimental) is stricter and more flexible&lt;/li&gt;
&lt;li&gt;SDK logging system upgraded to Go slog (modern logging)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These capabilities enable Helm charts to enter serious software supply chain governance.&lt;/p&gt;
&lt;h2 id="feature-comparison-helm-v3--v4"&gt;Feature Comparison (Helm v3 → v4)&lt;/h2&gt;
&lt;p&gt;The table below compares core features between Helm v3 and v4 for a quick understanding of the upgrade value.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Helm 3&lt;/th&gt;
&lt;th&gt;Helm 4&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Apply Model&lt;/td&gt;
&lt;td&gt;Three-way merge&lt;/td&gt;
&lt;td&gt;Default SSA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wait Behavior&lt;/td&gt;
&lt;td&gt;Fuzzy, not extensible&lt;/td&gt;
&lt;td&gt;kstatus + annotation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plugin System&lt;/td&gt;
&lt;td&gt;Script, uncontrollable&lt;/td&gt;
&lt;td&gt;WASM, typed plugins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Post-renderer&lt;/td&gt;
&lt;td&gt;External executable&lt;/td&gt;
&lt;td&gt;Plugin subsystem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build&lt;/td&gt;
&lt;td&gt;Not reproducible&lt;/td&gt;
&lt;td&gt;Reproducible build&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache&lt;/td&gt;
&lt;td&gt;name/version&lt;/td&gt;
&lt;td&gt;Content hash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chart API&lt;/td&gt;
&lt;td&gt;v2&lt;/td&gt;
&lt;td&gt;v2 + v3 (experimental)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SDK Logs&lt;/td&gt;
&lt;td&gt;stdlib log&lt;/td&gt;
&lt;td&gt;slog&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Helm v3 vs v4 Feature Comparison
&lt;/figcaption&gt;
&lt;p&gt;This is a release that &amp;ldquo;repays technical debt in bulk + aligns with contemporary Kubernetes semantics.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="why-is-helm-v4-a-paradigm-convergence-event"&gt;Why Is Helm v4 a Paradigm Convergence Event?&lt;/h2&gt;
&lt;p&gt;The release of Helm v4 is not just a feature upgrade, but a deep convergence of delivery paradigms, mainly in three aspects:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes Delivery Semantics Unified to SSA&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Previously: kubectl, GitOps, and Helm each had their own logic.
Now: All unified to SSA, consistent delivery behavior, smoother ecosystem collaboration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plugin System Enters the Platform Era&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;WASM (WebAssembly) brings a secure, universal, and controllable plugin runtime. Infrastructure projects widely adopt WASM: Envoy → WASM Filters, Kubernetes → WASM CRI/OCI, and now Helm joins the platform camp.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Charts Enter Supply Chain Governance&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Reproducible builds and digest verification allow Helm charts to be managed as seriously as container images, greatly enhancing supply chain security.&lt;/p&gt;
&lt;p&gt;The entire ecosystem moves to a unified capability baseline, driving cloud-native delivery standardization.&lt;/p&gt;
&lt;h2 id="my-helm-history-and-observations"&gt;My Helm History and Observations&lt;/h2&gt;
&lt;p&gt;As an early user from the Helm v2 era, I have experienced the following stages:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Tiller security controversies&lt;/li&gt;
&lt;li&gt;v3 migration (state stored in secrets)&lt;/li&gt;
&lt;li&gt;Large-scale chart consolidation in the community&lt;/li&gt;
&lt;li&gt;OCI adoption&lt;/li&gt;
&lt;li&gt;Today&amp;rsquo;s SSA / WASM / reproducible build&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each major Helm version upgrade is not about chasing trends, but proactively aligning with Kubernetes paradigms:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;v3 aligns with K8s &amp;ldquo;no cluster-side runtime&amp;rdquo; principle&lt;/li&gt;
&lt;li&gt;v4 aligns with SSA, kstatus, WASM, OCI, and other advances from the past five years&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Helm exemplifies the evolution rhythm of infrastructure projects: &lt;strong&gt;not by piling on features, but by evolving in semantic alignment with the platform.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;The release of Helm v4 marks a new paradigm for Kubernetes application delivery. SSA, WASM plugins, kstatus, and reproducible builds make Helm not just a templating tool, but a core for supply chain governance and platform extensibility. For cloud-native developers and platform teams, Helm v4 is a paradigm upgrade worth attention.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/helm/helm/releases/tag/v4.0.0" target="_blank" rel="noopener"&gt;Helm v4.0.0 Release - github.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://helm.sh/docs/overview/" target="_blank" rel="noopener"&gt;Helm Documentation Overview - helm.sh&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://artifacthub.io/" target="_blank" rel="noopener"&gt;ArtifactHub Charts Index - artifacthub.io&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Kimi K2 Thinking: The True Awakening of China's Thinking Model</title><link>https://jimmysong.io/blog/kimi-k2-thinking-cn-awakening/</link><pubDate>Fri, 14 Nov 2025 08:25:26 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/kimi-k2-thinking-cn-awakening/</guid><description>Kimi K2 Thinking&amp;#39;s open source marks China&amp;#39;s entry into thinking models. This article reviews its technical approach and compares it with Claude and Gemini.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;China&amp;rsquo;s large language models have finally moved from &amp;ldquo;writing like humans&amp;rdquo; to &amp;ldquo;thinking like humans.&amp;rdquo; The open-sourcing of Kimi K2 is a watershed moment for China&amp;rsquo;s AI trajectory.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The narrative around China&amp;rsquo;s large language models is shifting from &amp;ldquo;Chat-style models&amp;rdquo; to &amp;ldquo;Thinking models (Thinking Model, Thinking Model).&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Moonshot AI&amp;rsquo;s open-sourcing of &lt;strong&gt;Kimi K2 Thinking&lt;/strong&gt; marks the first real landing of this transition. K2 is not just another iteration like ChatGLM or Qwen; it&amp;rsquo;s the first time a Chinese team has unified &amp;ldquo;deep reasoning + long context + tool invocation continuity&amp;rdquo; in training. This is the core of the thinking model approach and the reason why models like Claude and Gemini have led the field.&lt;/p&gt;
&lt;h2 id="the-significance-of-k2s-open-source-china-enters-the-era-of-thinking-models"&gt;The Significance of K2&amp;rsquo;s Open Source: China Enters the Era of Thinking Models&lt;/h2&gt;
&lt;p&gt;Why is K2&amp;rsquo;s open source a turning point? Because it enables Chinese models to achieve the following capabilities for the first time:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Stable execution of 200–300 tool invocations (toolchain reasoning stability)&lt;/li&gt;
&lt;li&gt;Deep, multi-stage reasoning chain execution (CoT Consistency, Chain-of-Thought Consistency)&lt;/li&gt;
&lt;li&gt;256k context as a &amp;ldquo;working memory&amp;rdquo; (Working Memory, Working Memory)&lt;/li&gt;
&lt;li&gt;Native INT4 acceleration + MoE activation sparsity scheduling&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is a completely different path from &amp;ldquo;stacking parameters → stacking benchmarks,&amp;rdquo; emphasizing reasoning ability over parameter scale.&lt;/p&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;K2 is the first time a Chinese model has entered the sequence of thinking models (Thinking Model, Thinking Model).&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="dissecting-k2s-technical-approach"&gt;Dissecting K2&amp;rsquo;s Technical Approach&lt;/h2&gt;
&lt;p&gt;K2&amp;rsquo;s technical approach can be broken down into five key points, each directly impacting the model&amp;rsquo;s reasoning ability and ecosystem adaptability.&lt;/p&gt;
&lt;h3 id="moe-expert-division-cognitive-division-rather-than-parameter-expansion"&gt;MoE Expert Division: Cognitive Division Rather Than Parameter Expansion&lt;/h3&gt;
&lt;p&gt;K2&amp;rsquo;s MoE (Mixture of Experts, Mixture of Experts) design philosophy is distinct from previous models. The core is not about activating fewer parameters or running larger models more cheaply, but about assigning different cognitive sub-skills to different experts. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mathematical reasoning expert&lt;/li&gt;
&lt;li&gt;Planning expert&lt;/li&gt;
&lt;li&gt;Tool invocation expert&lt;/li&gt;
&lt;li&gt;Browser task expert&lt;/li&gt;
&lt;li&gt;Code generation expert&lt;/li&gt;
&lt;li&gt;Long-chain retention expert&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This division aligns directly with Claude 3.5&amp;rsquo;s cognitive layering (Cognitive Layering, Cognitive Layering) approach. K2&amp;rsquo;s MoE is about &amp;ldquo;dividing thinking among the model,&amp;rdquo; not just &amp;ldquo;making computation cheaper.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="256k-context-building-the-models-working-memory"&gt;256K Context: Building the Model&amp;rsquo;s Working Memory&lt;/h3&gt;
&lt;p&gt;K2&amp;rsquo;s ultra-long context is not just a parameter showcase; it&amp;rsquo;s designed to build the model&amp;rsquo;s &amp;ldquo;thinking buffer.&amp;rdquo; It allows the entire process to retain reasoning chains, tool invocation states, multi-stage reflection, and uninterrupted long tasks (such as research or code refactoring), stably executing multi-stage agent workflows. Long-term thinking requires long-term memory support, and K2&amp;rsquo;s long context is the &amp;ldquo;memory&amp;rdquo; for sustained reasoning chains.&lt;/p&gt;
&lt;h3 id="intertwined-training-of-tool-invocation-and-reasoning-chains"&gt;Intertwined Training of Tool Invocation and Reasoning Chains&lt;/h3&gt;
&lt;p&gt;K2 excels in the intertwined training of tool invocation and reasoning chains. Traditional open-source models typically follow this process:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Generate reasoning&lt;/li&gt;
&lt;li&gt;Output JSON function call&lt;/li&gt;
&lt;li&gt;Tool returns result&lt;/li&gt;
&lt;li&gt;Continue reasoning&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In this approach, the reasoning chain and invocation chain are separated. K2&amp;rsquo;s training allows the reasoning chain to invoke tools at any time and feed tool results back into the reasoning chain for the next stage of thinking. It supports 200–300 consecutive tool invocations without interruption, fully aligning with Claude 3.5&amp;rsquo;s Interleaved CoT + Tool Use.&lt;/p&gt;
&lt;h3 id="native-int4-quantization-ensuring-reasoning-chain-stability"&gt;Native INT4 Quantization: Ensuring Reasoning Chain Stability&lt;/h3&gt;
&lt;p&gt;K2&amp;rsquo;s INT4 (INT4, 4-bit Integer Quantization) approach is not ordinary post-quantization. Its purpose is not only to reduce memory usage and increase throughput, but more importantly, to ensure that deep reasoning chains do not break due to insufficient computing power. The biggest killer of deep thinking chains is timeout, freezing, or unstable workers. INT4 enables Chinese GPUs (non-H100) to run complete reasoning chains, which is highly significant for China&amp;rsquo;s ecosystem.&lt;/p&gt;
&lt;h3 id="moe--long-context--toolchain-unified-training-rather-than-module-stitching"&gt;MoE + Long Context + Toolchain: Unified Training Rather Than Module Stitching&lt;/h3&gt;
&lt;p&gt;K2&amp;rsquo;s most important feature is its holistic training approach: expert division, long context-driven consistency, tool invocation trained through real execution, browser tasks and long-step task reinforcement, and INT4 entering the training loop. It&amp;rsquo;s not a &amp;ldquo;ChatLLM + Memory + RAG + Tools&amp;rdquo; patchwork, but an integrated reasoning system.&lt;/p&gt;
&lt;h2 id="alignment-and-differences-between-k2-and-international-mainstream-approaches"&gt;Alignment and Differences Between K2 and International Mainstream Approaches&lt;/h2&gt;
&lt;p&gt;K2 is highly aligned with international mainstream models (such as Claude, Gemini, OpenAI) in cognitive reasoning, ultra-long context, and tool invocation mechanisms, but also has unique advantages for Chinese models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Native INT4 + adaptation to Chinese computing power is rare globally&lt;/li&gt;
&lt;li&gt;Toolchain continuity is more stable than most open-source models&lt;/li&gt;
&lt;li&gt;Higher degree of open source, stronger ecosystem reusability&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="collaborative-value-of-chinas-ai-infra-k2--rlinf--mem-alpha"&gt;Collaborative Value of China&amp;rsquo;s AI Infra: K2 × RLinf × Mem-alpha&lt;/h2&gt;
&lt;p&gt;A series of important open-source infrastructures have emerged in the K2 ecosystem. The table below summarizes these project types and their value to K2:&lt;/p&gt;
&lt;p&gt;Here is a comparison table of the collaborative value of each infrastructure with K2:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Project&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Value to K2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RLinf&lt;/td&gt;
&lt;td&gt;Reinforcement Learning&lt;/td&gt;
&lt;td&gt;Used to train stronger planning/browser task capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mem-alpha&lt;/td&gt;
&lt;td&gt;Memory Enhancement&lt;/td&gt;
&lt;td&gt;Can be combined with K2 to form long-term memory agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AgentDebug&lt;/td&gt;
&lt;td&gt;Agent Error Debugging&lt;/td&gt;
&lt;td&gt;Used to analyze K2&amp;rsquo;s toolchain errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UI-Genie&lt;/td&gt;
&lt;td&gt;GUI Agent Training&lt;/td&gt;
&lt;td&gt;Can serve as an experimental field for K2&amp;rsquo;s agent capability expansion&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: Collaborative Value of China&amp;rsquo;s AI Infra Ecosystem
&lt;/figcaption&gt;
&lt;p&gt;This combination is already forming a China AI Agent Infra Stack.&lt;/p&gt;
&lt;h2 id="personal-view-the-significance-of-k2s-approach"&gt;Personal View: The Significance of K2&amp;rsquo;s Approach&lt;/h2&gt;
&lt;p&gt;I believe the significance of K2 lies not in the model itself, but in its technical approach:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;K2 marks the first time Chinese models have shifted from &amp;ldquo;language generation competition&amp;rdquo; to &amp;ldquo;thinking ability competition.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For the past three years, the main line of China&amp;rsquo;s open-source models has been evaluation scores, parameter scale, instruction following, and alignment data. But K2 is the first to clearly take the path of deep reasoning, tool intertwining, cognitive division, long-term task chains, and native performance optimization. This means China&amp;rsquo;s model trajectory is now synchronized with the US, rather than chasing old paths.&lt;/p&gt;
&lt;h2 id="key-directions-to-watch-in-k2s-ecosystem-over-the-next-year"&gt;Key Directions to Watch in K2&amp;rsquo;s Ecosystem Over the Next Year&lt;/h2&gt;
&lt;p&gt;K2&amp;rsquo;s future ecosystem influence will depend on several key points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Whether it opens the tool registry (Tool Registry, Tool Registry)&lt;/li&gt;
&lt;li&gt;Whether it supports dynamic memory (Mem-alpha integration)&lt;/li&gt;
&lt;li&gt;Whether it opens the MoE expert structure&lt;/li&gt;
&lt;li&gt;Whether it can form a Chinese reasoning chain optimization path with vLLM / llm-d / KServe&lt;/li&gt;
&lt;li&gt;Whether it supports fault tolerance for multi-node continuous reasoning chains&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These capabilities will determine K2&amp;rsquo;s ecosystem influence and technical extensibility.&lt;/p&gt;
&lt;h2 id="k2-thinking-model-architecture-diagram"&gt;K2 Thinking Model Architecture Diagram&lt;/h2&gt;
&lt;p&gt;The following flowchart illustrates the core architecture of the K2 thinking model and its collaboration with external agents/applications:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/kimi-k2-thinking-cn-awakening/8883c2cf12acbe9362d56d664577b67c.svg" data-img="https://assets.jimmysong.io/images/blog/kimi-k2-thinking-cn-awakening/8883c2cf12acbe9362d56d664577b67c.svg" alt="Figure 3: K2 Thinking Model Architecture" data-caption="Figure 3: K2 Thinking Model Architecture"
width="1600"
height="1158"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: K2 Thinking Model Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;K2 is the first time China&amp;rsquo;s model trajectory is heading in the right direction:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;From &amp;ldquo;writing like humans&amp;rdquo; to &amp;ldquo;thinking like humans.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The era of thinking models is coming, and Chinese models are finally standing on the same roadmap as the international forefront.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://moonshotai.github.io/Kimi-K2/thinking.html" target="_blank" rel="noopener"&gt;Introducing Kimi K2 Thinking - moonshot.github.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/moonshotai/Kimi-K2-Thinking" target="_blank" rel="noopener"&gt;moonshotai/Kimi-K2-Thinking - huggingface.co&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Beyond Gateway: Inference Traffic Control Practice with Gateway API Inference Extension</title><link>https://jimmysong.io/blog/gateway-api-inference-extension-inference-traffic-control/</link><pubDate>Fri, 14 Nov 2025 08:19:16 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/gateway-api-inference-extension-inference-traffic-control/</guid><description>Exploring how Gateway API Inference Extension brings model-aware inference traffic control through InferencePool, InferenceObjective, and metrics-driven routing.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;AI inference traffic governance is undergoing a paradigm shift. The Gateway API Inference Extension makes &amp;ldquo;model awareness&amp;rdquo; the new 主线 of traffic control.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="brief-review-of-gateway-apis-current-state"&gt;Brief Review of Gateway API&amp;rsquo;s Current State&lt;/h2&gt;
&lt;p&gt;Kubernetes Gateway API has entered the stable v1 series and continued iterating after the 1.0 GA release, enhancing advanced traffic governance capabilities such as WebSocket, timeout and retry, Service Mesh integration, GRPCRoute, request mirroring, CORS, Retry Budget, and more. Major cloud providers and gateway implementations (such as Alibaba Cloud ACK, GKE Gateway, Envoy Gateway, NGINX Gateway Fabric) have all adopted Gateway API as the new generation north-south traffic model.&lt;/p&gt;
&lt;p&gt;Building on this foundation, the community has proposed an extension specification specifically for AI inference traffic - the &lt;strong&gt;Gateway API Inference Extension&lt;/strong&gt;. This specification is not about &amp;ldquo;reinventing an API&amp;rdquo; but rather supplementing the Gateway API core model with &lt;strong&gt;model-aware load balancing and traffic control capabilities&lt;/strong&gt; for Large Language Model (LLM) and other inference scenarios.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Early documentation often mentioned the InferenceModel + InferencePool CRD combination; the latest specification has evolved to InferenceObjective + InferencePool (with optional InferencePoolImport). This article consistently uses the latest terminology.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="typical-challenges-in-genai-inference-traffic"&gt;Typical Challenges in GenAI Inference Traffic&lt;/h2&gt;
&lt;p&gt;Traditional gateway, Ingress, and Service Mesh load balancing models are essentially &amp;ldquo;request-agnostic + endpoint-agnostic&amp;rdquo;: they distribute traffic across a group of static backends through algorithms like round-robin, least requests, and hashing.&lt;/p&gt;
&lt;p&gt;In GPU-powered Large Language Model (LLM) inference scenarios, this model exposes obvious problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Invisible GPU utilization and queuing&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A single LLM instance simultaneously maintains KV Cache, LoRA Adapter, and Token queues. Resource consumption varies significantly across the same batch of requests. Load balancing based solely on QPS or connection count can easily lead to extreme situations where &amp;ldquo;idle GPUs have no work while busy GPUs crash.&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Lack of semantic binding between models and requests&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;From the business perspective, one typically only sees a POST /v1/chat/completions endpoint, making it difficult to express intentions like &amp;ldquo;high-priority model / test version / canary weight&amp;rdquo; at the routing layer.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Difficult unified management of multiple models, versions, and LoRAs&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Each model service implements its own routing and A/B testing solutions, making it difficult for the platform to achieve unified governance and observability at the control plane level.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The goal of Gateway API Inference Extension is to introduce AI-specific semantics and metrics into load balancing and traffic control decisions while maintaining the existing Gateway model.&lt;/p&gt;
&lt;h2 id="core-concepts-and-resource-model-of-gateway-api-inference-extension"&gt;Core Concepts and Resource Model of Gateway API Inference Extension&lt;/h2&gt;
&lt;p&gt;This section introduces the overall architecture and key resource objects of the Inference Extension.&lt;/p&gt;
&lt;h3 id="overall-architecture"&gt;Overall Architecture&lt;/h3&gt;
&lt;p&gt;The Inference Extension uses Envoy External Processing (ext-proc) mechanism to upgrade Gateway API + ext-proc capable gateways (such as Envoy Gateway, kgateway, GKE Gateway) into &lt;strong&gt;Inference Gateways&lt;/strong&gt;. Requests still follow the standard Gateway + HTTPRoute path, but before being forwarded to the backend, they pass through an &amp;ldquo;Endpoint Picker&amp;rdquo; component that selects the most suitable backend instance based on real-time metrics exposed by the model server.&lt;/p&gt;
&lt;p&gt;The flowchart below shows the overall architecture:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/gateway-api-inference-extension-inference-traffic-control/777d6816e4df5acaf96f5e28b01bc441.svg" data-img="https://assets.jimmysong.io/images/blog/gateway-api-inference-extension-inference-traffic-control/777d6816e4df5acaf96f5e28b01bc441.svg" alt="Figure 3: Inference Extension Architecture Flow" data-caption="Figure 3: Inference Extension Architecture Flow"
width="1980"
height="140"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Inference Extension Architecture Flow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="inferencepool-platform-side-model-service-pool"&gt;InferencePool: Platform-Side &amp;ldquo;Model Service Pool&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;InferencePool is the core resource introduced by the Inference Extension, used to describe a group of inference Pods and routing plugin configurations. It is similar to a Service with a selector, responsible for selecting a group of model service Pods and specifying exposed ports, while also allowing attachment of Endpoint Picker plugins (such as Prefix-Cache-Aware, LoRA-aware, etc.).&lt;/p&gt;
&lt;p&gt;In the Gateway API model, InferencePool is treated as a type of &amp;ldquo;Backend&amp;rdquo; that can be referenced by HTTPRoute.backendRefs.&lt;/p&gt;
&lt;p&gt;The code block below shows a simplified example of InferencePool:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;apiVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;inference.networking.x-k8s.io/v1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;InferencePool&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;vllm-llama3-chat-pool&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;targetPortNumber&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;8000&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;selector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;app&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;vllm-llama3-chat&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;extensionRef&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;prefix-cache-aware-endpoint-picker&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The meaning of the above configuration is as follows:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Select all Pods with app=vllm-llama3-chat, port 8000.&lt;/li&gt;
&lt;li&gt;Use the plugin named prefix-cache-aware-endpoint-picker to make routing decisions based on metrics such as KV Cache hit rate, queue length, and GPU utilization.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="inferenceobjective-business-side-request-objective"&gt;InferenceObjective: Business-Side &amp;ldquo;Request Objective&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;InferenceObjective is used to express the goal and priority of a single request, decoupled from the model service pool. One request corresponds to one InferenceObjective, and the same InferencePool can serve multiple different InferenceObjectives.&lt;/p&gt;
&lt;p&gt;Typical fields include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Business criticality (Critical / High / BestEffort)&lt;/li&gt;
&lt;li&gt;Required model family / version preference&lt;/li&gt;
&lt;li&gt;Acceptable latency / cost upper limits, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Endpoint Picker can combine InferenceObjective with backend metrics to make decisions: prioritize Critical requests when resources are constrained, and shed BestEffort requests when necessary.&lt;/p&gt;
&lt;h3 id="inferencepoolimport-cross-cluster--gateway-reuse"&gt;InferencePoolImport: Cross-Cluster / Gateway Reuse&lt;/h3&gt;
&lt;p&gt;InferencePoolImport supports importing InferencePools defined in remote clusters into the local cluster, facilitating consistent governance of multi-cluster, multi-region inference services.&lt;/p&gt;
&lt;h3 id="maturity-and-implementation-ecosystem-of-inference-extension"&gt;Maturity and Implementation Ecosystem of Inference Extension&lt;/h3&gt;
&lt;p&gt;The current project version is v1.1.x, overall in the &lt;strong&gt;Alpha&lt;/strong&gt; stage. Official recommendation is not to use it directly in production yet; it&amp;rsquo;s more suitable for platform teams to experiment with the technology stack.&lt;/p&gt;
&lt;p&gt;Multiple implementations and integrations already exist:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Inference Gateway implementations for Envoy Gateway / kgateway&lt;/li&gt;
&lt;li&gt;GKE Inference Gateway: Enhanced capabilities based on GKE Gateway, including KV Cache-aware routing, LoRA reuse, priority scheduling, etc.&lt;/li&gt;
&lt;li&gt;NGINX Gateway Fabric, cloud provider ACK, and others are also following up on related extensions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Therefore, when designing practical solutions, it should be treated as a &amp;ldquo;forward-looking solution / future main path.&amp;rdquo; For production environments, prioritize managed implementations (such as GKE Inference Gateway) or commercial products from gateway vendors.&lt;/p&gt;
&lt;h2 id="practice-using-inference-extension-for-inference-traffic-control"&gt;Practice: Using Inference Extension for Inference Traffic Control&lt;/h2&gt;
&lt;p&gt;The following example uses a self-hosted LLM cluster on Kubernetes to demonstrate how to serve external traffic through an OpenAI-compatible interface and achieve:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Scheduling by request priority (real-time conversation vs. batch processing)&lt;/li&gt;
&lt;li&gt;Multi-version model canary and rollback&lt;/li&gt;
&lt;li&gt;Optimized routing using GPU metrics and KV Cache hit rates&lt;/li&gt;
&lt;li&gt;Unified platform-side observability and rate limiting entry point&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="deploying-gateway-api-and-inference-extension"&gt;Deploying Gateway API and Inference Extension&lt;/h3&gt;
&lt;p&gt;The deployment process is as follows:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Install a Gateway API-compatible gateway implementation (such as Envoy Gateway, kgateway, GKE Gateway) in the cluster.&lt;/li&gt;
&lt;li&gt;Install the Gateway API Inference Extension CRD and control plane components.&lt;/li&gt;
&lt;li&gt;Enable the metrics endpoints and plugin protocols required by Inference Extension on the model server side (such as vLLM, Triton, TGI).&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="defining-inferencepool-abstracting-llm-pods-as-an-inference-pool"&gt;Defining InferencePool: Abstracting LLM Pods as an &amp;ldquo;Inference Pool&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;The code block below shows a typical InferencePool configuration:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;apiVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;inference.networking.x-k8s.io/v1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;InferencePool&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;chat-llama3-pool&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;targetPortNumber&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;8000&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;selector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;app&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;chat-llama3&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;extensionRef&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;prefix-cache-aware&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="c"&gt;# Optional: Plugin configuration ConfigMap / CR&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Key points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;selector selects LLM Pods, targetPortNumber specifies the inference service port.&lt;/li&gt;
&lt;li&gt;extensionRef binds the Endpoint Picker plugin, implementing KV Cache prefix-aware routing, selecting replicas with lighter load based on metrics like queue_length/gpu_utilization, and triggering load shedding when necessary.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="defining-inferenceobjective-connecting-business-intent-to-routing-decisions"&gt;Defining InferenceObjective: Connecting &amp;ldquo;Business Intent&amp;rdquo; to Routing Decisions&lt;/h3&gt;
&lt;p&gt;The code block below shows a sample InferenceObjective configuration (fields can be adjusted according to the actual version and implementation):&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;apiVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;inference.networking.x-k8s.io/v1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;InferenceObjective&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;chat-critical&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;criticality&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;Critical &lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="c"&gt;# Real-time conversation&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;preferredModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;llama3-70b&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;fallbackModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;llama3-8b&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nn"&gt;---&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;apiVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;inference.networking.x-k8s.io/v1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;InferenceObjective&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;chat-batch&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;criticality&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;BestEffort &lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="c"&gt;# Batch analysis, log summarization&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;preferredModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;llama3-8b&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The Endpoint Picker can combine InferenceObjective with InferencePool metrics to make the following decisions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When GPU is constrained and queues are too long, prioritize chat-critical requests and discard chat-batch if necessary.&lt;/li&gt;
&lt;li&gt;For Critical requests, prioritize the large model; if the target pool is unavailable, fall back to the small model pool.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="importing-business-traffic-to-inferencepool-via-httproute"&gt;Importing Business Traffic to InferencePool via HTTPRoute&lt;/h3&gt;
&lt;p&gt;The code block below shows the HTTPRoute configuration on the Gateway API side:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;apiVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;gateway.networking.k8s.io/v1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;HTTPRoute&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;llm-route&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;parentRefs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;public-gateway&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;matches&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;PathPrefix&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;/v1/chat/completions&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;backendRefs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;- &lt;span class="nt"&gt;group&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;inference.networking.x-k8s.io&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;InferencePool&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;chat-llama3-pool&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Some implementations (such as GKE Inference Gateway) can route based on the model field in the request body, mapping OpenAI-style model names to different InferencePool / InferenceObjective combinations.&lt;/p&gt;
&lt;h3 id="fine-grained-inference-traffic-control-using-metrics"&gt;Fine-Grained Inference Traffic Control Using Metrics&lt;/h3&gt;
&lt;p&gt;Inference Extension provides platform teams with a unified metrics system, including kv_cache_hits, gpu_utilization, request_queue_length, per-request inference duration, token count, and more.&lt;/p&gt;
&lt;p&gt;Based on these metrics, multi-level traffic control strategies can be built:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Priority + Capacity&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Set priority and capacity upper limits for different InferenceObjectives, automatically guaranteeing critical business when resources are constrained.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Rate Limiting by Cost / Token&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Aggregate token / latency metrics exposed by Inference Extension to Prometheus, then add cost-based rate limiting logic at the gateway / API Gateway level (such as total tokens per minute per user / application). The specification itself doesn&amp;rsquo;t mandate &amp;ldquo;Token-level rate limiting,&amp;rdquo; but provides observability and hooks for easy policy extension.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Prefix Cache Aware Routing&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For requests with shared context (such as RAG, template generation), enable the Prefix Cache Aware plugin to route requests with the same prefix to the same replica, maximizing KV Cache hit rates and significantly reducing TTFT.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Auto-Scaling Integration&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Use metrics output by Inference Extension as HPA input to achieve &amp;ldquo;model-aware&amp;rdquo; auto-scaling, rather than relying solely on CPU / memory.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="relationship-and-trade-offs-with-traditional-gateway--service-mesh"&gt;Relationship and Trade-offs with Traditional Gateway / Service Mesh&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Control Plane&lt;/strong&gt;: Continue using Gateway API as the unified north-south / east-west traffic modeling specification. Service Mesh can still perform fine-grained circuit breaking, retry, mTLS, etc. within the cluster.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Plane&lt;/strong&gt;: Inference Extension pushes &amp;ldquo;model-aware routing&amp;rdquo; down to the ext-proc path implemented by the Gateway, avoiding redundant business-side wheel reinvention.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adoption Strategy&lt;/strong&gt;: In the current Alpha stage, prioritize productionized implementations (such as GKE Inference Gateway, commercial gateway vendor Inference Gateways), start with &amp;ldquo;bypass pilots&amp;rdquo; outside the critical path, and gradually migrate existing AI gateway routing rules to the Inference Extension model.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Combining official documentation and community implementations, we can see:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Gateway API has become the Kubernetes standard north-south traffic model and continues to enhance in the 1.x series.&lt;/li&gt;
&lt;li&gt;Gateway API Inference Extension introduces GPU metrics, KV Cache, LoRA, and other inference semantics into load balancing decisions through InferencePool, InferenceObjective, and Endpoint Picker.&lt;/li&gt;
&lt;li&gt;The project is still in Alpha stage, but has achieved experimental or productionized adoption in implementations such as GKE, kgateway, and NGINX Gateway Fabric. It is one of the important future directions for inference traffic control.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Previous descriptions of Inference Extension as &amp;ldquo;built-in Token rate limiting CRD, AIInferencePolicy, and other objects&amp;rdquo; are no longer accurate and should all be replaced with the design based on InferencePool / InferenceObjective + metrics-driven approach.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://kubernetes.io/blog/2023/10/31/gateway-api-ga/" target="_blank" rel="noopener"&gt;Gateway API v1.0: GA Release - kubernetes.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kubernetes.io/blog/2024/11/21/gateway-api-v1-2/" target="_blank" rel="noopener"&gt;Gateway API v1.2: WebSockets, Timeouts, Retries, and More - kubernetes.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gateway-api-inference-extension.sigs.k8s.io/" target="_blank" rel="noopener"&gt;Kubernetes Gateway API Inference Extension – Overview - gateway-api-inference-extension.sigs.k8s.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gateway-api-inference-extension.sigs.k8s.io/concepts/api-overview/" target="_blank" rel="noopener"&gt;API Overview – InferencePool / InferenceObjective / InferencePoolImport - gateway-api-inference-extension.sigs.k8s.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gateway-api-inference-extension.sigs.k8s.io/guides/metrics-and-observability/" target="_blank" rel="noopener"&gt;Metrics and Observability - gateway-api-inference-extension.sigs.k8s.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cncf.io/blog/2025/04/21/deep-dive-into-the-gateway-api-inference-extension/" target="_blank" rel="noopener"&gt;Deep Dive into the Gateway API Inference Extension – cncf.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kubernetes.io/blog/2025/xx/xx/gateway-api-inference-extension/" target="_blank" rel="noopener"&gt;Introducing Gateway API Inference Extension – kubernetes.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/kubernetes-engine/docs/concepts/about-gke-inference-gateway" target="_blank" rel="noopener"&gt;GKE Inference Gateway - cloud.google.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.nginx.com/nginx-gateway-fabric/" target="_blank" rel="noopener"&gt;NGINX Gateway Fabric – Kubernetes Gateway API and AI Inference - docs.nginx.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.alibabacloud.com/help/en/ack/product-overview/gateway-api" target="_blank" rel="noopener"&gt;Alibaba Cloud ACK – Gateway API Components and Versions - alibabacloud.com&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>TRAE SOLO vs VS Code: Rethinking Coding Tools from the Perspective of AI Engineering Entities</title><link>https://jimmysong.io/blog/trae-vs-vscode-insiders-agent-hq-and-ai-engineering-entity/</link><pubDate>Fri, 14 Nov 2025 07:14:39 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/trae-vs-vscode-insiders-agent-hq-and-ai-engineering-entity/</guid><description>A comparison of TRAE SOLO and VS Code (Copilot, Agent HQ) via the AI Engineering Entity framework, focusing on automation, collaboration, model transparency, and engineering roles.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Coding tools are evolving from &amp;ldquo;AI assistants&amp;rdquo; into true engineering entities. How can we reinterpret the roles of TRAE SOLO and VS Code from a pipeline perspective?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Recently, &lt;a href="https://www.trae.ai" target="_blank" rel="noopener"&gt;TRAE International Edition&lt;/a&gt; SOLO mode has been fully opened to overseas users. It claims to be a &amp;ldquo;responsive coding agent&amp;rdquo; and is now available for official trial, with token-based rate limiting.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve used early versions of TRAE (without SOLO access), and also tried Qoder and Kiro. The AI coding field is flourishing, each tool with its own strengths.&lt;/p&gt;
&lt;p&gt;This article compares TRAE SOLO and VS Code (with Copilot, Plan/Agent mode, and Agent HQ) from the perspective of AI engineering entities, combining personal experience to outline their differences in engineering automation, collaboration, and governance.&lt;/p&gt;
&lt;h2 id="three-engineering-role-abstractions-end-to-end-executor-contextual-collaborator-and-expert-orchestrator"&gt;Three Engineering Role Abstractions: End-to-End Executor, Contextual Collaborator, and Expert Orchestrator&lt;/h2&gt;
&lt;p&gt;From an engineering perspective, current mainstream AI coding tools can be abstracted into three roles:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;End-to-End Executor&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focuses on &amp;ldquo;requirement to deployment&amp;rdquo; workflows, capable of autonomous planning, task breakdown, coding, testing, previewing, and even deployment. Officially called &amp;ldquo;AI-Powered Context Engineer.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;User experience: like a &amp;ldquo;full-chain executor&amp;rdquo;—give it a requirement, and it handles the project, even if it&amp;rsquo;s slow or imperfect.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Contextual Collaborator&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;VS Code is a powerful editor. Copilot has evolved from line-level completion to Chat, Plan agent, Agent mode, supporting multi-step tasks and codebase analysis.&lt;/li&gt;
&lt;li&gt;It doesn&amp;rsquo;t take over the whole project, but efficiently handles local tasks under your guidance, acting as an automated unit for specific segments.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Expert Orchestrator / Specialist Engine&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GitHub&amp;rsquo;s Agent HQ is a &amp;ldquo;central platform for AI coding agents,&amp;rdquo; a unified control plane that can connect to OpenAI, Anthropic, Google, xAI, etc., run agents in parallel, and compare results.&lt;/li&gt;
&lt;li&gt;Functions as an &amp;ldquo;expert orchestrator&amp;rdquo; for key steps—planning, review, refactoring, or decision-making—providing high-quality output without taking over the entire project.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These three roles correspond to the structure in &amp;ldquo;AI Engineering Entity (AIEE)&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Single end-to-end executor (TRAE SOLO)&lt;/li&gt;
&lt;li&gt;Contextual collaborator residing in the IDE (VS Code + Copilot)&lt;/li&gt;
&lt;li&gt;Specialist orchestrator platform for multi-entity scheduling (Agent HQ)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="quick-product-status-check"&gt;Quick Product Status Check&lt;/h2&gt;
&lt;p&gt;To avoid memory bias, let&amp;rsquo;s clarify some key facts.&lt;/p&gt;
&lt;h3 id="trae--trae-solo"&gt;TRAE / TRAE SOLO&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;TRAE claims to be a &amp;ldquo;10x AI Engineer,&amp;rdquo; able to independently understand requirements and execute development tasks.&lt;/li&gt;
&lt;li&gt;SOLO mode is GA for international users, emphasizing full-chain automation, available directly but with token limits.&lt;/li&gt;
&lt;li&gt;Underlying open-source Trae Agent CLI can execute multi-step engineering tasks in real codebases.&lt;/li&gt;
&lt;li&gt;TraeIDE&amp;rsquo;s official page shows built-in Claude 3.5/3.7, DeepSeek, etc., but the community notes slow integration of new models like Claude Sonnet 4.5.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So, &amp;ldquo;TRAE does not support Claude&amp;rdquo; is now inaccurate—at least officially, Claude models are included. Which model is used in SOLO mode and whether it&amp;rsquo;s exposed to users remains unclear; the experience still needs improvement.&lt;/p&gt;
&lt;h3 id="vs-code--copilot--agent-hq"&gt;VS Code + Copilot + Agent HQ&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Copilot in VS Code now features Chat, Plan agent, Todo/multi-step execution:
&lt;ul&gt;
&lt;li&gt;Plan mode analyzes codebases, generates execution plans, splits into Todos, then implementation agents execute step by step.&lt;/li&gt;
&lt;li&gt;Agent mode provides a more automated &amp;ldquo;multi-step companion programmer&amp;rdquo; experience.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;GitHub launched Agent HQ at Universe 2025, integrating Copilot and third-party agents (Anthropic, OpenAI, Google, xAI, Cognition, etc.) into a unified control plane, supporting parallel runs and result comparison.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;TRAE is like &amp;ldquo;embedding an engineering entity into the IDE.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;VS Code + Copilot is &amp;ldquo;adding a set of engineering entities to a mature IDE.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Agent HQ is positioned as &amp;ldquo;headquarters for multiple engineering entities.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="reconstructing-comparison-dimensions-with-the-ai-engineering-entity-framework"&gt;Reconstructing Comparison Dimensions with the AI Engineering Entity Framework&lt;/h2&gt;
&lt;p&gt;In &amp;ldquo;AI Engineering Entity (AIEE),&amp;rdquo; the definition is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI has evolved from editor auto-completion to a formal node in the software supply chain, able to receive tasks, produce reviewable artifacts (PR/diff/report), pass tests/gates, and be replaced if it fails. It&amp;rsquo;s no longer just an &amp;ldquo;enhanced human developer,&amp;rdquo; but a &amp;ldquo;functional engineering unit&amp;rdquo; in the pipeline.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Based on this, we can reconstruct key comparison dimensions for TRAE and VS Code:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Existence as an Independent Functional Unit&lt;/strong&gt;&lt;br&gt;
Can it autonomously plan, implement, and produce PRs/reports from natural language requirements, without continuous human intervention?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Context Modeling Capability&lt;/strong&gt;&lt;br&gt;
Can it model across files, directories, terminal output, and browser content to form a stable engineering context?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Position in the Pipeline&lt;/strong&gt;&lt;br&gt;
Is it an enhancement layer within the IDE, or a formal node in CI/CD and code review flows?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Reviewability and Replaceability&lt;/strong&gt;&lt;br&gt;
Are its outputs standardized (PR, diff, report) and suitable for regular pipeline review and rollback?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Multi-Agent Collaboration Capability&lt;/strong&gt;&lt;br&gt;
Does it natively support multi-agent collaboration, or is it focused on single-agent enhancement?&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="trae-solo-vs-vs-code-engineering-entity-comparison-table"&gt;TRAE SOLO vs VS Code: Engineering Entity Comparison Table&lt;/h2&gt;
&lt;p&gt;The following table summarizes the main differences from the engineering entity perspective. Note: VS Code includes Copilot Chat + Plan/Agent mode by default and can mount the Agent HQ ecosystem.&lt;/p&gt;
&lt;p&gt;You can interpret this table as: &amp;ldquo;If AI is treated as an engineering entity in the pipeline, what roles do TRAE and VS Code play?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Table:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;TRAE SOLO&lt;/th&gt;
&lt;th&gt;VS Code + Copilot / Agent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Engineering Role&lt;/td&gt;
&lt;td&gt;Single strong entity, directly handles end-to-end tasks from idea to deployment&lt;/td&gt;
&lt;td&gt;IDE + multiple entities (Plan, Implementation, Review), IDE itself is the engineering base&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task Granularity&lt;/td&gt;
&lt;td&gt;Project/feature level: from PRD-style description to full project scaffold, implementation, testing, preview&lt;/td&gt;
&lt;td&gt;Function/file level mainly; Plan mode can scale to feature/subsystem level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Modeling&lt;/td&gt;
&lt;td&gt;Emphasizes &amp;ldquo;context engineering&amp;rdquo;: reads codebase, terminal output, browser content as unified input for SOLO&lt;/td&gt;
&lt;td&gt;Mainly codebase; Plan Agent generates plans based on code analysis, Agent mode schedules by plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automation&lt;/td&gt;
&lt;td&gt;Can proactively modify files, run commands, tests, start local services, forming a complete loop&lt;/td&gt;
&lt;td&gt;Plan/Agent can run commands, modify files, run tests, but is more dependent on your current project/workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Human Intervention&lt;/td&gt;
&lt;td&gt;More &amp;ldquo;post-review&amp;rdquo;: let it run first, then review and fine-tune&lt;/td&gt;
&lt;td&gt;More &amp;ldquo;in-process collaboration&amp;rdquo;: frequent intervention in planning, implementation, and review, with control points at each step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output Form&lt;/td&gt;
&lt;td&gt;Code changes, test results, preview; sometimes PRs/docs&lt;/td&gt;
&lt;td&gt;Code completion, refactoring, PR comments, CodeQL reports, Plan/Todo lists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-Agent&lt;/td&gt;
&lt;td&gt;Core is the SOLO agent; other capabilities (like Trae Agent CLI) are extensions&lt;/td&gt;
&lt;td&gt;Copilot itself is an agent; Agent HQ allows parallel competition among multiple agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model Transparency&lt;/td&gt;
&lt;td&gt;Product exposes specific models poorly; users can&amp;rsquo;t tell which model is used&lt;/td&gt;
&lt;td&gt;GitHub clearly marks Copilot&amp;rsquo;s model family; Agent HQ shows agent sources directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Strong automation but slow; complex projects may stall at &amp;ldquo;thinking&amp;rdquo; stage; hard token limits&lt;/td&gt;
&lt;td&gt;Stable response in familiar projects; mostly local changes, overall latency is controllable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Privacy &amp;amp; Compliance&lt;/td&gt;
&lt;td&gt;Official and third-party reviews mention extensive telemetry/data collection; enterprise adoption needs extra evaluation&lt;/td&gt;
&lt;td&gt;Copilot for Enterprise has clear data isolation/compliance, suitable for most enterprise governance needs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;figcaption class="text-center mb-3"&gt;
Table 3: TRAE SOLO vs VS Code Engineering Entity Comparison
&lt;/figcaption&gt;
&lt;p&gt;From the table:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you want &amp;ldquo;an AI engineering entity that takes full responsibility from requirement to deployment,&amp;rdquo; TRAE SOLO fits that role.&lt;/li&gt;
&lt;li&gt;If you want &amp;ldquo;a stable engineering base + a set of pluggable entities,&amp;rdquo; VS Code + Copilot + Agent HQ fits better.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="workflow-comparison-two-engineering-entity-pipelines"&gt;Workflow Comparison: Two Engineering Entity Pipelines&lt;/h2&gt;
&lt;p&gt;To clarify their engineering flows, the following diagram illustrates typical workflows for TRAE SOLO and VS Code.&lt;/p&gt;
&lt;p&gt;Before the diagram, here&amp;rsquo;s an introductory sentence:&lt;br&gt;
The following Mermaid diagram visually compares the engineering pipelines of TRAE SOLO and VS Code, highlighting their respective collaboration models.&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/trae-vs-vscode-insiders-agent-hq-and-ai-engineering-entity/85e9f25c80e9c24f2501f2089382295d.svg" data-img="https://assets.jimmysong.io/images/blog/trae-vs-vscode-insiders-agent-hq-and-ai-engineering-entity/85e9f25c80e9c24f2501f2089382295d.svg" alt="Figure 3: TRAE SOLO vs VS Code Engineering Entity Pipeline Comparison" data-caption="Figure 3: TRAE SOLO vs VS Code Engineering Entity Pipeline Comparison"
width="4366"
height="810"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: TRAE SOLO vs VS Code Engineering Entity Pipeline Comparison&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This diagram shows two typical collaboration models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;TRAE SOLO attempts to encapsulate &amp;ldquo;context aggregation → planning → implementation → testing → preview/deployment&amp;rdquo; within a single engineering entity, with user intervention only at requirement input and output review.&lt;/li&gt;
&lt;li&gt;VS Code + Copilot + Agent HQ uses the IDE as runtime, with Plan/Implementation/Review agents corresponding to different roles. Agent HQ supports parallel agent competition, allowing developers to select the best solution.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="model-transparency-speed-and-predictability"&gt;Model Transparency, Speed, and Predictability&lt;/h2&gt;
&lt;p&gt;Based on personal experience, here are the model transparency and speed issues from the engineering entity perspective:&lt;/p&gt;
&lt;h3 id="model-transparency"&gt;Model Transparency&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;TRAE currently exposes &amp;ldquo;which model is called&amp;rdquo; poorly; switching MAX mode only suggests &amp;ldquo;stronger model or higher quota,&amp;rdquo; but no clear feedback.&lt;/li&gt;
&lt;li&gt;Community feedback notes slow integration of new models; some strong models (like Claude series) are available elsewhere but not yet in TRAE.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;TRAE is hard to use as a &amp;ldquo;precisely configurable engineering unit,&amp;rdquo; more like a black box, making model change management in CI/CD or production pipelines difficult.&lt;/li&gt;
&lt;li&gt;VS Code + Copilot + Agent HQ is stronger in standardization; GitHub clearly marks Copilot&amp;rsquo;s model family, Agent HQ uses agent source as the abstraction boundary.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="speed-and-predictability"&gt;Speed and Predictability&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;TRAE SOLO&amp;rsquo;s &amp;ldquo;slowness&amp;rdquo; comes from executing more steps (reading files, analyzing, planning, testing) and insufficient engineering process visualization. The UI shows &amp;ldquo;Thinking…&amp;rdquo; prompts, making it hard to tell if it&amp;rsquo;s stuck or planning.&lt;/li&gt;
&lt;li&gt;VS Code&amp;rsquo;s Plan mode explicitly lists plans and Todos; Agent mode emphasizes &amp;ldquo;execution by plan,&amp;rdquo; letting users clearly see the entity&amp;rsquo;s work status, improving predictability.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="agent-hq-positioning-single-entity-vs-multi-entity-headquarters"&gt;Agent HQ Positioning: Single Entity vs Multi-Entity Headquarters&lt;/h2&gt;
&lt;p&gt;From a platform perspective, GitHub and TRAE differ as follows:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Agent HQ&amp;rsquo;s core idea: future development will rely on multiple specialized agents collaborating in parallel. GitHub is building &amp;ldquo;agent headquarters,&amp;rdquo; not a single engineering agent. Developers can schedule agents in a unified control plane, integrating with existing GitHub Flow (Issue, PR, Review, CI/CD).&lt;/li&gt;
&lt;li&gt;TRAE is more like &amp;ldquo;proprietary IDE + agent + full-stack context engineering,&amp;rdquo; delivering an integrated experience.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In terms of engineering entity organization:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GitHub is building &amp;ldquo;infrastructure and governance for multi-entity engineering systems.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;TRAE is building &amp;ldquo;vertically integrated engineering entity + private runtime.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;They&amp;rsquo;re not mutually exclusive, representing &amp;ldquo;broad platform + multi-entity scheduling&amp;rdquo; vs &amp;ldquo;single strong entity + proprietary toolchain.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="subjective-experience-and-engineering-framework-integration"&gt;Subjective Experience and Engineering Framework Integration&lt;/h2&gt;
&lt;p&gt;Translating personal experience into engineering language:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;VS Code is more accustomed to a &amp;ldquo;single IDE + multiple views&amp;rdquo; experience; TRAE splits IDE and SOLO modes, requiring mental switching.&lt;/li&gt;
&lt;li&gt;TRAE&amp;rsquo;s engineering entity capabilities surpass ordinary completion tools, able to take on tasks, but model transparency and context quality are unstable, and governance needs improvement.&lt;/li&gt;
&lt;li&gt;VS Code doesn&amp;rsquo;t take over the whole project, but local work is stable; Plan, Agent, and Review combinations enable multi-entity collaboration.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;According to the &amp;ldquo;AI Engineering Entity (AIEE)&amp;rdquo; framework:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;TRAE SOLO is a &lt;strong&gt;single AI engineering entity (AIEE) capable of handling complete engineering tasks&lt;/strong&gt;, but still has clear shortcomings in model transparency, engineering governance, and enterprise-level controllability.&lt;/li&gt;
&lt;li&gt;VS Code + Copilot + Agent HQ is an &lt;strong&gt;infrastructure platform for multiple engineering entities&lt;/strong&gt;, less aggressive in end-to-end outsourcing in the short term, but clearer in engineering consistency, model replaceability, and organizational governance.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;This article systematically compares TRAE SOLO and VS Code (with Copilot, Agent HQ) from the perspective of AI engineering entities, focusing on automation, collaboration, and model transparency. TRAE SOLO is better suited for individual developers or small teams seeking end-to-end automation, while VS Code + Copilot + Agent HQ provides stronger infrastructure for multi-entity collaboration, enterprise governance, and engineering consistency. In the future, AI engineering entities will become formal nodes in the software development pipeline, and tool selection should be based on engineering needs and governance requirements.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.webull.com/news/13844395366540288" target="_blank" rel="noopener"&gt;ByteDance&amp;rsquo;s AI programming tool TRAE announced the &amp;hellip; - webull.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/news-insights/company-news/welcome-home-agents" target="_blank" rel="noopener"&gt;Introducing Agent HQ: Any agent, any way you work - github.blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://traesolo.net" target="_blank" rel="noopener"&gt;TRAE SOLO - AI-Powered Context Engineer for Faster &amp;hellip; - traesolo.net&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/blogs/2025/02/24/introducing-copilot-agent-mode" target="_blank" rel="noopener"&gt;Introducing GitHub Copilot agent mode (preview) - code.visualstudio.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jimmysong.io/book/ai-handbook/infra/ai-engineering-entity/" target="_blank" rel="noopener"&gt;AI Engineering Entity | Jimmy Song - jimmysong.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.trae.ai" target="_blank" rel="noopener"&gt;TRAE - Collaborate with Intelligence - trae.ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/bytedance/trae-agent" target="_blank" rel="noopener"&gt;bytedance/trae-agent - github.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://traeide.com" target="_blank" rel="noopener"&gt;TraeIDE - traeide.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/chat/copilot-chat" target="_blank" rel="noopener"&gt;Get started with chat in VS Code - code.visualstudio.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://skywork.ai/blog/trae-ai-ide-review-2025-cursor-alternative" target="_blank" rel="noopener"&gt;Trae AI IDE Review 2025: ByteDance&amp;rsquo;s Free IDE vs Cursor - skywork.ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/changelog/2025-10-28-github-copilot-in-visual-studio-code-gets-upgraded" target="_blank" rel="noopener"&gt;GitHub Copilot in Visual Studio Code gets upgraded - github.blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.traesolo.org/blog/trae-solo-comprehensive-review" target="_blank" rel="noopener"&gt;TRAE 2.0 Preview: AI-Native Development Paradigm Leap | T&amp;hellip; - traesolo.org&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Closed-Source Flagships Accelerate, Open-Source Ecosystem Forced to 'Synchronize'</title><link>https://jimmysong.io/blog/closed-source-flagships-and-open-source-twins/</link><pubDate>Fri, 14 Nov 2025 04:21:07 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/closed-source-flagships-and-open-source-twins/</guid><description>Analysis of closed-source model acceleration and open-source ecosystem response, exploring core engineering contradictions and infrastructure evolution.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;Closed-source models are accelerating, while open-source ecosystems are forced to catch up. What engineers truly need to focus on is infrastructure and controllability—not just the surface-level &amp;ldquo;Twin&amp;rdquo; phenomenon.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Recently, I came across an email titled: &lt;strong&gt;&amp;ldquo;Every Big AI Model Now Has an Open-Source Twin&amp;rdquo;&lt;/strong&gt;. Literally translated, it means &amp;ldquo;Every major closed-source model now has an open-source sibling.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;From the perspective of media or venture capital, this is an easy story to tell: a closed-source flagship model is released, the community quickly produces an open-source counterpart, and the narrative becomes &amp;ldquo;open source is catching up with closed source&amp;rdquo;—the ecosystem is thriving, innovation is accelerating, and the future looks promising.&lt;/p&gt;
&lt;p&gt;But from the viewpoint of someone deeply involved in infrastructure, cloud native, and architecture, this narrative has several issues:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It equates &amp;ldquo;synchronized pace&amp;rdquo; with &amp;ldquo;matched capabilities.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;It overlooks the real factors that determine the ceiling: data, compute, and engineering systems.&lt;/li&gt;
&lt;li&gt;It blurs a key fact: the open-source ecosystem is fundamentally in a &lt;strong&gt;reactive state&lt;/strong&gt;, not leading in parallel.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This article breaks down the &amp;ldquo;Open-Source Twin&amp;rdquo; narrative from an engineering and infrastructure perspective, and shares the core issues I personally care about.&lt;/p&gt;
&lt;h2 id="from-theres-a-twin-to-forced-synchronization-how-the-narrative-changed"&gt;From &amp;ldquo;There&amp;rsquo;s a Twin&amp;rdquo; to &amp;ldquo;Forced Synchronization&amp;rdquo;: How the Narrative Changed&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s first outline the phenomenon.&lt;/p&gt;
&lt;p&gt;In the past two years, the industry has repeatedly seen the following pattern:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Big tech releases a closed-source flagship model (e.g., GPT-5 series, Claude 4/4.5, Gemini 2.5).&lt;/li&gt;
&lt;li&gt;Soon after, a batch of open-source counterparts emerge (e.g., Qwen, GLM, Yi, K2), aligning on parameter scale and benchmark metrics.&lt;/li&gt;
&lt;li&gt;Media and community start using terms like &amp;ldquo;open-source twin,&amp;rdquo; &amp;ldquo;replacement,&amp;rdquo; and &amp;ldquo;counterpart.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At this level, it&amp;rsquo;s easy to draw optimistic conclusions: &lt;strong&gt;open source has established full benchmarking capabilities; no matter how fast closed source runs, the community can keep up.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;But the more critical question is: &lt;strong&gt;Who sets the pace, who defines the rules, and who bears the real cost?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The current structure is clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Pace is set by closed-source giants&lt;/strong&gt;: They decide when to boost inference, extend context, push multimodality, or specialize reasoning (like the R1 series).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open-source ecosystem passively initiates response mechanisms&lt;/strong&gt;: Each closed-source upgrade triggers a new round of &amp;ldquo;open-source benchmarking.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In other words, the current model isn&amp;rsquo;t parallel innovation or mutual stimulation—it&amp;rsquo;s closed source constantly shifting gears at the front, with open source adjusting to avoid falling out of sight.&lt;/p&gt;
&lt;p&gt;From an engineering perspective, &amp;ldquo;Every Big AI Model Has an Open-Source Twin&amp;rdquo; is more accurately:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Every Big AI Model Now Forces an Open-Source Response.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="what-exactly-is-open-source-synchronizing-with"&gt;What Exactly Is Open Source &amp;ldquo;Synchronizing&amp;rdquo; With?&lt;/h2&gt;
&lt;p&gt;To understand the &amp;ldquo;synchronized response&amp;rdquo; phenomenon, we need to break it down into three categories.&lt;/p&gt;
&lt;p&gt;Before listing them, let&amp;rsquo;s add some context: every closed-source flagship update isn&amp;rsquo;t just &amp;ldquo;more parameters, higher scores&amp;rdquo;—it&amp;rsquo;s constantly &lt;strong&gt;rewriting constraints&lt;/strong&gt;, including inference cost, interaction patterns, context length, multimodal consistency, and explainability.&lt;/p&gt;
&lt;p&gt;In this context, open source isn&amp;rsquo;t just synchronizing &amp;ldquo;scores,&amp;rdquo; but increasingly complex &lt;strong&gt;objective functions&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="synchronizing-the-imagination-boundary-of-capability-ceilings"&gt;Synchronizing the &amp;ldquo;Imagination Boundary&amp;rdquo; of Capability Ceilings&lt;/h3&gt;
&lt;p&gt;Closed-source models essentially expand &amp;ldquo;what people think a model should be able to do,&amp;rdquo; such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;From pure text to text + image + audio + video.&lt;/li&gt;
&lt;li&gt;From single-turn Q&amp;amp;A to engineering-level reasoning, coding, debugging, fixing, and refactoring.&lt;/li&gt;
&lt;li&gt;From thousands of tokens of context to hundreds of thousands or more.&lt;/li&gt;
&lt;li&gt;From &amp;ldquo;black box output&amp;rdquo; to having chains of thought, reasoning traces, and verifiable outputs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open-source models then align their goals: &amp;ldquo;We also need long context, multimodality, coding ability, and agent workflow support.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="synchronizing-the-expectation-value-of-interfaces-and-usage-patterns"&gt;Synchronizing the &amp;ldquo;Expectation Value&amp;rdquo; of Interfaces and Usage Patterns&lt;/h3&gt;
&lt;p&gt;Once developers and enterprise users are educated by closed-source models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How low can interaction latency go?&lt;/li&gt;
&lt;li&gt;How long can context be extended without breaking?&lt;/li&gt;
&lt;li&gt;How smooth can multimodal input be?&lt;/li&gt;
&lt;li&gt;How &amp;ldquo;smart&amp;rdquo; can the reasoning process get?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Their expectations for &lt;strong&gt;any open-source model&lt;/strong&gt; are recalibrated.&lt;/p&gt;
&lt;p&gt;Thus, open source must:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Continuously optimize inference frameworks (e.g., vLLM, SGLang, TGI) to narrow the latency gap.&lt;/li&gt;
&lt;li&gt;Make serving experiences closer to closed-source, such as OpenAI API compatibility and better SDKs.&lt;/li&gt;
&lt;li&gt;Forcefully catch up on multimodality and long context, even if training costs are high.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="synchronizing-surface-metrics-not-complete-capabilities"&gt;Synchronizing &amp;ldquo;Surface Metrics,&amp;rdquo; Not &amp;ldquo;Complete Capabilities&amp;rdquo;&lt;/h3&gt;
&lt;p&gt;From a benchmark perspective, open source can indeed reach 80–90% on public test sets:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MMLU, GSM8K, HumanEval.&lt;/li&gt;
&lt;li&gt;Common reasoning, reading comprehension, code generation metrics.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But these metrics only reflect &lt;strong&gt;surface capabilities&lt;/strong&gt;, not:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Robustness to long-tail problems.&lt;/li&gt;
&lt;li&gt;Stability in complex, multi-step scenarios.&lt;/li&gt;
&lt;li&gt;Reliability and controllability in large-scale production systems.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Engineering health&amp;rdquo; over long-term evolution.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is why I&amp;rsquo;m skeptical of the &amp;ldquo;Twin&amp;rdquo; term: &lt;strong&gt;it uses superficial metric similarity to mask deep structural differences.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="why-open-source-twin-sounds-good-but-misses-the-core-contradiction"&gt;Why &amp;ldquo;Open-Source Twin&amp;rdquo; Sounds Good but Misses the Core Contradiction&lt;/h2&gt;
&lt;p&gt;From an infrastructure and engineering perspective, the real issue isn&amp;rsquo;t &amp;ldquo;can open source copy a similar architecture,&amp;rdquo; but:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Who can sustainably manage data, compute, scheduling systems, and engineering teams to build a long-term model production pipeline.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There are three key contradictions here.&lt;/p&gt;
&lt;h3 id="data-is-unavailable-training-recipes-cant-be-fully-reproduced"&gt;Data Is Unavailable, Training Recipes Can&amp;rsquo;t Be Fully Reproduced&lt;/h3&gt;
&lt;p&gt;Open source can replicate general network structures and optimization tricks, but can&amp;rsquo;t access:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Closed-source data sources and cleaning standards.&lt;/li&gt;
&lt;li&gt;Filtering strategies, detoxification, alignment details.&lt;/li&gt;
&lt;li&gt;Large-scale synthetic data generation and selection methods.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Result: &lt;strong&gt;Even if you match parameter scale and training steps, the effect may not truly align.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Many open-source projects have to use rougher data, limited compute budgets, and more conservative training strategies, ending up with a &amp;ldquo;usable but not truly stable&amp;rdquo; state.&lt;/p&gt;
&lt;h3 id="compute-gap-is-structural-not-solved-by-one-time-funding"&gt;Compute Gap Is Structural, Not Solved by One-Time Funding&lt;/h3&gt;
&lt;p&gt;Training flagship models requires compute that&amp;rsquo;s not just hundreds or thousands of GPUs—it&amp;rsquo;s a &lt;strong&gt;structural, long-term investment&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In the open-source camp, those approaching this scale usually have:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Backing from large companies or national labs.&lt;/li&gt;
&lt;li&gt;Funding from real business budgets, not community donations.&lt;/li&gt;
&lt;li&gt;Compute supply that can be planned long-term, not just a one-off &amp;ldquo;burn.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Reality:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Entities truly capable of &amp;ldquo;flagship open-source models&amp;rdquo; are essentially &amp;ldquo;institutions,&amp;rdquo; not loose personal communities.&lt;/li&gt;
&lt;li&gt;Most &amp;ldquo;open-source twins&amp;rdquo; are backed by enterprises, with product goals and commercial interests.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So, &lt;strong&gt;&amp;ldquo;open source vs closed source&amp;rdquo; is more like &amp;ldquo;many big companies vs a few giants,&amp;rdquo; not &amp;ldquo;community vs company.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id="reproducing-architecture--leading-architecture"&gt;&amp;ldquo;Reproducing&amp;rdquo; Architecture ≠ &amp;ldquo;Leading&amp;rdquo; Architecture&lt;/h3&gt;
&lt;p&gt;Many open-source models look architecturally similar to closed-source:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Transformer variants, MoE variants.&lt;/li&gt;
&lt;li&gt;Minor tweaks at the decision layer.&lt;/li&gt;
&lt;li&gt;Some inference optimizations.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But in terms of industry power, those truly pushing these architectures to production scale and validating feasibility are still on the closed-source side. Open source mainly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Validates closed-source approaches on weaker compute.&lt;/li&gt;
&lt;li&gt;Explores &amp;ldquo;smaller, cheaper&amp;rdquo; approximations.&lt;/li&gt;
&lt;li&gt;Prunes and adapts for specific scenarios.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So, &amp;ldquo;Twin&amp;rdquo; is more a marketing term than an engineering one.&lt;/p&gt;
&lt;h2 id="my-perspective-what-really-matters-in-this-game"&gt;My Perspective: What Really Matters in This Game&lt;/h2&gt;
&lt;p&gt;As an engineer in cloud native, service mesh, and distributed systems, my default thinking when looking at AI infrastructure is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Treat &amp;ldquo;models&amp;rdquo; as just one component in the system, and focus on the underlying infrastructure, scheduling systems, and engineering pipelines.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;From this angle, the real concern behind &amp;ldquo;every closed-source model has an open-source twin&amp;rdquo; is:&lt;/p&gt;
&lt;figure class="mx-auto text-center"&gt;
&lt;img src="https://assets.jimmysong.io/images/blog/closed-source-flagships-and-open-source-twins/4fb6b788319f45d2c5a4aedd1af26057.svg" data-img="https://assets.jimmysong.io/images/blog/closed-source-flagships-and-open-source-twins/4fb6b788319f45d2c5a4aedd1af26057.svg" alt="Figure 3: Mermaid Diagram" data-caption="Figure 3: Mermaid Diagram"
width="2616"
height="1359"
loading="lazy" decoding="async" class="image-loading"
onload="this.classList.remove('image-loading'); this.classList.add('image-loaded');"
onerror="handleImageError(this); this.classList.remove('image-loading');"&gt;
&lt;figcaption&gt;Figure 3: Mermaid Diagram&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="can-open-source-models-stand-firm-in-production-over-time"&gt;Can Open-Source Models Stand Firm in Production Over Time?&lt;/h3&gt;
&lt;p&gt;The focus isn&amp;rsquo;t whether it can run a demo, but:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Is there a clear upgrade cadence?&lt;/li&gt;
&lt;li&gt;Are rollback and compatibility strategies robust?&lt;/li&gt;
&lt;li&gt;Is there a sound evolution path for model weights, inference frameworks, and configurations?&lt;/li&gt;
&lt;li&gt;Is the entire stack observable and debuggable?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If these infrastructure layers aren&amp;rsquo;t mature, the so-called &amp;ldquo;Twin&amp;rdquo; is just &amp;ldquo;something that looks similar, but don&amp;rsquo;t ask if it can support your production workloads.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="has-training-and-inference-infrastructure-formed-a-replicable-engineering-paradigm"&gt;Has Training and Inference Infrastructure Formed a Replicable &amp;ldquo;Engineering Paradigm&amp;rdquo;?&lt;/h3&gt;
&lt;p&gt;The real value in open source is whether it can form a unified, teachable, and transferable engineering paradigm, such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Training pipeline: data preparation → preprocessing → training → evaluation → alignment → deployment.&lt;/li&gt;
&lt;li&gt;Inference infrastructure: how vLLM / SGLang / TGI maintain consistent performance across different GPU topologies.&lt;/li&gt;
&lt;li&gt;Scheduling and resource management: how to manage large-scale inference loads on Kubernetes and cloud-native infrastructure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If these can be established, &amp;ldquo;open-source twin&amp;rdquo; isn&amp;rsquo;t just &amp;ldquo;we also have a model,&amp;rdquo; but a reusable, transparent, and learnable engineering system.&lt;/p&gt;
&lt;h3 id="the-true-value-of-open-source-controllability-and-bargaining-power-not-absolute-performance"&gt;The True Value of Open Source: Controllability and Bargaining Power, Not Absolute Performance&lt;/h3&gt;
&lt;p&gt;Realistically, closed-source flagships will continue to lead in &lt;strong&gt;overall capability&lt;/strong&gt; for the foreseeable future: larger scale, more complex training, better data, richer scenario tuning.&lt;/p&gt;
&lt;p&gt;For enterprises and developers, the key value of open source isn&amp;rsquo;t &amp;ldquo;I want to fully replace closed source,&amp;rdquo; but:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maintaining controllability over technical direction.&lt;/li&gt;
&lt;li&gt;Gaining bargaining power, avoiding vendor lock-in.&lt;/li&gt;
&lt;li&gt;Building your own model stack in privacy-sensitive or compliance-heavy scenarios.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;From this perspective, the &amp;ldquo;Twin&amp;rdquo; term should be soberly rewritten as:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In many scenarios, open source can provide a controllable, more flexible alternative path—but it&amp;rsquo;s not a mirror of closed source, it&amp;rsquo;s a separate engineering decision space.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="practical-advice-for-engineers-and-teams-dont-worship-twins-see-the-structure-clearly"&gt;Practical Advice for Engineers and Teams: Don&amp;rsquo;t Worship &amp;ldquo;Twins,&amp;rdquo; See the Structure Clearly&lt;/h2&gt;
&lt;p&gt;Before the final summary, here are my actionable views on this topic.&lt;/p&gt;
&lt;p&gt;Premise: If you&amp;rsquo;re an engineer, architect, or technical leader, your real decision isn&amp;rsquo;t &amp;ldquo;choose open source or closed source,&amp;rdquo; but:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Under your business constraints, how do you combine closed-source APIs, open-source weights, and self-built infrastructure to create an &lt;strong&gt;evolvable, observable, and portable&lt;/strong&gt; system.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;With this in mind, &amp;ldquo;Every Big AI Model Has an Open-Source Twin&amp;rdquo; breaks down into several sober judgments:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When you see an &amp;ldquo;open-source twin,&amp;rdquo; first ask: can it run stably in your production environment over time, not just pass benchmarks?&lt;/li&gt;
&lt;li&gt;What you really need to understand: is there a clear story behind its training/inference infrastructure, not just a weight download link?&lt;/li&gt;
&lt;li&gt;Reframe &amp;ldquo;open source vs closed source&amp;rdquo; as &amp;ldquo;where do I need closed source (capability/cost), and where do I need open source (controllability/compliance)?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;If you&amp;rsquo;re working on infrastructure and platform layers, focus on:
&lt;ul&gt;
&lt;li&gt;How to run different models in a unified scheduling, monitoring, and logging system.&lt;/li&gt;
&lt;li&gt;How to treat large models as observable, governable services on Kubernetes/cloud-native stacks, not mysterious black boxes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Closed-source flagships keep accelerating, shifting gears, and adding dimensions, while open-source ecosystems are forced to develop increasingly mature synchronized response mechanisms. What truly determines the gap is data, compute, and engineering infrastructure—not just a single model release.&lt;/p&gt;
&lt;p&gt;Personally, my focus will remain on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The evolution of inference infrastructure (like vLLM, SGLang, TGI).&lt;/li&gt;
&lt;li&gt;Training and scheduling: how to stably manage model lifecycles in cloud-native environments.&lt;/li&gt;
&lt;li&gt;Engineering paradigm accumulation: moving from &amp;ldquo;can run&amp;rdquo; to &amp;ldquo;reproducible, maintainable, and evolvable.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Key Takeaways from Ingress NGINX Retirement: Managing Technical Debt in Cloud Native Migration</title><link>https://jimmysong.io/blog/ingress-nginx-retirement-insights/</link><pubDate>Thu, 13 Nov 2025 01:43:05 +0000</pubDate><author>Jimmy Song</author><guid>https://jimmysong.io/blog/ingress-nginx-retirement-insights/</guid><description>The retirement of Ingress NGINX reveals technical debt, migration paths, and the trend toward standardized traffic management in cloud native infrastructure.</description><content:encoded>
&lt;blockquote&gt;
&lt;p&gt;The evolution of cloud native infrastructure inevitably faces the reality of technical debt and governance. The retirement of Ingress NGINX is a profound reminder about standardization and sustainability.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Kubernetes officially announced: &lt;a href="https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/" target="_blank" rel="noopener"&gt;Ingress NGINX will be completely discontinued in March 2026&lt;/a&gt;. This is not just a typical project sunset, but a landmark event in the evolution of the Kubernetes networking model. It signals the inevitable shift of the tech stack from &amp;ldquo;flexible but fragile&amp;rdquo; to &amp;ldquo;controllable and governable.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;As someone who has long promoted Kubernetes and cloud native practices, I have witnessed both the golden age of Ingress NGINX and the gradual accumulation of its technical debt. Here are the clear insights this event has brought me.&lt;/p&gt;
&lt;h2 id="technical-debt-will-eventually-backfire-especially-for-infrastructure-components"&gt;Technical Debt Will Eventually Backfire, Especially for Infrastructure Components&lt;/h2&gt;
&lt;p&gt;The core issue with Ingress NGINX is not a decline in users, but that &amp;ldquo;maintenance costs permanently exceed the pace of contributions.&amp;rdquo; High flexibility leads to a huge attack surface, years of complex configuration legacy, and a shortage of community maintainers, ultimately making the project unsustainable.&lt;/p&gt;
&lt;p&gt;Once infrastructure components can no longer be securely updated, they cease to be assets and become liabilities.&lt;/p&gt;
&lt;div class="alert alert-note-container"&gt;
&lt;div class="alert-note-title px-2"&gt;
My Conclusion
&lt;/div&gt;
&lt;div class="alert-note px-2"&gt;
&lt;strong&gt;The threshold for future infrastructure will be higher, with stricter requirements for security and maintainability. The model of individual hero maintainers will continue to fail.&lt;/strong&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="kubernetes-officially-enters-the-gateway-api-era"&gt;Kubernetes Officially Enters the Gateway API Era&lt;/h2&gt;
&lt;p&gt;Before introducing Gateway API (Gateway API, Gateway Application Programming Interface), it&amp;rsquo;s important to review the design of Ingress. Ingress was once praised for its simplicity, but now it cannot meet modern needs for traffic management, scalability, security policies, and multi-team collaboration.&lt;/p&gt;
&lt;p&gt;Gateway API is designed with a more modern philosophy:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Governance model across roles (Infra / Dev / Ops)&lt;/li&gt;
&lt;li&gt;Strong CRD (Custom Resource Definition) extensibility&lt;/li&gt;
&lt;li&gt;Pluggable implementation&lt;/li&gt;
&lt;li&gt;Significantly improved observability and lifecycle management&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This means: &lt;strong&gt;The entire ecosystem is moving from &amp;ldquo;controller differentiation&amp;rdquo; to &amp;ldquo;API standardization&amp;rdquo; at the traffic layer.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="most-users-are-unprepared-for-the-complexity-of-the-underlying-network-stack"&gt;Most Users Are Unprepared for the Complexity of the Underlying Network Stack&lt;/h2&gt;
&lt;p&gt;Long-term community observation shows that most users treat Ingress NGINX as a black box. Now, migrating from Ingress to Gateway API or other Ingress controllers represents a &amp;ldquo;hidden migration wave&amp;rdquo; for many clusters.&lt;/p&gt;
&lt;p&gt;This announcement highlights two points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When a &amp;ldquo;default component&amp;rdquo; in a complex system stops being updated, it brings widespread invisible risks&lt;/li&gt;
&lt;li&gt;The cloud native ecosystem needs long-term, sustainable supply chain governance&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="security-is-the-final-straw"&gt;Security Is the Final Straw&lt;/h2&gt;
&lt;p&gt;The official announcement repeatedly emphasizes that security risks and vulnerabilities can no longer be continuously fixed. This once again proves: &lt;strong&gt;Flexibility and security are always a tradeoff, and the closer a component is to the data plane, the less compromise is acceptable.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-individual-maintainer-bottleneck-in-cloud-native-will-become-more-pronounced"&gt;The &amp;ldquo;Individual Maintainer Bottleneck&amp;rdquo; in Cloud Native Will Become More Pronounced&lt;/h2&gt;
&lt;p&gt;Ingress NGINX has long relied on just one or two maintainers, and ultimately had to retire. This exposes a long-standing issue in the open source world: &lt;strong&gt;Critical projects are heavily relied upon, but contributions are insufficient.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The future of infrastructure is clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Large companies will be more willing to invest in core open source infrastructure&lt;/li&gt;
&lt;li&gt;Individual maintainers cannot support critical foundational components&lt;/li&gt;
&lt;li&gt;The boundary between commercialization and open source will continue to tighten&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="my-personal-takeaway-gateway-api-l7-traffic-management-and-the-integration-with-ai-native-infra"&gt;My Personal Takeaway: Gateway API, L7 Traffic Management, and the Integration with AI Native Infra&lt;/h2&gt;
&lt;p&gt;The retirement of Ingress NGINX points to an underlying trend: &lt;strong&gt;Unified and extensible APIs will become the dominant paradigm for cloud native infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The AI-Native infrastructure I&amp;rsquo;m researching—such as inference routing, model gateways, AI Gateway, and Agent Orchestrator—will follow a similar path: from early flexible hacks to mature, standardized, and governed APIs.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;Ingress NGINX is arguably one of the most important control planes in Kubernetes history. Its retirement is not a failure, but an inevitable result of the system advancing to the next stage.&lt;/p&gt;
&lt;p&gt;For me, this is a strong reminder:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Technical debt cannot be avoided&lt;/li&gt;
&lt;li&gt;Infrastructure must be built for the long term&lt;/li&gt;
&lt;li&gt;Standardized APIs are the future&lt;/li&gt;
&lt;li&gt;Sustainable open source requires collective investment&lt;/li&gt;
&lt;li&gt;The convergence of AI and cloud native will follow the same evolutionary trajectory&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/" target="_blank" rel="noopener"&gt;Ingress NGINX Retirement - kubernetes.io&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gateway-api.sigs.k8s.io/guides/" target="_blank" rel="noopener"&gt;Gateway API Official Documentation - gateway-api.sigs.k8s.io&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item></channel></rss>