Meta launched Muse on September 8. Each Muse runs on a dedicated Muse Secure VM that holds both the agent and user data; longer tasks keep running after the user closes the app, and Muse only comes back when something changes or needs approval. The product took off fast and briefly topped the U.S. mobile download charts.
For Muse's product side and what it means for Meta's thesis, see Dolphin's note *Muse Blew Up—Real Inflection or False Peak for Meta?*. Below we look at what this changes for the upstream supply chain.
Muse's boom not only drove double-digit gains in $Meta.US—it also lifted names across the CPU chain. So: is Meta's Muse the ChatGPT moment for CPUs?
How Muse Pulls CPU Demand
Unlike a traditional chatbot, Muse is a subscription for “a 24/7 dedicated computer + an AI butler.”
Example: you want to “watch Mac mini prices”:
① Traditional chatbot: “Watch Mac mini prices for me” → GPU inference → “OK” → done → CPU idle → next day you ask again, everything restarts;
② Muse: “Watch Mac mini prices for me” → GPU inference → CPU sets a watch script → VM keeps running → you close the app → VM still runs in the cloud → checks a price API once a day → 30 checks a month, each trigger: GPU judgment + CPU tool execution → 30 days later price hits $899 — CPU script detects it → wakes GPU for a buy decision → CPU drives the browser to fill address, card, and click buy.
In short: traditional chatbots leave when the turn ends; Muse keeps running. Sounds a bit like OpenClaw? Muse was inspired by OpenClaw, but they differ.
The biggest difference: Muse gives each user a dedicated VM (Linux VM, 2 vCPUs), while OpenClaw defaults to local. So even after the app closes, Muse keeps working in the cloud.
Why CPUs Matter in Agents
Put CPU and GPU side by side and the silicon story diverges sharply.
GPUs pour almost all transistors into arithmetic units (green blocks below); CPUs spend most transistors on things that aren't pure math—control and management (red blocks below): branch predictors, out-of-order windows, prefetchers, multi-level caches, and so on.
Break an agent workflow down and you get roughly: Agent = perceive → plan → reason → tool execute → verify → re-plan → re-execute. The GPU only owns the “reason” slice; planning, execution, and verification must run on CPU (control/management).
In pure inference: the CPU mainly tokenizes requests, hands them to the GPU, and reassembles results. Heavy compute sits on the GPU.
In agentic AI: the CPU is the instruction layer—planning tasks, splitting goals into subtasks, scheduling parallel sub-agents, managing tool calls and APIs, monitoring token streams, merging results, and running reflection loops. What stays occupied is mainly agent CPU.
With agentic AI, datacenter CPUs fall into three buckets: traditional CPUs, AI head-node CPUs, and agentic AI CPUs:
a) Traditional / general-purpose CPU: runs “traditional non-AI workloads”—web servers, app servers, databases, caches, storage, message queues.
Main job: site serving, DB management, enterprise tasks—complex logic, low need for massive parallelism, mostly serial. Fully separate from AI GPU servers; no GPU dependency; deployed on traditional racks.
b) AI head-node CPU: “the CPU embedded in a GPU server that manages GPUs and devices”—feeds data to GPUs, coordinates communication, manages memory, runs the non-accelerated parts GPUs can't do.
Main job: manage attached GPUs and keep them fed. To cut tail latency you want high-performance single cores with large caches, high-bandwidth memory, and IO. Physically bound to GPUs in the same server. There is no standalone head-node CPU rack—it is the GPU server's “butler.”
c) Agentic AI CPU: the agent's brain and hands—standalone CPU racks for agent orchestration, tool execution, secure sandboxes, and RAG retrieval.
Main job: autonomous decision and planning for AI agents. Tool calls and the logic between them don't need high GPU utilization—mostly CPU.
Separate from but tightly coupled to AI GPU servers—agent CPU racks interconnect with GPU racks over high-speed fabric (Spectrum-X / Ethernet) into a full AI factory. Agent CPUs are not inside GPU servers; they are their own rack category.
Incremental Demand from Agentic AI
Traditional CPU demand mostly tracks general server refresh. For datacenter CPUs entering the agentic AI phase, the main incremental demand is AI head-node CPU and agentic AI CPU.
### **AI head-node CPU
Even for earlier LLMs, head-node CPUs were already required for GPUs. The CPU must feed the GPU at extreme speed while handling communication, sync, and non-accelerated app work—weight loading, GPU mode config, KV-cache management, data validation, and more.
Without enough AI head-node CPU, GPU efficiency suffers: head-node bandwidth short → GPU waits for data → utilization drops. Hence the shift from HGX's classic 4:1 toward NVL72's 2:1—more complex accelerators need more CPUs to manage and feed them.
Dolphin's view: this mix is already rising to “spend CPU to wring every dollar out of the GPU”—pure economics, not a direct pull from Muse-style agentic AI.
### **Agentic AI CPU
Agents are essentially CPU workloads: they lean on sequential task execution, not pure parallelism—something GPUs can't own. In the agentic AI phase, agent CPU is “must-have,” and that is the pure incremental Muse unlocks.
ARM published an AI-agent workflow map:
① User → AI agent: state a goal, not a question;
② AI agent → Cloud: the agent body runs on cloud CPU. No accelerator at this layer—pure CPU.
③ Cloud → Agents: one agent splits work into many sub-agents (12 in the diagram). This step is the demand multiplier.
④ Agents ⇅ AI data center (loop): note the bidirectional arrows, one pair per agent—the “think–act–see” loop. Not one call; repeated round trips.
⑤ Inside the AI datacenter: orchestrating CPUs (agentic CPUs—new demand) take requests, hold state, schedule; feed inference slices to accelerators; accelerators emit tokens; back to CPU for the next decision.
Finally Agents → Cloud → Answer → user: results merge in the cloud into an “Answer” for the user.
Note: “orchestrating” here means the agent main-loop orchestration—splitting work, spawning sub-agents, holding state, deciding when to reflect / branch / retry, dispatching tool calls—not GPU job orchestration inside inference serving (that is the head-node CPU near the GPU).
Entering agentic AI, datacenters add many standalone CPU racks for agentic orchestration, scheduling, and management. To meet that, Nvidia previously launched standalone Vera CPU racks—256 Vera CPU chips, 88 cores each. That is the true agentic AI incremental.
In Nvidia's AI-factory framing, beyond Vera in Vera Rubin NVL72 (AI head node) and Vera CPU racks (standalone), Vera also appears in Vera BlueField storage for storage management, DPU, and KV-cache persistence.
Unlike BF-4 DPUs inside NVL72 (Grace CPU), BF-4 DPUs in Nvidia's STX storage racks use Vera CPUs.
In the STX storage rack reference design, each BF-4 includes one Vera CPU, two CX-9 NICs, and two SOCAMM modules. A full STX rack has 16 chassis (two BF-4 units each), so one rack means 32 Vera CPUs, 64 CX-9 NICs, and 64 SOCAMMs.
BF-4 (DPU) actually spans two CPU types: one inside the NVL72 compute tray (Grace), one inside the STX storage rack (Vera).
So whether BF-4 (DPU) counts as AI head-node vs. agentic CPU depends on whether the chip mainly serves the GPU cluster or agent workloads (persisting agent state, KV cache, long-context storage). STX leans more agentic.
CPU's “ChatGPT Moment”?
People often mix up CPU, cores, and threads. Quick framing:
If a CPU is a factory, one core is one worker—one core handles one instruction stream. More cores on a processor means more instructions in the same time—e.g. Nvidia Vera CPU (88 cores).
Threads are like conveyor belts. Factory A: 8 cores / 8 threads—one belt per worker; Factory B: 8 cores / 16 threads—two belts per worker; Factory C: 16 cores / 16 threads.
C is fastest (most cores). Between A and B, B is relatively faster. With the same worker count (cores), more threads shrink idle gaps on the belt so cores stay busy.
So more cores (workers) and more threads (belts) both show up as CPU performance gains.
Against the three CPU types above, demand drivers differ completely:
a) Traditional CPU: tracks enterprise IT and cloud request volume; the new piece is mainly third-party sites/services agents hit. Unit counts stay relatively flat; core counts move with product cycles; agent impact is limited;
b) AI head node: mainly driven by GPU count and rack mix, proportional to AI accelerators (GPU/ASIC). Cores per GPU for AI head nodes rose ~3.75× in recent years. Unpack that and the lift is mostly CPU:GPU mix (doubled), cores per CPU (+22%), and new DPUs inside NVL72 (+36%).
For head nodes, per-GPU core growth is already in the numbers; the next gen doesn't yet look like another huge step. Agentic AI doesn't move head-node CPU much.
c) Agentic CPU: the biggest incremental from agents. Unit count is not set by GPUs—it tracks concurrent agents.
For agent demand, Nvidia's single Vera CPU rack packs 256 Vera CPUs (88 cores each, two threads per core)—22,528 cores per rack, matching the claim of supporting 22,500+ sandboxes (isolated rooms to prevent contamination).
Note: Muse's VM is 2 vCPUs per user. 2 vCPUs = two SMT threads of one physical core. SMT threads share execution units, so 2 vCPU throughput is not 2× 1 vCPU—roughly 1.2–1.3×.
Agentic CPU demand splits into fixed and elastic layers: ① Fixed: one user, pre-allocated—like Muse's 2 vCPUs / fixed sandbox; ② Elastic: runs split-out sub-agents in reusable temporary sandboxes. Both users and sub-agents stay in isolated environments.
So Agent = model (GPU side) + fixed sandbox (user-side agent decomposition) + temporary sandbox (sub-agents). Flow: user issues a command → fixed-layer CPU decomposes the agent → temporary sandboxes handle sub-agents → inference runs on the GPU side.
Both fixed (user) and elastic (sub-agent) layers need secure isolation (sandboxes). CPU isolation comes from the ISA and hardware itself. From day one CPUs solved sharing one machine among mutually untrusted programs; decades of circuits exist specifically for sandboxing.
From prior math, one VR NVL72 (incl. DPU) has 4,320 CPU cores (= 72 GPUs × 60 cores per GPU), while one CPU rack hits 22,528 cores (256 Vera × 88)—about 5× a VR NVL72.
Where Does Core-Count Growth Come From?
ARM's outlook: traditional AI datacenters need ~30M CPU cores per GW of compute; in the agentic AI era that jumps to 120M.
Market talk mixes GPU:CPU ratios, cores per GPU, and cores per GW. Dolphin's take: GPU:CPU is mainly a head-node story and matters less in the agentic phase. Cores per GPU will jump once CPU racks arrive, so cores per GW is the more meaningful metric.
① Traditional AI datacenter: take GB300 NVL72—rack power ~130 kW. If 1 GW were all GB300 NVL72, that's ~7,692 NVL72 racks.
One GB300 NVL72 has ~2,592 CPU cores (72 GPUs × 36 CPU cores per GPU), so 1 GW needs ~20M CPU cores (below ARM's 30M);
② Agentic AI datacenter: VR NVL72 and CPU racks both rise toward ~200 kW, so 1 GW is ~5,000 NVL72 racks, then how many of those slots go to CPU racks.
Today is still mostly pure GPU racks; the conservative / base / bull cases in the table below all imply pure incremental from CPU racks.
So within a single GW, ARM's 120M-core outlook sits above the theoretical ceiling of ~113M cores in the chart above.
Dolphin's reasonable estimate: with CPU racks, a more realistic lift in datacenter core demand is about 2–3×.
Nvidia CPU racks are pure CPU—no GPU. Introducing them in the agentic phase sharply raises cores per GPU across the AI datacenter (driven by CPU-rack share). In the base case (CPU racks = 25% of racks), cores per GPU can rise from ~60 to ~160—that is the direct CPU incremental.
How Much Demand Can Agentic AI Add?
As above, AI datacenters are already being built, and head-node demand rises with them (not new news). Muse-style agentic AI mainly adds CPU racks under agent demand—the main agentic CPU source. That is independent of GPUs and tracks agent users.
Formula: incremental CPU demand = (A head node) GPU shipments × head-node cores per GPU + (B fixed) registered users × DAU rate × (active hours/24) × 2 vCPUs per user × peak/avg ÷ oversell ÷ 2 (two threads per core) + (C elastic) active users × task concurrency × cores per task sandbox
When sizing agentic AI, focus on B (fixed) and C (elastic):
① B fixed: still user-side, but Muse's 2 vCPUs don't hold cores forever. a) When active, the user gets 1 physical core (2 threads); b) when idle, state goes to storage → VM pauses → physical core frees for others, then restores quickly on wake.
Muse's 2 vCPUs guarantee you can use up to 2 vCPUs, but idle use doesn't burn physical cores. One physical server (256 cores) can allocate hundreds or thousands of vCPUs because most VMs are idle at once—that's oversell.
At 100M Muse users, assume 30% DAU, 3 active hours/day, peak/avg 2, oversell 6 → B fixed needs ~1.25M CPU cores (= 100M × 30% × 3/24 × 2 × 2 / 6 / 2).
② C elastic: doesn't cover all users—only tasks from active users. Task concurrency = average sandboxes (sub-agents) per active user; also watch cores per task sandbox.
Still at 100M Muse (30% DAU), 3 hours/day, peak/avg 2, assume 4 CPU cores per task sandbox, and scenario on concurrency: peak active = 7.5M (= 100M × 30% × 3/24 × 2)
At 100M Muse users and the assumptions above, 100% concurrency (one sub-agent per active user on average) needs about 1 GW of factory. With 25% CPU architecture, that's ~1,332 CPU racks—C elastic alone hits ~30M CPU cores.
In that (base) case: A (GPU-side) ~17.28M cores (= 288k GPUs × 60 cores per GPU); B (user-side) ~1.25M; C (sub-agent) ~30M.
A+B+C ≈ 48.53M cores—about 2.8× the pure-GPU-rack plan (~17.28M). Adding Nvidia storage racks (DPU) and other extras, Dolphin estimates agentic CPU could lift core demand toward ~3× prior levels.
Opportunities Across the CPU Chain
AMD once guided server CPU market to ~$220B by 2030 at >50% CAGR—implying ~$29B in 2025.
Formula: server CPU market = total cores × ASP per core. From above, agentic CPU can 3× core demand; also assume ASP per core rises 18%/year.
Most vendors guide to 2030. Assume agentic CPU modes land at scale by 2030 (core demand 3×) and ASP/core +18%/year → 2030 market ~$200B (= 290 × 3 × 1.18^5), CAGR near 50%.
That lines up with vendor outlooks. ARM management has also said its prior $100B view was conservative; ~$200B by 2030 is now roughly consensus for the CPU chain.
Today's server CPU market is mainly Intel and AMD. On 2025 related revenue, shares are ~58% Intel and ~35% AMD, with others under 10% combined.
Because market growth is driven by agent CPU while traditional server CPU grows slower (mostly refresh), Dolphin expects Intel's share to keep slipping while AMD, Nvidia, and peers regain—especially Nvidia, ARM, and Qualcomm as pure incremental.
Assume by 2030 Intel and AMD each hold ~36%. As “new entrants,” Nvidia / ARM / Qualcomm get ~15% / 5% / 2%.
In total, agentic AI could add ~$55B and ~$62B of annual server-CPU revenue for Intel and AMD. Nvidia, ARM, and Qualcomm could see ~$30B / $10B / $4B of pure annual incremental from agent CPU.
That roughly lifts 2030 Intel/AMD server-CPU revenue expectations by $10–20B vs. prior—about a ~10% bump to 2030 earnings power.
Nvidia's ~$30B agent-CPU revenue would be under 3% of total company revenue—limited company-wide impact.
Among pure incremental, agent CPU matters more for ARM and Qualcomm. At 5% and 2% share in 2030, that's ~$10B and ~$4B revenue. At 25% OPM on agent CPU, ~$2.5B and ~$1B of 2030 operating profit uplift.
Overall, agentic AI needs CPU racks and directly raises core demand—supportive for the whole CPU chain. Structurally, incremental demand is mostly agent CPU and head-node CPU; traditional server CPU stays flat.
Ordering of agent pull across the chain: ARM > Intel/AMD > Qualcomm > Nvidia—also visible in last week's price moves. Muse-driven agent-CPU demand is largely in the stock; watch whether agent demand keeps growing the pie (server CPU market) and who captures more incremental share.