An analyst I respect posted something last week that most crypto traders will scroll past. Jukan, who covers the semiconductor space for Citrini, published a short note on August 8 arguing that memory prices peak within two quarters and that NVIDIA's Rubin Ultra platform is quietly reducing its per-GPU HBM configuration. Two sentences. No blockchain mention. No token tickers. Zero engagement from our corner of the internet.
That is a mistake, and I will tell you why.
That note describes a re-wiring of the exact hardware assumptions your AI token portfolio is built on. Whether you hold GPU-scarcity tokens, decentralized storage coins, or AI agent plays, the profit pool underneath those narratives is moving. Not shrinking. Moving. From local memory stacked onto a single chip to pooled memory stitched together across racks by optical interconnects.
I have spent enough time on the ground — surviving the 2018 ICO purge, the DeFi Summer minefield, the Terra collapse — to know that narratives detach from fundamentals months before they detach from price. Jukan's note is a fundamental signal. Let's break down what it actually says and what it means for your holdings.
The Context
First, the technical backdrop. For the last three AI buildout cycles, the playbook was simple: buy the hottest GPU, stuff it with as much High Bandwidth Memory as possible, and charge accordingly. HBM sits directly beside the compute die, delivering the bandwidth that large language models need to keep tens of thousands of GPUs busy. SK Hynix, Samsung, and Micron turned HBM into the most valuable memory product in history. CoWoS advanced packaging became the gatekeeper, and whoever controlled that stack controlled the AI supply chain. The industry called this the memory wall, and for two years every scaling roadmap was really just a plan to build a taller wall.
NVIDIA's Rubin Ultra changes that equation. According to Jukan's read — and consistent with independent industry signals — Rubin Ultra leans toward a reduced HBM configuration per GPU, compensating at the system level with aggressive optical interconnect linking multiple racks. Instead of one GPU with maximum memory density, you get clusters where memory is pooled across machines. The GPU is no longer the compute island. The rack, or the cluster, becomes the new unit of AI performance.
For anyone who has lived through distributed systems, this pattern is familiar. It is the same philosophical fork that split the blockchain world years ago: monolithic architectures that try to do everything in one place versus modular stacks that specialize and connect. Ethereum's rollup-centric roadmap, Celestia's data availability layers, the entire multi-chain thesis — that is what NVIDIA is doing at the silicon level. Pooling resources, fragmenting workloads, and building optical bridges between islands.
Jukan's credentials matter here. Citrini's semiconductor desk has a track record of catching memory cycle turns before the crowd does. When an analyst with that history flags a top within two quarters, the market usually listens — eventually. The gap between the analyst note and the market's adjustment is exactly where traders get caught, and exactly where patient communities build advantage.
Why should a crypto trader care? Because the token market has not repriced for this yet. Bittensor rewards GPUs doing training and inference. Render and io.net price compute by the hour based on scarcity. GPU tokens are still priced in the old assumption: a single GPU's HBM count is the measure of value. If that measure shifts to cluster-level interconnects, the economics underneath those tokens shift with it.
What the HBM Cut Actually Tells Us
First, let me be precise about what "reduced HBM configuration" means. It does not mean NVIDIA is abandoning memory. It means treating memory as a network resource rather than a chip attachment. Pooled memory via low-latency optical interconnect lets multiple GPUs access a shared memory fabric. In practice, that is distributed shared memory at data-center scale — a design pattern supercomputing has used for decades but never hit commercial AI economics. Jukan's note says that day is arriving. For investors, that requires unlearning the habit of measuring AI power by HBM gigabytes glued to a single chip.
Why now? Three reasons, in my view.
The first reason is supply risk. HBM remains structurally tight. Ramping HBM3E and HBM4 requires TSV, advanced hybrid bonding, and yield learning curves that have been brutal for every memory maker. One flawed product generation and you burn billions. NVIDIA can reduce that dependency by shifting the performance burden to optics, which have more suppliers and faster capacity expansion.
The second reason is cost. HBM pricing has climbed with AI demand. From my audit work in this sector, memory is the fastest-rising line item in AI server bills of materials. A platform that cuts HBM per GPU while maintaining cluster performance is a platform protecting its gross margins. Jukan is right to read this as NVIDIA sending a soft price signal to its memory partners: do not chase the premium forever.
The third reason is the most interesting. The optics pivot is a moat-building move. NVIDIA already owns NVLink, InfiniBand, and the Quantum switch family. If Rubin Ultra becomes a system-level product where optics do the heavy lifting, NVIDIA locks in the full stack. Competitors who cannot match the interconnect layer end up stranded at the single-GPU level.
The packaging wrinkle is worth naming, too. Reduced HBM per GPU eases pressure on CoWoS capacity — the same 2.5D packaging that has been the AI supply chain's narrowest bottleneck. But co-packaged optics and silicon photonics introduce their own packaging challenges. The constraint does not disappear; it migrates. That is a critical nuance for anyone trying to identify the next bottleneck before the market does.
This matters for token communities because the "GPU scarcity" narrative is being deliberately softened by the most powerful buyer in the world. The same narrative underpins a dozen compute-token projects. NVIDIA is telling you the scarcity premium has peaked. Trust the hands, not just the charts.
Two Quarters to Peak
Now the cycle. Jukan's call is that memory prices top within two quarters. Sit with that. The memory cycle historically runs three to four years. If this upcycle peaks in early 2026, we are looking at one of the shorter AI-driven memory booms on record. That tells me something important: the AI buildout is being engineered for efficiency gains, not indefinite hardware scarcity.
Here is the part crypto traders understand better than traditional equity traders. Jukan's note references Korean leveraged ETFs that suffered an unwind, with LP redemptions adding sell pressure to memory equities. Retail reads that as "memory demand is collapsing." Smart money reads it as a capital structure event — a levered instrument breaking, not the underlying demand breaking.
We have watched this movie. In May 2022, when Luna and UST collapsed, the instant reaction was "stablecoins are dead, DeFi is dead." But the underlying protocols with real revenue and real users survived, and many later thrived. The collapse was a capital structure failure, not a demand failure. Same logic applies here. A leveraged ETF unwinding does not cancel a single HBM order from a hyperscaler. It just means someone borrowed too much against a volatile asset. My community survived 2022 by learning to separate price mechanics from fundamental reality. That discipline is exactly what this moment requires.
There is also a hidden feedback loop worth naming. If the consensus becomes "memory prices peak in two quarters," memory makers may hold back on capacity expansion. That could keep supply tighter for longer, producing a peak that does not crash. This is why my longer-term view stays structurally constructive even while the short-term picture wobbles. Short-term pain from position unwinds; long-term strength from real AI order books. The middle is where most people lose money.
And there is a timing trap on top of that. HBM capacity expansion, from equipment move-in to volume production, typically takes twelve to eighteen months. That means real supply release lags the price-peak consensus by at least two quarters. If Jukan is right about the peak but capacity arrives late, we get a strange outcome: prices peaking on paper while physical supply stays tight. Understanding that lag separates readers who nod at headlines from traders who actually manage risk.
Mapping the Value Shift to Tokens
Now let us map this to tokens carefully, because narratives get slippery around hardware cycles.
The value chain is splitting into two pools. One is shrinking relative to expectations: local memory attached to GPUs. The HBM scarcity premium, which has bled into everything from GPU tokens to crypto AI infrastructure wings, is going to compress. If you hold tokens whose entire thesis is "GPUs are scarce and memory is scarce," re-examine that thesis. Based on my audit experience inside copy-trading communities, I can tell you most retail portfolios are overweight this exact narrative.

The second pool is expanding: optical interconnect and pooled memory infrastructure. Silicon photonics, co-packaged optics, optical engines, high-speed transceivers. The total addressable market grows every time a rack needs to talk to another rack at speeds copper cannot sustain. The commercial supply chain — Broadcom, Marvell, Coherent on one side, aggressive Chinese module makers on the cost side — is already shifting. But the more interesting development for our space is structural: it strengthens the argument for decentralized compute networks.
Think it through. If the unit of AI performance becomes the interconnected cluster, then geographically distributed GPU clusters with high-bandwidth optical links become economically viable for training and inference workloads. That is the DePIN thesis — Render, Akash, io.net and others — but with a twist. The bottleneck moves from "who owns the most HBM" to "who provides the best network fabric." Decentralized networks that can prove low-latency, verifiable interconnection gain a stronger value proposition than they had three months ago.
Storage tokens deserve attention here too. The shift from local memory to pooled memory is a shift from owning more hardware to accessing shared resources through a network. That is the founding ethos of Filecoin, Arweave, and the data availability layers. If NVIDIA is adopting network-attached memory as its architecture, the old criticism of decentralized storage — that it is slower than local disks — becomes less relevant. The whole industry is moving toward network-dependent performance models. Community first, coins second. Always. That means we evaluate storage projects on their ability to serve a networked memory era, not on their ability to mirror obsolete local-storage economics.
Signals I'm Watching
What am I actually watching on the ground? Three signals.
The first is HBM order books. The real test of Jukan's thesis is whether hyperscaler purchase orders for HBM keep growing even after Rubin Ultra's final configuration is revealed. Strong orders mean the HBM reduction is a system-design choice, not a demand warning. Paused orders mean something darker for the whole AI complex.
The second is optical module procurement. 800G and 1.6T transceiver order data is the cleanest leading indicator for the interconnect shift I know. When module makers report 1.6T ramping into hyperscaler data centers, the narrative is confirmed by actual money movement. Follow the people, follow the profit — and right now, the people placing the biggest orders are buying optics, not stacking more HBM.
The third is token-level revenue signals. For compute DePIN projects, I am looking at whether utilization and revenue per GPU respond to cluster-level demand rather than single-GPU scarcity. A project demonstrating revenue growth from distributed, interconnected compute is positioned for the new architecture. A project still pricing scarcity is a short thesis waiting to happen.
I want to be careful here. I am not telling anyone to dump memory-adjacent tokens and buy optical names. That binary thinking is exactly how retail loses to smart money. I am asking you to inspect the assumptions underneath every AI-related token in your portfolio. What is the hardware thesis? Does it assume per-GPU memory scarcity? Or does it assume network-level performance? The answer determines whether you are holding an asset aligned with the next architecture — or a relic of the last one.
The Contrarian Read
Here is where I will frustrate people on both sides. The bearish read — "HBM is doomed" — is just as wrong as the naive bullish read. NVIDIA is not abandoning memory. It is renegotiating its cost structure while capturing the interconnect narrative from would-be competitors. That is not a bearish story for the AI infrastructure complex. It is a power-shift story. Aggregate memory demand across data centers is still growing. What changes is who captures the premium and how the performance narrative gets told.
And here is the genuinely contrarian crypto angle: the market is treating this as bad for memory suppliers and good for optics suppliers. But the biggest beneficiary may be the decentralized compute layer — the projects ridiculed for being "too distributed" to compete with centralized AI clouds. If AI performance becomes a network property rather than a chip property, the DePIN thesis strengthens rather than weakens. Meanwhile, the centralized operators running HBM-stuffed superclusters are the ones carrying the heaviest depreciation burden when the memory cycle turns. That is not an opinion; that is the accounting reality of owning the most expensive inventory at the top of the cycle.
One more thing nobody is saying. Software eats the memory wall too. If NVIDIA reduces HBM per GPU, inference workloads will need algorithmic efficiency — quantization, distillation, sparsity, speculative decoding. AI agent tokens that genuinely improve inference efficiency, rather than the ones simply slapping "AI" on a ticker, become more valuable when memory per GPU gets dearer. The market has not priced that distinction yet. It will.
And keep the Korean ETF unwind in perspective. That event is being used as evidence of a broken memory narrative. It is not. It is a leverage cleanup. The same crowd that screamed "DeFi is dead" in 2022 is now screaming "memory is dead" in 2025. Both are wrong for the same reason: they confuse a capital structure event with a fundamental demand event. We have been trained by this market to spot that difference. Use it. The data will tell you who is right, but only if you are watching the right data.
The Takeaway
Two quarters. That is the window Jukan gives memory prices. Use it to stress-test your AI token thesis before the market does it for you.
The architecture of AI is moving from chip-level dominance to network-level coordination. The token market is still pricing the old architecture. That gap is where the next cycle's winners and losers get decided. I cannot tell you which token wins; that is not my job. My job is making sure we ask the right questions. Where is the profit pool flowing? Whose hands are moving real product volume? Get that right, and price eventually follows. Trust the hands, not just the charts.