
Bypassing the Memory Wall — $CBRS and $QCOM: Two Listed Ways to Route Around HBM
Internal Deep Dive · Published: 2026-07-27 · Tickers: $CBRS, $QCOM
The fastest-growing constraint in AI isn't compute — it's the speed of feeding it. In inference's decode phase, generating each token means dragging the whole model back out of memory, and memory bandwidth, not FLOPs, sets the pace. Every dollar HBM's price climbs raises the payoff for designing around it — and there are now exactly two public, liquid ways to own that trade: $CBRS, which puts the memory inside the chip, and $QCOM, which stacks cheap memory onto it. They sit at opposite ends of the same bet and, not coincidentally, at opposite ends of the risk spectrum.
Thesis
The memory wall is an access-speed problem, and HBM — the incumbent fix — is expensive, supply-constrained, and made by three companies. That premium is the catalyst: the pricier HBM gets, the more valuable any architecture that needs less of it. Two listed names express the "route around HBM" trade from opposite corners:
- $CBRS (Cerebras) — the SRAM-centric extreme: 44GB of memory printed on a wafer-scale chip, no HBM at all. The purest expression, the fastest decode, and the most concentrated risk (~80% of a $24.6B backlog is one customer, OpenAI).
- $QCOM (Qualcomm) — the stacked-LPDDR approach: cheap mobile memory stacked on top of the accelerator via TSVs (its HBC architecture). A ~$188B cash machine at ~14× earnings, where the entire data-center effort is a free call option.
- The one risk that governs both: if HBM prices ever break, both trades lose urgency together.
Backdrop — How the Memory Wall Became the Trade
Start with the physics, because the whole basket rides on it. Training a model is a one-time job; inference is the standing one — every query, every agent step, burns cycles forever. And inference's decode phase has a peculiar bottleneck: to generate a single token it must read the entire active weight set back from memory. Produce ten tokens a second and you're pulling terabytes per second, per user, while the expensive compute sits mostly idle. As @JKeynesAlpha relayed Cerebras CEO Andrew Feldman putting it this month: every token forces the weights out of memory and into compute, and "the weights are the intelligence." No amount of FLOPs helps if the pipe from memory is too narrow. That pipe is the memory wall.
HBM — stacked DRAM on a silicon interposer — is the industry's answer, and a good one: it's what drove the DRAM makers' record cycle. But it carries three taxes: interposer/advanced-packaging cost, power per stack, and, above all, supply — only a few lines on earth make it, so when demand piles in, price doesn't hold. Boil two years of DRAM earnings calls into one word and it's sold out. The rule in this neighborhood is that the more successful the answer, the thicker its price tag — and a thick enough price tag is, to an engineer, an invitation to route around it.
That's the demand shift. Through 2026 it stopped being a conference-paper idea and showed up as shipping hardware and real contracts: Qualcomm put HBM-free silicon on an annual data-center roadmap, and a wafer-scale startup that wants HBM gone went public at one of the year's largest IPOs. When the incumbent fix gets this expensive, the market pays up for the escape routes — the only question is which routes are real, and what you're paying for each.
Mechanism — Two Roads Around the Wall
There are really only two ways to beat a slow round-trip to memory: shrink the trip, or make the memory cheap enough that you can afford a lot of it close by. $CBRS and $QCOM take one road each.

$CBRS shrinks the trip to zero — memory inside the die. Cerebras doesn't dice its wafer into hundreds of chips; it uses the whole wafer as one processor (the WSE-3), which lets it print 44GB of SRAM directly onto the silicon, delivering ~21 petabytes/second of on-chip bandwidth. SRAM is the fastest memory there is and it sits millimeters from the compute, so there's essentially no round-trip. For models too big to fit, weights stream in from an external "MemoryX" appliance. The payoff is blistering decode: independent benchmarks put the WSE-3 at ~2,522 tokens/second per user on Llama 4 Maverick — roughly double NVIDIA's published B200-system figure. The tradeoff is the whole story: SRAM is unmatched on speed and tiny on capacity, so this design wins ultra-low-latency decode and gives up flexibility. There's a second-order tell here worth holding onto — as Damnang notes, SRAM-centric compute is silicon printed in a logic fab, so the more this camp grows, the more memory demand quietly shifts from the DRAM makers to the foundry.
$QCOM makes the memory cheap and stacks it close — LPDDR onto the die. Qualcomm's High Bandwidth Compute (HBC) skips HBM and the interposer entirely: it stacks LPDDR — the low-power memory it perfected over 20+ years in phones — directly on top of the accelerator through TSVs, so data drops vertically instead of crossing sideways. LPDDR is cheaper and lower-power than HBM and you can stack a lot of it: Qualcomm claims 768GB per card and 133TB/s of "effective" bandwidth on the AI250, an 18× jump over the LPDDR5X AI200. In the inference-chip taxonomy, this is the "route around HBM with cheap DRAM" camp — the same instinct as d-Matrix's LPDDR5X or Tenstorrent's GDDR6, which is telling, because Qualcomm was reported in June to be circling Tenstorrent ($8–10B) on top of building HBC. Both roads exploit the same wrinkle Nvidia is now attacking from inside its own stack (Rubin's 3-bit weights): decode is bandwidth-bound, so the winner is whoever serves that workload more cheaply than HBM does.
Two roads, one destination — and neither is trying to win frontier training, where HBM stays king. This is a fight for the fast-growing, cost-sensitive, latency-hungry middle of inference.
Basket & Positioning
Same trade, opposite risk profiles. One is a pre-profit pure-play priced for a backlog to convert; the other is a profit machine where the theme is upside you're not paying for.
| Node on the memory map | Name | Memory architecture | The bet | Theme exposure | Rough valuation |
|---|---|---|---|---|---|
| Memory inside the die (SRAM-centric) | $CBRS | 44GB wafer-scale SRAM + MemoryX streaming; no HBM | Fastest decode, ultra-low latency | ~Pure-play — but ~80% of backlog is one customer | ~$42B cap · ~50× FY26 rev · pre-GAAP-profit |
| Memory stacked onto the die (cheap-DRAM) | $QCOM | LPDDR on accelerator via TSV (HBC); no interposer | Cheap, capacity-friendly inference | Data center ≈ 0% of revenue today — a call option on a $44B business | ~$188B cap · ~14× non-GAAP EPS |
$CBRS — the purest bet, and the most concentrated. The numbers finally caught up to the technology. Q1'26 revenue hit $193.4M, up 94%, split between hardware ($110.6M, +59%) and a cloud/services line that jumped 178% to $82.8M — and the company ran essentially breakeven (core net loss just $2.5M; GAAP loss $14.0M on stock comp). FY25 revenue was $510M (+76%). The moat is real and nearly uncopyable: nobody else ships a working wafer-scale processor. The catch is who's buying it. UAE-linked entities (G42 + MBZUAI) were ~86% of 2025 revenue; the January OpenAI deal — $10B+, 750MW of inference through 2028 — diversifies the name but not the concentration, since it's now ~80% of a $24.6B backlog. As one analyst put it, concentration rotated; it didn't go away. Watch how fast that backlog converts to recognized revenue, and watch the margin: management guided full-year core gross margin down to 38–41% from Q1's 46.5%, which is what knocked the stock.


Positioning matches the profile: a 10-week-old IPO with violently wide dealer walls (call wall $340, +80%; put wall $115, −39%) and near-zero net GEX — an un-pinned, high-volatility name that will trade on backlog-conversion headlines and lockup supply, not levels.
$QCOM — the theme as a free option on a cash machine. Here the memory-wall trade comes wrapped in a business that already prints money: FY25 revenue $44.3B (+14%), non-GAAP EPS $12.03, chip-segment revenue $38.4B (+16%). (GAAP EPS was a scary-looking $5.01, but that's a one-time $7.1B tax charge; operating income was a healthy $12.4B.) At $170 that's ~14× earnings — cheap enough you double-check the ticker — because the market prices in Apple in-sourcing its modem and handset cyclicality. What it doesn't price is the data center, which is ~0% of revenue today. The June Investor Day put a number on the ambition: >$15B of data-center revenue by FY29, with the ramp now pulled forward to FY27. And it's not a slide-deck fantasy — Qualcomm closed Alphawave (the 800G/1.6T SerDes/optical connectivity in the Dragonfly line) and acquired Modular, bringing in Chris Lattner (of Swift and Google's TPU compiler stack) to fix the software problem that kills most accelerator challengers. CFO Akash Palkhiwala flagged the economics plainly:
"We also expect custom silicon gross margin to be slightly below our overall Qualcomm gross margin, but it'll be accretive at the operating margin level."
Akash Palkhiwala, EVP & CFO, Qualcomm Investor Day (6/24/26)


Dealer positioning frames the range cleanly: put wall $150 (−12%) as support, max pain $190, call wall $220 (+29%) — a name that grinds rather than gaps, with the data-center ramp as the catalyst that could walk it toward that upper wall.
Risks & What Breaks It
- The whole trade is short HBM's price (both names). These architectures exist because HBM is expensive. If DRAM prices break — a capacity glut, a demand air-pocket — the incentive to stack LPDDR or print SRAM softens, and both theses lose urgency at once. This is the single assumption underneath everything; it's also the least likely to break before 2027 given how sold-out the DRAM makers are.
- $CBRS customer concentration is the thesis-killer. ~80% of a $24.6B backlog is OpenAI; the rest still leans on UAE entities. A renegotiation, a slip in OpenAI's own capex cadence, or a CFIUS-style complication (the 2024 IPO was pulled over exactly this) hits revenue disproportionately. This gets the most airtime for a reason — it's the difference between a hyper-growth compounder and a single-contract story.
- $CBRS margin + supply, and a fresh IPO's mechanics. Guiding core gross margin down to 38–41% says pricing power is capped while it scales; a wafer-scale part is also a yield and packaging bet. Layer on post-IPO lockup expiries and heavy insider Form-4 activity, and the near-term supply of stock is a real overhang independent of the business.
- $QCOM's core can wobble while the option ripens. The data-center payoff is FY27+, and in the meantime the handset franchise faces Apple's modem in-sourcing and memory-cost inflation (Qualcomm just notified OEMs of double-digit price hikes as sub-$100-phone memory BoM spiked). The option is close to free precisely because the base business carries known risks.
- Both are pre-benchmark on the load-bearing claim. Qualcomm's 133TB/s is a vendor-defined "effective" figure with no third-party benchmark yet, and Cerebras's backlog is signed but not yet recognized. What flips the memo: verifiable HBC benchmarks + a second marquee HBC customer disclosed in Microsoft's own words (not just a vendor-stage intro) for $QCOM; the pace of backlog-to-revenue conversion and any margin stabilization for $CBRS.
Valuation & House View
The two names force a genuine choice between cheap-with-optionality and expensive-with-purity — which is why owning both, sized differently, is the honest answer.
$QCOM trades at ~14× trailing non-GAAP earnings and ~13× forward — a mid-teens multiple on a business growing low-double-digits, versus AI-semi peers at 30–40×. You are paying handset-cyclical prices for a company handing you the data-center ramp for free.
| $QCOM scenario | Prob. | 2Y target | Implied | Trigger |
|---|---|---|---|---|
| Bull | 30% | $255 | +50% | HBC benchmarks land, FY27 data-center ramp visible, re-rate toward 18× |
| Base | 50% | $205 | +21% | Core holds, data center begins contributing, modest re-rate to ~15× |
| Bear | 20% | $135 | −21% | Apple modem loss bites, data center slips to FY28, multiple stays ~11× |
$CBRS is a different instrument — ~$42B cap on ~$860M of FY26 core revenue (~50×), pre-GAAP-profit, priced for the OpenAI backlog to convert on schedule. The technology moat justifies a premium; the concentration and margin trajectory justify sizing it small.
| $CBRS scenario | Prob. | 2Y target | Implied | Trigger |
|---|---|---|---|---|
| Bull | 30% | $300 | +59% | Backlog converts fast, OpenAI Sol scales, margin stabilizes, customers diversify |
| Base | 45% | $200 | +6% | Revenue compounds ~60–70%, margin drifts to guide, concentration lingers |
| Bear | 25% | $95 | −50% | Backlog conversion slips or a key customer wavers; lockup supply + de-rate |
House View: own the theme through both, weighted to the cash machine. $QCOM is the risk-adjusted way to be long the memory-wall trade — a profitable, cheap franchise where the data center is asymmetric upside you're barely paying for; start a position and add as HBC benchmarks and a real hyperscaler contract confirm. $CBRS is the high-conviction pure-play — the deepest technology moat on the board and a credible path to profitability, but a name to size small and accumulate on backlog-conversion proof and margin stabilization, not on the OpenAI headline alone. The elegant part: they hedge each other. If SRAM-centric designs win, $CBRS captures it directly and $QCOM barely notices; if cheap-DRAM stacking wins the cost-sensitive middle, $QCOM's free option prints. Either way, someone is routing around the wall.
Figures independently verified against SEC filings ($QCOM FY25 10-K, $CBRS Q1'26 10-Q), the Qualcomm Investor Day transcript, and live market data. h/t PhotonCap and Damnang for the memory-wall framing, and P Equity Research for the accelerator-cycle backdrop. See Three Routes Around the Memory Wall and Accelerators, ABF Substrate & InP.