As smartphones get smarter with on-device AI, Nvidia is building the massive cloud infrastructure that powers the world’s largest AI models. Here’s why both approaches matter.
Anamika Dey, editor
By TechSun News Desk | techsunnews.com | July 20, 2026 | Tech / AI / Trending | 8 min read
Speaking in Tokyo on July 15, Nvidia CEO Jensen Huang confirmed that the company’s next-generation Rubin AI chips are already in production and heading toward what he called “giant” volumes, brushing aside earlier reports of manufacturing trouble with a specialized circuit board. It is a small comment with a big backdrop: Rubin is the platform Nvidia is betting will power the next phase of the AI boom, and the world’s largest cloud providers are already lined up to use it.
If that sounds like it belongs in the same conversation as our recent explainer on Edge AI, it does — but not in the way the headlines might suggest. Rubin has nothing to do with the AI running on your phone. It is built for the opposite end of the spectrum: the data centers doing the heaviest lifting in AI. Understanding that distinction is the key to understanding where AI is actually headed.
What Is Nvidia Rubin?
Rubin is Nvidia’s newest AI computing platform, first unveiled at CES 2026 and now, according to Nvidia’s own newsroom, in full production with seven integrated chips designed to function as a single AI supercomputer. The lineup includes a new Vera CPU, Rubin GPUs, and a set of networking and storage chips — Nvidia calls the approach “extreme codesign,” engineering every piece together rather than selling chips separately.
The headline number is cost. Nvidia says Rubin delivers up to a 10x reduction in the cost of generating each AI token compared with its current Blackwell chips, and can train large “mixture-of-experts” models using roughly four times fewer GPUs. In plain terms: the same AI query that costs a company a certain amount today could cost a fraction of that once Rubin is fully deployed.
This is not aimed at consumers. Rubin’s confirmed first customers are the companies that run the internet’s AI backbone — Microsoft, AWS, Google Cloud and Oracle Cloud Infrastructure are named as among the first to deploy Rubin-based systems, alongside traditional server makers Dell, HPE and Lenovo, who are building it into their own enterprise hardware. Microsoft has said it will use Rubin systems in its next-generation “Fairwater” AI data centers as part of its cloud buildout.
Why Rubin Isn’t Edge AI
It would be an easy mistake to lump this in with the on-device AI trend — both stories are technically about “new AI chips.” But they sit at opposite ends of the same industry.
Edge AI, which we covered in our recent guide, means AI models running directly on your phone, laptop or watch — small, efficient, and built to fit inside a battery-powered device. Rubin is the reverse: chips built for data centers the size of warehouses, consuming enormous amounts of power, designed to train and run the largest, most demanding AI models that could never fit on a phone. Nvidia does not make phone chips at all; that market belongs to Apple’s Neural Engine, Qualcomm’s Hexagon NPU, and Google’s Tensor silicon, all covered in our GPU vs NPU vs TPU explainer.
So when a headline connects Nvidia to “the future of edge AI,” read it carefully — Rubin is not going into your next phone. It is going into the data centers your phone’s AI features sometimes quietly call back to.
Cloud AI vs Edge AI: Where Rubin Actually Fits
| Factor | Cloud AI (Nvidia Rubin) | Edge AI (phone/laptop chips) |
|---|---|---|
| Where it runs | Data centers and AI factories | Your device, on-hand |
| Built by | Nvidia, for hyperscalers | Apple, Qualcomm, Google, etc. |
| Scale | Warehouse-sized server racks | A single chip in your pocket |
| Power use | Extremely high, industrial-scale | Battery-constrained, low-power |
| Best for | Training and running the largest models | Fast, private, everyday tasks |
| You interact with it | Indirectly, through apps like ChatGPT | Directly, on your own device |
For the fuller breakdown of the edge side of this table, our What Is Edge AI? guide covers Apple Intelligence, Copilot+ PCs and Gemini Nano in detail.
Why Both Will Keep Growing Together
It’s tempting to treat this as a competition — edge versus cloud, phone versus data center. In practice, the two are growing side by side because they solve different problems.
- Phones handle the fast, private, everyday work — photo edits, live translation, quick replies — the moment-to-moment tasks people want to feel instant.
- The cloud handles what a phone never could — training frontier models in the first place, and running the heaviest reasoning tasks that need far more computing power than any device could hold.
Nvidia’s own technical team frames this as the industry moving into an “industrial phase,” where always-on AI factories continuously convert power and silicon into intelligence at scale — the exact opposite job description from a phone’s NPU, but a necessary partner to it. Every on-device feature that feels lightweight today, from Face ID to live captions, exists because a cloud-trained model was built somewhere first.
What This Means for Businesses and AI Costs
The 10x inference cost reduction Nvidia is promising matters well beyond Silicon Valley balance sheets. Independent analysis of Rubin-era pricing suggests that a task costing roughly $0.003 per 1,000 output tokens on today’s Blackwell-based infrastructure could fall to $0.0003–$0.0006 once Rubin capacity comes online — a drop that would ripple down into cheaper AI subscriptions, cheaper customer-support bots, and cheaper AI features baked into ordinary software.
That matters because demand has been outpacing supply. Microsoft told investors in a Q1 2026 earnings call that it was facing an AI compute capacity shortage affecting its entire fiscal year, and a Flexential industry report found that roughly 80% of organizations are now planning their AI data center needs a full year in advance. Rubin is Nvidia’s answer to that squeeze — more computing power per watt, and a lower cost for every AI response generated. First deployments are expected in the second half of 2026, so the effect on everyday AI pricing will likely show up gradually rather than overnight.
| A QUESTION FOR YOU
Do you think about where your AI queries are actually processed — a phone chip or a data center on the other side of the world — or does it not cross your mind as long as the answer shows up fast? Tell us in the comments. |
Nvidia Rubin FAQ
1. Is Nvidia Rubin an edge AI chip?
No. Rubin is built for data centers and “AI factories” run by cloud providers like Microsoft, AWS, Google Cloud and Oracle, not for phones or laptops. Edge AI chips — the ones that run AI directly on your device — come from companies like Apple, Qualcomm and Google instead.
2. When will Nvidia Rubin actually be available?
Nvidia says Rubin is now in full production, with partner deployments from cloud providers expected in the second half of 2026. As a data-center platform, it won’t appear as a product you buy directly — you’ll experience its effects indirectly, through cheaper or faster responses from the AI apps and services that run on it.
3. Will cheaper AI chips make on-device AI less necessary?
Not really — they solve different problems. Cheaper cloud inference makes heavy AI tasks more affordable to run at scale, but it doesn’t remove the reasons people want edge AI: working offline, keeping data on the device, and getting instant responses without a network round trip. Expect both to keep improving in parallel rather than one replacing the other.
| EDITOR’S OBSERVATION
It’s easy to read every new AI chip announcement as part of the same story, but Rubin and edge AI are answers to two different questions. Edge AI asks: how do we make AI fast and private enough to live in your pocket? Rubin asks: how do we make AI cheap and powerful enough to run at planetary scale? Neither side is winning at the other’s expense — the phone in your hand and the data center on the other side of the world are becoming more dependent on each other, not less. The future of AI is not a fork in the road. It’s a supply chain, and most of us are only ever going to see one end of it. |
techsunnews.com | Tech / AI / Trending | © 2026




