GPUs for Local LLM

Intel Arc Pro B60: 192GB of VRAM the Cheap Way, and What It Really Costs

Four dual-GPU Intel cards promise 192GB of VRAM for a fraction of NVIDIA's price. We checked what that card actually costs in the US, and why two single B60s are the better buy.

Intel Arc Pro B60: 192GB of VRAM the Cheap Way, and What It Really Costs
Editor checklist: verify all prices and specs against primary sources before publishing (GPU pricing is moving weekly on the memory shortage); swap Amazon search links for exact product/ASIN links; confirm the video embed renders; add any real testing notes; delete this box before publishing.

If you want to run big models at home, the wall you hit is memory. Not compute, memory. So when a single card shows up carrying 48GB for a fraction of what NVIDIA charges, and four of them promise 192GB in one chassis, it deserves a hard look.

Alex Ziskind (@AZisk) built exactly that machine and asked the practical question: for 48GB, do you buy two Intel Arc Pro B60s, or one Maxsun card with two B60 GPUs on it? His answer is genuinely useful. But the buying advice underneath it has a problem he does not cover, and it changes the math by thousands of dollars.

What Ziskind found

The lineup he tested: Intel's Arc Pro B60 (24GB), the Arc Pro B70 (32GB), and the Maxsun Arc Pro B60 Dual 48G Turbo, a single card carrying two B60 GPUs. Running Intel's LLM Scaler stack (a fork of vLLM), he found:

WorkloadResult
Qwen3 4B, FP8B70 at 92 tok/s, ahead of two B60s in tensor parallel
Qwen 3.5 27B, int4B70 at 28.3 tok/s, but two B60s beat it. A single B60 cannot load it at all (~30GB)
30B MoE, int4B70 at 45 tok/s, single B60 at 44.5 tok/s
Maxsun dual on the 27B~26.7 tok/s, prompt processing around 1,800 to 2,200 tok/s
LTX2 video generationB70 at 167s vs ~225s for both dual-GPU setups (compute-bound, single-GPU bound)
Power, same workloadMaxsun ~129W vs ~134W for two separate B60s

The headline result: splitting one PCIe slot between two GPUs cost him nothing. The Maxsun card matched two separate B60s within margin of error, while drawing slightly less power. His conclusion was that you buy the dual card for density, not speed: a four-slot board fits four of them for 192GB, where four single B60s would give you 96GB.

Getting there was not smooth. The card did not appear at all on first boot, no device detected, empty PCI list. The Maxsun dual is switchless, meaning its two GPUs have no bridge chip and talk through the CPU, so the motherboard has to split that x16 slot into x8/x8. That is PCIe bifurcation, and until he enabled it in the BIOS the second GPU was invisible. Wendell from Level1Techs turns up to confirm bifurcation is broadly supported on modern boards, with the caveat to check your manual.

That warning deserves more weight than the video gives it. In LTT Labs' May 2026 review, the required x8/x8 option "wasn't available on many of our test benches," including the ASUS ROG MAXIMUS Z890 HERO and the Pro WS TRX50-SAGE WIFI. Without it, as they put it, you will only be using one of the GPUs.

The price problem

Here is where the buying case falls apart. The video prices the Maxsun dual at $1,750 and gets to $7,000 for 192GB, against roughly $24,000 for the NVIDIA equivalent. The top comment on the video, with over 350 likes, tells a different story: "the price range goes from 'request a quote' to 'contact us'." Others report $2,200 to $3,500, or being unable to find one at all.

We went looking. As of July 19, 2026:

  • Maxsun publishes no price. Its product page offers only a "Request a Quote" link.
  • No mainstream US retail. The card is not listed at Newegg, Amazon US, or B&H.
  • The grey market is thin and expensive. An eBay search filtered to US-located sellers returns zero results. Eight listings existed worldwide, all shipping from China, Australia, or Israel, at roughly $2,599 to $3,609 delivered before import duties.
  • The famous $1,200 figure was never real. It traces to a single August 2025 Reddit post relaying a bulk-order quote for ten units to Singapore. VideoCardz reported Maxsun did not confirm it. The tech press repeated it anyway.

So the 192GB build is not $7,000 in the US. At real delivered prices it lands closer to $10,400 to $14,400, before duties, from overseas sellers, with no domestic warranty. Still under NVIDIA, but a different proposition entirely.

And there is a cheaper answer sitting right there. Single B60s are ordinary US retail stock: $599.99 for the ARKN, $649.99 for ASRock's B60 CT on Newegg. Two of those give you the same 48GB for around $1,250, with a warranty and a returns window. One owner, Faux_Grey, skipped the dual for exactly that reason: "hard 8/8 requirement, and having a ~30% price overhead ... drove me away from it."

Cheap VRAM is real. Cheap fast VRAM is not

Ziskind says early on that "cheap VRAM and useful VRAM are different things," which is the right instinct. The numbers make it concrete.

Intel's B60 datasheet lists 24GB of GDDR6 on a 192-bit interface, rated at up to 456 GB/s. The B70 moves 608 GB/s. For comparison, a used RTX 3090 with the same 24GB does about 936 GB/s, and NVIDIA's RTX PRO 5000 Blackwell does 1,344 GB/s.

Token generation is memory-bandwidth-bound, so that ratio sets your ceiling. Intel doubled capacity relative to the consumer B580 without touching bandwidth. You get room for bigger models and longer context, not more speed. As one r/LocalLLaMA commenter put it, "each GPU has the same memory as a 3090, but half the memory bandwidth."

This is why MoE models are the workload where these cards earn their keep. A Mixture-of-Experts model only reads the experts that fire for each token, so it sidesteps the bandwidth deficit while still needing the capacity. Owner numbers line up: dense Mistral 14B lands at a rough 10 to 15 tok/s, while GPT-OSS 20b hits 60+ and Qwen3.6 35B-A3B around 51.

The 48GB is not one pool

Worth being blunt about, because the marketing blurs it: a "48GB" dual-B60 card is two B60 GPUs with 24GB each. Two separate memory pools, no fast link between them, routed through the CPU like two discrete cards. A model still has to fit in 24GB per GPU unless your runtime explicitly splits it across both.

Never treat it as 48GB of unified memory, and never add the bandwidth figures together. This is also the weakest part of Intel's stack: Intel's "Linux multi-GPU LLM support" claim had no independent confirmation on the B70 as of July 2026, with Phoronix noting the multi-card configuration is still being worked on.

What owners are saying

Sentiment splits on a clean line: people who bought for VRAM-per-dollar and expected to tinker are satisfied. People who wanted a drop-in CUDA replacement are not. Almost nobody disputes the hardware value. The complaints are about software.

The most upvoted critical thread is "Don't buy b60 for LLMs", where damirca lays out a rough ride: a custom-compiled kernel to fix ffmpeg crashes, a trip to a Windows machine to flash firmware and calm the fans, and "the speed is about 10-15tks at best in models like mistral 14b. The noise level is just unbearable."

Others in that thread go further. "I warned people about this. The B60 is about the same speed as the A770. Which makes it the slowest GPU I have," wrote fallingdowndizzyvr.

The counterweight comes from owners who changed runtimes. In the same thread, FortyFiveHertz posted llama-bench output from llama.cpp Vulkan on Windows showing Ministral-3 14B at 877 tok/s prefill and 24.7 tok/s generation, and said the card is "a low power, warrantied option depending on your local GPU market and whether you're happy to tinker." On fans: "Mine stays at 48 decibels under full, sustained load," though "the default fan curve is SO LOUD."

That runtime gap is the most actionable finding anywhere in this research. One tester on the Level1Techs forums documented the same card and model going from 328 tok/s prefill to 1,634 tok/s across one month of llama.cpp builds. Default backends are the slow path. OpenVINO, ipex-llm, Intel's LLM Scaler, and Vulkan under Windows are where the good numbers live, a spread of roughly 2x to 5x on identical silicon.

In the small r/IntelArcPro community, one B70 owner reported running Qwen 3.5 122B (IQ3_S, heavy offload) at 10 to 13 tok/s with 24k context, concluding that "your best return on investment will be running MOEs."

What viewers are saying

Ziskind's comment section leaned positive on the hardware and skeptical on buying it:

  • "I have dual b70s in my server. They are awesome!!" @MrMcWitt on YouTube
  • "I want to see a big model being loaded into those 192gb!" @typoerror177 on YouTube
  • "Double the GPUs in one single slot can save space." @SkyWarrior2000 on YouTube
  • "Maxsun's website says 'Request a Quote'. Sounds like unobtanium." @TheStanglehold on YouTube

So what should you buy?

If you want cheap VRAM and enjoy tinkering: two single Arc Pro B60s at $600 to $650 each. Same 48GB as the dual, roughly half the delivered price, US retail, warranty, and no bifurcation requirement. Budget a weekend for runtime setup and run MoE models.

If you mostly chat with small-to-mid models: one Arc Pro B70 (32GB, 608 GB/s, around $999 street). Ziskind's advice here is sound: one strong card is simpler and snappier, and it drops into any board.

If you want speed over capacity: a used RTX 3090 still does roughly twice the bandwidth per 24GB, with a software stack that works on the first try. That remains the default recommendation for a reason.

If you want the Maxsun dual: only if slot density is a hard constraint, your board documents x8/x8 bifurcation, and you accept importing a card at $2,600 to $3,600 with no domestic warranty. For most buyers that is a worse deal than two separate cards.

Ziskind's technical finding stands up: two GPUs sharing one slot cost nothing in performance, and that is a genuinely clever bit of engineering at a price nobody else is matching. The catch is that at real US prices, the card that makes 192GB possible is the most expensive way to buy Intel VRAM.

Sources and how we researched this

  • Video: "192GB of VRAM in One PC… The Cheap Way" by Alex Ziskind, July 15, 2026. All benchmark figures attributed to him are from that video. We have not tested any of this hardware first-hand.
  • Specs: Intel's Arc Pro B60 datasheet and B70 spec page; Igor's Lab teardown for electrical PCIe width; NVIDIA RTX PRO 5000 for the comparison bandwidth.
  • Bifurcation behavior: LTT Labs' review of the Maxsun B60 Dual 48G Turbo.
  • Pricing and availability checked July 19, 2026 against Maxsun's product page, Newegg listings, and live eBay searches. Prices on this class of card are moving weekly during the memory shortage, so re-check before buying.
  • Owner sentiment: linked r/LocalLLaMA, r/IntelArc, r/IntelArcPro threads and the Level1Techs forums, quoted verbatim and attributed. We include the most upvoted critical thread rather than only favorable reports.
  • Multi-GPU software status: Phoronix and the vLLM project's Intel Arc Pro B benchmarks.

Related: Intel Arc B580 buyer's guide · Why active parameters decide speed · Bandwidth, not TFLOPS · Used GPU price index

Get the Vetted Consumer newsletter

Reviews, buying advice, and field notes. Delivered monthly.

Almost there, check your inbox and click the confirmation link. ✓

Something went wrong, please try again, or email [email protected].