I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements. When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set. This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open. The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right. The hardware: one card is really two devices My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices. Current system / Planned system Physical cards: 2 / 3 Ascend devices/npus/AIcpus: 4 / 6 Nameplate device memory: 192 GB / 288 GB Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB Combined maximum accelerator-board power: 300 W / 450 W Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e). Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X. This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs. https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player Passive cooling is not a deal-breaker The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement. I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it. So I would not bat an eye at adding another passive card. The actual checklist is mundane: • Keep unobstructed airflow through the heatsink fins. • Make sure the chassis fans have enough static pressure. • Budget another 150 W of board power per card, plus the rest of the host. • Exhaust the heat from the room instead of recirculating it through the rack. • Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading. Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me. Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths. ECC, nameplate memory, and what is actually usable My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation. Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration. That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?” Why I forked vLLM as well as vLLM Ascend The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel. I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA. I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware. Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators. The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch. What took Qwen from incoherent ~1 tok/s to where it is now These are the architectural changes that moved the needle. Make the hybrid model correct before making it fast The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package. One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time. Shard the model at load time instead of loading everything everywhere The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU. I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout. Keep low-bit weights packed and do the work on the NPU My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case. The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation. This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency. Build caches for the model I actually have Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics. I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak. Remove synchronization and launch overhead from the token loop On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python. Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches. Optimize the service, not just an isolated kernel Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary. That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle. Qwen performance today These are milestones from different stages and workloads, not one controlled single-variable benchmark: Qwen3.8 Flash-Next milestone / Measured result Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent Correct eager W4 dequant reference: 0.229 tok/s Stable W8 service baseline: About 18–19 tok/s Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak Same optimized service, four requests: 60.88 tok/s aggregate median Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate 40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story. https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ. On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense. GLM is my next hard(er) model I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile. It now loads and generates on the same machine. The speed progression so far has been: GLM four-chip milestone / One request / Four-request aggregate Initial full-service profile / 0.764 tok/s / 1.474 tok/s Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s NZ-packed code layout / 1.386 tok/s / 2.946 tok/s Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation. I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.” The third card: 50% more memory and cores, not just a spare I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks. The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance. The extra capacity gives me several options: • Keep larger expert sets or higher-precision layers resident. • Fit models that are just over the four-chip limit without host offload. • Spend more memory on long-context state and cache. • Reduce aggressive quantization where the quality trade is not worthwhile. • Run a large distributed model while retaining room for another smaller service or evaluation workload. The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership. Known issues and active work This is what is still on my bench: • I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks. • Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization. • Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag. • GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen. • GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result. • Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology. • Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic. Should you buy one? I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM. I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following: You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers should expect a less polished setup than CUDA/mainline vLLM and be willing to follow pinned versions, read runbooks, report failures, and occasionally help diagnose a new hardware or model combination. If you need a supported appliance that runs arbitrary models without touching the stack, this is not the product I would recommend today. If you want a lot of inference memory at this price point, are comfortable working through rough edges, and see value in helping expand an open Ascend serving stack, it can be a very capable development platform. What I think these cards are good at The Atlas 300I Duo is compelling when used capacity is more valuable than a polished CUDA ecosystem and when you are willing to own the systems work. Two single-slot, 150 W cards give four accelerator ranks and a lot of local model memory. MoE models are an especially interesting fit because expert weights can be partitioned while only the selected experts run for each token. The weak point is not that the cards are passive or limited to 150 W. The weak point is the amount of model-specific runtime and kernel work still required on 310P: layouts, supported dtypes, graph-safe metadata, custom operators, sparse attention, state ownership, tooling, and documentation. If you expect to point stock vLLM at an arbitrary new Hugging Face model, this is not that experience. If you enjoy bringing hardware up from first principles, though, they are a lot of machine in a very interesting form factor. Going from incoherent output at around 1 tok/s to coherent Qwen near 30 tok/s per stream and 61 tok/s aggregate has made me substantially more optimistic about what the platform can do. The third card should let me push both model capacity and expert parallel throughput further, and I will publish the numbers—good or bad—once they are measured.   submitted by   /u/matteiuspi [link]   [comments]