This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast. This has turned from a weekend experiment into something I think I will use in the future. I am running 4 BC-250 boards connected by 2.5gb ethernet adapters to a 2.5gb switch using llama with vulkan and RPC. All the boards are inside an asrock 4u12g bc250 case that they originally came in. After starting out at 30 tok/s on flash next Iq2 and 80 tok/s with 3.6 35B I pointed Claude Code at the cluster and we are now running a better quant next flash as well as maintaining 60-70 tok/s (~150 ppt) at context. The 4 boards can run 2 instances of 3.6 35B at 145tok/s (~450 ppt) starting and going below 100 tok/s at 150k context. These boards cost me less than $100usd each when i bought them but i have seen deals recently on aliexpress for $150usd. I now have outlet power monitors and under load with next flash it pulls 600 - 700 W with the 2 instances of 3.6 35B it pulls 800 - 900 W. The next step for me is to use AIO coolers for each board as now I have been seeing some thermal throttling as the boards are better utilized . I will also 3d printing and building a AIO cooler lid for the case (upgrading from the cardboard plenum lol) now that I am going to keep the cluster. Before anyone suggests upgrading to a m.2 networking solution to get more bandwidth the limiting constraint is latency as its not sending a lot of data between boards at a time. BC-250 Info: https://elektricm.github.io/amd-bc250-docs/ Here is the Claude summary of optimizations: Next Flash The current build runs a 100K context: 59.4 / 65.2 tok/s on that request (first / second run) 66.6-68.3 tok/s behind a 5K-token prompt (sampled, 512 tokens) 72.6 tok/s on one browser generation of 50,357 tokens Most of the work was done on the Q2_0 file first (27.9 → 68-69 tok/s on a 5K prompt) and then carried over to IQ3: Speculative decoding with the model’s own MTP draft head, run as a chain of passes. Pass A verifies the known token. Passes of two drafts each follow one board behind (six drafts, four passes). A round ends at the first rejected draft. Adding the third pass took Q2_0 from 61 to 68 tok/s, and six drafts took an IQ3 temperature-0 edit request from 84.0 to 90.1-90.5. Recorded and replayed graphs. Workers record each stage graph’s Vulkan commands once and replay them (saves 3.7-4.5 ms of host time per graph, +2%). Remote graphs’ nodes are regrouped so a layer’s independent products share a barrier group (+5.7%). Exact kernels for this architecture (routed-expert top-k, two-token expert passes, recurrent-state fusion) and for the parts that grow with context: +7.3% at 62K, then 66.1 → 68.9-69.6 tok/s. Context handling. Context checkpoints are copied on the boards instead of through the coordinator (665 → 18 ms each, prompt speed +60%). Prompt batches are pipelined through the boards. Context went from 65,536 to 100,096 tokens with 32-token prompt batches. Qwen3.6-35B-A3B Q4_K_M (HauhauCS GGUF), two boards, K/V cache q8_0 . Step one (the Flash-Next kernels, graph replay and transport, two drafts): +39% to +66% over the baseline from 120 to 99.6K prompt tokens. Chain of passes with the MTP head (four drafts), plus attention that decodes a K/V tile once into shared memory: first request 105.8 → 116.0 tok/s at a short prompt, 73.0 → 95.2 at 44K. Longer context. The attention mask is one limit per token instead of a cells × tokens tensor (768 MiB less on the coordinator), and prompt batches read the q8_0 K/V in place. Usable context went from about 100K to 258K (of 262,144). Replay on the coordinator too. The coordinator’s own half is replayed from recorded command buffers (+2 to +4%). Draft-head checkpoints keep positions only, so a repeated request’s prompt time at 99.6K went from 932-953 ms to 142-163 ms. Startup and saving. A pair starts in 32-35 s instead of 105 s, and a conversation can be saved and restored bit for bit after a restart. Latest ( verified ). The draft head’s catch-up pass now only writes its K/V (it was computing attention nobody read), and RPC sends are batched. At 258K: 51.1 → 53.5 tok/s on a repeated request, 51.7 → 54.4 sampled. Plain decoding at that depth is 36.4. Exactness Weights, quantization and sampler settings are untouched. At temperature 0, the speculative rounds are verified bit for bit against plain decoding of the same build by hashing every logits row (five depths up to 99.6K on the 35B). Against stock llama.cpp, the logits differ at float-rounding level (about 0.03) where kernels were rewritten. On the 35B that includes summing attention in fixed-size chunks, so a result no longer depends on how many tokens a pass carries. Stock’s own drafted rounds differ from its plain decoding by the same order.   submitted by   /u/Ok-Breadfruit-3523 [link]   [comments]