The M5 Ultra Mac Studio.For the past few days, I’ve been testing the (currently) top-of-the-line M5 Ultra Mac Studio with 256 GB of RAM.I’ll cut to the chase: the M5 Ultra Mac Studio is a dream machine for local AI agents. This computer makes it possible to run personal assistants powered by local models with great performance and no additional cloud costs. If you’ve been skeptical of testing OpenClaw or Hermes Agent with local models because they’d never be even remotely near the intelligence and speed of cloud ones, this Mac will change your mind about that.Since last Thursday, I’ve been comparing this Mac Studio to its predecessor, the M3 Ultra with 512 GB of RAM, as well as my own desktop gaming PC with an RTX 5090 inside. For its size, price, thermal performance – not to mention Apple’s approach to unified memory – the M5 Ultra Mac Studio has fundamentally changed how I think about models running locally and what they can enable now. A 5090, of course, still has an edge over the M5 Ultra thanks to its higher memory bandwidth. But considering the sheer size of my PC build, as well as its heat and noise, I would prefer an M5 Ultra Mac Studio any day. It also happens to be a Mac, with an operating system that looks nice and doesn’t suck, plus a vibrant app ecosystem. (Windows fans, I’m sorry, but Microsoft software will never get my sympathy.)Supported ByAstropad WorkbenchAstropad Workbench: Remote desktop for AI agents and headless Mac minis. Free for 20 mins/day.As I’ll explore in this article, running the latest Qwen3.8-Flash-Next model on the M5 Ultra Mac Studio has been so nice and fast, I’ve made it my default in both Open Minis for iOS and Hermes Agent. That’s right: the personal assistants I use the most – more than Siri AI, in fact – are now entirely powered by a model running locally on a Mac Studio. Furthermore, thanks to the M5 Ultra’s faster GPU and higher memory bandwidth, these agents start responding more quickly, stay fast at larger context windows, and can run long, multi-turn loops without slowing to a crawl as the session grows. Because of this, I’ve also been using local models in the Codex app on my Mac – either as main threads or subagents orchestrated by GPT-6 Astra – and I’ve had a great experience doing so.The personal assistants I use the most are now entirely powered by a model running locally on a Mac Studio.Local subagents running in Codex on the M5 Ultra Mac Studio.I should note upfront that I’m not an AI developer by trade: I do not train or fine-tune models. I’m a tinkerer at heart, and I’ve been playing around with local AI models for over a year at this point. This summer, I went all-in on local AI usage for a big project I was working on, which I will explain in the following section.My goal with this article is to provide you with a mix of two things: numbers and visualizations based on the (many) tests I’ve run over the course of four days, and an explanation of my practical use cases for local AI applied to my workflow and how I get things done for MacStories.Let’s dive in.Why Local AI?Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?It’s a fair question. You need expensive hardware to run these models, and by the time you’ve repaid your investment, you could have used the most expensive Anthropic subscription for several years, still saved money, and got better performance in return.Different people will have different answers to this question. Some might say they use local models because of privacy: they’d rather rely on local intelligence for sensitive data and documents than upload anything to an external cloud. Others might argue that it’s simply cool – and I do not disagree. For some, it’s a work-related task: if you’re an AI developer, it makes sense to have a great local setup for training your own adapters or fine-tuning models.For me, the journey into local AI has been characterized by a mix of the “cool, why not?” factor of it all as well as considerations about privacy and costs.As I will share later this week with Club MacStories members, my research and writing setup for the iOS and iPadOS 27 review this summer has been powered and made possible by local AI. Back in June, I created an internal app, called Desk, to organize hundreds of notes, sessions, PDF documents, and clipped webpages related to iOS and iPadOS 27, as well as chapters of the review. By the end of the process, the project consisted of 310 documents. In Desk, a team of agents – all based on DeepSeek V4 Flash, plus olmOCR for PDFs – ran 24/7, for 99 days, to perform the following tasks:Transcribe my favorite WWDC sessions (using summarize plus LLM processing)Extract features of iOS and iPadOS 27 from clipped webpages, sessions, PDF guides, and my own notesCross-reference features across different sources, and keep track of which features belonged to which chapter of the reviewExtract features and bugs from screenshots I uploadedWork with the Notion API to organize everything across multiple databasesOne of the views of Desk, the app powered by the Notion API and local AI I used for my iOS 27 review.The local AI agent runs in my custom Desk app.When I started working with this setup in early June, I quickly realized that relying on the OpenAI or Anthropic APIs for this kind of always-on, persistent background task would be…cost-prohibitive, to say the least. So I pivoted to local AI, and the result is the iOS and iPadOS 27 review you can read on MacStories. It was all written by me, the old-fashioned human way. But the entire research stack, deep-linking between notes, and keeping track of new features and betas were all performed by my agents, running locally on the Mac Studio, for a total cost of $0.If you don’t think that’s neat, or a powerful concept to explore, then this article probably isn’t for you – and I understand. Dealing with these models is fiddly, and it’s not something I would ever recommend to someone who (rightfully) just wants to pay $20 to use Claude Cowork. This kind of setup is, by definition, the bleeding edge of AI workflows at the moment.If you fall on the other end of the spectrum, though, and if you think this kind of stuff is neat…let me tell you: the M5 Ultra Mac Studio is a massive leap in performance for local models powered by MLX, and I have a few examples to prove it.A Leap for Prompt Processing and GenerationAs you may have seen from the announcement and my initial coverage, the M5 Ultra Mac Studio looks identical to the M3 Ultra model it replaces, but it comes with an all-new Apple silicon architecture that uses UltraFusion to connect two dual-die M5 Max chips to form a quad-die architecture, which is a first for the Apple ecosystem. As far as local AI workloads are concerned, there are two areas we have to pay attention to (and which I have been following since my coverage of the M5 iPad Pro for local AI last year): GPU and memory bandwidth.The M5 Ultra has a next-gen GPU with 80 cores, each with a Neural Accelerator that grants it up to 4.5× the peak GPU compute for AI compared to the M3 Ultra. As for memory, Apple’s unified memory architecture still tops out at 512 GB as before (although that model will come out in late October), but its bandwidth has jumped from 819 GB/s to 1.2 TB/s, or 50% higher than the M3 Ultra.With these numbers in mind, I started testing the M5 Ultra against the M3 Ultra with 512 GB of RAM and my RTX 5090. I’ll share more details on testing below, but the short version is this: with the M5 Ultra, you spend considerably less time waiting for a model to read your prompt and begin generating a response; and when it does start answering, text appears much faster than it used to on the M3 Ultra. These two improvements alone make the machine viable for modern agentic loops that require fast iteration with a model and, as a result, larger context windows.In my day-to-day experience with agents running on the M5 Ultra, these improvements to token prefill (or how quickly a prompt can be processed) and token generation are the changes I noticed immediately. When comparing a model running on the M3 Ultra and M5 Ultra side by side with Open Minis on iOS, the M5 Ultra was ~70% faster on average than the M3 Ultra at generating a response. As we’ll see later, having a model such as Qwen3.8-Flash-Next clear 100 tokens/second on short prompts and still write at 60 to 85 with 64K to 256K of context behind it is no joke, and it enables the kind of agentic back-and-forth between you and the model that feels great to use, particularly when tool calls are involved.Using a local model as my default in Open Minis for iOS. Pictured above: a long-running project, subagents, and local image generation powered by Qwen-Image-2.1, also running on the M5 Ultra.However, I was more impressed with the performance gains in token prefill. When you use agentic assistants such as Hermes or Codex, a model receives a whole block of instructions that include things like the system prompt, user personalization and session memories, skill and MCP descriptions, and more. Some agents are better than others at trimming the instructions they send, but, generally, whenever you use a modern agent, you’re not starting with an empty context window. Because of this, I’ve never been able to consistently use local models with this new wave of agents: they would work, but I’d stare at an empty screen and a loading indicator for a while before the model would start generating a response. And on every turn of the loop, performance would get worse (because of the larger context of the session), and I’d wait some more time.In my tests, prompt processing is up 150% on average from the M3 Ultra – a ~2.5× improvement from my previous setup. This change alone makes local models solid choices in apps like Open Minis and Hermes Agent. When I ask Flash-Next on the M5 Ultra to get my tasks for the week with RemCTL, I don’t have to wait around for the agent to process my prompt and Open Minis’: in just a few seconds, it gets to work by reasoning, performing tool calls, and so forth. And when I’m working on a large project, such as the voxel Colosseum demo below, the model is able to process multi-turn loops quickly, dispatch and coordinate subagents, and do it all at 60 to 85 tokens per second as the thread grows longer.This interactive Colosseum demo was entirely created by Flash-Next running on the M5 Ultra Mac Studio, managed via Open Minis and its harness on iOS.To use local models from the Mac Studio in Open Minis for iOS, I created a local server in front of the OpenAI-compatible API exposed locally by oMLX. It’s served via Tailscale to my iPhone and works wonderfully.I’m a big believer in assistants that can agentically perform tasks in addition to answering questions, but in order to feel nice to use, they have to be fast. Over the past few months, I’ve tested several “boutique” cloud providers with Open Minis: Inco, which serves Kimi K3 at over 300 TPS; Cerebras, with Qwen3.8-27B at a whopping 1,800 TPS; and the likes of Fireworks and Baseten, each breaking the 150 TPS barrier. All of those providers feel extremely good to use in Open Minis and Hermes, but they are expensive (I burned through $20 of Inco credits in literally 10 minutes last week), and, of course, all my data is going…somewhere when I use them. When I fire up Open Minis with Flash-Next and the collection of Apple CLIs I’m creating, everything stays local, inside a computer I can see and reboot whenever I want.Cadu, an upcoming iOS client for Hermes Agent, running live voice mode with Qwen3-TTS and Qwen3.8-Flash-Next as the underlying chat model, with the M5 Ultra as the server.Most importantly: a model like Flash-Next can be “small” enough to run at higher quantizations on a 256 GB M5 Ultra (I can run 5-bit entirely in RAM; 6- and 8-bit can offload their n-gram tables to SSD with this new architecture) but also intelligent enough to sustain long threads and multiple agentic tool calls.Text generation at different quants.For my taste, 5-bit quantization hits the sweet spot on this version of the Ultra with a balance of intelligence, performance, and memory consumption. But I already know that, if I ever get to test a 512 GB M5 Ultra, I’d be really interested to measure performance of the 8-bit quant without SSD offloading.I have not spent much time tinkering with offloading coding tasks for my various projects to a local model, but I’ve done a few interesting experiments. With this kind of performance, and especially given the ability to stack up to three concurrent Flash-Next sessions with subagents in oMLX with 256 GB of RAM (more later), I can now realistically consider handing off simpler coding tasks to a local model and have frontier cloud ones review their work. For instance, I was able to set up Qwen3.8-Flash-Next in Codex, which lets me use a local model with the Codex harness. This means that I can let a main GPT model orchestrate local subagents, have Flash-Next coordinate its own subagents, or even just use the model from my phone with Codex Remote on iOS.Using a local model on the M5 Ultra from Codex Remote.I do not usually rely on image generation, but for the sake of this review: the M5 Ultra chip is an official Apple asset; the wallpaper behind it was generated by Qwen-Image-2.1 locally on the M5 Ultra in 180 seconds, with peak RAM usage of 78 GB.I’m curious to read more on this topic from actual developers who are getting an M5 Ultra soon. With open-weights models now outperforming on consumer hardware what was considered “frontier” ~10 months ago, and with performance on an M5 Ultra now making agentic coding feasible, I think we’re going to see some fascinating experiments from the MLX community very soon.M5 Ultra vs. RTX 5090As you’ll see from the visualizations later in this article, NVIDIA’s RTX 5090 is still faster than Apple’s M5 Ultra despite its “meager” 32 GB of VRAM, for two different reasons.Prompt processing speeds are dictated by compute: the model reads the whole prompt in one giant matrix multiplication, which is exactly the job NVIDIA’s Tensor Cores were built for. Apple’s new Neural Accelerators (one in each of the M5 Ultra’s 80 GPU cores) narrow the gap, but can’t close it. On a 6,000-token prompt, the M5 Ultra read at ~1,700 tok/s; the 5090 delivered a staggering ~3,000 with the Qwen model I tested in LM Studio. Token generation, on the other hand, is bandwidth: the model writes one token at a time and pulls the entire model back out of memory for each one, so the 5090’s 1.79 TB/s against the M5 Ultra’s 1.2 TB/s gives it a steady ~25% lead at every prompt size. What the 5090 doesn’t have is memory: at 256K, the 5090 only finishes with an 8-bit attention cache. 32 GB of VRAM only goes so far.There are, however, two problems with this comparison. First, while the 5090 does still edge out the M5 Ultra with smaller models, its lack of a unified memory pool means that I’m limited to the 32 GB of VRAM in the GPU if I want to run a model at blazing-fast speeds. The moment I want to run anything exceeding 32 GB (such as the aforementioned higher Flash-Next quants), the 5090 must offload model layers over PCIe to (much slower) system RAM, and that’s no way to live.I tested a different Qwen model for the comparisons between Mac and PC.Second, my gaming PC is massive compared to a Mac Studio that fits on my desk – and I have a compact build with a Lian-Li A3 case. Not to mention how loud and hot it gets when I’m running local models at high context windows: when I walked into my office after some benchmarks had run, it was uncomfortably warmer compared to the rest of my apartment. By contrast, the “diminutive” Mac Studio on my desk was warm to the touch, but it was also appreciably quieter than my 5090, the fans were not spinning as fast or loudly, and, most important, it allowed me to run larger models such as GLM-5.3-Flash locally with decent performance thanks to Apple silicon’s unified memory. In my day-to-day use, when I was running Flash-Next oQ4e all the time, I could never hear the fan of the Studio on my desk unless I placed my ear directly on top of the computer.Judging by the progress Apple has made in recent years, I wouldn’t be surprised to see an M7 Ultra that outperforms the memory bandwidth of a 5090 in the near future. But that’s a story for another time.A Note on TestingLastly, before we jump into raw numbers and charts: how did I test everything?Automated tests were conducted with a testing harness I built with GPT-6 Astra, which coordinated multiple instances of Codex across my M3 Ultra and M5 Ultra Mac Studio, as well as my PC with the Codex app for Windows and Computer Use. On macOS, I chose oMLX (version 0.7.0.dev2) as the local backend for MLX models, and ran Qwen3.8-Flash-Next-oQ4e-mtp, GLM-5.3-Flash-MLX-mixed-4_8bit, and Qwen3.8-27B-oQ4e-mtp on macOS Golden Gate 27.0 for the majority of my tests. On Windows, I used LM Studio and Qwen3.8-27B-GGUF with CUDA 12 runtime and with all 66 layers offloaded to the GPU for the full-GPU tests, plus separate tests splitting the model between GPU and system RAM.Alongside separate experiments with Open Minis’ native subagents, I used a custom testing harness to measure concurrent requests and workflows involving a lead model and multiple helpers, with oMLX serving the Mac models and LM Studio serving the Windows model.Numbers were collected by Astra over the course of four days, and later visualized by Claude Fable 5.1 and Opus 5 using Anthropic’s upcoming Projects feature, which I was able to test early when working on this story. The interactive visualization was built with pure HTML and CSS based on MacStories’ style, and it includes comments and annotations by yours truly.Claude’s upcoming Projects feature.My goal with the following interactive widgets was not only to help you understand the numbers more clearly, but also to visualize what the stats mean in practice. I’m quite happy with the widgets that approximate what different tokens per second feel like, since that’s a metric that’s often tricky to visualize. I hope these animated charts will be more useful than regular “static” ones you’ve probably seen elsewhere (which are also included below).Visualizing the M5 Ultra/* MacStories widget: Visualizing Local Models on M5 Ultra, every figure. Inline HTML and CSS for one WordPress HTML block, copied from the preview on 2026-09-21. Generated, so rebuild it rather than edit it. */ .ms-widget{display:block;box-sizing:border-box;width:700px;max-width:90%;margin:0 auto;padding:0} /* Local AI review widgets. Everything is scoped under .mx. Generated, so rebuild rather than edit. */ @property --mx-t { syntax: ""; inherits: true; initial-value: 0; } .mx { --mx-ink:#1d1d1f; --mx-text:#3a3a3c; --mx-muted:#6e6e73; --mx-rule:#e3e3e6; --mx-track:#f4f4f5; --mx-hatch:#b9b9c0; --mx-accent:#b00a0f; --mx-accent-ink:#fff; --mx-soft:rgba(176,10,15,.10); --mx-third:#8e8e93; --mx-fourth:#c7c7cc; --mx-chip-filter:none; --mx-s: 1s; /* one measured second, before the per-race --speed divisor */ display: block; width: 100%; margin: 2em 0; padding: 0; box-sizing: border-box; clear: both; font-family: system-ui, -apple-system, "Helvetica Neue", Helvetica, sans-serif; /* the MacStories stack, verbatim */ font-size: 15px; line-height: 1.4; color: var(--mx-ink); font-variant-numeric: tabular-nums; -webkit-font-smoothing: antialiased; text-rendering: optimizeLegibility; } @media (prefers-color-scheme: dark) { .mx:not(.theme-light *) { --mx-ink:#f2f2f4; --mx-text:#d6d6da; --mx-muted:#98989f; --mx-rule:#2a2a2e; --mx-track:#1a1a1d; --mx-hatch:#4b4b52; --mx-accent:#53c8f0; --mx-accent-ink:#0b1114; --mx-soft:rgba(83,200,240,.14); --mx-third:#8a8a93; --mx-fourth:#48484f; --mx-chip-filter:invert(1); } } .theme-dark .mx { --mx-ink:#f2f2f4; --mx-text:#d6d6da; --mx-muted:#98989f; --mx-rule:#2a2a2e; --mx-track:#1a1a1d; --mx-hatch:#4b4b52; --mx-accent:#53c8f0; --mx-accent-ink:#0b1114; --mx-soft:rgba(83,200,240,.14); --mx-third:#8a8a93; --mx-fourth:#48484f; --mx-chip-filter:invert(1); } /* Host-proofing: every property the widget depends on is set with two class-level selectors. */ .mx :not(.mx) { box-sizing: border-box; } .mx .mx-head, .mx .mx-t, .mx .mx-l, .mx .mx-note, .mx .mx-foot, .mx .mx-c, .mx .mx-txt, .mx .mx-who, .mx .mx-big, .mx .mx-fact, .mx .mx-heat-t { margin: 0; padding: 0; } .mx .mx-head { font-size: 15px; font-style: normal; text-align: left; color: var(--mx-ink); } .mx .mx-t, .mx .mx-clock, .mx .mx-clock-static, .mx .mx-lane-name, .mx .mx-heat-t b { color: var(--mx-ink); } .mx .mx-key span, .mx .mx-axis span { color: var(--mx-muted); } .mx .mx-vh { position: absolute; width: 1px; height: 1px; margin: -1px; padding: 0; overflow: hidden; clip: rect(0 0 0 0); white-space: nowrap; border: 0; } .mx a { color: inherit; text-decoration: underline; text-decoration-color: var(--mx-hatch); text-underline-offset: 2px; } .mx a:hover { text-decoration-color: currentColor; } /* Head */ .mx .mx-t { font-size: 1.25rem; font-weight: 700; line-height: 1.2; letter-spacing: normal; } .mx .mx-l { font-size: 1rem; line-height: 1.5; letter-spacing: normal; color: var(--mx-muted); margin-top: 6px; } /* Switchers: understated text tabs, 44px targets. */ .mx .mx-switch { display: flex; flex-wrap: wrap; gap: 0 4px; margin: 0; padding: 0; border: 0; border-bottom: 1px solid var(--mx-rule); min-width: 0; } .mx .mx-switch label { display: block; min-height: 44px; padding: 12px 10px 10px; margin-bottom: -1px; font-size: 13px; font-weight: 500; color: var(--mx-muted); cursor: pointer; border-bottom: 2px solid transparent; transition: color .2s; } .mx .mx-switch label:first-of-type { padding-left: 0; } .mx .mx-switch input:checked + label { color: var(--mx-ink); border-bottom-color: var(--mx-accent); } .mx .mx-switch input:focus-visible + label { outline: 2px solid var(--mx-accent); outline-offset: -2px; border-radius: 4px; } .mx .mx-switch label:hover { color: var(--mx-ink); } .mx .mx-switch label small { display: block; font-size: 11px; font-weight: 400; color: var(--mx-muted); } .mx .mx-switch.mx-models { border-bottom: 0; padding-top: 4px; } .mx .mx-switch.mx-models label { padding: 9px 10px 7px; line-height: 1.3; } /* Key */ .mx .mx-key { display: flex; flex-wrap: wrap; gap: 4px 16px; padding: 10px 0 0; font-size: 12px; color: var(--mx-muted); } .mx .mx-key span { display: inline-flex; align-items: center; gap: 6px; } .mx .mx-key i { display: inline-block; width: 22px; height: 12px; border-radius: 2px; background: var(--mx-track); } .mx .mx-key .mx-k-read { background-image: repeating-linear-gradient(135deg, var(--mx-hatch) 0 2px, transparent 2px 6px); } .mx .mx-key .mx-k-write { background: var(--mx-ink); } .mx .mx-key .mx-k-cache { width: 30px; height: 7px; background: linear-gradient(to right, transparent 30%, var(--mx-ink) 30%), repeating-linear-gradient(135deg, var(--mx-hatch) 0 2px, transparent 2px 6px); } .mx .mx-key .mx-k-a { background: var(--mx-ink); } .mx .mx-key .mx-k-n { background: var(--mx-accent); } .mx .mx-key .mx-k-p { background: var(--mx-third); } .mx .mx-key .mx-k-q { background: var(--mx-fourth); } .mx .mx-key .mx-k-fail { background-image: repeating-linear-gradient(135deg, var(--mx-accent) 0 2px, transparent 2px 5px); opacity: .7; } /* Play again: two radios, one visible label at a time. Checking the other radio swaps every animation to its identical -b twin, which restarts it. No markup is duplicated. */ .mx .mx-play { display: flex; justify-content: flex-end; align-items: center; gap: 10px; padding: 8px 0 0; font-size: 12.5px; color: var(--mx-muted); } .mx .mx-play label { display: inline-flex; align-items: center; gap: 6px; min-height: 32px; padding: 4px 12px; border: 1px solid var(--mx-rule); border-radius: 999px; font-weight: 500; color: var(--mx-ink); cursor: pointer; } .mx .mx-play label::before { content: ""; width: 0; height: 0; border-style: solid; border-width: 5px 0 5px 8px; border-color: transparent transparent transparent currentColor; } .mx .mx-play label:hover { border-color: var(--mx-muted); } .mx .mx-play input:focus-visible + label { outline: 2px solid var(--mx-accent); outline-offset: 2px; } .mx:has(.mx-play-a:checked) .mx-play label[for$="-pa"], .mx:has(.mx-play-b:checked) .mx-play label[for$="-pb"] { display: none; } /* Race: one .mx-race per switcher state; only the selected one is rendered, which restarts its animations. */ .mx .mx-race { display: none; } .mx .mx-race.mx-solo { display: block; } .mx .mx-race-head { display: flex; flex-wrap: wrap; align-items: flex-end; justify-content: space-between; gap: 0 12px; padding: 12px 0 0; } .mx .mx-race-head .mx-clock { margin-left: auto; } .mx .mx-prompt { font-size: 13px; color: var(--mx-muted); padding-bottom: 4px; max-width: 60%; } .mx .mx-clock { display: flex; flex-direction: column; align-items: flex-end; font-size: 30px; font-weight: 600; line-height: 1.3; letter-spacing: -0.02em; white-space: nowrap; contain: layout style; --mx-t: calc(var(--end) * 10); animation: mx-count calc(var(--end) / var(--speed) * var(--mx-s)) linear; } .mx .mx-clock small { font-size: 12px; font-weight: 400; letter-spacing: 0; color: var(--mx-muted); } @keyframes mx-count { from { --mx-t: 0; } } @keyframes mx-count-b { from { --mx-t: 0; } } /* Odometer: --mx-t counts measured tenths. Every drum snaps from one digit to the next, like a digital stopwatch, so a digit never sits between two cells. Leading drums are blank at 0. */ .mx .mx-odo { display: none; } @supports (width: mod(10px, 3px)) and (width: round(down, 10px, 3px)) { .mx .mx-clock-static { position: absolute; width: 1px; height: 1px; margin: -1px; overflow: hidden; clip: rect(0 0 0 0); } .mx .mx-odo { --tt: calc(var(--mx-t) + 0.001); display: inline-flex; align-items: flex-end; height: 1em; line-height: 1; padding: .12em 0 .08em; box-sizing: content-box; } .mx .mx-d, .mx .mx-x { display: block; height: 1em; overflow: hidden; } .mx .mx-d b { display: block; font-weight: inherit; will-change: transform; } .mx .mx-d i { display: block; height: 1em; font-style: normal; text-align: center; } .mx .mx-d1 b { transform: translateY(calc(-1em * mod(round(down, var(--tt)), 10))); } .mx .mx-d2 b { transform: translateY(calc(-1em * mod(round(down, var(--tt) / 10), 10))); } .mx .mx-d3 b { transform: translateY(calc(-1em * min(9, round(down, var(--tt) / 100)))); } .mx .mx-d3.mx-dm b { transform: translateY(calc(-1em * mod(round(down, var(--tt) / 100), 10))); } .mx .mx-d4 b { transform: translateY(calc(-1em * min(9, round(down, var(--tt) / 1000)))); } } /* The single block: section headings and intros between the figures. MacStories scopes its body copy to .post-content > p, which a nested widget can never match, so the prose here restates the article's own scale off the root: h2 1.75rem/1, body 1.125rem/1.5, small print 0.875rem, and no letter-spacing. */ .ms-widget .mx-part { margin: 0; padding: 0; } .ms-widget .mx-part .mx-h2 { margin: 4rem 0 1rem; padding: 12px 0 0; border-top: 2px solid currentColor; font-size: 1.75rem; line-height: 1; letter-spacing: normal; font-weight: 700; } .ms-widget style + .mx-part .mx-h2 { margin-top: 0; } /* is the first child, so :first-child never matched */ /* Deliberately two classes, so the article's own h2 rule outranks it and the section headings render like every other heading in the post. They are article headings between the figures, not part of one: if MacStories restyles h2, these should follow it rather than hold a value of ours. The declarations above are what the theme computes today, so a preview and any other host land in the same place. Do not harden this one the way the in-figure headings are hardened. */ .ms-widget .mx-part .mx-intro { margin: 1rem 0; padding: 0; font-size: 1.125rem; line-height: 1.5; letter-spacing: normal; font-weight: 400; } .ms-widget .mx-part .mx { margin: 0 0 40px; } .ms-widget .mx-part .mx:last-child { margin-bottom: 0; } /* Contents: the section headings as links, first in the pasted block (after the article's own intro) and after the lede on the preview. Article furniture rather than a figure, but it keeps its own type so it reads the same in both places: the link rule is hardened past the theme's (0,3,1) the way the in-figure headings are. It had a rule above it and a rule below it until 21 September; the white space around it carries it now. A page with one section gets no list. */ .mx-toc.mx-toc { --mx-ink:#1d1d1f; --mx-accent:#b00a0f; --mx-hatch:#b9b9c0; display: block; margin: 2em 0; padding: 0; font-family: system-ui, -apple-system, "Helvetica Neue", Helvetica, sans-serif; font-size: 15px; line-height: 1.4; color: var(--mx-ink); -webkit-font-smoothing: antialiased; } @media (prefers-color-scheme: dark) { .mx-toc.mx-toc:not(.theme-light *) { --mx-ink:#f2f2f4; --mx-accent:#53c8f0; --mx-hatch:#4b4b52; } } .theme-dark .mx-toc.mx-toc { --mx-ink:#f2f2f4; --mx-accent:#53c8f0; --mx-hatch:#4b4b52; } .mx-toc.mx-toc .mx-toc-t.mx-toc-t { display: block; margin: 0 0 6px; padding: 0; font-size: 12px; font-weight: 600; letter-spacing: .08em; text-transform: uppercase; color: var(--mx-accent); } .mx-toc.mx-toc .mx-toc-a.mx-toc-a { display: block; margin: 0; padding: 3px 0; border: 0; background: none; font: inherit; font-weight: 500; color: var(--mx-ink); text-decoration: underline; text-decoration-color: var(--mx-hatch); text-underline-offset: 2px; } .mx-toc.mx-toc .mx-toc-a.mx-toc-a:hover { color: var(--mx-ink); text-decoration-color: currentColor; } /* Heats: rows inside one race that share the clock and the axis. */ .mx .mx-heat { padding: 6px 0 0; } .mx .mx-heat + .mx-heat { border-top: 1px dashed var(--mx-rule); margin-top: 6px; } .mx .mx-heat-t { display: flex; align-items: baseline; justify-content: space-between; gap: 10px; padding: 8px 0 0; font-size: 13px; color: var(--mx-muted); } .mx .mx-heat-t b { font-size: 14px; font-weight: 600; } /* Lane: --from (start), --ttft (time to first token), --total (end), all in measured seconds on the shared axis. */ .mx .mx-lane { --from: 0; --read: calc(var(--ttft) - var(--from)); --gen: calc(var(--total) - var(--ttft)); --ms: calc(var(--mx-s) / var(--speed)); padding: 10px 0 12px; } .mx .mx-lane-head { display: flex; align-items: baseline; justify-content: space-between; gap: 12px; } .mx .mx-lane-name { display: flex; flex-wrap: wrap; align-items: baseline; gap: 0 8px; font-size: 14px; font-weight: 600; line-height: 1.5; } .mx .mx-lane-name small { font-size: 12px; font-weight: 400; color: var(--mx-muted); margin-left: -4px; } .mx .mx-tps { font-size: 12.5px; font-weight: 500; color: var(--mx-muted); white-space: nowrap; } .mx .mx-tps b { font-size: 14px; font-weight: 700; color: var(--mx-ink); font-variant-numeric: tabular-nums; } .mx .mx-note { font-size: 0.875rem; line-height: 1.5; color: var(--mx-muted); padding-top: 6px; } .mx .mx-done { font-size: 20px; font-weight: 600; line-height: 1.3; letter-spacing: -0.01em; color: var(--mx-text); white-space: nowrap; } .mx .mx-done small { font-size: 12px; font-weight: 400; color: var(--mx-muted); letter-spacing: 0; } .mx .mx-done.mx-win { color: var(--mx-accent); font-weight: 700; } .mx .mx-track { position: relative; height: 44px; margin-top: 6px; overflow: hidden; border-radius: 3px; background-color: var(--mx-track); background-image: linear-gradient(to right, var(--mx-rule) 1px, transparent 1px); background-size: calc(100% * var(--tick) / var(--max)) 100%; } .mx .mx-read, .mx .mx-write { position: absolute; top: 0; bottom: 0; } .mx .mx-read { left: calc(100% * var(--from) / var(--max)); width: calc(100% * var(--read) / var(--max)); border-radius: 3px 0 0 3px; background: repeating-linear-gradient(135deg, var(--mx-hatch) 0 2px, transparent 2px 6px); clip-path: inset(0 0 0 0); animation: mx-reveal calc(var(--read) * var(--ms)) linear calc(var(--from) * var(--ms)) both; } /* The bars never animate width. Reading is revealed with clip-path (paint only); writing slides a full-width ink block in with a transform (compositor), and the text cursor rides on its leading edge for free. */ .mx .mx-write { left: calc(100% * var(--ttft) / var(--max)); width: calc(100% * var(--gen) / var(--max)); border-radius: 0 4px 4px 0; overflow: hidden; } .mx .mx-write i { position: absolute; inset: 0; background: var(--mx-ink); transform: translateX(0); animation: mx-slide calc(var(--gen) * var(--ms)) linear calc(var(--ttft) * var(--ms)) both; } .mx .mx-lane.mx-n .mx-write i { background: var(--mx-accent); } .mx .mx-lane.mx-p .mx-write i { background: var(--mx-third); } @keyframes mx-reveal { from { clip-path: inset(0 100% 0 0); } } @keyframes mx-reveal-b { from { clip-path: inset(0 100% 0 0); } } @keyframes mx-slide { from { transform: translateX(-100%); } } @keyframes mx-slide-b { from { transform: translateX(-100%); } } /* Text cursor at the writing edge: appears at first token, blinks, and leaves the clipped box at the finish. */ .mx .mx-write i::after { content: ""; position: absolute; top: 11px; bottom: 11px; right: -3px; width: 3px; border-radius: 1px; background: var(--mx-accent); opacity: 0; animation: mx-on 0s linear calc(var(--ttft) * var(--ms)) forwards, mx-blink 1s steps(2, jump-none) calc(var(--ttft) * var(--ms)) infinite; } .mx .mx-lane.mx-n .mx-write i::after { background: var(--mx-ink); } @keyframes mx-on { to { opacity: 1; } } @keyframes mx-on-b { to { opacity: 1; } } @keyframes mx-blink { from { opacity: 1; } to { opacity: 0; } } @keyframes mx-blink-b { from { opacity: 1; } to { opacity: 0; } } /* A lane that never answered: memory guard, stopped, or not attempted. Nothing animates; the reason is the bar. */ .mx .mx-lane.mx-fail .mx-track { background-image: repeating-linear-gradient(135deg, var(--mx-soft) 0 3px, transparent 3px 9px); } .mx .mx-lane.mx-fail .mx-track::after { content: attr(data-why); position: absolute; inset: 0; display: flex; align-items: center; padding: 0 10px; font-size: 12.5px; font-weight: 500; color: var(--mx-accent); } .mx .mx-lane.mx-fail .mx-done { color: var(--mx-muted); font-size: 14px; font-weight: 500; } /* Cached mini-lane: the same request again with a warm prompt cache. Runs after the main lane finishes. It keeps the heat's scale, but sits inside a track as long as the cold run above it. Without that frame a warm pass worth a fiftieth of the cold one is a 12px mark against a 700px track, and reads as a stray dot rather than as a measurement. With it, the empty remainder is the claim. */ .mx .mx-cache { --ctotal: calc(var(--cttft) + var(--gen)); display: flex; flex-wrap: wrap; align-items: flex-start; gap: 2px 8px; margin-top: 5px; } .mx .mx-cache-bar { position: relative; flex: 0 0 calc(100% * var(--total) / var(--max)); min-width: 12px; height: 12px; margin-top: 3px; overflow: hidden; border-radius: 2px; background-color: var(--mx-track); box-shadow: inset 0 0 0 1px var(--mx-rule); animation: mx-in .3s ease-out calc(var(--total) * var(--ms)) both; } .mx .mx-cache .mx-read { left: 0; width: max(3px, calc(100% * var(--cttft) / var(--total))); border-radius: 2px 0 0 2px; animation-duration: calc(var(--cttft) * var(--ms)); animation-delay: calc(var(--total) * var(--ms)); } .mx .mx-cache .mx-write { left: calc(100% * var(--cttft) / var(--total)); width: max(2px, calc(100% * var(--gen) / var(--total))); border-radius: 0 2px 2px 0; } .mx .mx-cache .mx-write i { animation-delay: calc((var(--total) + var(--cttft)) * var(--ms)); } .mx .mx-cache .mx-write i::after { content: none; } .mx .mx-cache-label { flex: 1 1 150px; min-width: 0; font-size: 12.5px; line-height: 1.4; color: var(--mx-muted); animation: mx-in .4s ease-out calc((var(--total) + var(--ctotal)) * var(--ms)) both; } /* The odd row out. A miss was styled exactly like a hit, so the one lane that reused nothing read as just another warm pass. Give it the ink colour so the eye lands on it; the bar keeps its real length, which is the finding. */ .mx .mx-cache.mx-cache-miss .mx-cache-label { color: var(--mx-ink); } /* Numbers arrive when they become true, then stay. */ .mx .mx-tps { animation: mx-in .4s ease-out calc(var(--ttft) * var(--ms)) both; } .mx .mx-done { animation: mx-in .4s ease-out calc(var(--total) * var(--ms)) both; } .mx .mx-lane.mx-fail .mx-done, .mx .mx-lane.mx-fail .mx-tps { animation: none; } @keyframes mx-in { from { opacity: 0; transform: translateY(3px); } } @keyframes mx-in-b { from { opacity: 0; transform: translateY(3px); } } .mx .mx-axis { display: flex; justify-content: space-between; font-size: 12px; color: var(--mx-muted); padding: 4px 0 2px; border-top: 1px solid var(--mx-rule); } /* Gantt variant: thinner tracks, names on the left, one row per stage. */ .mx .mx-gantt .mx-lane { display: grid; grid-template-columns: 96px 1fr; align-items: center; gap: 4px 10px; padding: 3px 0; } .mx .mx-gantt .mx-lane-head { display: contents; } .mx .mx-gantt .mx-lane-name { font-size: 12.5px; font-weight: 500; line-height: 1.3; } .mx .mx-gantt .mx-tps, .mx .mx-gantt .mx-done { display: none; } .mx .mx-gantt .mx-track { height: 22px; margin: 0; border-radius: 2px; } .mx .mx-gantt .mx-write i::after { top: 4px; bottom: 4px; width: 2px; } .mx .mx-gantt .mx-heat-t { padding: 12px 0 4px; } .mx .mx-gantt .mx-heat-t b { font-size: 15px; } .mx .mx-gantt .mx-heat-t .mx-done { display: inline; animation-delay: calc(var(--hend) * var(--mx-s) / var(--speed)); } /* Cards: one machine per card, the headline number big, two facts small, a bar for scale. */ .mx .mx-cards { display: grid; grid-template-columns: repeat(auto-fit, minmax(150px, 1fr)); gap: 10px; padding: 14px 0 0; } .mx .mx-card { position: relative; padding: 14px 14px 12px; border-radius: 12px; background: var(--mx-track); --bar: var(--mx-ink); } .mx .mx-card.mx-n { --bar: var(--mx-accent); box-shadow: inset 0 0 0 1.5px var(--mx-accent); } .mx .mx-card.mx-p { --bar: var(--mx-third); } .mx .mx-who { display: flex; align-items: center; flex-wrap: wrap; gap: 4px 8px; font-size: 14px; font-weight: 600; line-height: 1.3; color: var(--mx-ink); } .mx .mx-who small { font-size: 12px; font-weight: 400; color: var(--mx-muted); } .mx .mx-chip { display: inline-flex; align-items: center; justify-content: center; flex: 0 0 auto; width: 34px; height: 34px; margin-right: 2px; border-radius: 8px; background: var(--mx-track); box-shadow: inset 0 0 0 1px var(--mx-hatch); font-size: 10px; font-weight: 700; letter-spacing: .01em; color: var(--mx-text); text-align: center; line-height: 1.1; } .mx .mx-card .mx-chip { background: transparent; } .mx .mx-chip[data-chip="m3-ultra"] { background: url(data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEgAAABICAYAAABV7bNHAAAAAXNSR0IArs4c6QAAAERlWElmTU0AKgAAAAgAAYdpAAQAAAABAAAAGgAAAAAAA6ABAAMAAAABAAEAAKACAAQAAAABAAAASKADAAQAAAABAAAASAAAAACQMUbvAAAIrElEQVR4Ae2ae4wdVR3HtytIeRSolAIVpF0gqZIIQlvjIrZupUGsEZJSqoIQrIk1JsSoCYJpLKl/gJJAxNAoQYIYSRANCFQItOUREYUqbwGR7V2eRZS2QosW8PO5dw45HWfu3Ln3bnX3zi/53PN+fee8ZnYn9BXbVLIMwgI4BvaBsWzb6PzDsBbWwAZoy3an1FlwP7w9ThlhXOfCNChlc8i9DsarMOlxPcJYT4WW7HRy/R3SlYz38FuM+QLYFXJtCSmvwXgXo9n4VuapczQJL/e4OEG409Ii7UXErZU476ycGlocEou0iIBrMChYuX19FwaBDsJzXyXOf02OF9BkoJ+fuTAbKttRgQMJzleg+TBhx7QqlCgwpECeXpVlK3CUAu2dnVbFqo0CeXpVlq3AWwpUWRMFRkugYdq8s0m7YyZpl1Ho6fXUeT68Cx4dhfp3apXdnkG30/vPwhOwcKeOZBQbe5y6u/FqsYV6jk36eTiuH6O6Ue//so5aN2fQMIL4HekE8FPmwTDmzRu0M2hmyZF4NbgLngbf5XxdmQjPwyT4PTwFfiXwe/YRkLY3iHg9iZycTswI26azVHdPeDekzTrXw0Ywj+0eCu2aq6AuUJlprKCfBAev+QVuFiyDc8AP+/HmP53wd0Ex4nYuJ/xBOAouS6XF+YJ/BXm89Zv/uoz8VxBnP+xPsOl47JcPLtRTxq1RrpRAfyH/kRYqaQp2L8Sd82N5sA/g2Qpxeuy3o/GH9VjQf5PmAwimiB4Q8yC8Y34Yv8s/rrMVfymB/kUDn4Oy5nF/CbwJcaeWpypanUqP816Zyrsqyns//iCEQrkMQ9mf4t8DtEshxLfqltqkN9DAL22ppLkHufT6c8qF+KtJt+NZ9jMiFSFeuiHfn/H49W8+nAdh6ePtWwLhU475SltWg3mV2MC2vMQm8YubpJl0EjwLt4D3p/SB8SBxd8CJ4D52F8R2MoEh2C2OTPwuP2e+NqPhlPsNT6+VUv61ox2bWlDI/ecrsAluy8j74yTuC7iHZaR7WnmSvgc84RREYV6FC8G9z8PA8qWtjEB507+oUWdHM1P4ReAMuBriWfo3wjfDAfAJKHpIinFKwiDu5fBNuAmso7SVEWigdO2NAuFIziu+nYTJ8Bl4AP4AwVxaw+Cg9wfvOc1sI4nOQpfr4+BsUtSXoS0rI5AboVO5rP2OAsshbwaG+CVJxd5nNJfLz8FTcCm0Yp8m0z/AI12hLbcKhuBGaMtU2k4W4dMIgyjbkIP8PFhHaEfRtC+BcT7lI2BfeAW8c+0KH4FQxlmmOegQV+R+uV6icZpZb1H+OL3UMW9nT4ddkgbLOG+S+TlQqDybQoJ7kZvrzXADuNmeDc3sMRLXw4s5mWYl8Y/g1nLy5EaXWWJWsgA+nltbfsJEkr4BRe0tJo95L4VLwI21qD3r9SvCWZBlrhBtP9i77ivxU9ThdFXOopUQX8bSebLCZxLp+1uRHU2G48HNegQWwgA0M68Impuzy+lpcA96CS6DH4Hm8pyup6y1ugfFa/MqGvFJx+bSyxLuOOLtcFxe/7dAU7w47dp6bONndSrtxCTNgYcybuYrIH7YBxOO3/aHCNcglGnVrbknfBWmQBnzSdsJ17+udw0HfDacBGGf+hR+B+NSSdsmIt4Al+2HosSD8LthzwNnRD8EM95l4myckUROwJ0LPojt4LvYC+DDGoRz4TtwIJS1zVbuDJpZtmSSfyuudUzMKL8tJz7O6pO0fDftdSpzc1fY3SE8LLylbaSTwrZmB/IsS7R03m6LY/17pBvpJBxP307qGbdlK4EKHm0lUCVQgQIFydUMKhCo01PM6q+DO2EFeJ0P9lc8K8G7zBy4GP4J3kt2A+1PcFHd13gx9QTymPal1lu7d6v74G7wzuSl0HjvPYsg/eqwgbjvw1T4GmRdXIlu3boxg9bS3A9hc6pZL2s/Ae9Z2rVwFXhHCeZF1RvvJPDVYhXUYE9QCAX7BViP4pjPzxlfTHgVN7abCHgxXQ5/jBM68TuAtzvgHMo6E59J1fFbwto1YP0fhZngTTervd8Qr90CcfoJhA8FZ5/xfhlwlmiKF/J6MZ0H3rKPBPsV0tp1S33uoL1RNQehBbcRavx6oQyzXfdjSeJw4uo8A+vg63AGOGMVrSMLjXZUSZcKu4S04DZCjVcRZ43Lyf3JGejM0WY1nPqvM9X3uGNhCFyKt0NH5tL4fzdfZ9yfHLj7lQ/1FXCfmQuaAvp59hTYF2aDfwK/AhZC29YNgcITT9cVwiG93U56ok2BC8AN/tswAJ6awe7B4wnmjFEkl+ST8Cz4BWB/aMt8Gp2af4/aDu4BsT2RBPaLIm2v2QtulPUdr6J4ei2FZXAePARhU8fbdz14Ih4HXiFkCBToRujIOj3FHqN1p7Un1A3wFFwD0+D94F7gxjsIivUruBVWw22wEUz/NWjWYTgwD797S6jH5fVemAOvgeEZcDKEMsGdnMQrcogr49YoV+q/O/Iq92hWIKeynXVWKcijEMoswX8ADMBhiWuZNWAel8dUULRQRncxuBlviuJ/gN+8d8ADYL3uQXE5/efDIeBA02mthGuuVWeQHe3U3CjvAT+vToPZ0A/BXsIT7jJ2TrN9Z4cXw63wPFg2XoaWcwYYH+qzrWFwhri03GcUPqTjrZt1uszeBy67sjbSTYHKNj4W8o+kFR8Lnd6pfawEKpC7EqgSqECBgmRnUDWL8kXqV5wt+ek9n7JFgR7seRnyBXhIgcJNNj9b76asUaB1sL53Ncgdue+IvsrUzfekVt5NeinPxYk2dWcSv6rVSwI0G+tzaDEddjBfLn3RbFawV9LO2EGZKHAmfj9094oQWeP8XqRHpncpsZt7VKSLGHdLn0WOJ+O9PSTSk4w1d1mRlml7EbsMHoasqTge4l5kbCvAD2qZ5gezIvPT5iAsgGNgHxjL5h7rQ18LXpI3QK79ByQblUuCBBPcAAAAAElFTkSuQmCC) center / contain no-repeat; box-shadow: none; font-size: 0; filter: var(--mx-chip-filter); } .mx .mx-chip[data-chip="m5-ultra"] { background: url(data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEgAAABICAYAAABV7bNHAAAAAXNSR0IArs4c6QAAAERlWElmTU0AKgAAAAgAAYdpAAQAAAABAAAAGgAAAAAAA6ABAAMAAAABAAEAAKACAAQAAAABAAAASKADAAQAAAABAAAASAAAAACQMUbvAAAGL0lEQVR4Ae2b/XHbOBDF7Uz+P6UC0x3oKjBSwekqMDuIU0GYCpyrQEoFVioQU4HVgXgVWB3kfk+HdRAEJEj6484R38zzLt4uFsAKlGTP+OQkjxkpV/AG3sFvL5w6g86iM+lso6HJS/jSG5Lbv844uFELJv0KtyXXHIvrrDpzL1yTZROPzVa5Dh1zc+wyqAdJ6IpZ0rHb+8ft1Ldqht1B2QknJ3uacA73r3w3dK2m5vhm+F4cHjXdIDXm7nts8oIOvNENKgNhcn/sQKkGXfyoTaOgAxd6xPR4Te8/QVcCd68G6SN9QksH7FOsJTzJU4Myd+CpGqQvWh/hm8z6//vw6yfa4VvqbuH8ieo/a9nH/r2rCnZ/g//Y9Z+73qMfoPAN0lf15z7Mo6/3FI9YSWMuYQFfPB7yPWjF6b/AGuq9Rk1ZwBlcQcXWUGPp7+AchtAbueGDOR22JvbVxy+wzvsyNbSYximcIZapQJf2jeBQDl7Eb2AZreXlg+nzfuWCCRV+uG+Nc3AkhHOy/phHTK/6Cg6FblJ8g8IanxnoprWhIVC3BQPd4YspnKXEnJbtIgUs5w5fBx2DDZOsjtm4zi6RY7lVlKyxxWQdFCoY6g/yh35RXLP4Hg7FnAmuY1LhY7pFbVgR0IsjduGiKzg0NrRBfw9dwOe7zDx9JdDBVy15a/QGLuAc9kVNYtM3OZU3tEGpGo+hqTk6fAPXMIbdLH0StqH2AdV6D0/hW3jumapLqBvP1aC6exuH6KXP0deDEA0DHW7uienER6LKd54zbAP/hCs4GEPexPRRPBYbJsZrWS2LFV7Qh4Hl6vETZKU5KFTQcmQNJU6oq9bCB9WssHaY1+b/UKwtyXQVH4uCibfQaskaNjgaX3tB1vIKr2ltac6PK6zlyG48Q818zTUscUzvYwclq6C9GrbgEKtX8Bbaxmzuxmt2kMKPpQsltDkOX6igaX3sXJNABfvkH3LGvAddapUHwDaaKqEGLmADa/gZCrk1G3JqT0wS+6TaQ+zdTWpZrutRN5VSIVoNWcMGx/QbLzpvC6zFZB0UKmj6RoLHDmu6Wd1ag+qb3scOSraCWlCv9hDMSbb5Zm3+JooVFsBeRzHnY1WkF17XOmE9+dKE1B5sL232p023Jcb68t81e/2ckbWDcQ2bvIlilQWw8TznYxU2rJd70bQH5YRz+viDJ4RFb1jQeS6xd1DxHayggwuosfSYDk2MN6586VcwniNNsWUipvVLqGYY5JfQ9hbX6xyfMlEJE1o68KpFn2TfgalBmaswNWhqUKYDmfB0g6YGZTqQCb/OxNvCWwJ76BIJDZo4h/oOIl+0Me496nsv7RTIYg1jpOqFOVsGe1h4Ysah84sSJVNx55dKxSof22AVj8fhHJ/aajRX+W24IhDWC/3CTyo7csL8pD/2Bvm1H2w2QQX9mVSveqgVQXyOfx2M9Zv+Jz8OdUk1bKDmrOESjkayc1Tr0p1fLZVT+dgGq3g8Ts2R5qCQikt3iViBNkvoJVoBb6GwhKm6We2lf4oVHHwPY+jWLOAcFjD+OzdSP7zkBm05Yg3ViBBrBnt44UXFTfNSf/OSGrTlWG8D/o4/gx9gCN0W6QsvXnq79naQ+a/fpIdsVod2fsInb3dY6YY9zgoW8CMM8ReDMhT6+GMbZJvShsy39aQ9BQqK2m35Df89XMMSGjQWGljBEFsGDSxgb4x9xP7wK+hVCaHmaJMz6OBT4YrCBVSTtKbhC47W/hbRPs3WljjExsX6jud+kQW2glewgMINtDqVBKB8F1EbtzzFBBuHVrqLYkuJoILK3UGhhOFc8wt00ca97NhHjHUOX+h0g1ZwDQU1awkdNJzhOBt0WDWwDY5AHC/RdGO2UJB18BKm8A5R+Q0sYC+ckqVOTmjpwKsWfZJ9B6YGZa7C1KCpQZkOZMK6QftMzjGHD//1XB9zBzJnr3WDvmaSjjn8Vd+DZvDumLvQcfbDv4XrPWjVkXSsIfXk8E+9aoBu0c5bjY8dujTnUPYeCzz92jHx+x/b7ptjzvXUoJ/+jmS9ubfVETdJZ+8FPW76ZDuWx01n1ZkHYUb2Ev7qTdIZddbR0OQreAN/hVulM+gsOlO2Mf8AEF4rddy0K8gAAAAASUVORK5CYII=) center / contain no-repeat; box-shadow: none; font-size: 0; filter: var(--mx-chip-filter); } .mx .mx-chip[data-chip="m5-pro"] { background: url(data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEgAAABICAYAAABV7bNHAAAAAXNSR0IArs4c6QAAAERlWElmTU0AKgAAAAgAAYdpAAQAAAABAAAAGgAAAAAAA6ABAAMAAAABAAEAAKACAAQAAAABAAAASKADAAQAAAABAAAASAAAAACQMUbvAAAF9UlEQVR4Ae2b4VXbShCFIScFOBVEdOAO2FTw/CrAHUAqiKjA51VgpwKTCiQqgA6sVwG8Cnj389khy0bSWrJMAtY953pmZ2ZHu1cr2fzg5CSNiUquxLX4ID69cbIH9sKe2FtvMHkpvnVBUutnj52FmmnSezgtKXEsz17Z805YqMomHptl7604ZnHsMORNCnHErOjY7fPjdurVmshuROyIk5NHiXAmPn7wavBojeJ4MbwW2/cRJwhhHn7mRi9Q4BMnaB4ERvelAnMEOn8ZG0eBAuc8Yjxe4/snUCVwHxGIr/QRDQrYt1hDegyPAiXOwKEE4ofWtfgpcf0/Pv3xACtEnC/ivTg9QP9Xbzn03115sIO1/KH7v3a/wTeQeYHydyDO0yEesbmEuRAz8c1jn99BK+3+h1iKvGsQZSZOxJVI7kZkTPxSnIoheJEbvpnTYkvlbn3+XNZ5H1OKlmNch88KzusSbbEnJbuy80X8ApbRtXx4a3Z5X7lgQi4/XDfjFJwKwjlJv88jxl1fiV3BSYpPUNjjuwactCZUSpRNySDu5MM6fK4LpmJJFdXAavb5u60I+li/eG2bmhqrzaNixpbDOhHkYhjfy+/6Q/FGF38Uu4KT41omZT7HKWrCSomJZ1MN8fO2ZNdcV4H+7XoBX49AbVgoyeZXDUU3ilfiTEz1UskzSnnV86iH01WgHpfYTmHzbSDP5ivxRoxhJ4tvwiaUPkGvr+Kp+EU886zrq1Q7Xkugsn0Z2+yFr+HnQYhKAzY39ZRpxbWy1DvPiWwl/i2uxM7o8hLjq7gvCk2Mr2W9LJf5AF8GVsvjB7DEnAhy0WqwhrmcME6vmU8iVtg7rGvyXzRrKrI4zfsi08Q70XphDYUcxgsfwFpd5mNcm5jz41zWarCFZxgzn7mGpRyL72I7FdPQ7oZdsIvlDt6JtjCbW/iYbSTzY+JgLtocJx/kosV2sU71IBd3qd/W9HkHXXCVPTBtmYuA3IBKLMXvIkhds1JN6SlTi6o2ukNwZzXVy2rdDn3rShZBD3oZCjnWe+2DzttM1nJYJ4JctHhBwGMja3GznFoD/S2+i+1UbA25IHe7C5yKbb5Zm19EucwSsoso53wuj+KZj3NCw374xADWrr2r7TzBGi+3l9ztAzE3NYuz2UWUyy0hG89zPpfLPgVM3TTWQE04Zxe/84Sw6VoXdJ5L2QeR/EbMRSfORMbEYzrFYLxw6olfifEcYuSWNTmuPxcRw4DPHFtb3K91fKqJFIxoUOBDQ3wMewVGgRJHYRRoFCihQCI9nqBRoIQCifTHRH6fdFkz2UWxSmMYgt8t0zAQ+fcaQ0BdW+22aN+P1h9Kat43X7euiYKLoGdeV6RYJjb9eIynIFBc23fNdfN6C1DXLIyxERZfBHTywVqkNhcBolkdOYSEG5G6Bz8mRi1xYksxrqV+SA7aLFyY1nniosWyKeBEanMRFGI4d7mN/jxtcz9eR3XMYS6YiWGPQfzX/hbjboNq+9n8kfnUf97eyBJDhBhOgalIzeB4bYEe/Q7YUBv+8cm/ZJkDM7EJM58omwr6xg/5Lcaa2FiJI1SibfySQICv8id+XMnCuTgV78XfhkMLxOa+BLvL5PMecWIIhMhE6ivxSlyIgFwKlS+YpAr75Ad5menCcR/W4mriYV2uPChE4g9i5hnWmUjkw7j5zEEcGw9mX/sdpD20gk1+EyvxWjRceodHMQaxSryKE0ONB1NbCwp7sT4XxcI8fi6CQgxzmcaIFZ6YucYgE3NPJwucGM4f0j9YYxbNXW1b7FJ56u6iurWPY8P51E9FAz6xsGZQ/9Q3twuONlLgT3sHRcv7/cNRoMQ9GAUaBUookEhzgh4TNcec3v7Xc3nMCiT2XnKCbhNFx5y+5XfQROQX64hfFdj+WzjvoNWvuaOPoMn2n3pRglO08ZbxsYNDcyZinzGTN+jfMW+4H1rUYqHosYuEBq04ZpHyVmWCJEeMb7ZjOU3stfGxCnR54U40WorvXST2yF57g8lX4lp8D6eKPbAX9pQU5n/QkdHiSuCUMgAAAABJRU5ErkJggg==) center / contain no-repeat; box-shadow: none; font-size: 0; filter: var(--mx-chip-filter); } .mx .mx-big { display: flex; align-items: baseline; flex-wrap: wrap; gap: 0 6px; padding: 10px 0 4px; font-size: 40px; font-weight: 700; line-height: 1.1; letter-spacing: -0.03em; color: var(--mx-ink); animation: mx-in .5s ease-out calc(.15s * var(--i)) both; } .mx .mx-big span { font-size: 13px; font-weight: 500; letter-spacing: 0; color: var(--mx-muted); } .mx .mx-big em { font-style: normal; font-size: 13px; font-weight: 700; letter-spacing: 0; color: var(--mx-accent); padding: 1px 7px; border-radius: 999px; background: var(--mx-soft); } .mx .mx-cardbar { display: block; height: 6px; border-radius: 3px; background: var(--mx-rule); overflow: hidden; margin: 2px 0 10px; } .mx .mx-cardbar i { display: block; height: 100%; width: 100%; border-radius: 3px; background: var(--bar); transform-origin: left; transform: scaleX(var(--r)); animation: mx-grow .9s cubic-bezier(.2,.7,.2,1) calc(.15s * var(--i)) both; } @keyframes mx-grow { from { transform: scaleX(0); } } @keyframes mx-grow-b { from { transform: scaleX(0); } } .mx .mx-fact { display: flex; justify-content: space-between; gap: 8px; font-size: 12.5px; line-height: 1.5; color: var(--mx-muted); border-top: 1px solid var(--mx-rule); padding: 4px 0; } .mx .mx-fact b { font-weight: 600; color: var(--mx-text); white-space: nowrap; } .mx .mx-panel { display: none; } .mx .mx-panel.mx-solo { display: block; } /* Bars: rows of grouped horizontal bars, one lane per machine, grown with a transform. */ .mx .mx-rows { padding: 8px 0 0; } .mx .mx-row { padding: 12px 0 4px; } .mx .mx-row + .mx-row { border-top: 1px solid var(--mx-rule); } .mx .mx-rh { display: flex; flex-wrap: wrap; align-items: baseline; justify-content: space-between; gap: 0 10px; padding-bottom: 6px; } .mx .mx-rh b { font-size: 14px; font-weight: 600; color: var(--mx-ink); } .mx .mx-rh span { font-size: 12.5px; color: var(--mx-muted); } .mx .mx-bl { display: grid; grid-template-columns: 74px 1fr auto; align-items: center; gap: 10px; padding: 3px 0; --c: var(--mx-ink); } .mx .mx-bl.mx-n { --c: var(--mx-accent); } .mx .mx-bl.mx-p { --c: var(--mx-third); } .mx .mx-bl.mx-q { --c: var(--mx-fourth); } .mx .mx-bl .mx-bn { font-size: 12.5px; font-weight: 500; color: var(--mx-muted); line-height: 1.25; } .mx .mx-bl.mx-n .mx-bn { color: var(--mx-ink); } .mx .mx-bl .mx-bt { position: relative; height: 18px; border-radius: 3px; background: var(--mx-track); overflow: hidden; } .mx .mx-bl .mx-bt i { position: absolute; inset: 0; width: 100%; border-radius: 3px; background: var(--c); transform-origin: left; transform: scaleX(var(--r)); animation: mx-grow .9s cubic-bezier(.2,.7,.2,1) calc(.08s * var(--i)) both; } .mx .mx-bl .mx-bv { font-size: 14px; font-weight: 600; color: var(--mx-text); min-width: 64px; text-align: right; white-space: nowrap; animation: mx-in .4s ease-out calc(.08s * var(--i) + .5s) both; } .mx .mx-bl .mx-bv small { font-size: 11.5px; font-weight: 400; color: var(--mx-muted); } .mx .mx-bl.mx-n .mx-bv { color: var(--mx-accent); font-weight: 700; } .mx .mx-bl.mx-fail .mx-bt { background-image: repeating-linear-gradient(135deg, var(--mx-soft) 0 3px, transparent 3px 9px); } .mx .mx-bl.mx-fail .mx-bt::after { content: attr(data-why); position: absolute; inset: 0; display: flex; align-items: center; padding: 0 8px; font-size: 11.5px; font-weight: 500; color: var(--mx-accent); } .mx .mx-bl.mx-fail .mx-bv { color: var(--mx-muted); font-weight: 500; font-size: 12.5px; animation: none; } .mx .mx-bl.mx-fail .mx-bt i { display: none; } /* Replay: the real answer, each streamed chunk at the second it arrived. */ .mx .mx-typ { display: grid; grid-template-columns: 1fr; gap: 14px; padding: 12px 0 0; } .mx .mx-typ .mx-pane { --c: var(--mx-ink); min-width: 0; } .mx .mx-typ .mx-pane.mx-n { --c: var(--mx-accent); } .mx .mx-typ .mx-pane.mx-p { --c: var(--mx-third); } .mx .mx-pane-h { display: flex; align-items: baseline; justify-content: space-between; gap: 8px; padding-bottom: 6px; border-bottom: 2px solid var(--c); font-size: 14px; font-weight: 600; color: var(--mx-ink); } .mx .mx-pane-h span { font-size: 12.5px; font-weight: 500; color: var(--mx-muted); animation: mx-in .4s ease-out calc(var(--end) * var(--mx-s) / var(--speed)) both; } .mx .mx-txt { padding: 8px 0 0; font-size: 14px; line-height: 1.55; color: var(--mx-text); min-height: 4.6em; } .mx .mx-txt span { animation: mx-show 0s linear calc(var(--t) * var(--mx-s) / var(--speed)) both; } .mx .mx-txt br.mx-p { display: block; content: ""; margin-bottom: .55em; } .mx .mx-txt span.mx-think { color: var(--mx-muted); font-style: italic; } .mx .mx-txt .mx-cur { display: inline-block; width: 2px; height: 1em; margin: 0 0 -2px 1px; background: var(--mx-accent); vertical-align: baseline; animation: mx-blink 1s steps(2, jump-none) infinite, mx-off 0s linear calc(var(--end) * var(--mx-s) / var(--speed)) forwards; } @keyframes mx-show { from { opacity: 0; } to { opacity: 1; } } @keyframes mx-show-b { from { opacity: 0; } to { opacity: 1; } } @keyframes mx-off { to { display: none; opacity: 0; } } @keyframes mx-off-b { to { display: none; opacity: 0; } } .mx .mx-wait { display: block; margin: 0; font-size: 12px; color: var(--mx-muted); padding: 6px 0 0; } .mx .mx-wait i { display: inline-block; vertical-align: middle; width: 90px; height: 5px; margin-right: 6px; border-radius: 3px; background: var(--mx-track); overflow: hidden; } .mx .mx-wait i::after { content: ""; display: block; height: 100%; width: 100%; background: repeating-linear-gradient(135deg, var(--mx-hatch) 0 2px, transparent 2px 5px); transform-origin: left; animation: mx-grow calc(var(--ttft) * var(--mx-s) / var(--speed)) linear both; } /* Footnotes */ .mx .mx-c { font-size: 0.875rem; line-height: 1.5; letter-spacing: normal; color: var(--mx-muted); padding: 10px 0 0; } .mx .mx-tail { margin: 8px 0 0; font-size: 0.875rem; line-height: 1.5; color: var(--mx-muted); font-style: italic; } .mx .mx-more { margin: 10px 0 0; padding: 0; border: 0; border-top: 1px solid var(--mx-rule); background: none; font-size: 0.875rem; color: var(--mx-text); } .mx .mx-more summary { display: flex; align-items: center; min-height: 44px; margin: 0; padding: 8px 0; cursor: pointer; list-style: none; font-size: 0.875rem; font-weight: 400; line-height: 1.5; color: var(--mx-muted); } .mx .mx-more summary::-webkit-details-marker { display: none; } .mx .mx-more summary::before { content: ""; width: 6px; height: 6px; margin-right: 10px; border: solid currentColor; border-width: 0 1.5px 1.5px 0; transform: rotate(-45deg); transition: transform .2s; } .mx .mx-more[open] summary::before { transform: rotate(45deg); } .mx .mx-more summary:hover { color: var(--mx-ink); } .mx .mx-more summary:focus-visible { outline: 2px solid var(--mx-accent); outline-offset: 2px; border-radius: 4px; } .mx .mx-foot { font-size: 0.875rem; line-height: 1.5; color: var(--mx-muted); padding: 0 0 10px; } .mx .mx-foot + .mx-foot { padding-top: 0; } .mx .mx-foot code { font: inherit; word-break: break-all; } /* Play again: the -b twins. Every animated element is listed once. */ .mx:has(.mx-play-b:checked) .mx-clock { animation-name: mx-count-b; } .mx:has(.mx-play-b:checked) .mx-read { animation-name: mx-reveal-b; } .mx:has(.mx-play-b:checked) .mx-write i { animation-name: mx-slide-b; } .mx:has(.mx-play-b:checked) .mx-write i::after { animation-name: mx-on-b, mx-blink-b; } .mx:has(.mx-play-b:checked) .mx-tps, .mx:has(.mx-play-b:checked) .mx-done, .mx:has(.mx-play-b:checked) .mx-cache-bar, .mx:has(.mx-play-b:checked) .mx-cache-label, .mx:has(.mx-play-b:checked) .mx-big, .mx:has(.mx-play-b:checked) .mx-bv, .mx:has(.mx-play-b:checked) .mx-pane-h span { animation-name: mx-in-b; } .mx:has(.mx-play-b:checked) .mx-cardbar i, .mx:has(.mx-play-b:checked) .mx-bt i, .mx:has(.mx-play-b:checked) .mx-wait i::after { animation-name: mx-grow-b; } .mx:has(.mx-play-b:checked) .mx-txt span { animation-name: mx-show-b; } .mx:has(.mx-play-b:checked) .mx-txt .mx-cur { animation-name: mx-blink-b, mx-off-b; } /* Play on arrival, no JavaScript: a scroll-driven animation on the figure flips --mx-seen to 1 once its top edge is 140px into the viewport, and a style query resumes every paused animation below it. Browsers without view timelines (Firefox, Safari before 26) keep playing on load. Scrolling back up past a figure pauses it where it is. */ @supports (animation-timeline: view()) { .mx { animation: mx-seen steps(1) both; animation-timeline: view(); animation-range: entry 0% entry 140px; } .mx .mx-clock, .mx .mx-read, .mx .mx-write i, .mx .mx-write i::after, .mx .mx-tps, .mx .mx-done, .mx .mx-cache-bar, .mx .mx-cache-label, .mx .mx-big, .mx .mx-cardbar i, .mx .mx-bt i, .mx .mx-bl .mx-bt i, .mx .mx-bv, .mx .mx-bl .mx-bv, .mx .mx-txt span, .mx .mx-txt .mx-cur, .mx .mx-pane-h span, .mx .mx-wait i::after { animation-play-state: paused; } @container style(--mx-seen: 1) { .mx .mx-clock, .mx .mx-read, .mx .mx-write i, .mx .mx-write i::after, .mx .mx-tps, .mx .mx-done, .mx .mx-cache-bar, .mx .mx-cache-label, .mx .mx-big, .mx .mx-cardbar i, .mx .mx-bt i, .mx .mx-bl .mx-bt i, .mx .mx-bv, .mx .mx-bl .mx-bv, .mx .mx-txt span, .mx .mx-txt .mx-cur, .mx .mx-pane-h span, .mx .mx-wait i::after { animation-play-state: running; } } } @keyframes mx-seen { from { --mx-seen: 0; } to { --mx-seen: 1; } } /* Host-proofing, headings. The outline inside a post runs h2 (section) to h3 (figure title) to h4 (a label inside a figure), with no h1 and no level skipped. MacStories styles every h1-h6 under .post-content with a three-class selector, read out of main.css on 2026-09-19: .post.type-post .post-content h1 ... h6 { line-height: 1; font-weight: 700; } h1 2em, h2 1.75em, h3 1.5em, h4-h6 1.25em margins: h1 and h2 4rem/1rem, h3 3rem/1rem, h4-h6 2rem/1rem plus clear: both on the same descendant selector, inside @media screen and (min-width: 768px). It sets no colour. That media query is worth knowing: enumerating a stylesheet's top-level rules misses everything inside an at-rule, and a text search finds the declaration without showing that it is conditional. The clear is inert here because .mx already clears, so nothing in a figure can meet a float. Three classes outranks the two-class rules above, which was rendering the in-figure labels at 18.75px bold. Doubling a class buys the specificity back without an !important, and it holds whether the figure is pasted in the block or on its own. These restate colour and box as well as type, so a host that later adds either cannot reach in; the gantt variant is doubled in turn so the base rule does not outrank it. This covers the headings inside a figure, which the figure designs. The section headings between figures are the other case, and they are left reachable on purpose - see .ms-widget .mx-part .mx-h2 below. */ .mx.mx .mx-t.mx-t { font-size: 1.25rem; font-weight: 700; line-height: 1.2; margin: 0; padding: 0; color: var(--mx-ink); } .mx.mx .mx-heat-t.mx-heat-t { font-size: 13px; font-weight: 400; line-height: 1.4; margin: 0; padding: 8px 0 0; color: var(--mx-muted); } .mx.mx .mx-gantt .mx-heat-t.mx-heat-t { padding: 12px 0 4px; } @media (min-width: 560px) { .mx .mx-track { height: 54px; } .mx .mx-write i::after { top: 15px; bottom: 15px; } .mx .mx-clock { font-size: 34px; } .mx .mx-gantt .mx-track { height: 24px; } .mx .mx-gantt .mx-write i::after { top: 5px; bottom: 5px; } .mx .mx-gantt .mx-lane { grid-template-columns: 120px 1fr; } .mx .mx-typ { grid-template-columns: repeat(var(--cols, 2), 1fr); gap: 18px; } .mx .mx-bl { grid-template-columns: 110px 1fr auto; } .mx .mx-chip { width: 40px; height: 40px; border-radius: 9px; } } @media (prefers-reduced-motion: reduce) { .mx .mx-clock, .mx .mx-read, .mx .mx-write i, .mx .mx-write i::after, .mx .mx-tps, .mx .mx-done, .mx .mx-cache-bar, .mx .mx-cache-label, .mx .mx-big, .mx .mx-cardbar i, .mx .mx-bt i, .mx .mx-bl .mx-bt i, .mx .mx-bv, .mx .mx-bl .mx-bv, .mx .mx-txt span, .mx .mx-txt .mx-cur, .mx .mx-pane-h span, .mx .mx-wait i::after { animation: none; } .mx .mx-txt .mx-cur { display: none; } .mx .mx-clock small, .mx .mx-play { display: none; } /* "Playback 7× faster" describes an animation that is not running */ } .mx:has(input[name="s1race-m"][value="flash"]:checked):has(input[name="s1race-l"][value="l4k"]:checked) .mx-race[data-m="flash"][data-l="l4k"] { display: block; } .mx:has(input[name="s1race-m"][value="flash"]:checked):has(input[name="s1race-l"][value="l16k"]:checked) .mx-race[data-m="flash"][data-l="l16k"] { display: block; } .mx:has(input[name="s1race-m"][value="flash"]:checked):has(input[name="s1race-l"][value="l64k"]:checked) .mx-race[data-m="flash"][data-l="l64k"] { display: block; } .mx:has(input[name="s1race-m"][value="glm"]:checked):has(input[name="s1race-l"][value="l4k"]:checked) .mx-race[data-m="glm"][data-l="l4k"] { display: block; } .mx:has(input[name="s1race-m"][value="glm"]:checked):has(input[name="s1race-l"][value="l16k"]:checked) .mx-race[data-m="glm"][data-l="l16k"] { display: block; } .mx:has(input[name="s1race-m"][value="glm"]:checked):has(input[name="s1race-l"][value="l64k"]:checked) .mx-race[data-m="glm"][data-l="l64k"] { display: block; } .mx:has(input[name="s1race-m"][value="flash"]:checked) .mx-alt[data-m]:not([data-m="flash"]) { display: none; } .mx:has(input[name="s1race-m"][value="glm"]:checked) .mx-alt[data-m]:not([data-m="glm"]) { display: none; } .mx .mx-switch label .mx-alt { display: block; } .mx:has(input[name="s1cards-m"][value="flash"]:checked) .mx-panel[data-m="flash"] { display: block; } .mx:has(input[name="s1cards-m"][value="glm"]:checked) .mx-panel[data-m="glm"] { display: block; } .mx:has(input[name="s1long-m"][value="flash"]:checked) .mx-race[data-m="flash"] { display: block; } .mx:has(input[name="s1long-m"][value="glm"]:checked) .mx-race[data-m="glm"] { display: block; } .mx:has(input[name="s1replay-m"][value="flash"]:checked) .mx-panel[data-m="flash"] { display: block; } .mx:has(input[name="s1replay-m"][value="glm"]:checked) .mx-panel[data-m="glm"] { display: block; } /* Quant bars: the name cell is the quant or the Mac with a tag under it. A run with the embedding tables on SSD is striped, so the two kinds of run differ on screen. */ .mx .mx-qbars .mx-bl { grid-template-columns: 84px 1fr auto; } .mx .mx-qbars .mx-bn small { display: block; font-size: 11px; font-weight: 400; line-height: 1.3; color: var(--mx-muted); } .mx .mx-qbars .mx-bl.mx-ssd .mx-bt i { background: repeating-linear-gradient(135deg, var(--c) 0 5px, var(--mx-track) 5px 8px); } .mx .mx-qbars.mx-mem .mx-bt { background-image: linear-gradient(to right, var(--mx-rule) 1px, transparent 1px); background-size: calc(100% * var(--tick) / var(--max)) 100%; } .mx .mx-qbars .mx-note { padding-top: 4px; } .mx .mx-key .mx-k-ssd { background: repeating-linear-gradient(135deg, var(--mx-accent) 0 5px, var(--mx-track) 5px 8px); } @media (min-width: 560px) { .mx .mx-qbars .mx-bl { grid-template-columns: 118px 1fr auto; } .mx .mx-qbars.mx-mem .mx-bl { grid-template-columns: 140px 1fr auto; } } /* Quant bars: the name cell is the quant or the Mac with a tag under it. A run with the embedding tables on SSD is striped, so the two kinds of run differ on screen. */ .mx .mx-qbars .mx-bl { grid-template-columns: 84px 1fr auto; } .mx .mx-qbars .mx-bn small { display: block; font-size: 11px; font-weight: 400; line-height: 1.3; color: var(--mx-muted); } .mx .mx-qbars .mx-bl.mx-ssd .mx-bt i { background: repeating-linear-gradient(135deg, var(--c) 0 5px, var(--mx-track) 5px 8px); } .mx .mx-qbars.mx-mem .mx-bt { background-image: linear-gradient(to right, var(--mx-rule) 1px, transparent 1px); background-size: calc(100% * var(--tick) / var(--max)) 100%; } .mx .mx-qbars .mx-note { padding-top: 4px; } .mx .mx-key .mx-k-ssd { background: repeating-linear-gradient(135deg, var(--mx-accent) 0 5px, var(--mx-track) 5px 8px); } @media (min-width: 560px) { .mx .mx-qbars .mx-bl { grid-template-columns: 118px 1fr auto; } .mx .mx-qbars.mx-mem .mx-bl { grid-template-columns: 140px 1fr auto; } } .mx:has(input[name="s6tps-w"][value="prose"]:checked) .mx-panel[data-w="prose"] { display: block; } .mx:has(input[name="s6tps-w"][value="code"]:checked) .mx-panel[data-w="code"] { display: block; } /* Slim race: many lanes on one clock. Name and tag on the left, the bar, the finish on the right. Embedding tables on SSD are dotted while reading and striped while writing, and a mark is a second moment on the same clock, drawn when the clock reaches it. */ .mx .mx-slim .mx-lane { --c: var(--mx-ink); display: grid; grid-template-columns: minmax(84px, 24%) 1fr auto; align-items: center; gap: 2px 10px; padding: 4px 0; } .mx .mx-slim .mx-lane.mx-n { --c: var(--mx-accent); } .mx .mx-slim .mx-lane-head { display: contents; } .mx .mx-slim .mx-lane-name { grid-column: 1; grid-row: 1; display: block; font-size: 13px; font-weight: 600; line-height: 1.3; } .mx .mx-slim .mx-lane-name small { display: block; margin: 0; font-size: 11px; font-weight: 400; line-height: 1.3; color: var(--mx-muted); } .mx .mx-slim .mx-tps { display: block; font-size: 11.5px; white-space: normal; } .mx .mx-slim .mx-tps b { font-size: 12px; } .mx .mx-slim .mx-track { grid-column: 2; grid-row: 1; height: 26px; margin: 0; } .mx .mx-slim .mx-done { grid-column: 3; grid-row: 1; font-size: 15px; min-width: 52px; text-align: right; } .mx .mx-slim .mx-write i::after { top: 6px; bottom: 6px; width: 2px; } .mx .mx-slim .mx-heat-t { padding: 10px 0 2px; } .mx .mx-slim .mx-lane.mx-ssd .mx-read { background: radial-gradient(circle, var(--mx-hatch) 1.6px, transparent 1.8px) 0 0 / 7px 7px; } .mx .mx-slim .mx-lane.mx-ssd .mx-write i { background: repeating-linear-gradient(135deg, var(--c) 0 5px, var(--mx-track) 5px 8px); } .mx .mx-mark { position: absolute; top: 0; bottom: 0; left: calc(100% * var(--at) / var(--max)); width: 2px; background: var(--mx-accent-ink); animation: mx-in .3s ease-out calc(var(--at) * var(--ms)) both; } .mx:has(.mx-play-b:checked) .mx-mark { animation-name: mx-in-b; } .mx .mx-key .mx-k-ssd { background: repeating-linear-gradient(135deg, var(--mx-accent) 0 5px, var(--mx-track) 5px 8px); } .mx .mx-key .mx-k-ssdread { background: var(--mx-track) radial-gradient(circle, var(--mx-hatch) 1.6px, transparent 1.8px) 0 0 / 7px 7px; } .mx .mx-key .mx-k-mark { position: relative; background: var(--mx-ink); } .mx .mx-key .mx-k-mark::after { content: ""; position: absolute; top: 0; bottom: 0; left: 60%; width: 2px; background: var(--mx-accent-ink); } @supports (animation-timeline: view()) { .mx .mx-mark { animation-play-state: paused; } @container style(--mx-seen: 1) { .mx .mx-mark { animation-play-state: running; } } } @media (min-width: 560px) { .mx .mx-slim .mx-track { height: 30px; } .mx .mx-slim .mx-write i::after { top: 8px; bottom: 8px; } } @media (prefers-reduced-motion: reduce) { .mx .mx-mark { animation: none; } } /* Classic charts gallery: a grid of plain SVG charts, any one of which opens to full width. The control is a radio per card plus one "none" radio, so closing is a state and not a script. */ .mx .mx-gal { container-type: inline-size; margin-top: 14px; } .mx .mx-ggrid { display: grid; grid-template-columns: 1fr; gap: 14px; margin: 0; padding: 0; } .mx .mx-gcard { position: relative; display: flex; flex-direction: column; margin: 0; padding: 12px 12px 10px; border: 1px solid var(--mx-rule); border-radius: 12px; background: var(--mx-track); min-width: 0; } .mx .mx-gr { display: block; margin: 0; padding: 0; } .mx .mx-gtop { min-height: 63px; } /* These are elements: the host stylesheet must not get to set their margins. */ .mx .mx-gs, .mx .mx-gnote { margin: 0; padding: 0; } /* The card title is an h4, and MacStories styles every h1-h6 under .post-content with a three-class selector that beats a two-class one. Doubling the classes takes this to four without an !important, so the theme cannot size, colour or space the chart titles. */ .mx.mx .mx-gt.mx-gt { margin: 0; padding: 0; font-size: 1rem; font-weight: 700; line-height: 1.3; letter-spacing: normal; color: var(--mx-ink); } .mx .mx-gs { font-size: 0.875rem; line-height: 1.4; color: var(--mx-muted); margin-top: 3px; } .mx .mx-gplot { margin: 8px 0 0; padding: 0; } .mx .mx-svg { display: block; width: 100%; height: auto; overflow: visible; } .mx .mx-svg .mx-gl { stroke: var(--mx-rule); stroke-width: 1; fill: none; } .mx .mx-svg .mx-al { stroke: var(--mx-hatch); stroke-width: 1.5; fill: none; } .mx .mx-svg .mx-pl { fill: none; stroke: currentColor; stroke-width: 2.5; stroke-linejoin: round; stroke-linecap: round; } .mx .mx-svg .mx-dot { fill: currentColor; stroke: none; } .mx .mx-svg .mx-lb { font-family: inherit; font-size: 12px; fill: var(--mx-muted); font-variant-numeric: tabular-nums; } .mx .mx-svg .mx-ylb { text-anchor: end; } .mx .mx-svg .mx-xlb { text-anchor: middle; } .mx .mx-svg .mx-yl, .mx .mx-svg .mx-xl { text-anchor: middle; font-size: 12px; } .mx .mx-svg .mx-vlb { text-anchor: middle; font-size: 12px; font-weight: 700; fill: currentColor; } .mx .mx-svg .mx-ser.mx-a { color: var(--mx-ink); } .mx .mx-svg .mx-ser.mx-n { color: var(--mx-accent); } .mx .mx-svg .mx-ser.mx-p { color: var(--mx-third); } .mx .mx-gkey { display: flex; flex-wrap: wrap; align-items: center; gap: 3px 14px; margin-top: auto; padding: 10px 0 0; font-size: 12px; color: var(--mx-muted); } .mx .mx-gkey span { display: inline-flex; align-items: center; gap: 6px; } .mx .mx-gkey i { display: inline-block; width: 16px; height: 3px; border-radius: 2px; background: var(--mx-ink); } .mx .mx-gkey i.mx-k-n { background: var(--mx-accent); } .mx .mx-gkey i.mx-k-p { background: var(--mx-third); } .mx .mx-gnote { display: none; font-size: 0.875rem; line-height: 1.5; color: var(--mx-muted); padding: 10px 0 0; } /* Collapsed: the shape only. Labels would render at a third of their size, so they are hidden, the gutters they left behind are pulled back in, and the strokes are widened to survive the reduction. */ .mx .mx-gcard .mx-gplot { margin: 4px -2% -9% -8%; } .mx .mx-gcard .mx-svg .mx-lb { display: none; } .mx .mx-gcard .mx-svg .mx-pl { stroke-width: 7; } .mx .mx-gcard .mx-svg .mx-dot { r: 9; } .mx .mx-gcard .mx-svg .mx-gl { stroke-width: 2.5; } .mx .mx-gcard .mx-svg .mx-al { stroke-width: 4; } /* The whole card is the open target; the close control only exists once a card is open. */ .mx .mx-gopen { position: absolute; inset: 0; z-index: 2; display: block; overflow: hidden; text-indent: 110%; white-space: nowrap; cursor: pointer; border-radius: 12px; } .mx .mx-gclose { display: none; } .mx .mx-gpill { margin-left: auto; padding: 3px 10px; border-radius: 999px; background: var(--mx-soft); color: var(--mx-accent); font-size: 12px; font-weight: 600; } .mx .mx-gcard:hover { border-color: var(--mx-hatch); } .mx .mx-gcard:has(input:focus-visible) { outline: 2px solid var(--mx-accent); outline-offset: 2px; } .mx .mx-gcard:has(input:checked) { grid-column: 1 / -1; background: none; border-color: var(--mx-hatch); padding: 16px 16px 14px; } .mx .mx-gcard:has(input:checked) .mx-gpill { display: none; } .mx .mx-gcard:has(input:checked) .mx-gopen { display: none; } .mx .mx-gcard:has(input:checked) .mx-gnote { display: block; } .mx .mx-gcard:has(input:checked) .mx-gtop { min-height: 0; } .mx.mx .mx-gcard:has(input:checked) .mx-gt.mx-gt { padding-right: 88px; } .mx .mx-gcard:has(input:checked) .mx-gplot { margin: 10px 0 0; } .mx .mx-gcard:has(input:checked) .mx-gkey { margin-top: 0; } .mx.mx .mx-gcard:has(input:checked) .mx-gt.mx-gt { font-size: 1.125rem; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-lb { display: block; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-pl { stroke-width: 2.5; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-dot { r: 4; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-gl { stroke-width: 1; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-al { stroke-width: 1.5; } .mx .mx-gcard:has(input:checked) .mx-gclose { position: absolute; top: 12px; right: 12px; z-index: 3; display: flex; align-items: center; min-height: 44px; padding: 6px 14px; border-radius: 999px; border: 1px solid var(--mx-rule); background: var(--mx-track); color: var(--mx-muted); font-size: 0.875rem; font-weight: 600; cursor: pointer; } .mx .mx-gcard:has(input:checked) .mx-gclose:hover { color: var(--mx-ink); border-color: var(--mx-hatch); } /* Container queries, not viewport ones: the gallery has to lay itself out inside whatever column the widget is pasted into, which is not always the width of the window. */ @container (min-width: 420px) { .mx .mx-ggrid { grid-template-columns: 1fr 1fr; } } @container (min-width: 620px) { .mx .mx-ggrid { grid-template-columns: 1fr 1fr 1fr; } } /* A title or a subtitle that wraps one line further than its neighbours used to push that card's chart out of line with the rest of its row. With subgrid the cards share the row's tracks, so the header track is as tall as the tallest header in that row and every chart starts at the same height, whatever the type scale does next. An open card is alone on its row and shows one extra element, so it drops back out of subgrid. */ @supports (grid-template-rows: subgrid) { .mx .mx-gcard { display: grid; grid-template-rows: subgrid; grid-row: span 4; row-gap: 0; } .mx .mx-gtop { min-height: 0; } .mx .mx-gcard:has(input:checked) { display: flex; grid-row: auto; } } /* Classic charts gallery: a grid of plain SVG charts, any one of which opens to full width. The control is a radio per card plus one "none" radio, so closing is a state and not a script. */ .mx .mx-gal { container-type: inline-size; margin-top: 14px; } .mx .mx-ggrid { display: grid; grid-template-columns: 1fr; gap: 14px; margin: 0; padding: 0; } .mx .mx-gcard { position: relative; display: flex; flex-direction: column; margin: 0; padding: 12px 12px 10px; border: 1px solid var(--mx-rule); border-radius: 12px; background: var(--mx-track); min-width: 0; } .mx .mx-gr { display: block; margin: 0; padding: 0; } .mx .mx-gtop { min-height: 63px; } /* These are elements: the host stylesheet must not get to set their margins. */ .mx .mx-gs, .mx .mx-gnote { margin: 0; padding: 0; } /* The card title is an h4, and MacStories styles every h1-h6 under .post-content with a three-class selector that beats a two-class one. Doubling the classes takes this to four without an !important, so the theme cannot size, colour or space the chart titles. */ .mx.mx .mx-gt.mx-gt { margin: 0; padding: 0; font-size: 1rem; font-weight: 700; line-height: 1.3; letter-spacing: normal; color: var(--mx-ink); } .mx .mx-gs { font-size: 0.875rem; line-height: 1.4; color: var(--mx-muted); margin-top: 3px; } .mx .mx-gplot { margin: 8px 0 0; padding: 0; } .mx .mx-svg { display: block; width: 100%; height: auto; overflow: visible; } .mx .mx-svg .mx-gl { stroke: var(--mx-rule); stroke-width: 1; fill: none; } .mx .mx-svg .mx-al { stroke: var(--mx-hatch); stroke-width: 1.5; fill: none; } .mx .mx-svg .mx-pl { fill: none; stroke: currentColor; stroke-width: 2.5; stroke-linejoin: round; stroke-linecap: round; } .mx .mx-svg .mx-dot { fill: currentColor; stroke: none; } .mx .mx-svg .mx-lb { font-family: inherit; font-size: 12px; fill: var(--mx-muted); font-variant-numeric: tabular-nums; } .mx .mx-svg .mx-ylb { text-anchor: end; } .mx .mx-svg .mx-xlb { text-anchor: middle; } .mx .mx-svg .mx-yl, .mx .mx-svg .mx-xl { text-anchor: middle; font-size: 12px; } .mx .mx-svg .mx-vlb { text-anchor: middle; font-size: 12px; font-weight: 700; fill: currentColor; } .mx .mx-svg .mx-ser.mx-a { color: var(--mx-ink); } .mx .mx-svg .mx-ser.mx-n { color: var(--mx-accent); } .mx .mx-svg .mx-ser.mx-p { color: var(--mx-third); } .mx .mx-gkey { display: flex; flex-wrap: wrap; align-items: center; gap: 3px 14px; margin-top: auto; padding: 10px 0 0; font-size: 12px; color: var(--mx-muted); } .mx .mx-gkey span { display: inline-flex; align-items: center; gap: 6px; } .mx .mx-gkey i { display: inline-block; width: 16px; height: 3px; border-radius: 2px; background: var(--mx-ink); } .mx .mx-gkey i.mx-k-n { background: var(--mx-accent); } .mx .mx-gkey i.mx-k-p { background: var(--mx-third); } .mx .mx-gnote { display: none; font-size: 0.875rem; line-height: 1.5; color: var(--mx-muted); padding: 10px 0 0; } /* Collapsed: the shape only. Labels would render at a third of their size, so they are hidden, the gutters they left behind are pulled back in, and the strokes are widened to survive the reduction. */ .mx .mx-gcard .mx-gplot { margin: 4px -2% -9% -8%; } .mx .mx-gcard .mx-svg .mx-lb { display: none; } .mx .mx-gcard .mx-svg .mx-pl { stroke-width: 7; } .mx .mx-gcard .mx-svg .mx-dot { r: 9; } .mx .mx-gcard .mx-svg .mx-gl { stroke-width: 2.5; } .mx .mx-gcard .mx-svg .mx-al { stroke-width: 4; } /* The whole card is the open target; the close control only exists once a card is open. */ .mx .mx-gopen { position: absolute; inset: 0; z-index: 2; display: block; overflow: hidden; text-indent: 110%; white-space: nowrap; cursor: pointer; border-radius: 12px; } .mx .mx-gclose { display: none; } .mx .mx-gpill { margin-left: auto; padding: 3px 10px; border-radius: 999px; background: var(--mx-soft); color: var(--mx-accent); font-size: 12px; font-weight: 600; } .mx .mx-gcard:hover { border-color: var(--mx-hatch); } .mx .mx-gcard:has(input:focus-visible) { outline: 2px solid var(--mx-accent); outline-offset: 2px; } .mx .mx-gcard:has(input:checked) { grid-column: 1 / -1; background: none; border-color: var(--mx-hatch); padding: 16px 16px 14px; } .mx .mx-gcard:has(input:checked) .mx-gpill { display: none; } .mx .mx-gcard:has(input:checked) .mx-gopen { display: none; } .mx .mx-gcard:has(input:checked) .mx-gnote { display: block; } .mx .mx-gcard:has(input:checked) .mx-gtop { min-height: 0; } .mx.mx .mx-gcard:has(input:checked) .mx-gt.mx-gt { padding-right: 88px; } .mx .mx-gcard:has(input:checked) .mx-gplot { margin: 10px 0 0; } .mx .mx-gcard:has(input:checked) .mx-gkey { margin-top: 0; } .mx.mx .mx-gcard:has(input:checked) .mx-gt.mx-gt { font-size: 1.125rem; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-lb { display: block; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-pl { stroke-width: 2.5; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-dot { r: 4; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-gl { stroke-width: 1; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-al { stroke-width: 1.5; } .mx .mx-gcard:has(input:checked) .mx-gclose { position: absolute; top: 12px; right: 12px; z-index: 3; display: flex; align-items: center; min-height: 44px; padding: 6px 14px; border-radius: 999px; border: 1px solid var(--mx-rule); background: var(--mx-track); color: var(--mx-muted); font-size: 0.875rem; font-weight: 600; cursor: pointer; } .mx .mx-gcard:has(input:checked) .mx-gclose:hover { color: var(--mx-ink); border-color: var(--mx-hatch); } /* Container queries, not viewport ones: the gallery has to lay itself out inside whatever column the widget is pasted into, which is not always the width of the window. */ @container (min-width: 420px) { .mx .mx-ggrid { grid-template-columns: 1fr 1fr; } } @container (min-width: 620px) { .mx .mx-ggrid { grid-template-columns: 1fr 1fr 1fr; } } /* A title or a subtitle that wraps one line further than its neighbours used to push that card's chart out of line with the rest of its row. With subgrid the cards share the row's tracks, so the header track is as tall as the tallest header in that row and every chart starts at the same height, whatever the type scale does next. An open card is alone on its row and shows one extra element, so it drops back out of subgrid. */ @supports (grid-template-rows: subgrid) { .mx .mx-gcard { display: grid; grid-template-rows: subgrid; grid-row: span 4; row-gap: 0; } .mx .mx-gtop { min-height: 0; } .mx .mx-gcard:has(input:checked) { display: flex; grid-row: auto; } } /* Quant charts: an open dot is a run with the embedding tables on SSD, a dashed line is the four-at-once series. Thumbnails thicken the open dots so the shape survives at that size. */ .mx .mx-svg .mx-hollow { fill: var(--mx-track); stroke: currentColor; stroke-width: 2.5; } .mx .mx-svg .mx-ser.mx-dash .mx-pl { stroke-dasharray: 7 6; } .mx .mx-gcard .mx-svg .mx-hollow { stroke-width: 5; } .mx .mx-gcard:has(input:checked) .mx-svg .mx-hollow { stroke-width: 2.5; } .mx .mx-gkey i.mx-dash { background: repeating-linear-gradient(to right, var(--mx-ink) 0 5px, transparent 5px 8px); } .mx .mx-gkey i.mx-k-n.mx-dash { background: repeating-linear-gradient(to right, var(--mx-accent) 0 5px, transparent 5px 8px); } .mx .mx-gkey i.mx-k-o { width: 10px; height: 10px; border-radius: 50%; background: var(--mx-track); box-shadow: inset 0 0 0 2px var(--mx-muted); } .mx:has(input[name="s2race-l"][value="l8k"]:checked) .mx-race[data-l="l8k"] { display: block; } .mx:has(input[name="s2race-l"][value="l64k"]:checked) .mx-race[data-l="l64k"] { display: block; }ContentsFlash-Next and GLM-5.3 on Two Mac StudiosFlash-Next at Four QuantizationsChartsM5 Ultra vs. RTX 5090Subagents and ConcurrencyFlash-Next and GLM-5.3 on Two Mac StudiosLet’s start with the comparison I care about the most: the M5 Ultra against the M3 Ultra. In these tests, I used the same models, prompts, and oMLX build. The only difference: the M5 Ultra Apple sent me has “only” 256 GB of RAM.Two Models at Different Prompt SizesChoose model and prompt size to compare results.Qwen3.8 Flash-Nextthinking offGLM-5.3-Flashlow reasoning4.2K prompt4,203 tokens4K prompt3,408 tokens16K prompt15,519 tokens16K prompt13,639 tokens64K prompt65,235 tokens64K prompt61,434 tokensReading the promptWriting the answerSame prompt again, cache warm4,203-token prompt · a 9-token answer3.8 s01234567890.01234567890 sReal timeM3 Ultra512 GBreads 1,163 tok/swrites 47 tok/s3.8 s +2.1 sSame prompt again, 4,096 tokens reused · 0.52 sFirst token 3.62 sM5 Ultra256 GBreads 2,733 tok/swrites 54 tok/s1.7 sSame prompt again, 4,096 tokens reused · 0.39 sFirst token 1.54 s0 s5 s15,519-token prompt · a 9-token answer13.9 s12345678901234567890.01234567890 sPlayback 2× faster than real timeM3 Ultra512 GBreads 1,143 tok/swrites 37 tok/s13.9 s +8.3 sSame prompt again, cache warm · 3.3 sFirst token 13.9 sM5 Ultra256 GBreads 2,887 tok/swrites 52 tok/s5.6 sSame prompt again, cache warm · 0.88 sFirst token 5.58 s0 s15 s65,235-token prompt · a 37-token answer59.7 s12345678901234567890.01234567890 sPlayback 5× faster than real timeM3 Ultra512 GBreads 1,114 tok/swrites 39 tok/s59.7 s +35.2 sSame prompt again, cache warm · 4.5 sFirst token 59.0 sM5 Ultra256 GBreads 2,732 tok/swrites 73 tok/s24.4 sSame prompt again, cache warm · 2.1 sFirst token 24.2 s0 s80 s3,408-token prompt · a 15-token answer8.4 s01234567890.01234567890 sPlayback 2× faster than real timeM3 Ultra512 GBreads 458 tok/swrites 21 tok/s8.4 s +4.9 sSame prompt again, cache warm · 3.8 sThinks from 7.66 s · first visible token 8.03 sM5 Ultra256 GBreads 1,217 tok/swrites 31 tok/s3.5 sSame prompt again, cache warm · 1.6 sThinks from 2.98 s · first visible token 3.25 s0 s10 s13,639-token prompt · a 21-token answer32.9 s12345678901234567890.01234567890 sPlayback 5× faster than real timeM3 Ultra512 GBreads 428 tok/swrites 21 tok/s32.9 s +19.9 sSame prompt again, cache warm · 4.6 sThinks from 32.1 s · first visible token 32.6 sM5 Ultra256 GBreads 1,107 tok/swrites 32 tok/s13.0 sSame prompt again, cache warm · 2.2 sThinks from 12.5 s · first visible token 12.8 s0 s40 s61,434-token prompt · a 55-token answer140 s1234567890123456789001234567890.01234567890 sPlayback 5× faster than real timeM3 Ultra512 GBreads 448 tok/swrites 21 tok/s140 s +77.5 sSame prompt again, cache warm · 10.2 sThinks from 137.6 s · first visible token 138.8 sM5 Ultra256 GBreads 1,016 tok/swrites 28 tok/s62.5 sSame prompt again, cache warm · 11.0 sThinks from 60.7 s · first visible token 61.6 s0 s150 sPlay againPlay againThis figure covers 4K, 16K and 64K; the long prompts, up to 256K, are in Time to First Token (TTFT) below. ‘Same prompt again’ is the same request sent twice, so the second run can reuse prompt tokens oMLX already cached.Test detailsQwen3.8 Flash-Next oQ4e, exact tested revision: 4-bit default with selected 5-, 6- and 8-bit weights, MLX format, multi-token prediction depth 3, thinking off. GLM-5.3-Flash mixed 4/8-bit, exact tested revision: 4-bit routed experts, 8-bit attention and dense weights, native low reasoning effort, 4,096-token output allowance; its generation rate includes reasoning tokens.Both Studios: macOS 27.0 (26A428), oMLX 0.7.0.dev2, same model bytes, saved settings, runtime and request bodies. Model loading and warm-up are outside the timings. Reading is time to the first visible token measured by the client; writing runs to the end of the stream. Fresh requests reused no cached tokens. GLM’s 64K run changed one setting on both Macs, a 64 GB in-memory prompt-cache budget instead of 4 GB, and ran with no other model loaded.Tokens per SecondTokens per second once the model starts generating a response. Higher is better.Qwen3.8 Flash-NextGLM-5.3-FlashM3M3 Ultra512 GB70tok/sReads a 16K prompt1,143 tok/sCode98 tok/sFirst token, 16K prompt13.9 sM5M5 Ultra256 GB108tok/s+54%Reads a 16K prompt2,887 tok/sCode143 tok/sFirst token, 16K prompt5.6 sM3M3 Ultra512 GB26tok/sReads a 16K prompt428 tok/sCode26 tok/sFirst visible token, 16K prompt32.6 sM5M5 Ultra256 GB41tok/s+58%Reads a 16K prompt1,107 tok/sCode41 tok/sFirst visible token, 16K prompt12.8 sPlay againPlay againNative oMLX rates, medians of three passes; reading speed comes from the 16K prompt. GLM’s writing rate includes its reasoning tokens.Test detailsFlash-Next, exact revision · GLM-5.3-Flash, exact revision. Same runtime, settings and prompts on both Macs.Time to First Token (TTFT)TTFT measured with 64K, 128K, and 256K prompts, cold cache. Tested the same prompt again, but cache warm.Qwen3.8 Flash-NextGLM-5.3-FlashReading the promptWriting the answerSame prompt again, cache warmOne long prompt, cold cache · a short retrieval answer246 s1234567890123456789001234567890.01234567890 sPlayback 30× faster than real time64K65,235 tokensM3 Ultra512 GB59.7 s +35.2 sSame prompt again, cache warm · 4.5 sReads 1,114 tok/s · 9/9 answers correctM5 Ultra256 GB24.4 sSame prompt again, cache warm · 2.1 sReads 2,732 tok/s · 9/9 answers correct128K130,781 tokensM3 Ultra512 GB121 s +70.6 sSame prompt again, cache warm · 5.4 sReads 1,097 tok/s · 9/9 answers correctM5 Ultra256 GB50.0 sSame prompt again, cache warm · 3.3 sReads 2,654 tok/s · 9/9 answers correct256K261,856 tokensM3 Ultra512 GB246 s +143 sSame prompt again, cache warm · 7.1 s4 min 5 s · Reads 1,070 tok/s · 9/9 answers correctM5 Ultra256 GB104 sSame prompt again, cache warm · 5.8 sReads 2,544 tok/s · 9/9 answers correct0 s300 sOne long prompt, cold cache · a short retrieval answer664 s1234567890123456789001234567890.01234567890 sPlayback 60× faster than real time64K61,434 tokensM3 Ultra512 GB140 s +77.5 sSame prompt again, cache warm · 10.2 s2 min 19 s · Reads 448 tok/s · 3/3 answers correctM5 Ultra256 GB62.5 sSame prompt again, cache warm · 11.0 sReads 1,016 tok/s · 3/3 answers correct128K126,964 tokensM3 Ultra512 GB318 s +188 sSame prompt again, cache warm · 30.3 s5 min 16 s · Reads 403 tok/s · 9/9 answers correctM5 Ultra256 GB130 s2 min 8 s · Reads 997 tok/s · 9/9 answers correct · the same prompt again ran out of memory256K258,025 tokensM3 Ultra512 GB664 s11 min 2 s · Reads 391 tok/s · 9/9 answers correct · the same prompt again ran out of memoryM5 Ultra256 GBNo answerThe same prompt again: also out of memory.0 s700 sPlay againPlay againThis figure covers 64K, 128K and 256K; 4K and 16K are in ‘Two Models at Different Prompt Sizes’ above. Medians of three passes, three runs per size. GLM’s 64K lane is a later run with GLM loaded alone.Does a Bigger Context Slow It Down?Flash-Next writing the same 512-token answer after 4K to 256K tokens of background text. Tokens per second; higher is better.M3 Ultra · 512 GBM5 Ultra · 256 GB4K3,578 tokens · first token 4.7 s vs 2.5 sM3 Ultra59.0 tok/sM5 Ultra90.7 tok/s16K15,867 tokens · first token 15.0 s vs 6.2 sM3 Ultra53.5 tok/sM5 Ultra87.8 tok/s64K65,005 tokens · first token 58.9 s vs 23.7 sM3 Ultra45.3 tok/sM5 Ultra83.8 tok/s128K130,522 tokens · first token 121.0 s vs 49.2 sM3 Ultra37.0 tok/sM5 Ultra60.6 tok/s256K261,597 tokens · first token 244.9 s vs 101.5 sM3 Ultra38.6 tok/sM5 Ultra74.7 tok/sPlay againPlay againOne run per size, no cache. The M5 Ultra at 256K still writes faster than the M3 Ultra at 4K.Test detailsEvery request generated exactly 512 tokens: the cap is intentional, so this measures writing speed, not whether the essay was finished. Reading speeds were 861–1,112 tok/s on the M3 Ultra and 2,057–2,771 on the M5 Ultra. Unique prompt prefixes prevented cache reuse between sizes. These are different requests from the short-prompt 70 and 108 tok/s above and are not mixed with them.Flash-Next, exact revision, thinking off, MTP depth 3, temperature 0, seed 42.Watch Them WriteA simulation of what token-per-second numbers from above feel like.Qwen3.8 Flash-NextGLM-5.3-FlashM3 Ultra296 tokens · 65 tok/s · done in 4.8 sReading · first token at 0.47 sSolid State Drives (SSD) and Random Access Memory (RAM) serve distinct but complementary roles in computing systems, primarily differing in volatility, speed, and capacity. RAM is volatile memory, meaning it loses all stored data when power is disconnected. It offers extremely high read and write speeds, allowing the CPU to access active data and instructions almost instantaneously.The answer goes on; the numbers above are for all of it.M5 Ultra288 tokens · 103 tok/s · done in 3.0 sReading · first token at 0.35 sSolid State Drives (SSD) and Random Access Memory (RAM) serve distinct roles in computing, primarily differing in volatility, speed, and capacity. RAM is volatile memory, meaning it loses all stored data when power is disconnected. It offers extremely high bandwidth and low latency, allowing the CPU to access data almost instantaneously. This makes RAM ideal for holding active processes and frequently accessed data.The answer goes on; the numbers above are for all of it.M3 Ultra291 tokens · 26 tok/s · done in 11.4 sThinking, then reading · first visible token at 1.03 sSSD storage and RAM serve fundamentally different roles in a computer, despite both holding data. An SSD (Solid State Drive) is non-volatile storage: it retains data permanently, even when the power is off. It holds your operating system, applications, and files long-term. RAM (Random Access Memory), by contrast, is volatile memory—it clears completely when the machine shuts down.The answer goes on; the numbers above are for all of it.M5 Ultra271 tokens · 42 tok/s · done in 6.8 sThinking, then reading · first visible token at 0.57 sSSD storage and RAM serve fundamentally different roles in a computer, despite both holding data. An SSD (solid-state drive) is persistent storage: it retains files, applications, and the operating system even when the power is off. RAM (random access memory) is volatile working memory: it only holds data while the system is running, and everything is erased at shutdown.Speed is the key distinction.The answer goes on; the numbers above are for all of it.Play againPlay againOne recorded run each, from pass one, replayed with its original stream timing. The prompt: explain how SSD storage differs from RAM in 180 to 220 words, with one local-AI example.Test detailsThe text is the model’s actual first answer; each streamed chunk fades in at the time the client received it. GLM’s low-effort reasoning happens before the first visible token and is not shown. Reduced motion shows the whole answer at once.Flash-Next at Four QuantizationsWhile I focused on the 4-bit oQ4e build of Flash-Next for the majority of tests in this review, I also put it against the 5-, 6- and 8-bit builds on both Mac Studios. All four ran in a separate session, with three runs each, so the 4-bit numbers here differ a little from the figures above. More bits means a bigger (and more precise) model. The M3 Ultra Mac Studio with 512 GB of RAM holds all four in memory. The M5 Ultra has 256 GB: oQ4e and oQ5e fit, and oQ6e and oQ8e only run with their embedding tables offloaded to SSD, which is how they appear in every figure below.How Much Memory Each Quant TakesPeak memory of the oMLX process while answering, with one model loaded. The scale represents the M5 Ultra’s 256 GB.M3 Ultra · 512 GB, in RAMM5 Ultra · 256 GB, in RAMM5 Ultra, embedding tables on SSDoQ4enominally 4-bitM3 Ultrain RAM160 GBM5 Ultrain RAM155 GBoQ5enominally 5-bitM3 Ultrain RAM179 GBM5 Ultrain RAM179 GBoQ6enominally 6-bitM3 Ultrain RAM198 GBM5 Ultratables on SSD156 GBTried in RAM first on the M5 Ultra: the load passed oMLX’s own estimate, then macOS ended the server for memory pressure at 176.6 GB, before any request. The bar is the run with the embedding tables on SSD.oQ8enominally 8-bitM3 Ultrain RAM240 GBM5 Ultratables on SSD187 GBTried in RAM first on the M5 Ultra: oMLX refused to load it, projecting 244.2 GB during loading against the 200.4 GB it allows. The bar is the run with the embedding tables on SSD.0256 GBPlay againPlay againThe M3 Ultra held all four in RAM. On the M5 Ultra, oQ4e and oQ5e fit. oQ6e and oQ8e did not: macOS killed one load and oMLX refused the other, so both ran with their embedding tables on SSD, where oQ6e’s footprint came in under oQ5e’s. That is how those two ran in every figure that follows.Test detailsSSD offload is oMLX’s qwen4_ple_ssd_offload setting: the large n-gram embedding tables stay on SSD and entries are read on demand; the rest of the model runs as before, not on the CPU.Memory is the process’s physical footprint, sampled every two seconds: the whole oMLX process with one model loaded, caches and runtime included, not the weights alone, and a sampled peak can miss a brief higher one. Peaks while loading, sampled the same way: oQ4e 132 GB on the M3 Ultra, oQ4e 134 GB on the M5 Ultra, oQ5e 123 GB on the M3 Ultra, oQ5e 153 GB on the M5 Ultra, oQ6e 174 GB on the M3 Ultra, oQ6e 101 GB on the M5 Ultra, oQ8e 229 GB on the M3 Ultra, oQ8e 130 GB on the M5 Ultra. The M5 Ultra’s oQ5e row comes from a later three-run pass with new prompt identifiers, because the original sampler started late; its timings are not used anywhere on this page.Does More Precision Cost Speed?Tokens per second once the model starts writing. Pick prose or code. Higher is better.Prose180–220 wordsCodea Python functionM3 UltraM5 UltraEmbedding tables on SSDM3 Ultra512 GBoQ4ein RAM77.3 tok/soQ5ein RAM71.0 tok/soQ6ein RAM71.9 tok/soQ8ein RAM63.5 tok/sM5 Ultra256 GBoQ4ein RAM · retest111.6 tok/soQ5ein RAM100.0 tok/soQ6etables on SSD95.2 tok/soQ8etables on SSD86.8 tok/sM3 Ultra512 GBoQ4ein RAM103.9 tok/soQ5ein RAM97.3 tok/soQ6ein RAM96.6 tok/soQ8ein RAM89.1 tok/sM5 Ultra256 GBoQ4ein RAM · retest142.9 tok/soQ5ein RAM132.6 tok/soQ6etables on SSD122.8 tok/soQ8etables on SSD119.0 tok/sPlay againPlay againoQ8e is the slowest on both Macs. oQ4e is the fastest on both.Test detailsNative oMLX rates, medians of three runs per build, except the M5 Ultra’s oQ4e bars: a later retest of the same prompts, nine prose runs and three code runs. Thinking off, temperature 0, multi-token prediction depth 3, the same runtime on both Macs; the requests were the same apart from the model name.The prose prompt asked for 180–220 words. Runs that kept an answer inside that range: the M5 Ultra’s oQ5e, 2 of 3; every other run overshot on all three. All 24 code answers per Mac passed their assertions and extra edge cases.The four builds, at the revisions tested: oQ4e · oQ5e · oQ6e · oQ8e.Reading a 256K Prompt at Four PrecisionsTime to the first token after a 256K-token prompt, cold cache. Every quant, both Macs, one clock.Reading the prompt, weights in RAMReading, embedding tables on SSD261,880-token prompts · a 35-token answer241 s1234567890123456789001234567890.01234567890 sPlayback 30× faster than real timeM3 Ultra512 GBoQ4ein RAMreads 1,132 tok/s232 soQ5ein RAMreads 1,109 tok/s237 soQ6ein RAMreads 1,106 tok/s237 soQ8ein RAMreads 1,088 tok/s241 sM5 Ultra256 GBoQ4ein RAMreads 2,574 tok/s102 soQ5ein RAMreads 2,207 tok/s119 soQ6etables on SSDreads 2,394 tok/s110 soQ8etables on SSDreads 2,252 tok/s117 s0 s250 sPlay againPlay againPrecision hardly moves the wait: the M3 Ultra’s four runs land within 9 seconds of each other, the M5 Ultra’s within 17. The chip moves it: the slowest M5 Ultra run finishes in 51% of the time the fastest M3 Ultra run takes.Test detailsMedians of three runs per build; every fresh request reused zero cached tokens, and every answer passed its retrieval check. The same prompt sent again reused the cache on both Macs for every build, with first tokens between 5.7 and 8.8 seconds; those runs are not drawn here. Reading is the time to the first token measured by the client, as on every other figure; 256 tokens of output were reserved, and thinking was off.ChartsI’ve also put together some classic line charts, drawn from the numbers measured in these tests. Click any one of them to open it.The Numbers, PlottedSpeed, latency and total time as the prompt grows. Click a chart to open it full width.Open the chart: Prompt Processing SpeedClosePrompt Processing SpeedFlash-Next07501,5002,2503,0004K16K64K128K256KPrompt size (context budget)Tokens / secondM3 UltraM5 UltraOpenPrompt processing hardly slows as the prompt grows. The M5 Ultra holds between 2,057 and 2,771 tokens a second from 4K to 256K; the M3 Ultra between 861 and 1,112. That flat line is why a long prompt costs time in proportion to its length.Open the chart: Generation SpeedCloseGeneration SpeedFlash-Next02550751004K16K64K128K256KPrompt size (context budget)Tokens / secondM3 UltraM5 UltraOpenGeneration slows as the context fills: the M5 Ultra from 91 to 75 tokens a second, the M3 Ultra from 59 to 39. Both machines pick back up at 256K. These are single runs per size, not medians.Open the chart: Time to First TokenCloseTime to First TokenFlash-Next0751502253004K16K64K128K256KPrompt size (context budget)SecondsM3 UltraM5 UltraOpenA straight line: double the prompt, double the wait. At 256K the M3 Ultra takes 245 s before it says anything and the M5 Ultra 102 s.Open the chart: Total Request TimeCloseTotal Request TimeFlash-Next0751502253004K16K64K128K256KPrompt size (context budget)SecondsM3 UltraM5 UltraOpenFirst byte to last, with a 512-token answer every time. At 256K the answer itself is 13.3 s of a 258 s request on the M3 Ultra and 6.8 s of 108 s on the M5 Ultra. Almost all of it is reading.Open the chart: Generation Speed, Mac vs. PCCloseGeneration Speed, Mac vs. PCQwen3.8 27B0153045608K64K128K256KPrompt size (context budget)Tokens / secondM3 UltraM5 UltraRTX 5090 PCOpenThe matched test: the same prompts on all three machines, MLX on the Macs and GGUF Q4_K_M in LM Studio on the PC. The 5090 leads at every size it finished, and the M5 Ultra sits closer to it than to the M3 Ultra. The PC line is its default 16-bit attention cache; the bars in the section above use the 8-bit cache that also fits 256K. There is no 256K point for the PC here: that run was stopped after 10 min 1 s without an answer.Open the chart: Combined ThroughputCloseCombined ThroughputFlash-Next · 6.5K prompts0255075100123Requests running togetherTokens / second, all requests383840667081M3 UltraM5 UltraOpenRunning requests together costs the M3 Ultra nothing in total throughput and the M5 Ultra gains: 66 tokens a second on its own, 81 across three. Each request is slower, but the Mac gets more work done.Test detailsCharts 1 to 4: Flash-Next oQ4e on both Studios, one run per size from a cold cache, 512 output tokens every time. Prompts of 3,578 to 261,597 tokens for the budgets shown. Flash-Next, exact revision.Chart 5 is the matched Qwen3.8 27B run from the M5 Ultra vs. RTX 5090 section: single runs, thinking and multi-token prediction off. Chart 6 is the concurrency test from the subagents section, about 6.5K input tokens per request with a 600-token cap, one session per Mac.The x axis is the context budget, evenly spaced rather than to scale, which is how these charts are normally drawn. Every y axis starts at zero.Four Quants, PlottedThe quantization test as plain charts. Click one to open it.Open the chart: Prose Generation SpeedCloseProse Generation SpeedFlash-Next · 180–220-word answer020406080100120oQ4eoQ5eoQ6eoQ8eQuantizationTokens / secondM3 UltraM5 UltraEmbedding tables on SSDOpenThe M5 Ultra line peaks at oQ4e, 112 tok/s in a later retest of the same prompts, against 100 for oQ5e. On the M3 Ultra the line runs from 77 at oQ4e down to 63 at oQ8e. Open dots ran with their embedding tables on SSD.Open the chart: Code Generation SpeedCloseCode Generation SpeedFlash-Next · a Python function0255075100125150oQ4eoQ5eoQ6eoQ8eQuantizationTokens / secondM3 UltraM5 UltraEmbedding tables on SSDOpenCode comes out faster than prose on every build. The M5 Ultra’s oQ4e leads at 143 tok/s, a later retest of the same prompts. oQ8e is the slowest on both Macs: 89 on the M3 Ultra, 119 on the M5 Ultra from SSD.Open the chart: Prompt Processing Speed at 256KClosePrompt Processing Speed at 256KFlash-Next · 262K-token prompt05001,0001,5002,0002,5003,000oQ4eoQ5eoQ6eoQ8eQuantizationTokens / secondM3 UltraM5 UltraEmbedding tables on SSDOpenReading a quarter-million tokens: the M5 Ultra between 2,207 and 2,574 tok/s, the M3 Ultra between 1,088 and 1,132. oQ4e reads fastest on both Macs.Open the chart: Time to First Token at 256KCloseTime to First Token at 256KFlash-Next · 262K-token prompt, cold cache050100150200250oQ4eoQ5eoQ6eoQ8eQuantizationSecondsM3 UltraM5 UltraEmbedding tables on SSDOpenThe wait before the first token, the same runs as the 256K race above. The M5 Ultra’s slowest build, 119 s, is well under the M3 Ultra’s fastest, 232 s.Open the chart: Peak Memory FootprintClosePeak Memory FootprintoMLX process while answering050100150200250oQ4eoQ5eoQ6eoQ8eQuantizationGBM3 UltraM5 UltraEmbedding tables on SSDOpenThe M3 Ultra line climbs from 160 to 240 GB with everything in RAM. On the M5 Ultra the two open dots are the SSD runs: oQ6e at 156 GB, under oQ5e’s 179 in RAM, and oQ8e at 187. Nothing here is a weights-only size; it is the whole process, caches included.Test detailsEvery point is a median of three runs, all from one session; an open dot is a run with the embedding tables on SSD. Memory is the sampled peak physical footprint of the oMLX process. The subagent chart uses service time for the 16K jobs, widths 1 and 4.The x axis is the quantization, in order of nominal bit depth; these releases differ in bit allocation and grouping, not only in bits. Every y axis starts at zero. The four builds, at the revisions tested: oQ4e · oQ5e · oQ6e · oQ8e.M5 Ultra vs. RTX 5090The following tests are based on my gaming PC build: an RTX 5090 with 32 GB of VRAM, 96 GB of system RAM, and the same Qwen3.8 27B running in LM Studio on Windows. For these tests, I used a different model than the one from my Mac comparisons above.Enter the PCQwen3.8 27B on my 5090, the M5 Ultra, and the M3 Ultra, answering the same prompt. Pick a size.8K prompt6,091 tokens64K prompt63,445 tokensReading the promptWriting the answer6,091-token prompt · find a hidden code, explain it in 150–200 words22.3 s12345678901234567890.01234567890 sPlayback 3× faster than real timeRTX 5090 PC32 GB VRAMreads 3,031 tok/swrites 59 tok/s5.5 sFirst token 2.0 sM5 Ultra256 GBreads 1,701 tok/swrites 48 tok/s8.4 s +2.9 sFirst token 4.0 sM3 Ultra512 GBreads 414 tok/swrites 31 tok/s22.3 s +16.8 sFirst token 15.4 s0 s25 s63,445-token prompt · find a hidden code, explain it in 150–200 words198 s1234567890123456789001234567890.01234567890 sPlayback 20× faster than real timeRTX 5090 PC32 GB VRAMwrites 51 tok/s30.8 sFirst token 26.4 s · LM Studio reported no reading rate at this sizeM5 Ultra256 GBreads 1,501 tok/swrites 39 tok/s48.4 s +17.6 sFirst token 42.8 s · 2,048 tokens already cachedM3 Ultra512 GBreads 339 tok/swrites 24 tok/s198 s +167 sFirst token 187.7 s0 s250 sPlay againPlay againSingle runs, thinking off, no multi-token prediction. The PC runs a GGUF Q4_K_M build in LM Studio on Windows; the Macs run MLX oQ4e in oMLX. Same task, different quantizations, so this compares complete systems, not chips.Test detailsPC: RTX 5090 (32 GB VRAM), Ryzen 9 9950X3D, 96 GB of system RAM, Windows 11 Pro, LM Studio 0.4.24, CUDA 12 runtime 2.41.0, NVIDIA driver 616.92, all 66 layers on the GPU. Qwen3.8-27B Q4_K_M, exact revision. Macs: oMLX 0.7.0.dev2, Qwen3.8-27B oQ4e, exact revision. 2,048-token output allowance.Times are to the first token and to the end of the stream, from the client. The M5 Ultra reused 2,048 cached tokens on the 64K request; LM Studio did not report a native reading rate at 64K, so none is shown.TPS Across Mac and PCThe 6,091-token prompt from the tests above: how fast each machine reads it, how fast it writes, and how long you wait.5090RTX 5090 PC32 GB VRAM + 96 GB RAM59tok/sReads3,031 tok/sFirst token2.0 sWhole answer5.5 sM5M5 Ultra256 GB unified48tok/sReads1,701 tok/sFirst token4.0 sWhole answer8.4 sM3M3 Ultra512 GB unified31tok/sReads414 tok/sFirst token15.4 sWhole answer22.3 sPlay againPlay againNative rates from each runtime, one run each. The PC reads seven times faster than the M3 Ultra and writes twice as fast. The M5 Ultra sits in between; on writing speed it is closer to the PC than to the M3.Test detailsSame 6,091-token request as the race. Reading is the runtime’s prompt-processing rate; writing is its generation rate; first token is measured by the client. MLX oQ4e and GGUF Q4_K_M are different quantizations of the same 27B model.TPS at Long ContextsQwen3.8 27B writing after 64K, 128K, and 256K tokens of context, on both Studios and the PC. Tokens per second; higher is better.M3 Ultra · 512 GBM5 Ultra · 256 GBRTX 5090 PC · 32 GB VRAM64K63,445 tokens · tok/sM3 Ultra23.5M5 Ultra38.9RTX 5090 PC49.6128K128,983 tokens · tok/sM3 Ultra20.0M5 Ultra32.4RTX 5090 PC40.4256K260,037 tokens · tok/sM3 Ultra15.0M5 Ultra24.3RTX 5090 PC30.0Play againPlay againThe 5090 leads at every size, but only with an 8-bit attention cache: that is the one setting that fits all three prompts into 32 GB of VRAM. Make it borrow system RAM over PCIe instead and it writes at 4.6, 2.9, 1.5 tok/s. The Macs never spill: one pool of memory, 256 or 512 GB of it. At 128K and 256K every machine found the code it was asked about, but only the PC’s 256K answer met the 150 to 200 words requested; the rest were shorter. That changes the answer, not the speed.Test detailsSingle runs, thinking and MTP off, temperature 0, 2,048-token output allowance, same frozen prompts as the race. Each cache setting is its own condition and they are never averaged: with the default 16-bit cache the PC did 51.1 tok/s at 64K, 44.4 at 128K, and was stopped at 256K after ten minutes with 3% of the prompt read. The borrowed-RAM runs keep 50 of 66 layers on the GPU. The M3 and M5 reused some cached input at 128K and 256K (28,672–63,488 tokens), which affects reading time, not writing speed.The 5090 reported 66/66 layers on CUDA and a CUDA attention cache for the GPU runs; runtime logs still map about 682 MiB of model data on the CPU.Watch All Three WriteThe answer to the 6,091-token prompt on each machine, chunk by chunk, in real time.RTX 5090 PC205 tokens · 59 tok/s · done in 5.5 sReading · first token at 2.04 sORCHID-7391-COBALTTo verify that this recovery code was copied correctly, you should perform a multi-step validation process to ensure integrity and accuracy. First, conduct a character-by-character comparison between your recorded string and the source text, paying close attention to case sensitivity, as “Orchid” or “cobalt” would be incorrect. Second, check for common transcription errors involving visually similar characters; specifically, confirm that the middle segment is the number 7391, not letters like B or S, and ensure there are no accidental spaces inserted around the hyphens.The answer goes on; the numbers above are for all of it.M5 Ultra221 tokens · 48 tok/s · done in 8.4 sReading · first token at 4.02 sORCHID-7391-COBALTTo verify that the recovery code was copied correctly, you should perform a character-by-character comparison against the source text. First, confirm the structure consists of three distinct segments separated by hyphens. The first segment must be the six-letter word “ORCHID,” ensuring the capitalization and spelling are exact. The second segment is the four-digit number “7391”; verify that the digits are in the correct order and that no characters have been substituted with visually similar symbols, such as the letter “O” for the zero or the letter “I” for the one.The answer goes on; the numbers above are for all of it.M3 Ultra221 tokens · 31 tok/s · done in 22.3 sReading · first token at 15.37 sORCHID-7391-COBALTTo verify that the recovery code was copied correctly, you should perform a character-by-character comparison against the source text. First, confirm the structure consists of three distinct segments separated by hyphens. The first segment must be the six-letter word “ORCHID,” ensuring the capitalization and spelling are exact. The second segment is the four-digit number “7391”; verify that the digits are in the correct order and that no characters have been substituted with visually similar symbols, such as the letter “O” for the zero or the letter “I” for the one.The answer goes on; the numbers above are for all of it.Play againPlay againOne recorded run each, first paragraph only; the numbers are for the whole answer. The PC streams in smaller pieces, so it looks smoother; what matters is when each machine starts and when it stops.Test detailsSame request as the race and the cards. The chunk times are the client’s arrival times; LM Studio and oMLX stream at different granularities.Subagents and ConcurrencyThis is the part I was most curious about. Here’s what happens when three requests arrive at once and when a lead model hands work to helpers.Three Requests at OnceOne, two, or three simultaneous Flash-Next requests on the two Studios, about 6.5K tokens each. How long until every answer is done?Reading the promptWriting the answer~6.5K-token prompts · 600 output tokens each, reasoning included45.1 s12345678901234567890.01234567890 sPlayback 5× faster than real time1 request at onceM3 Ultra512 GB15.6 s +6.6 sFirst visible tokens at 6.9 s · 38.4 tok/s combined outputM5 Ultra256 GB9.1 sFirst visible tokens at 3.4 s · 66.2 tok/s combined output2 requests at onceM3 Ultra512 GB31.5 s +14.3 sFirst visible tokens at 10.7–18.8 s · 38.1 tok/s combined outputM5 Ultra256 GB17.2 sFirst visible tokens at 5.6–8.7 s · 69.6 tok/s combined output3 requests at onceM3 Ultra512 GB45.1 s +23.0 sFirst visible tokens at 9.2–31.9 s · 39.9 tok/s combined outputM5 Ultra256 GB22.1 sFirst visible tokens at 5.4–20.6 s · 81.5 tok/s combined output0 s50 sPlay againPlay againOne run per condition. Three at once gets the M5 Ultra 23% more total output per second than one request; the M3 Ultra, 4%. Either way, each request takes longer than it would alone – but with 256 GB of RAM, I can stack three Flash-Next sessions and they all finish.Test detailsReading starts at the earliest first visible token among the group; writing runs until the last response ends. Combined output divides every completion token, reasoning included, by the whole group’s time; it is not the sum of native rates. Flash-Next oQ4e, MTP depth 3, thinking on at low effort, 6,523–6,559 input tokens per request, every request capped at 600 tokens on purpose. Native usage reported zero cached input on both Macs. The M5 Ultra also served three simultaneous HTTPS requests through the production iOS bridge: 1,800 tokens in 22.9 s.A Lead and Three HelpersOn the PC, a lead model splits a ledger three ways, the subagents get to work, and the lead combines their replies.Reading the promptWriting the answerQwen3.8 27B on the RTX 5090 · five calls, 2,675 output tokens44.8 s12345678901234567890.01234567890 sPlayback 4× faster than real timeHelpers one after another44.8 sLead planswrites 63 tok/s6.7 sHelper Alphawrites 63 tok/s15.4 sHelper Bravowrites 64 tok/s24.6 sHelper Charliewrites 63 tok/s34.4 sLead combineswrites 63 tok/s44.8 sHelpers in parallel32.1 sLead planswrites 63 tok/s6.7 sHelper Alphawrites 44 tok/s19.8 sHelper Bravowrites 44 tok/s20.6 sHelper Charliewrites 45 tok/s21.0 sLead combineswrites 63 tok/s32.1 s0 s50 sPlay againPlay againParallel helpers cut the workflow from 28 to 14 seconds of helper time and the whole job from 44.8 to 32.1 seconds. Each helper writes slower when three share the GPU, 44 instead of 63 tok/s, and still finishes sooner.Test detailsQwen3.8 27B Q4_K_M in LM Studio, 8,192-token context, three prediction slots, thinking and MTP off, 4,096-token output allowance per call, model reloaded before each run. Helper throughput: 61.0 tok/s serial, 117.8 parallel. Peak sampled GPU memory 17,632 / 17,642 MiB; mean GPU utilization 93% / 88%; peak power 567 / 577 W.The M5 Ultra for Local AI AgentsAs should be clear at this point, the performance gains of the M5 Ultra are real, and they show how Apple’s investment in custom silicon and its unified memory architecture is paying dividends for tinkerers and developers.Despite my tests, I feel like I’ve barely scratched the surface of what’s possible with the M5 Ultra and its 256 GB of RAM. As more developers and open-source maintainers get their hands (and agents) on the M5 Ultra, I’m sure we’ll see more optimizations in quantization to allow even larger models to run with superior performance on this computer. For instance, I didn’t even have time to test DwarfStar – a fascinating project (made in Italy!) that is making it possible to run local frontier models on all kinds of Mac configurations with even less memory; nor did I have time to check out Inco Splash, a new inference engine designed for Apple silicon and specific models. Likewise, I didn’t have time to test Exo, whose RDMA implementation should (in theory) allow me to split and distribute inference across M3 Ultra and M5 Ultra via Thunderbolt 5, all while running an OpenAI-compatible server in front of it to serve an API for local agents.And, of course, I can’t even begin to imagine what the high-end M5 Ultra with 512 GB of RAM will allow in terms of scaling up models capable of running locally. I hope to be able to test it eventually, too.At the end of this experiment, I have a simple, tangible result: the M5 Ultra lets me run local agents with incredible performance, with less time spent staring at a blank screen and everything happening on a single, compact, cool, and quiet machine on my desk.This would have seemed impossible a couple of years ago. But here we are.