GPT-6.1 Sol looped "leak" hints at nested models serving architecture

Wait 5 sec.

Hello llamas. I am posting this because I believe that, despite it being closed source models, the discussion will bring value to the local AI community. As many of you probably heard, GPT-6 Astra is speculated to be a looped transformer architecture that outputs a token after multiple forward passes instead of one. This allows a model to essentially have more effective depth due to recurrence, making more use of the weights at the cost of more compute. Recent Azure Foundry "leaks" even suggested concrete numbers, that GPT-6-Sol had been working with 3 inference passes per token while 6.1-Sol only needs 2. Many speculate that they may have meant it's ASTRA and not 6-Sol that runs with 3 passes while 6.1-Sol is essentially the same model with 2 passes instead. So I did some back-of-the-envelope math to see if the numbers add up. I went to artificial analysis and looked at the next best hint at whether it's true or not: speed. I know it doesn't prove it, but hear me out. If you look at the image, it shows something interesting: - GPT-6-Sol and 5.6-Sol: ~100 tok/s - GPT-6.1-Sol: ~60 tok/s - GPT-6-Astra: ~60 tok/s This may suggest that, if they're essentially the same model weights, that they may be running with batched inference and that Sol may have to wait an extra cycle for Astra requests to finish a token, which caps both models at around the same speed. Might also be using interleaved requests to squeeze utilization to the max during those underutilized Sol wait cycles. But then I also realized that 5.6 Luna was between 126-137 tok/s and then 6.0-Luna dropped to around 110-115. Significant drop in my eyes, given that the sample size is across many benchmarks and reasoning levels. Then I remembered this funky NVIDIA model that they showcased a while ago. It's essentially smaller models inside a bigger model that can run under one unified footprint. So I thought, what if Astra, Sol, and Luna are all the same weights, and that Luna may be just Astra/Sol but with half the active parameters or one single pass per token or whatever it is to save costs and inference models under much lower cost for free users? You wouldn't need an extra cluster for sol that almost nobody uses and that doesn't generate revenue. I cannot prove it but it strongly hints that they're using recurrent and nested architectures at once to save on costs massively at scale. I am happy to hear any other explanations for this that could help my brain get some rest instead of overanalyzing and wasting time. Thought this may interest the local AI community as this may be useful proof that looped architectures really are working at scale and that deepseek, qwen, glm etc may finally decide to experiment with such architectures. Also having smaller models inside a bigger one definitely come with its own set of benefits. PS: fully human generated text. 0.7 tokens per second. ~100T parameter wetware model. Running on two coffees and a muesli bar.   submitted by   /u/QuackerEnte [link]   [comments]