2 months (meme-d) research, Do People Notice the Difference Between AI Models?

Wait 5 sec.

Preamble and disclaimer, Sample size: 8, at any size of form this research is just screwing around being writing it down. This was just a little,(well... big), curiosity test I wanted to run to see whether everyday folks could actually differentiate between high end AI models. In my country, the general reception toward AI is fairly neutral. People are neither strongly anti-AI nor boot licking about it, except for a noticeable distaste (rightfully so), toward lazy, overburnt style use of gpt AI gen images for product listings. After the experiment, I asked the participants for permission to publish the results. ------ Anyway ------ For the experiment, I initially used both my local models and OpenRouter. However, after the first two weeks, I dropped OpenRouter because, surprisingly, my RTX 3090 was barely being hit. Usage mostly came in short bursts of around 5-6 reqs/hour. Although i use 96G of my ram for warm KV store, and the rest of cold KV are on SSD I told the 13 participants, translated roughly: "I got access to the latest chatgpt model for free, but only for a limited time. You no longer have to deal with things like 'memory is low,' but you have to use my website because I had to wire it up to OpenAI. Also, don't ask anything weird. I don't want to, but technically I can see your activity logs in the database." Based on the IP traffic patterns, I think they (8 people) somehow bought it. The web UI was basically a gray-themed OpenWebUI instance modified by Qwen 3.6 27B, with the admin panels hidden. The experiment itself was simple. I rotated between Qwen 3.6 27B, Qwen 3.6 35B-A3B, Gemma 26B-A4B, Qwen 3.5 9B, Gemma 12B, Gemma E4B, and Gemma E2B. Reasoning effort was parameterized into four levels: instant, low, medium, and high. All ran on VLLM except for 35B The goal was to find out whether people would notice meaningful differences between the models, complain about quality, or develop preferences without knowing which model they were actually using. Complaints started appearing at the 9B level and below. The most common complaint was basically that the model "does not get it." Unsurprisingly, everyone preferred the responses from the Gemma models. While the other models ,where 3 participants even said that sometimes the model "thinks too much" and ends up sounding like a confused robot. We probably know which model they were talking about. last pic is from Q3.6 27B OpenWebui restyle Across 7,912 requests, only 51 used high reasoning. Around 4,588 used instant, 2,263 used medium, and the rest used low. So, yes, the overwhelming majority of usage was nowhere near high reasoning, i already told them there is a toggle to set it highest thinking mode. Interestingly, some participants still described the models as top of the line because they could see the reasoning process. One comment was roughly: "Wow, this model is really observant about its own behavior because it thinks very carefully." At the end of the experiment, I ran an LLM judge using Qwen 3.8 27B to categorize the requests. Around half were related to writing documents, including things like drafting documents and generating excel style tables. Next is were grammar-related requests. This category overlapped somewhat with document writing, and the LLM judge reported fairly high uncertainty when classifying them. Roughly 1/3 of the grammar-checking requests could reasonably have been placed in the document-writing category instead. Most of the remaining requests were basically "Google search" type questions, welp i paid for sonar credit for the most overkill cooking recipe question..... Almost nobody used it for coding. There was only one notable coding-related request, when a friend wanted to showcase one of their projects using a simple HTML-only landing-page hero section. About context length, although i install hook to auto prune+summarized old message, it kinda never being used. most of the request sit arround 64K Model: Q 3.6 27B INT4 https://huggingface.co/Lorbus/Qwen3.6-27B-int4-AutoRound Q 3.6 35B Unsloth UD Q4 https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/main/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf Gemma 26B A4B https://huggingface.co/cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit Gemma 12B QAT https://huggingface.co/google/gemma-4-12B-it-qat-w4a16-ct Ornith Q 3.5 9B https://huggingface.co/ornith-ai/Ornith-1.0-9B Gemma E4B https://huggingface.co/google/gemma-4-E4B-it Gemma E2B https://huggingface.co/google/gemma-4-E2B-it - 12B and above use FP8/Q8_0 KV. - 9B and below use BF16 KV. - VLLM ran on AOT with small batch tok to increase ctx - For Gemma 26B and Q 27B have image sub LLM (E4B) that ran on my processor. Although image request is also rare   submitted by   /u/Altruistic_Heat_9531 [link]   [comments]