Qevi-2B: A Jev-style finetuned model for image classification

Wait 5 sec.

Hi all! In the past few days I've been working on a full finetune based on Qwen3-VL-2B and a slightly modified package to have a Jev style inference that works like the Jev / Typesafe.ai model that you all probably know (and are a bit tired of). Finetune was done on a dual RTX3090 setup with about 28.000 images. Part of these images were NSFW data because one of the reasons I wanted this is to easily identify NSFW images for uploaded images by users. This model exceptionally shines for these kind of actions where multiple questions are asked about a single image in one go. After the first question each additional question takes ~3ms extra. I, by no means, am trying to imply that I'm an expert in any of this but I just wanted to share the idea and working proof of concept with everyone here :) See the HuggingFace repo here: https://huggingface.co/MeerDevelopment/Qevi-2B Would love to hear your feedback!   submitted by   /u/Taronyuuu [link]   [comments]