SenseNova-Vision-7B-MoT is a fine-tune of ByteDance's Bagel, 14B total, 7B active. Boxes and OCR come back as plain text with coordinates, depth maps and masks as images it generates. The paper puts it first on object detection and OCR (one tie) against the models they picked, and it beats Depth Anything V2 on every depth set. It trails MoGe-2 on most of them though, and dedicated segmentation models still win most mask tests. Running it locally looks hard. The README's only tested setup is one 80GB A800, and smaller cards haven't been fully validated. No GGUF, and llama.cpp can't load Bagel yet. Weights are non-commercial too, even though Bagel itself is Apache 2.0. If a GGUF got this onto a 24GB card, would you run one model for all four or stick with Depth Anything, SAM and a small VLM?   submitted by   /u/mostlyired12 [link]   [comments]