I’m releasing OpenJudgement-4B-Preview, an experimental Qwen-based model fine-tuned on custom datasets for classification, scoring and true/false judgments. It scores answer options directly, and Python formats the results into JSON with probabilities. It still uses an LLM backbone, but doesn’t generate the response token by token. It’s unfinished and isn’t at Jev’s level yet. I’d love feedback, especially examples where it gets things wrong. Use it via api at: https://kitani.ai/models/kitani/OpenJudgement-4B-Preview (paid) Model and inference code: https://huggingface.co/kitaniai/OpenJudgement-4B-Preview   submitted by   /u/bakatristan [link]   [comments]