Really need someone to help me run tests on my benchmark, any model

Wait 5 sec.

I've had some updates to this benchmark, meaning it should run smoothly compared to before. I have an umm, measly GPU with 8 GB of VRAM, so I can't run models like Qwen3.8-27B on it. Any result would be good from you guys; although this is more suited to frontier models, smaller models should still work. I'll credit you for any results if you'd like. The primary reason I thought it would be interesting to run: benchmarks are nearly all pass or fail on a single-dimention graph, so I thought it might be worth shaping a new one up. BinkBench measures video quality and video compression rate, which gives you two things to plot on. The agent also can't score 100% - there isn't an end, which makes it progressively harder as the agents get smarter, because they need to implement more novel techniques. I also thought video encoding would be good as a benchmark, since it's not something we've tested agents on before and is pretty hard. It's like the kernel/compiler optimisation things we've seen other labs show tests on. More info is on GitHub, Old post: https://www.reddit.com/r/LocalLLaMA/comments/1vn6nlr/looking_for_people_to_help_me_run_a_benchmark/   submitted by   /u/-MaskNinja- [link]   [comments]