Hi everyone, I'm Julian, and I just launched the first model on my inference startup, ObitMC. I wrote a custom scheduler and inference engine in Rust which allows the model Qwen 3.8 to run entirely on interruptible/unreliable compute providers like AWS/GCP Spot and Vast Interruptible, saving ~65% vs OpenRouter. This is the sustainable/breakeven point for us, not a promotional deal. Our stats show that expected reliability, throughput and time to first token are excellent, but we need real users hammering the service to get an idea of how we hold up in production. This is why I'm giving everyone who responds to this post 100 million free tokens of Qwen 3.8 27B (~$12 of platform/API credit) to test out the service. No credit card required. After that, our token rates are $0.10/M in, $1/M out, $0.01/M cached. You can use this via our OpenAI-compatible API and setup takes about 5 minutes. We're going to expand to Qwen 3.6 35B-A3B (MoE) and Gemma 4, which we're piloting internally right now. Reply in a comment if you want the credit promotion and I'll PM you, or you can sign up directly through our website (promotion does not apply if you sign up this way) Also happy to answer any questions about the scheduler, engine or other infra if anyone's curious.   submitted by   /u/Charming_Car_504 [link]   [comments]