I built MOLT: a local fine-tuning system with fit tests, checkpoints, and deployment tracing

Wait 5 sec.

I’ve been building MOLT, a local fine-tuning system for consumer NVIDIA GPUs. The goal is not only to start a QLoRA run. It is to make the whole process from model and dataset to a working local adapter less fragile. MOLT currently handles: - dataset detection, preparation, and validation - GPU, VRAM, system-RAM, storage, and thermal checks before a run - automatic microbatch fit testing - 4-bit NF4 QLoRA training with BF16 adapters - safe checkpoints with integrity checks and proper resume state - telemetry for VRAM, temperature, energy, clocks, and throughput - base-vs-adapter evaluation - local adapter chat, export/GGUF workflows, and runtime diagnostics Under the hood, I’m also experimenting with custom CUDA/Triton kernel and runtime paths, memory planning, CUDA graphs, replay modes, and fused optimization where the hardware and shape are qualified. On my RTX 4060 Laptop 8GB, a recent Qwen 3B training run reached about 1,005 training targets/sec end to end very close to my historical Unsloth run at about 1,007 targets/sec. That is not a broad “MOLT beats Unsloth” claim yet; I still need repeated, matched quality/energy/memory tests. What I want MOLT to become is a reliable local fine-tuning environment: you bring the model and dataset; it checks the plan, runs safely, preserves the evidence, and helps verify the adapter actually works after deployment. I’m looking for people who genuinely fine-tune on a single NVIDIA GPU and would be willing to test it. What would MOLT need to support before you would try it?   submitted by   /u/MKP_Nimilka [link]   [comments]