Update #5: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Wait 5 sec.

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wxlytt/comment/pdz728b/?screen_view_count=1 Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress. Last update explained underfitting and next steps. Training has begun again! I've synthesized about 5M more tokens for the SFT, this time across a much larger general instruct trajectory to try to reduce the underfitting. Dropped the learning rate about 4x over my original LoRA adapter. I'm live streaming training again: https://geological-estimate-fifth-pct.trycloudflare.com/ This one should last 12-14 hours, and I plan to run another epoch if this isn't sufficient. Stay tuned! Thanks for following along.   submitted by   /u/jjusko20 [link]   [comments]