guys is it a dumb idea to use one of the smaller jev knockoffs to decide which speculative draft is better

Wait 5 sec.

as in like generated so far: A B C draft 1: D E F draft 2: G H I draft 3: J K L then some small model that runs locally really fast decides which speculative draft is best but only when the token entropy is high this would be in the generation loop itself   submitted by   /u/Aggravating-Push-207 [link]   [comments]