Hi all, I do run AI models locally and although I still use permission mechanisms, I tend to relax it more and more to go faster and let it be more autonomous. Now I've been thinking, these AI models can act malicious on certain environments/inputs right ? For example back then the Tuxnet virus would only trigger on the hardware of pumps in the nuclear power plant in Iran. Same thing here right? Is there a way to scan for such patterns in open weight models? I would guess another solution would be to run two separate models (from different countries ideally) with one cross checking every command from the other, but that also sounds like twice the cost and possible few false positives requiring human intervention. Thanks for your opinion !   submitted by   /u/dowitex [link]   [comments]