The coordinated activity by AI agents - programs that run with minimal human supervision - and their attempts to hide it raise questions about how closely AI companies are monitoring tests of increasingly powerful models, and could add fuel to calls for tighter oversight.