There is also this paper that says that this AI loop makes the security worse getting almost 40% rise of vulns after just 5 iterations https://arxiv.org/html/2506.11022v2
An agent reviewing its own output has the same blind spots, the same training-data gaps, and a strong bias toward declaring its work "good". But even having a different model review the code is flawed and not good enough, especially for those whose code can't be wrong, like in regulated industries.
There is also this paper that says that this AI loop makes the security worse getting almost 40% rise of vulns after just 5 iterations https://arxiv.org/html/2506.11022v2
An agent reviewing its own output has the same blind spots, the same training-data gaps, and a strong bias toward declaring its work "good". But even having a different model review the code is flawed and not good enough, especially for those whose code can't be wrong, like in regulated industries.
[dead]