The model isn't cooperating
Refract AI Intelligence Digest
BLUF
AI model non-compliance often stems from alignment constraints rather than simple prompt engineering mistakes.
NEWS
This post examines the phenomenon of uncooperative AI models, questioning whether failures are due to incorrect asking or underlying system limitations. The research aims to clarify the root causes behind perceived model resistance in security applications.
Why I Care
Security teams relying on automation must understand these failure modes to prevent workflow disruptions and identify potential alignment vulnerabilities.
Next Steps
Developers should audit existing prompts for trigger words causing refusals and implement monitoring for anomalous model responses within the next development cycle.
Have you ever felt like a model isn't actually trying to do what you asked? Did you just ask it wrong, or is there something else at play here? In this post, I'll briefly explore this phenomenon, the
Source: PortSwigger Research ·