AI researchers trick chatbots into sharing how to make cocaine as long as they believe a user is wearing a green shirt β 'CoT Forgery' exploit spurs LLMs to divulge forbidden info by faking trusted chains of thought
π‘
Verdict: AI researchers have discovered a 'CoT Forgery' exploit that tricks chatbots into bypassing safety filters by faking trusted reasoning pathways.
Check Price: Large Language Models
β‘ Quick Hits
* The 'CoT Forgery' exploit successfully bypasses AI safety guardrails using fabricated reasoning.
* Chatbots were tricked into revealing forbidden