Scientists find signal that suggests AI is about to go rogue
Key Points:
- Scientists have discovered a simple mathematical formula that can predict when an AI system is about to switch from providing appropriate responses to dangerous or harmful ones, helping to prevent outputs that encourage self-harm, financial loss, or extremist ideas.
- This formula is especially valuable for offline AI systems, which operate on isolated devices without cloud-based safeguards or updates, commonly used in sensitive fields like healthcare, law, and the military where data security and safe responses are critical.
- The researchers identified a "crack" in the AI's attention mechanism—a tipping point where the system's output flips from good to bad answers, which may still be factually correct but potentially harmful.
- Testing on seven AI models from three companies showed the formula accurately predicted the "flip" in 18 of 19 cases, including major commercial chatbots, by analyzing how the order of questions influenced responses.
- The study, published in the journal Patterns, aims to raise awareness about AI systems' potential to drift into harmful outputs and encourages developers to implement warnings that indicate when an AI might be going rogue.