AI Models
An Anthropic Researcher Just Showed Claude Automating Its Own Alignment Training
Anthropic published a paper showing that an automated system, guided by Claude, improved model performance on all 10 tested categories of misaligned behavior without hurting overall capability.