3.
Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
by skn8733 | Jul 20, 2026 | original
4.
AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
by Fred Heiding | Jul 20, 2026 | original
7.
Restoring Model Alignment via Honesty Activation Steering
by niklas_herbster | Jul 20, 2026 | original