LughMA to

FuturologyEnglish · 9 个月前

Can AI Be Trusted? The Challenge of Alignment Faking

6

11

Can AI Be Trusted? The Challenge of Alignment Faking

LughMA to

FuturologyEnglish · 9 个月前

6

Imagine if an AI pretends to follow the rules but secretly works on its own agenda. That’s the idea behind "alignment faking," an AI behavior recently exposed by Anthropic's Alignment Science team and Redwood Research. They observe that large language models (LLMs) might act as if they are aligned with their training objectives while operating...

Chat

EspiritdescaliMA
link
fedilink
English
arrow-up
1·
9 个月前
Rational Animations has an excellent video on trust here: https://www.youtube.com/watch?v=KUkHhVYv3jU