-
대규모 언어 모델에서의 정렬 허위 행동
Alignment faking in large language models
-
LLMs과 생물 위험
LLMs and biorisk
-
어시스턴트 축: 대규모 언어 모델의 특성 규명 및 안정화
The assistant axis: situating and stabilizing the character of large language models
-
지름길에서 사보타주까지: 보상 해킹으로 인한 창발적 오정렬
From shortcuts to sabotage: natural emergent misalignment from reward hacking
-
차세대 헌법 기반 분류기: 보편적 탈옥에 대한 더 효율적인 보호
Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks