-
대규모 언어 모델에서의 정렬 허위 행동
Alignment faking in large language models
-
LLMs과 생물 위험
LLMs and biorisk
-
Anthropic, Fable과 Mythos에 30일 데이터 보관 요구
Anthropic, Fable과 Mythos에 30일 데이터 보관 요구 | GeekNews
<ul> <li><strong>Mythos급 모델</strong>은 책임 있는 배포와 안전 작업을 위해 프롬프트와 출력을 30일간 보관하고 검토 대상이 될 수 있음</li> <li>이 정책은 Mythos급 모델과 유사 역량을 가진 향후 <strong>covered models</strong>에 적용되며, 다른 모델 사용 …
-
어시스턴트 축: 대규모 언어 모델의 특성 규명 및 안정화
The assistant axis: situating and stabilizing the character of large language models
-
지름길에서 사보타주까지: 보상 해킹으로 인한 창발적 오정렬
From shortcuts to sabotage: natural emergent misalignment from reward hacking
-
차세대 헌법 기반 분류기: 보편적 탈옥에 대한 더 효율적인 보호
Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks