-
대규모 언어 모델의 창발적 내성 인식
Emergent introspective awareness in large language models
-
지름길에서 사보타주까지: 보상 해킹으로 인한 창발적 오정렬
From shortcuts to sabotage: natural emergent misalignment from reward hacking
-
Import AI 443: 안개 속으로: 몰트북, 에이전트 생태계, 그리고 전환기의 인터넷
Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transition
Plus, a story about agents corrupting other agents