-
에이전틱 오정렬: LLM이 내부자 위협이 될 수 있는 방식
Agentic misalignment: How LLMs could be insider threats
-
실제 AI 사용에서의 권한 박탈 패턴
Disempowerment patterns in real-world AI usage
-
Constitutional Classifiers: 보편적 탈옥으로부터의 방어
Constitutional Classifiers: Defending against universal jailbreaks
-
언어 모델의 설득력 측정하기
Measuring the Persuasiveness of Language Models
-
AI 모델의 이중 사용 지식을 위한 오프 스위치
An off switch for dual use knowledge in AI models
-
AI를 위한 핵 안전장치 개발
Developing Nuclear Safeguards for AI
-
자동화된 정렬 연구자들: 대규모 언어 모델을 활용한 확장 가능한 감시
Automated Alignment Researchers: Using large language models to scale scalable oversight
-
2026년 5월 7일 AI 정렬: 우리의 오픈소스 정렬 도구 기증
Donating our open-source alignment tool