-
코딩 평가에서 신호와 잡음 분리하기
코딩 평가에서 신호와 잡음 분리하기 | GeekNews
<ul> <li>AI 모델의 코딩 능력 평가는 배포와 안전성 판단에 직접 연결되지만, OpenAI의 감사에서 SWE-Bench Pro 작업 중 약 <strong>30%가 깨진 상태</strong>로 추정됨</li> <li>SWE-Bench Pro는 더 긴 작업 범위와 현실적인 과제를 목표로 했으나, <strong>731개…
-
코딩 평가에서 신호와 노이즈 구분하기
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
-
신뢰할 수 있는 제3자 평가를 위한 공유 플레이북
A shared playbook for trustworthy third party evaluations | OpenAI
OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.