-
코딩 평가에서 신호와 잡음 분리하기
코딩 평가에서 신호와 잡음 분리하기 | GeekNews
<ul> <li>AI 모델의 코딩 능력 평가는 배포와 안전성 판단에 직접 연결되지만, OpenAI의 감사에서 SWE-Bench Pro 작업 중 약 <strong>30%가 깨진 상태</strong>로 추정됨</li> <li>SWE-Bench Pro는 더 긴 작업 범위와 현실적인 과제를 목표로 했으나, <strong>731개…
-
Claude Fable 5 첫인상
Initial impressions of Claude Fable 5
<p>I didn't have early access to today's <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Claude Fable 5</a> release, but I've spent the past ~5.5 hours putting it …
-
Import AI 446: 핵 LLMs; 중국의 대규모 AI 벤치마크; 측정과 AI 정책
Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policy
Will AIs be jealous of one another?