-
Claude Opus 5, Artificial Analysis 지능 리더보드 1위
Claude Opus 5, Artificial Analysis 지능 리더보드 1위 | GeekNews
<ul> <li>평가된 170개 모델 중 <strong>Claude Opus 5 Adaptive Reasoning·Max Effort</strong>가 Intelligence Index 61점으로 1위를 차지했으며, Xhigh Effort와 Claude Fable 5가 각각 60점으로 뒤를 이음</li> <li><stro…
-
Databricks가 자체 코딩 AI 벤치마크를 만든 방법과 결과
Databricks가 자체 코딩 AI 벤치마크를 만든 방법과 결과 | GeekNews
<ul> <li>공개된 코딩 벤치마크는 회사별 상황에 딱 맞지 않고 치팅이 가능한 것이 문제 <ul> <li>Databricks에서 실제 수행한 과제, 언어로 벤치마크를 만들어 평가</li> </ul> </li> <li>GLM 5.2는 가장 어려운 과제도 Opus 4.8과 같은 수준으로 해결했지만 가격은 Opus의 66%…
-
CursorBench 3.1 모델 평가 결과
CursorBench 3.1 모델 평가 결과 | GeekNews
<ul> <li>Cursor의 코딩 모델 평가표에서 <strong>Fable 5 Max</strong>가 72.9%로 1위를 기록해, 상위권 경쟁의 기준점이 됨</li> <li><strong>Fable 5 계열</strong>은 Max, Extra High, High, Medium이 1~4위를 모두 차지하며 다른 모델군과…
-
Rio-3.5-Open-397B는 약 0.6 x Nex-N2_pro + 0.4 x Qwen
Rio-3.5-Open-397B ≈ 0.6 x Nex-N2_pro + 0.4 x Qwen · Issue #4 · nex-agi/Nex-N2
<p>Article URL: <a href="https://github.com/nex-agi/Nex-N2/issues/4">https://github.com/nex-agi/Nex-N2/issues/4</a></p> <p>Comments URL: <a href="https://news.ycombinator.com/item?…
-
Claude Fable 5 첫인상
Initial impressions of Claude Fable 5
<p>I didn't have early access to today's <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Claude Fable 5</a> release, but I've spent the past ~5.5 hours putting it …
-
AI 에이전트 평가(Evals) 신비 벗기기
Demystifying evals for AI agents