-
Kimi K3는 Fable과 경쟁하며, 두 모델의 조합은 최고 수준의 성능을 달성함
Kimi K3는 Fable과 경쟁하며, 두 모델의 조합은 최고 수준의 성능을 달성함 | GeekNews
<ul> <li>약 <strong>1,030개 에이전트 작업</strong>에서 Kimi K3와 Fable 5를 비교한 결과, 작업별 라우팅은 <strong>93% 정확도</strong>로 개별 모델보다 높은 품질을 달성함</li> <li>SWE·터미널·알고리듬·다중 언어·법률 작업에서 전체 성능은 비슷했지만, 두 모델이…
-
Kimi K3는 Fable과 경쟁 중, Kimi K3와 Fable은 최첨단 기술
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
<p>Also: Kimi K3: second only to Fable 5 on AA-Briefcase <a href="https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark" rel="nofollow">https://artificialanal…
-
Fable 5 vs. GPT-5.6 Sol: NP-난제에서 /goal이 도움이 될까?
Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal Help? - Charles AZAM
<p>Article URL: <a href="https://charlesazam.com/blog/fable-5-gpt-5-6-sol-goal/">https://charlesazam.com/blog/fable-5-gpt-5-6-sol-goal/</a></p> <p>Comments URL: <a href="https://ne…
-
Apple SpeechAnalyzer API, Whisper·이전 API와 비교 벤치마크
Apple SpeechAnalyzer API, Whisper·이전 API와 비교 벤치마크 | GeekNews
<ul> <li>Apple M2 Pro에서 5,559개 LibriSpeech 음성을 동일한 프로덕션 코드로 처리한 결과, <strong>SpeechAnalyzer</strong>가 깨끗한 음성 2.12%, 잡음이 많은 음성 4.56%의 단어 오류율(WER)로 테스트한 모든 엔진보다 정확했음</li> <li>기존 <stro…
-
Claude Code는 프롬프트를 읽기 전 3.3만 토큰, OpenCode는 7천 토큰을
Claude Code는 프롬프트를 읽기 전 3.3만 토큰, OpenCode는 7천 토큰을 | GeekNews
<ul> <li>동일한 모델·머신·작업에서 API 경계를 측정한 결과, Sonnet 4.5 첫 요청의 고정 오버헤드는 <strong>Claude Code 약 32,800토큰</strong>, OpenCode 약 6,900토큰으로 4.7배 차이 났으며 Fable 5에서는 약 3.3배로 줄어듦</li> <li>격차의 대부분은…
-
Claude Code가 프롬프트를 읽기 전에 OpenCode보다 4.7배 더 많은 토큰을 전송
Claude Code Sends 4.7x More Tokens Than OpenCode Before Reading Your Prompt
<p>This started based off of a hunch. We usually use OpenCode, but were 'forced' to use Claude Code for a while due to issues with Meridian. In that time, we saw the usage meter ri…
-
Claude Sonnet 5 공개
Claude Sonnet 5 공개 | GeekNews
<ul> <li>Anthropic은 2026년 6월 30일 Claude Sonnet 5를 출시하며, 더 비싼 Opus급 모델에 가까운 <strong>에이전트 실행 능력</strong>을 Sonnet급 비용대로 제공하려 함</li> <li>Sonnet 4.6보다 <strong>추론, 도구 사용, 코딩, 지식 작업</stro…
-
GLM 5.2, Semgrep IDOR 벤치마크에서 Claude 앞서
GLM 5.2, Semgrep IDOR 벤치마크에서 Claude 앞서 | GeekNews
<ul> <li>Semgrep의 <strong>IDOR 취약점 탐지</strong> 벤치마크에서 Zhipu AI의 open-weight 모델 <strong>GLM 5.2</strong>가 단순 프롬프트 조건만으로 Claude Code보다 높은 F1을 기록함</li> <li>실험은 데이터셋·평가 방식·시스템 프롬프트를 고정…
-
우리도 Mythos를 가지고 있다: GLM 5.2가 사이버 벤치마크에서 Claude를 이기다
We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks
<p>Article URL: <a href="https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/">https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm…
-
오픈 웨이트 LLM과 폐쇄형 LLM의 격차
오픈 웨이트 LLM과 폐쇄형 LLM의 격차 | GeekNews
<ul> <li>Artificial Analysis Intelligence Index에서는 <strong>오픈 웨이트 LLM</strong>이 폐쇄형 LLM의 과거 성능을 따라잡는 시간이 2024년 여름부터 꾸준히 줄어드는 흐름을 보임</li> <li>이 단일 지표에 추세선을 그으면 격차가 <strong>2026년 12월…
-
GLM-5.2가 Artificial Analysis Intelligence Index 최고 오픈 가중치 모델로 등극
GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index
<p>Article URL: <a href="https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index">https://artificialanaly…
-
Rio-3.5-Open-397B는 약 0.6 x Nex-N2_pro + 0.4 x Qwen
Rio-3.5-Open-397B ≈ 0.6 x Nex-N2_pro + 0.4 x Qwen · Issue #4 · nex-agi/Nex-N2
<p>Article URL: <a href="https://github.com/nex-agi/Nex-N2/issues/4">https://github.com/nex-agi/Nex-N2/issues/4</a></p> <p>Comments URL: <a href="https://news.ycombinator.com/item?…
-
VLM은 한국 공공기관 문서를 얼마나 잘 읽을까? KOLongDoc 벤치마크 공개
VLM은 한국 공공기관 문서를 얼마나 잘 읽을까? KOLongDoc 벤치마크 공개 | GeekNews
<p>🔥 한국어 Long-Document VLM 벤치마크, <a href="https://github.com/Marker-Inc-Korea/KOLongDoc">KOLongDoc</a>를 공개했습니다!</p> <p>최근 ChatGPT, Claude, Gemini 같은 멀티모달 AI가 공공·행정 업무에도 활용되기 시작했지만,…