더 작은 모델로 에이전트 비용을 줄이려면, 모델보다 하네스를 먼저 최적화해야 한...
Better Harnesses, Smaller Models는 작은 언어 모델의 실패 trajectory를 진단해 context·tools·agent loop를 자동 적응시키고, 반복 업무에서 frontier LL...
Tag
Better Harnesses, Smaller Models는 작은 언어 모델의 실패 trajectory를 진단해 context·tools·agent loop를 자동 적응시키고, 반복 업무에서 frontier LL...
WeiboAI의 VibeThinker-3B는 Qwen2.5-Coder-3B 위에 Spectrum-to-Signal post-training, 다중 도메인 RL, offline self-distillation, C...
arXiv 2509.24945의 MobileLLM-R1은 140M·360M·950M reasoning model을 공개하면서, 초대형 말뭉치보다 능력별 데이터 선별·재혼합·지식 압축이 작은 모델의 reasonin...
SOD는 tool-integrated reasoning에서 학생 모델의 잘못된 tool call이 만든 상태 드리프트를 step-level divergence로 감지하고, 온폴리시 증류 신호를 단계별로 재가중해...