OPD는 답보다 교사의 정책을 옮기고, 그래서 다중 교사에서 줄다리기가 된다
Every Coin Has Two Sides는 on-policy distillation이 교사의 문제별 정답보다 reasoning behavior를 전이하며, 같은 계보 teacher에서는 범위를 넓히지만 mul...
Tag
Every Coin Has Two Sides는 on-policy distillation이 교사의 문제별 정답보다 reasoning behavior를 전이하며, 같은 계보 teacher에서는 범위를 넓히지만 mul...
U-OPSD는 여러 on-policy rollout의 합의를 pseudo-solution으로 만들고, 그 합의와 갈린 trajectory에만 token-level distillation을 적용해 ground tr...
Direct-OPD는 작은 RL teacher의 최종 정책을 모방하지 않고, 사전·사후 체크포인트의 로그비에 남은 policy shift를 큰 student의 온폴리시 학습용 dense reward로 바꿔 wea...
WeiboAI의 VibeThinker-3B는 Qwen2.5-Coder-3B 위에 Spectrum-to-Signal post-training, 다중 도메인 RL, offline self-distillation, C...
arXiv 2509.24945의 MobileLLM-R1은 140M·360M·950M reasoning model을 공개하면서, 초대형 말뭉치보다 능력별 데이터 선별·재혼합·지식 압축이 작은 모델의 reasonin...
SU-01은 30B-A3B reasoning backbone에 reverse-perplexity SFT, two-stage RL, test-time verification/refinement를 얹어 IMO·USA...