Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

Published in COLM, 2026

Recommended citation: Gu, S., Chen, J., Zhou, S., Cohan, A., & Ying, R. (2026). "Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation." COLM. https://arxiv.org/abs/2606.19327

Rubric-Conditioned Self-Distillation uses structured, criterion-level feedback for on-policy self-distillation. The framework first learns to generate task-specific rubrics, then conditions a teacher model on those rubrics to provide token-level guidance for a student’s sampled trajectories.

Read the paper · View the code