A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
-
Updated
Oct 6, 2026
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
🔥 DanceOPD: On-Policy Generative Field Distillation
Source code of paper "RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation"
Official implementation of On-Policy Delta Distillation (OPD2)
Tiny-R2: A hybrid architecture integrating SWA, CSA, HCA, mHC, and DSMoE under the DeepSeek V4 design paradigm, enabling single-GPU OPD post-training.
The implementation for our paper: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning.
H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
A repo for distilling a large teacher into a small vision-language model for efficient embodied spatial reasoning and action planning.
Simplifying Hospital OPD System by bridging the gap between patients and hospitals.
Основы профессиональной деятельности (ОПД) Программная инженерия (ПИиКТ) Нейротехнологии и программирование ИТМО
Official implementation of the paper "STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models. "
🔥 Learning When to Trust via Selective Context Preference Optimization
Open Source Health Information System. Get your Frappe Health instance ready within a few seconds and uncover the potentials of digitising your operations.
Open Source Health Information System. Get your Frappe Health instance ready within a few seconds and uncover the potentials of digitising your operations.
To associate your repository with the opd topic, visit your repo's landing page and select "manage topics."