STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
Structured pruning of self-distilled reasoning traces reduces token use while largely preserving accuracy in low-data fine-tuning.
•
1 min read
Read more
Structured pruning of self-distilled reasoning traces reduces token use while largely preserving accuracy in low-data fine-tuning.