🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Training

Tiny Reasoning Models

First page
Tiny Reasoning Models
Paper summary

Tina is a family of 1.5B parameter reasoning models trained using LoRA-based reinforcement learning (RL) to achieve high reasoning accuracy at very low cost. It outperforms or matches full fine-tuned models on reasoning tasks like AIME and MATH with only ~$9 post-training cost, demonstrating that efficient reasoning can be instilled via minimal updates to a tiny model.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack