🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Memory

LongLoRA

First page
LongLoRA
Paper summary

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.

Ask this paper

Key points
01

Shift short attention: Uses "shift short attention" during training, a pattern-shifted sparse approximation that mimics full attention while cutting cost.

02

LoRA-compatible: Works with standard LoRA, making it compatible with the existing parameter-efficient fine-tuning ecosystem.

03

Lower GPU cost: Dramatically reduces GPU memory and training time compared to full fine-tuning for context extension.

04

No accuracy compromise: Achieves comparable accuracy to full fine-tuning at extended context lengths, despite using a much cheaper approximation.

Every Monday
Get next week’s papers.
Subscribe on Substack