LongLoRA
Free while signed in. Answers cite the passages they came from.

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.
Shift short attention: Uses "shift short attention" during training, a pattern-shifted sparse approximation that mimics full attention while cutting cost.
LoRA-compatible: Works with standard LoRA, making it compatible with the existing parameter-efficient fine-tuning ecosystem.
Lower GPU cost: Dramatically reduces GPU memory and training time compared to full fine-tuning for context extension.
No accuracy compromise: Achieves comparable accuracy to full fine-tuning at extended context lengths, despite using a much cheaper approximation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack