🚀NEW LABGetting Started with Claude AgentsStart lab
Retrieval

ModernBERT

First page
ModernBERT
Paper summary

a new encoder-only transformer model that achieves state-of-the-art performance on classification and retrieval tasks while being more efficient than previous encoders; it was trained on 2T tokens with 8192 sequence length and incorporates modern optimizations that represent a significant improvement over BERT; the model is specifically designed for practical deployment, offering superior speed and memory efficiency on common GPUs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack