🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

ShieldGemma

First page
ShieldGemma
Paper summary

offers a comprehensive suite of LLM-based safety content moderation models built on Gemma 2; includes classifiers for key harm types such as dangerous content, toxicity, hate speech, and more.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack