🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Safety

TrustLLM (Trustworthiness in LLMs)

First page
TrustLLM (Trustworthiness in LLMs)
Paper summary

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

Ask this paper

Key points
01

Eight dimensions of trustworthiness: Principles span truthfulness, safety, fairness, robustness, privacy, machine ethics, transparency, and accountability.

02

Six-dimension benchmark: TrustLLM evaluates the first six dimensions across 30+ datasets, giving a single, comparable trustworthiness score per model.

03

16 LLMs compared: Tests both proprietary (GPT-4, Claude, PaLM 2) and open-source models (Llama 2, Vicuna, others), finding proprietary leads on average but open-source closing the gap on several dimensions.

04

Practitioner takeaway: Offers a standardized framework for evaluating trustworthiness claims, shifting the conversation from anecdote to measurable comparison.

Every Monday
Get next week’s papers.
Subscribe on Substack