TrustLLM (Trustworthiness in LLMs)
Free while signed in. Answers cite the passages they came from.

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.
Eight dimensions of trustworthiness: Principles span truthfulness, safety, fairness, robustness, privacy, machine ethics, transparency, and accountability.
Six-dimension benchmark: TrustLLM evaluates the first six dimensions across 30+ datasets, giving a single, comparable trustworthiness score per model.
16 LLMs compared: Tests both proprietary (GPT-4, Claude, PaLM 2) and open-source models (Llama 2, Vicuna, others), finding proprietary leads on average but open-source closing the gap on several dimensions.
Practitioner takeaway: Offers a standardized framework for evaluating trustworthiness claims, shifting the conversation from anecdote to measurable comparison.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack