Benchmark results

Model performance

Score by model release

SOTA at release

Release date uses the model repository publication date on Hugging Face.

Rank Run

Reference

Cite the benchmark

Hager et al., Nature Medicine (2024). Evaluation and mitigation of the limitations of large language models in clinical decision-making.