A patient-derived benchmark for evaluating large language models in connective tissue diseases: blinded multi-stakeholder assessment and guideline comparison
To co-develop disease-specific patient questions for connective tissue diseases (CTDs), compare patient/rheumatologist ratings of answers from language models (LLMs) versus Google Search, and quantify EULAR coverage of these patient-prioritized questions. In this prospective sin…