Bench to Bed
The system

AI Tool Prof. Valmed Speeds Up Rheumatology Diagnosis

A randomized trial found the AI system Prof. Valmed cut diagnosis time by more than half compared to physicians working alone, but did not improve

The System: A randomized trial found the AI system Prof

The large language model Prof. Valmed helped physicians diagnose rheumatologic conditions more than twice as fast, but did not make them more accurate, according to a new trial. The study, led by Johannes Knitza, MD, PhD, MHBA, of Philipps-Universität Marburg in Germany, was reported in a preprint on medRxiv and has not yet undergone peer review.

Doctors using the AI system took an average of 94 seconds to reach a diagnosis across three case scenarios. Physicians relying on their own resources required 206 seconds per case, a statistically significant difference. Diagnostic accuracy, however, was nearly identical between the two groups.

How the Trial Was Conducted

The trial, named ALLIANCE, involved 82 physicians from seven institutions in Germany and Norway. They were randomly assigned to diagnose three clinical scenarios either with or without the support of Prof. Valmed. The cases, drawn from published reports, involved Cogan syndrome, dermatomyositis, and familial Mediterranean fever. Participants listed up to three possible diagnoses with probabilities and rated their confidence in each.

Prof. Valmed is an EU-certified AI system designed specifically for medical use. Founded by attorney Vera Roedel and neuroimmunologist Heinz Wiendl, MD, it received the CE mark in March 2025. The tool is billed as an "AI copilot" to assist, not replace, doctors. Only about a quarter of the participants in the trial were rheumatology specialists.

Accuracy and Confidence Findings

The primary measure of accuracy was whether a doctor's top-listed diagnosis matched the actual published diagnosis. This occurred in 33.3% of cases for the AI-assisted group and 35.0% for the conventional group, a non-significant difference. A less strict measure-having the correct diagnosis appear anywhere in the doctor's top three possibilities-also showed no significant benefit from the AI.

A notable finding was a disconnect between accuracy and confidence. "An interesting finding was that confidence exceeded observed accuracy across all evaluated settings and increased further after [Prof." the investigators wrote. Confidence ratings often far surpassed actual accuracy scores.

MetricAI-Assisted GroupConventional Group
Average Diagnosis Time94 seconds206 seconds
Top Diagnosis Accuracy33.3%35.0%
Accuracy in Top 3 Diagnoses49.2%39.2%

User Experience and Safety Concerns

Physicians generally gave Prof. Valmed positive reviews. Over 80% said they would use it again, and about two-thirds found it trustworthy. The system received high marks for ease of use and interface design. However, only 36% said it was easy to correct mistakes while using the tool.

The researchers highlighted potential safety issues related to over-reliance on the AI. "Exploratory analyses also suggested substantial AI over-reliance in the intervention group, whereas under-reliance was uncommon," they noted. This suggests diagnostic support may boost confidence more readily than correctness.

When operating alone without a doctor's input, Prof. Valmed's performance mirrored that of the assisted physicians. Its top diagnosis was correct 33.3% of the time, and the correct condition appeared in its top three suggestions 47.6% of the time. The study concluded that such AI tools may be most valuable for improving efficiency and broadening the range of considered diagnoses, but real-world trials are needed to define their role in routine care.

Related coverage

More from The system