Published research: benchmark, not demonstrated clinical impact
Arkangel AI: A conversational agent for real-time, evidence-based medical question-answeringAuthors as registered for the article: Maria Camila Villa; Natalia Castano-Villegas; Isabella Llano; Julian Martinez; Maria Fernanda Guevara; Jose Zea; Laura Velásquez.
The full article describes five LLMs and predefined retrieval and answer workflows, evaluated with MedQA and PubMedQA. Accuracy used the 1,273-question MedQA test set without fine-tuning. Retrieval and quality assessment used 127 MedQA questions (10%) and 500 Human-Evaluated PubMedQA questions (50%), with RAGAS context, relevance and faithfulness metrics, statistical tests and 95% confidence intervals.
- 90.26% accuracy on MedQA (1,273 questions).
- 80% retrieval of expected articles.
- 82% relevant answers on PubMedQA.
- Table 8: 401/500 expected abstracts retrieved (80.20%).
- Table 9: 414/500 relevant responses (82.80%).
These benchmarks evaluate the studied system, not patient outcomes, diagnostic accuracy in practice or a guarantee for the current service. External validation and real-world physician scenarios are still needed. Specialties are unevenly represented in MedQA. PubMedQA Table 9 also lists 86 non-relevant and 23 empty responses alongside 414 relevant responses; the categories require clarification before interpreting exclusions. The authors disclose relationships with Arkangel AI; this is not independent validation.
Reference checked on October 8, 2026: DOI metadata and the full 12-page editorial PDF publicly linked by Arkangel AI. The abstract rounds retrieval and relevance to 80% and 82%; Tables 8 and 9 report 80.20% and 82.80%.
Indexed published abstractFull editorial article (PDF)