Health Interoperability
6 Min Read

Top LLM insights from FHIR DevDays 2025

Subscribe to our newsletter

Subscribe

We usually keep HL7 FHIR DevDays session recordings exclusive to attendees for six months before releasing them publicly on YouTube. But let’s face it: when it comes to generative AI, six months is a lifetime.

That’s why we’re unlocking three standout LLM sessions from DevDays 2025 ahead of schedule—because the insights are too good to keep under wraps.

Whether you’re building with LLMs, benchmarking their performance, or just exploring how GenAI fits into the FHIR ecosystem, these talks offer practical tips, real-world use cases, and valuable takeaways. Find our top takeaways and the full sessions below.

Unstructured data—like scanned lab reports—is a major interoperability hurdle. Max explores how GenAI can automate the transformation from messy text to structured FHIR resources. 

Two approaches are tested: 

  1. Manual FHIR mapping from extracted key-value pairs
  2. AI-generated resources based on predefined profiles

    Key takeaways

    • Reinforcement learning can improve AI’s performance on mapping to valid FHIR 
    • GenAI shows real promise in healthcare ETL workflows by automating data extraction and transformation— and could eventually replace ETL pipeline, shifting from assistant to end-to-end solution 
    • AI needs to act as both reader and formatter: interpreting unstructured inputs and generating valid FHIR 
    • Validation is still essential—AI can get close, but not perfect

    LLMs are everywhere—but how well do they actually perform on FHIR-specific tasks? 

    Joshua introduces an open-source community eval framework designed to measure how well LLMs handle key FHIR-related challenges in real-world implementations, and covers the benchmarking of two unique tasks:

    • FHIR “tool use” & $validate: Can an LLM use $validate to improve resource generation?
    • FHIRPath evaluation: How reliable are LLMs when handling complex FHIRPath logic? 

    Key takeaways

    • Use evals for scalable, repeatable testing of LLM performance on FHIR 
    • Detailed prompts and tool-calling are essential 
    • Some evaluations even use one LLM to score another 
    • Open-source benchmarks make results transparent and reproducible 
    • LLM performance is showing signs of saturation—we may be approaching a point where things “just work” for certain FHIR tasks

    Authoring jurisdictional IGs is notoriously complex. Alex shows how LLMs can streamline IG harmozation and profile generation, using the Canadian healthcare landscape as a case study. 

    He walks through building an AI stack for FHIR IG development and offers practical tips to avoid common pitfalls—especially validation errors. 

    Key takeaways

    • LLMs helped harmonize multiple IGs across Canada into a unified national standard 
    • Detailed prompts = better outputs 
    • LLMs excel at FSH generation and applying rule sets (less so at creating them)—Chris Moesel still beats the bots at debugging FSH errors 
    • Surprisingly good at generating ValueSets and FQL tables

    As the FHIR community dives deeper into GenAI, these early-access sessions offer a view of what’s working now, what still needs refinement, and how to make the most of these tools today. From benchmarking LLMs and authoring IGs to automating data extraction, the talks highlight practical strategies—grounded in real-world use cases. 

    Whether you’re already building with LLMs or just exploring what’s possible, these videos will give you a solid foundation to experiment, validate your approaches, and start building smarter solutions. 

    The rest of the DevDays 2025 recordings will be available in December, but in the meantime, you can catch up on all past sessions over on the DevDays YouTube channel

    Recommendations for you

    Explore more topics

    Post a comment

    Your email address will not be published. Required fields are marked *